Information processing device, information processing method, and recording medium
The information processing device enhances image-based authentication accuracy by converting and updating image parameters across different wavelength regions, addressing the challenge of inconsistent light sources in existing systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2025-01-16
- Publication Date
- 2026-07-23
AI Technical Summary
Existing image-based authentication systems face challenges in improving accuracy when using different combinations of light wavelength ranges for capturing registered and matching images, particularly when one is infrared and the other is visible light.
An information processing device and method that acquires images in different wavelength regions, converts one image using trained parameters based on features from another, and updates these parameters to enhance feature matching accuracy, regardless of wavelength combinations.
Improves the accuracy of image-based authentication by transforming images to extract features suitable for matching, irrespective of the wavelength ranges used for capture, thus enhancing robustness to lighting conditions and ensuring accurate authentication.
Smart Images

Figure JP2025001160_23072026_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and recording medium
[0001] This disclosure relates to an information processing device, an information processing method, and a recording medium.
[0002] For example, Patent Document 1 discloses a technique for authenticating a subject by converting a visible light image, which includes the iris of the subject, into a monochrome image, and then comparing the features extracted from the converted image with the features generated from an infrared image of the iris.
[0003] According to Patent Document 1, an image is selected from among multiple converted images obtained by converting a visible light image D1 to monochrome using different image conversion algorithms, to be used for extracting features for matching. The document states that the image used for extracting features is one in which the features used for iris authentication are clearly represented, specifically an image with a wide range from which features can be extracted.
[0004] International Publication No. 2021 / 192134
[0005] According to Patent Document 1, the accuracy of authentication can be improved when the registered image is an infrared image and the matching image is a visible light image. However, when the registered image is a visible light image and the matching image is an infrared image, it is difficult to improve the accuracy of authentication by applying the technology described for infrared images.
[0006] Furthermore, even when the registered image is an infrared image and the matching image is a visible light image, it is desirable to further improve the accuracy of authentication.
[0007] One of the challenges of this disclosure is to improve the accuracy of image-based authentication, regardless of the combination of light wavelength ranges used to capture each image being matched.
[0008] An information processing device in one aspect of this disclosure includes: a first image acquisition means for acquiring a first image obtained by capturing light in a first wavelength region; a second image acquisition means for acquiring a second image obtained by capturing light in a second wavelength region; an image conversion means for converting the second image to generate a converted image; a feature extraction means for extracting feature quantities from an image; and an update means for updating image conversion parameters for converting the second image based on a first feature quantity which is the feature quantity of the first image and a second feature quantity which is the feature quantity of the converted image.
[0009] An information processing device in one aspect of this disclosure includes: a second image acquisition means for acquiring a second image obtained by capturing light in a second wavelength region; an image conversion means for acquiring a converted image obtained by converting the second image; a feature extraction means for acquiring feature quantities of an image; and a matching means for comparing a first feature quantity, which is the feature quantity of a first image obtained by capturing light in a first wavelength region, with a second feature quantity, which is the feature quantity of the converted image. The image conversion means acquires a converted image obtained by converting the second image using image conversion parameters obtained in prior training using a first training feature quantity extracted from a first training image and a second training feature quantity extracted from the converted image of the second training image.
[0010] An information processing method in one aspect of this disclosure includes one or more computers acquiring a first image of light in a first wavelength region, acquiring a second image of light in a second wavelength region, transforming the second image to generate a transformed image, extracting feature quantities from the image, and updating image transformation parameters for transforming the second image based on a first feature quantity which is the feature quantity of the first image and a second feature quantity which is the feature quantity of the transformed image.
[0011] An information processing method in one aspect of this disclosure includes one or more computers acquiring a second image obtained by capturing light in a second wavelength region, acquiring a transformed image obtained by transforming the second image, acquiring image features, and comparing a first feature, which is the features of a first image obtained by capturing light in a first wavelength region, with a second feature, which is the features of the transformed image. In acquiring the transformed image, the computer acquires a transformed image obtained by transforming the second image using image transformation parameters obtained in prior training using a first training feature extracted from a first training image and a second training feature extracted from the transformed image of the second training image.
[0012] One aspect of this disclosure is a recording medium which contains a program that causes one or more computers to perform the following actions: acquire a first image by capturing light in a first wavelength region; acquire a second image by capturing light in a second wavelength region; convert the second image to generate a converted image; extract feature quantities from the image; and update image conversion parameters for converting the second image based on the first feature quantities, which are the feature quantities of the first image, and the second feature quantities, which are the feature quantities of the converted image.
[0013] In one aspect of this disclosure, the recording medium contains a program that causes one or more computers to acquire a second image obtained by capturing light in a second wavelength region, acquire a converted image obtained by transforming the second image, acquire image features, compare a first feature, which is the features of a first image obtained by capturing light in a first wavelength region, with a second feature, which is the features of the converted image, and acquire the converted image using image transformation parameters obtained in prior training using a first training feature extracted from a first training image and a second training feature extracted from the converted image of the second training image.
[0014] This is a block diagram showing an example configuration of the first information processing device according to this disclosure. This is a flowchart showing an example processing operation of the first information processing device according to this disclosure. This is a block diagram showing a detailed configuration example of the first information processing device according to this disclosure. This is a diagram showing an example of the solar spectrum on the ground. This is a block diagram showing an example configuration of the first update unit according to this disclosure. This is a flowchart showing an example processing operation of the first update unit according to this disclosure. This is a block diagram showing a physical configuration example of the first information processing device according to this disclosure. This is a block diagram showing an example configuration of the second information processing device according to this disclosure. This is a block diagram showing an example configuration of the second update unit according to this disclosure. This is a flowchart showing an example processing operation of the second information processing device according to this disclosure. This is a flowchart showing an example processing operation of the second update unit according to this disclosure. This is a block diagram showing an example configuration of the third information processing device according to this disclosure. This is a block diagram showing an example configuration of the third information processing device according to this disclosure. This is a flowchart showing an example processing operation of the third information processing device according to this disclosure. This is a block diagram showing an example configuration of the fourth information processing device according to this disclosure. This is a flowchart showing an example processing operation of the fourth information processing device according to this disclosure. This is a block diagram showing an example configuration of the first information processing system according to this disclosure. This is a block diagram showing a detailed configuration example of the fourth information processing device according to this disclosure. This is a flowchart showing a detailed processing operation example of the fourth information processing device according to this disclosure.
[0015] The embodiments relating to this disclosure will be described below with reference to the drawings. In this disclosure, the drawings are associated with one or more embodiments. In all drawings, similar components are denoted by the same reference numerals, and their descriptions are omitted where appropriate.
[0016] <Embodiment 1> As shown in Figure 1, the information processing device 100 according to this disclosure comprises a first image acquisition unit 110, a second image acquisition unit 120, an image conversion unit 130, a feature extraction unit 140, and an update unit 150.
[0017] The first image acquisition unit 110 acquires a first image captured of light in the first wavelength region.
[0018] The second image acquisition unit 120 acquires a second image obtained by photographing light in the second wavelength region.
[0019] The image conversion unit 130 converts the second image to generate a converted image.
[0020] The feature amount extraction unit 140 extracts the feature amount of the image.
[0021] The update unit 150 updates the image conversion parameters for converting the second image based on the first feature amount that is the feature amount of the first image and the second feature amount that is the feature amount of the converted image.
[0022] The information processing apparatus 100 executes the information processing shown in the flowchart of FIG. 2.
[0023] The first image acquisition unit 110 acquires a first image obtained by photographing light in the first wavelength region (step S110).
[0024] The second image acquisition unit 120 acquires a second image obtained by photographing light in the second wavelength region (step S120).
[0025] The image conversion unit 130 converts the second image to generate a converted image (step S130).
[0026] The feature amount extraction unit 140 extracts the feature amount of the image (step S140).
[0027] The update unit 150 updates the image conversion parameters for converting the second image based on the first feature amount that is the feature amount of the first image and the second feature amount that is the feature amount of the converted image (step S150). The information processing may be executed repeatedly.
[0028] According to this information processing apparatus 100, image conversion parameters for converting a second image are updated using feature amounts respectively derived from a first image and a second image. Therefore, the second image can be converted so that feature amounts suitable for matching the first image and the second image are extracted. Accordingly, it is possible to improve the accuracy of authentication using the first image and the second image regardless of the combination of the first wavelength region and the second wavelength region of light for capturing each of the first image and the second image to be matched. That is, it is possible to improve the accuracy of authentication using an image regardless of the combination of the wavelength regions of light for capturing each of the images to be matched.
[0029] Hereinafter, a detailed example of the information processing apparatus 100 will be described.
[0030] (Detailed Example) The information processing apparatus 100 is an apparatus that constructs an image conversion unit 130. As shown in FIG. 3, for example, the information processing apparatus 100 includes a training data storage unit 160 in addition to a first image acquisition unit 110, a second image acquisition unit 120, an image conversion unit 130, a feature amount extraction unit 140, and an update unit 150.
[0031] Note that the training data storage unit 160 may be provided inside the information processing apparatus 100 as shown in the figure, or may be provided in an external apparatus not shown. The external apparatus is, for example, an apparatus connected so as to be able to transmit and receive information to and from the training data storage unit 160 via a communication network configured by wire, wirelessly, or a combination thereof.
[0032] (Regarding the Training Data Storage Unit 160) The training data storage unit 160 is a storage unit for storing training data. The training data may be prepared in advance and stored in the training data storage unit 160. The training data includes a first image, a first label, a second image, and a second label.
[0033] The first image is an image captured with light in the first wavelength region as described above. The first label is a label for identifying the object included in the first image. The second image is an image captured with light in the second wavelength region as described above. The second label is a label for identifying the object included in the second image.
[0034] For example, the training data includes at least one pair of first and second images. For example, the first image is associated with the first label, and the second image is associated with the second label. If the first and second labels represent the same object, then the objects contained in the associated first and second images are the same. If the first and second labels represent different objects, then the objects contained in the associated first and second images are different.
[0035] The training data may include one or more combinations of the first and second images that contain the same subject. The training data may also include one or more combinations of the first and second images that contain different subjects.
[0036] (Regarding the example of the first and second images) Both the first and second images are, for example, images of a specific part of a target, and are used for biometric authentication of the target. When the first and second images are used for human facial recognition, the target is a person, and the specific part is the face. In this case, both the first and second images include the face of the person, which is the specific part of the target.
[0037] In other words, the first image is an image of a specific part of the target taken using light in the first wavelength range. The second image is an image of a specific part of the target taken using light in the second wavelength range. When the first and second images are used for human face recognition, the first image is an image of light in the first wavelength range traveling from the human face to the sensor. In this case, the second image is an image of light in the second wavelength range traveling from the human face to the sensor.
[0038] The wavelength range of light used for capturing the first image and the second image may be different. The first wavelength range and the second wavelength range may contain at least some different wavelength ranges. The first wavelength range and the second wavelength range may be different wavelength ranges from each other. The first wavelength range and the second wavelength range may be wavelength ranges in which one encompasses the other.
[0039] Each of the first wavelength region and the second wavelength region may include part or all of one or more of the following: (1) the wavelength range of visible light, (2) the wavelength range of short-wave infrared light, and (3) the wavelengths at which sunlight is blocked by the atmosphere.
[0040] (1) The wavelength range of visible light is, for example, 360 nm (nanometers) to 400 nm at the lower limit and 760 nm to 830 nm at the upper limit. Visible light images, which are captured using visible light, can be easily obtained using commonly available imaging equipment.
[0041] (2) The wavelength range of short-wave infrared radiation is, for example, 1000 nm or more and less than 2500 nm.
[0042] Figure 4 shows an example of the solar spectrum on the ground. Short-wave infrared radiation has a low energy density on the ground. By using short-wave infrared radiation for imaging, it is possible to obtain images that are less affected by backlighting, that is, images that are highly robust to the lighting environment during imaging.
[0043] (3) The wavelengths at which sunlight is blocked by the atmosphere are, for example, 940 nm, 1450 nm, and 1940 nm.
[0044] As shown in the figure, among the light emitted from the sun, light with wavelengths such as 940 nm, 1450 nm, and 1940 nm has a lower energy density on the ground than light with wavelengths in the vicinity of those wavelengths. This is because light with wavelengths of 940 nm, 1450 nm, and 1940 nm is blocked by the atmosphere before it reaches the ground.
[0045] By using light wavelengths that are blocked by the atmosphere when sunlight is captured, it is possible to obtain images that are less affected by backlighting, that is, images that are more robust to the lighting environment during shooting.
[0046] Here, the “wavelength region” in this disclosure includes not only a wavelength region of a predetermined width, but also a region consisting substantially of only one or more wavelengths. For example, each of the first wavelength region and the second wavelength region may consist only of wavelengths to which sunlight is blocked by the atmosphere. In this case, each of the first wavelength region and the second wavelength region may consist only of one wavelength to which sunlight is blocked by the atmosphere, or of two or more wavelengths to which sunlight is blocked by the atmosphere.
[0047] Furthermore, the uses of the first and second images are not limited to biometric authentication. Also, when the first and second images are used for biometric authentication, the biometric authentication is not limited to human facial recognition. The target is not limited to humans; for example, it may be an animal such as a dog, cow, pig, or snake. The specific body part is not limited to the face; for example, it may be the iris, veins, fingerprints, palm print, etc. The first wavelength region and the second wavelength region are not limited to the examples given above.
[0048] (Regarding the first image acquisition unit 110) The first image acquisition unit 110 acquires a first image from the training data storage unit 160, for example. The first image acquisition unit 110 may also acquire a first label along with the first image from the training data storage unit 160.
[0049] The source from which the first image acquisition unit 110 acquires the first image is not limited to the training data storage unit 160, but may also be a camera or other imaging device (not shown). Similarly, the source from which the first image acquisition unit 110 acquires the first label is not limited to the training data storage unit 160, but may also be the said imaging device (not shown), various devices used in conjunction with the said imaging device, etc.
[0050] (Regarding the second image acquisition unit 120) The second image acquisition unit 120 acquires a second image from the training data storage unit 160, for example. The second image acquisition unit 120 may also acquire a second label along with the second image from the training data storage unit 160.
[0051] Furthermore, the source from which the second image acquisition unit 120 acquires the second image is not limited to the training data storage unit 160, but may also be a camera or other imaging device (not shown). Also, the source from which the second image acquisition unit 120 acquires the second label is not limited to the training data storage unit 160, but may also be the said imaging device (not shown), various devices used in conjunction with the said imaging device, etc.
[0052] (Regarding the image conversion unit 130) As described above, the image conversion unit 130 converts the second image to generate a converted image.
[0053] For example, the image conversion unit 130 includes an image conversion model.
[0054] An image transformation model is a model that transforms an input image to generate a transformed image, and includes image transformation parameters for transforming the image. The image transformation parameters are typically multiple, but may be one.
[0055] An image transformation model is, for example, a machine learning model composed of a neural network. However, an image transformation model is not limited to machine learning models composed of neural networks; it may also be a model that applies general techniques used for image transformation.
[0056] The image input to the image conversion model includes a second image. The converted image may have a number of channels equal to or less than that of the second image, or it may have a number of channels greater than that of the second image. The number of channels is the number of channels in an image, which is the number of elements used to represent the color information of the image. For example, a monochrome image has 1 channel, and an RGB color image has 3 channels.
[0057] The image conversion unit 130 generates a converted image, for example, using the second image and an image conversion model. More specifically, the image conversion unit 130 may generate a converted image by inputting the second image into the image conversion model and converting the input second image.
[0058] (Regarding the feature extraction unit 140) The feature extraction unit 140 extracts features from the image as described above.
[0059] For example, the feature extraction unit 140 includes a feature extraction model. In this embodiment, we will describe an example in which the feature extraction unit 140 includes only one feature extraction model.
[0060] A feature extraction model is a model that extracts features from an input image, and includes feature extraction parameters for extracting image features. While there are typically multiple feature extraction parameters, there may be only one.
[0061] The feature extraction unit 140 extracts a first feature using, for example, the first image and a feature extraction model. The first feature is a feature of the first image. More specifically, the feature extraction unit 140 may extract the first feature from the input first image by inputting the first image into the feature extraction model.
[0062] The feature extraction unit 140 extracts a second feature using, for example, the transformed image and the feature extraction model. The second feature is a feature of the transformed image. More specifically, the feature extraction unit 140 may input the transformed image into the feature extraction model and extract the second feature from the input transformed image.
[0063] A feature extraction model is, for example, a machine learning model composed of a neural network. However, a feature extraction model is not limited to machine learning models composed of neural networks; it may also be a model that applies general techniques used for extracting features from images.
[0064] The feature extraction model may be a machine learning model that has been pre-trained to extract features from an image. In this case, the feature extraction parameters are already set. The first image may be used for training. That is, the features may be extracted using a machine learning model that has been pre-trained using the first image to extract features from an image. In addition, images other than the first image, such as the second image, and images other than the first and second images, may be used for training.
[0065] The image input to the feature extraction model includes the first image and the transformed image. Features are represented, for example, by an N-dimensional tensor, where N is a non-negative integer. For instance, 0-dimensional, 1-dimensional, and 2-dimensional tensors are scalars, vectors, and matrices, respectively.
[0066] Note that the training data may include the first feature instead of, or together with, the first image. In this case, the feature extraction unit 140 does not need to extract the first feature.
[0067] (Regarding the update unit 150) As described above, the update unit 150 updates the image transformation parameters based on the first feature and the second feature.
[0068] The update unit 150 may calculate a loss based on the first feature and the second feature, and update the image transformation parameters based on the loss. The update unit 150 may also update the image transformation parameters using the loss calculated from the first feature, the second feature, the first label, and the second label.
[0069] Figure 5 shows an example of the functional configuration of the update unit 150.
[0070] The update unit 150 includes a loss calculation unit 151 and a parameter update unit 152.
[0071] The loss calculation unit 151 calculates the loss based on the first feature and the second feature.
[0072] The parameter update unit 152 updates the image conversion parameters based on the calculated loss.
[0073] Figure 6 is a flowchart showing an example of the processing operation of the update unit 150, specifically a detailed example of the update process (step S150).
[0074] The loss calculation unit 151 calculates the loss based on the first feature and the second feature (step S151).
[0075] The parameter update unit 152 updates the image conversion parameters based on the calculated loss (step S152).
[0076] (Regarding the loss calculation unit 151) As described above, the loss calculation unit 151 calculates a loss based on the first feature amount and the second feature amount.
[0077] For example, the loss calculation unit 151 calculates a loss using the first feature amount, the second feature amount, the first label, the second label, and a predetermined loss function.
[0078] The loss function L1 used by the loss calculation unit 151 is represented by, for example, the following equation (1).
[0079] In equation (1), x 1 , x 2 are data to be compared for calculating the loss function L1, and are the first feature amount and the second feature amount, respectively. l becomes either 0 or 1 depending on the combination of the labels of the data x 1 , x 2 . When the labels of the data x 1 , x 2 match, that is, when the data x 1 , x 2 are derived from the same object, l becomes 1. When the labels of the data x 1 , x 2 do not match, that is, when the data x 1 , x 2 are derived from different objects, l becomes 0. m is a margin parameter for increasing the distance between data with different labels when the labels of x [[ID= thirty-five ]] 1 , x 2 are different. m is preset in order to train the image conversion model so that the distance between data with different labels is separated by a certain amount or more during training. m may be, for example, a value predetermined to be -1 or more and 1 or less.
[0080] The loss function L1 in equation (1) is an example of a function in which the loss becomes smaller as the first feature amount and the second feature amount derived from the same object are more similar, and the loss becomes larger as the first feature amount and the second feature amount derived from different objects are more similar. As a result, the first feature amount and the second feature amount derived from the same object can be brought closer together, and the first feature amount and the second feature amount derived from different objects can be separated.
[0081] Note that the loss function L1 is not limited to the one expressed by equation (1). For example, the loss function L1 may include only terms that bring the first and second features originating from the same object closer together, or it may include only terms that move the first and second features originating from different objects further apart.
[0082] For example, the loss function L1 may also be expressed by the following equation (2).
[0083] In equation (2), x target x is the data to be used in training. 1 is x target Same label data, x 2 is x target These are data with different labels. α is a parameter that sets the combined ratio of the first and second terms.
[0084] The similarity between the first and second features is a value corresponding to the degree of similarity between the first and second features. The similarity between the first and second features may be a large value as the similarity between the first and second features increases, and may be a value based on the L1 norm (e.g., a constant minus the L1 norm, the reciprocal of the L1 norm), cosine similarity, etc. Distance may be used instead of similarity, and the same applies below. The distance between the first and second features is a value corresponding to the degree of similarity between the first and second features, similar to the similarity. The distance between the first and second features may be a small value as the similarity between the first and second features increases, and may be, for example, the L1 norm. The similarity and distance between the first and second features are not limited to those exemplified herein.
[0085] The first and second features originating from the same subject means that the first image from which the first feature was extracted and the second image from which the second feature originates contain the same subject, for example, the same person.
[0086] The second image from which the second feature is derived is the second image before the transformation of the transformed image from which the second feature was extracted. The fact that the first image and the second image contain the same object may mean that the first image and the second image contain a specific part of the same object, for example, the face of the same person.
[0087] The first and second features originating from different subjects mean that the first image from which the first feature is extracted and the second image from which the second feature is derived contain different subjects, such as different people. The inclusion of different subjects in the first and second images may also mean that the first and second images contain specific parts of different subjects, such as the faces of different people.
[0088] Whether the first and second features originate from the same object or different objects may be determined depending on whether the first and second labels indicate the same object or different objects.
[0089] More specifically, it may be determined whether the first and second features originate from the same object or different objects based on the first label associated with the first image from which the first feature was extracted, and the second label associated with the second image from which the second feature originates. In this case, if the first and second labels indicate the same object, the first and second features are determined to originate from the same object. If the first and second labels indicate different objects, the first and second features are determined to originate from different objects.
[0090] (Regarding the parameter update unit 152) As described above, the parameter update unit 152 updates the image transformation parameters based on the calculated loss. This enables the training of the image transformation model.
[0091] For example, the parameter update unit 152 updates the image conversion parameters to reduce the loss.
[0092] As a result, if the first and second features used to calculate the loss originate from the same object, the image transformation parameters are updated to increase the similarity between the first and second features originating from the same object. In this case, an image transformation model can be constructed that transforms the second image so that the second feature becomes closer to the first feature in the feature space.
[0093] Here, the feature space is the space representing the first and second features, and the same applies below.
[0094] Furthermore, if the first and second features used to calculate the loss originate from different objects, the image transformation parameters are updated so that the similarity between the first and second features originating from different objects decreases. In this case, an image transformation model can be constructed that transforms the second image so that the second feature is farther from the first feature in the feature space.
[0095] In this way, the image transformation parameters and feature extraction parameters are updated using features derived from the first image and the second image, respectively. The image transformation parameters may be repeatedly updated using multiple combinations of the first and second images. As a result, the second image can be transformed so that features suitable for matching the first and second images are extracted, that is, features that can clearly distinguish whether or not the objects contained in the first and second images are the same. Therefore, it becomes possible to improve the accuracy of image-based authentication regardless of the combination of light wavelength ranges used to capture each image being matched.
[0096] (Example of physical configuration of the information processing device 100) Figure 7 shows an example of the physical configuration of the information processing device 100. Physically, the information processing device 100 includes, for example, a bus 1010, a processor 1020, a memory 1030, a storage device 1040, a network interface 1050, an input interface 1060, and an output interface 1070.
[0097] Bus 1010 is a data transmission path for the processor 1020, memory 1030, storage device 1040, network interface 1050, input interface 1060, and output interface 1070 to send and receive data to and from each other. However, the method of connecting the processor 1020 and the other components to each other is not limited to bus connection.
[0098] Processor 1020 is a processor implemented using components such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit).
[0099] Memory 1030 is a main memory device implemented as RAM (Random Access Memory), etc.
[0100] The storage device 1040 is an auxiliary storage device implemented as an HDD (Hard Disk Drive), SSD (Solid State Drive), memory card, or ROM (Read Only Memory). The storage device 1040 stores program modules for realizing the functions of the device equipped with it. The processor 1020 reads each of these program modules into the memory 1030 and executes them, thereby realizing the functions corresponding to those program modules.
[0101] The network interface 1050 is an interface for connecting a device equipped with it to a communication network.
[0102] The input interface 1060 is an interface for the user to input information. The input interface 1060 consists of, for example, a touch panel, a keyboard, a mouse, and the like.
[0103] The output interface 1070 is an interface for presenting information to the user. The output interface 1070 is composed of, for example, a liquid crystal panel, an organic EL (Electro-Luminescence) panel, and the like.
[0104] Embodiment 1 has been described above.
[0105] (Function and Effects) According to this embodiment, the information processing device 100 includes a first image acquisition unit 110, a second image acquisition unit 120, an image conversion unit 130, a feature extraction unit 140, and an update unit 150.
[0106] The first image acquisition unit 110 acquires a first image captured of light in the first wavelength region. The second image acquisition unit 120 acquires a second image captured of light in the second wavelength region. The image conversion unit 130 converts the second image to generate a converted image. The feature extraction unit 140 extracts the features of the image. The update unit 150 updates the image conversion parameters for converting the second image based on the first features, which are the features of the first image, and the second features, which are the features of the converted image.
[0107] This updates the image transformation parameters for transforming the second image using features derived from the first and second images, respectively. Therefore, the second image can be transformed in such a way that features suitable for matching the first and second images are extracted. Consequently, it becomes possible to improve the accuracy of image-based authentication regardless of the combination of light wavelength ranges used to capture each image being matched.
[0108] According to this embodiment, the update unit 150 may update the image transformation parameters using a loss calculated from a first feature, a second feature, a first label for identifying an object included in the first image, and a second label for identifying an object included in the second image.
[0109] This updates the image transformation parameters for transforming the second image using the features derived from the first and second images, as well as the first and second labels, respectively. As a result, the second image can be transformed so that features suitable for matching the first and second images are extracted, that is, features that can clearly distinguish whether the objects contained in the first and second images are the same or not. Consequently, it becomes possible to improve the accuracy of image-based authentication regardless of the combination of light wavelength ranges used to capture each of the images being matched.
[0110] According to this embodiment, the second image may be an image of light in a wavelength range where sunlight is blocked by the atmosphere.
[0111] This allows the use of a second image that is highly robust to the lighting conditions during capture. Therefore, regardless of the shooting environment, especially the lighting conditions, a second feature suitable for matching can be extracted. Consequently, regardless of the combination of light wavelength ranges used to capture each image being matched, the accuracy of image-based authentication can be further improved.
[0112] According to this embodiment, the light in the first wavelength region is visible light.
[0113] This allows the easily obtainable image to be transformed so that the second image can be accurately authenticated using the first image and the second image. Therefore, accurate authentication can be easily performed regardless of the combination of light wavelength ranges used to capture each of the images being compared.
[0114] According to this embodiment, the features are extracted using a machine learning model that has been pre-trained using the first image to extract image features.
[0115] This allows the second image, which is the other image used for authentication, to be transformed using a machine learning model built with the first image, which is one of the images used for authentication, so that authentication can be performed with high accuracy. Therefore, it becomes possible to easily perform accurate authentication regardless of the combination of light wavelength ranges used to capture each of the images being matched.
[0116] According to this embodiment, the converted image has more channels than the second image.
[0117] This allows the number of channels in the second image to be increased and transformed so that features suitable for matching the first and second images are extracted. Therefore, it becomes possible to improve the accuracy of image-based authentication regardless of the combination of light wavelength ranges used to capture each image being matched.
[0118] According to this embodiment, the first wavelength region and the second wavelength region include at least a portion of different wavelength regions.
[0119] This allows the second image to be transformed so that features suitable for matching the first and second images can be extracted, even if the wavelength ranges of light used to capture the first and second images are different. Therefore, it becomes possible to improve the accuracy of image-based authentication regardless of the combination of wavelength ranges of light used to capture each image being matched.
[0120] <Modification 1> In Embodiment 1, an example was described in which the training data included a first image, a second image, a first label, and a second label. The training data includes a first image and a second image, and it is possible to determine whether or not the objects included in the first image and the second image are the same, but it is not necessary to include one or both of the first label and the second label.
[0121] For example, the training data includes at least one pair of first and second images. The training data may include one or more combinations of first and second images that contain the same object. The training data may include one or more combinations of first and second images that contain different objects. Each pair of first and second images may be associated with information indicating whether or not the first and second images in that pair contain the same object.
[0122] The update unit 150 may update the image transformation parameters using a loss calculated only from the first and second features. Whether the first and second features originate from the same object or different objects can be determined depending on whether the first image from which the first feature was extracted and the second image from which the second feature originates are a combination of images containing both the same object and different objects.
[0123] In this case, the loss calculation unit 151 may calculate the loss based only on the first and second features.
[0124] According to this embodiment, the update unit 150 updates the image transformation parameters using a loss calculated only from the first and second feature quantities.
[0125] This also allows the second image to be transformed in a way that, similar to Embodiment 1, features suitable for matching the first and second images can be extracted. Therefore, it becomes possible to improve the accuracy of image-based authentication regardless of the combination of light wavelength ranges used to capture each of the images being matched.
[0126] <Embodiment 2> In Embodiment 1, we described an example in which the feature extraction model is a machine learning model that has been pre-trained using a first image to extract features from an image. In this embodiment, we will describe an example in which the feature extraction model is trained together with the image transformation model.
[0127] In this embodiment, in order to simplify the explanation, descriptions that overlap with other embodiments will be omitted as appropriate.
[0128] As shown in Figure 8, the information processing device 200 includes a first image acquisition unit 110, a second image acquisition unit 120, an image conversion unit 130, and a training data storage unit 160, similar to those in Embodiment 1, as well as a feature extraction unit 240 and an update unit 250.
[0129] (Regarding the feature extraction unit 240) The feature extraction unit 240 extracts image features in the same way as the feature extraction unit 140 in Embodiment 1. The feature extraction unit 240 includes one feature extraction model, similar to the feature extraction unit 140 in Embodiment 1. This feature extraction model is a machine learning model that has not yet been trained. Except for this point, the feature extraction model included in the feature extraction unit 240 may be the same as the feature extraction model described in Embodiment 1.
[0130] (Regarding the update unit 250) The update unit 250 updates the image transformation parameters based on the first feature and the second feature, similar to the update unit 150 in Embodiment 1. The update unit 250 further updates the feature extraction parameters for extracting features from the image based on the first feature and the second feature.
[0131] (Example of processing operation of the information processing device 200) The information processing device 200 executes the information processing shown in the flowchart of Figure 9.
[0132] Steps S110, S120, and S130 are performed in the same manner as in Embodiment 1.
[0133] The feature extraction unit 240 extracts features from the image (step S240).
[0134] The feature extraction unit 240 inputs the first image to the feature extraction model and extracts the first feature. The feature extraction unit 240 inputs the transformed image to the feature extraction model and extracts the second feature. The feature extraction model in this embodiment is a machine learning model that has not yet been fully trained.
[0135] The update unit 250 updates the image transformation parameters and feature extraction parameters based on the first and second features (step S250). The information processing may be performed repeatedly.
[0136] (Details of the update unit 250) As described above, the update unit 250 updates the image transformation parameters and feature extraction parameters based on the first and second features.
[0137] The update unit 250 may calculate a loss based on the first and second features and update the image transformation parameters and feature extraction parameters based on the calculated loss.
[0138] Figure 10 shows an example of the functional configuration of the update unit 250.
[0139] The update unit 250 includes a loss calculation unit 251 and a parameter update unit 252.
[0140] The loss calculation unit 251 calculates the loss based on the first feature and the second feature.
[0141] The parameter update unit 252 updates the image conversion parameters and feature extraction parameters based on the calculated loss.
[0142] Figure 11 is a flowchart showing a detailed example of the processing operation of the update unit 250, i.e., the update process (step S250).
[0143] The loss calculation unit 251 calculates the loss based on the first feature and the second feature (step S251).
[0144] The parameter update unit 252 updates the image transformation parameters and feature extraction parameters based on the calculated loss (step S252).
[0145] (Regarding the loss calculation unit 251) As described above, the loss calculation unit 251 calculates the loss based on the first feature and the second feature.
[0146] For example, the loss calculation unit 251 calculates the loss using the first feature, the second feature, the first label, the second label, and a predetermined loss function.
[0147] The loss function L2 used by the loss calculation unit 251 is expressed, for example, by the following equations (3) to (5).
[0148] The loss function L2 is a loss function for updating the image transformation parameters and feature extraction parameters, and is, for example, a function that integrates the first loss function L2a and the second loss function L2b. The loss function L2 is also a function in which the loss becomes smaller the more similar the first and second features originating from the same object are, and the loss becomes larger the more similar the first and second features originating from different objects are. Equation (3) shows an example of integrating the first loss function L2a and the second loss function L2b by adding them together.
[0149] Both the first loss function L2a and the second loss function L2b are functions in which the loss decreases as the first and second features originating from the same object are similar, and the loss increases as the first and second features originating from different objects are similar.
[0150] Equation (4) shows an example in which the loss function L1 expressed in equation (1) above is used as the first loss function L2a. In equation (4), x 1 , x 2 These are the data used for comparison to update the image transformation parameters, and are the first and second features, respectively. l is the data x 1 , x 2The combination of labels will result in either a value of 0 or 1. Data x 1 , x 2 If the labels match, i.e., data x 1 , x 2 If they originate from the same object, l will be 1. Data x 1 , x 2 If the labels do not match, i.e., data x 1 , x 2 If they originate from different objects, l is 0. m is, as described above, x 1 , x 2 This is a margin parameter used to increase the distance between data with different labels when the labels are different. m is pre-set to train the image transformation model to maintain a certain distance between data with different labels. m may be a predetermined value, for example, between -1 and 1.
[0151] Equation (5) shows an example in which a loss function called Arc Face (Additive Angular Margin Loss), etc., is used as the second loss function L2b. In equation (5), y i represents the correct label, indicating the label when the data being compared (i.e., the first and second features) originate from the same object. n represents the number of classes, i.e., the number of label variations in training. N represents the batch size, i.e., the number of data points used for parameter updates in a single training run. θ represents the angle between the data features (i.e., the first and second features) and the parameters of the feature extraction model in the feature space. 2 This represents the margin parameter applied when the data being compared (i.e., the first and second features) originate from different sources.
[0152] Here, we have described an example where the first loss function L2a and the second loss function L2b are different functions, and the data compared in each are the first and second features. However, the first loss function L2a and the second loss function L2b may be the same function.
[0153] Furthermore, the data compared in the first loss function L2a and the second loss function L2b may be different. For example, the first loss function L2 may compare the first feature and the second feature, while the second loss function L2b may compare the first image and the transformed image. In this case, the loss calculation unit 151 further uses the first image and the transformed image to calculate the loss.
[0154] If the first loss function L2a and the second loss function L2b are the same loss function and the data being compared are the same, the loss function L2 may include only either the first loss function L2a or the second loss function L2b.
[0155] The loss functions in equations (3) to (5) are examples of functions in which the loss decreases as the first and second features originating from the same object become similar, and the loss increases as the first and second features originating from different objects become similar. The loss function is not limited to those expressed in equations (3) to (5). For example, the loss function L1 expressed in equation (2) above may be used for the second loss function L1b. For example, softmax, SphereFace, CosFace, etc. may be used for the second loss function L2b.
[0156] (Regarding the parameter update unit 252) The parameter update unit 252 updates the image transformation parameters and feature extraction parameters based on the calculated loss. This allows for simultaneous training of the image transformation model and the feature extraction model.
[0157] For example, the parameter update unit 252 updates the image conversion parameters and feature extraction parameters so that the loss is reduced.
[0158] As a result, if the first and second features used to calculate the loss originate from the same object, the image transformation parameters and feature extraction parameters are updated to increase the similarity between the first and second features originating from the same object. In this case, an image transformation model can be constructed that transforms the second image so that the second feature is close to the first feature in the feature space. Along with this, a feature extraction model can be constructed that extracts features so that the first and second features are close to each other in the feature space.
[0159] Furthermore, if the first and second features used in calculating the loss originate from different objects, the image transformation parameters and feature extraction parameters are updated so that the similarity between the first and second features originating from different objects is reduced. In this case, an image transformation model can be constructed that transforms the second image so that the second feature is farther from the first feature in the feature space. Along with this, a feature extraction model can be constructed that extracts features so that the first and second features are farther from each other in the feature space.
[0160] In this way, the image transformation parameters and feature extraction parameters are updated using features derived from the first image and the second image, respectively. The image transformation parameters and feature extraction parameters may be repeatedly updated using multiple combinations of the first image and the second image. As a result, the image transformation model and the feature extraction model can be constructed in association so that features suitable for matching the first image and the second image are extracted, that is, features that can clearly distinguish whether or not the objects contained in the first image and the second image are the same.
[0161] In other words, the image transformation model is trained to generate a transformed image suitable for a simultaneously trained feature extraction model. Furthermore, the feature extraction model is trained to extract features suitable for the transformed image generated by the simultaneously trained image transformation model. "Suitable for the model" means that the first and second features originating from the same object are close in the feature space, while the first and second features originating from different objects are far apart in the feature space.
[0162] Therefore, regardless of the combination of light wavelength ranges used to capture each image being matched, it becomes possible to further improve the accuracy of image-based authentication.
[0163] The information processing device 200 may be physically configured in the same way as the information processing device 100.
[0164] Embodiment 2 has been described above.
[0165] (Function / Effect) According to this embodiment, the update unit 250 further updates the feature extraction parameters for extracting features from the image based on the first feature and the second feature.
[0166] This updates the image transformation parameters and feature extraction parameters for transforming the second image using features derived from the first and second images, respectively. Therefore, the image transformation model and feature extraction model can be constructed in association so that features suitable for matching the first and second images—that is, features that clearly distinguish whether the objects contained in the first and second images are the same—are extracted. Consequently, the accuracy of image-based authentication can be further improved regardless of the combination of light wavelength ranges used to capture each image being matched.
[0167] <Modification 2> In Embodiment 2, an example was described in which the feature extraction model is an untrained machine learning model. However, the feature extraction model may be a machine learning model that is further trained. That is, in this case, the feature extraction model may be trained together with the image transformation model as described in Embodiment 2, using the feature extraction parameters determined in prior training as initial values.
[0168] According to this modified example, the image transformation model and the feature extraction model can be constructed in association with each other so that features suitable for matching the first and second images can be extracted with less training data than when the feature extraction model is untrained, as described in Embodiment 2. Therefore, regardless of the combination of light wavelength ranges used to capture each of the images to be matched, it becomes possible to further easily improve the accuracy of authentication using images.
[0169] <Embodiment 3> Embodiment 1 described an example using one feature extraction model. This embodiment describes an example using multiple feature extraction models.
[0170] In this embodiment, in order to simplify the explanation, descriptions that overlap with other embodiments will be omitted as appropriate.
[0171] As shown in Figure 12, the information processing device 300 includes a first image acquisition unit 110, a second image acquisition unit 120, an image conversion unit 130, and a training data storage unit 160, similar to those in Embodiment 1, as well as a feature extraction unit 340 and an update unit 350.
[0172] (Regarding the feature extraction unit 340) The feature extraction unit 340 extracts multiple features from an image. The feature extraction unit 340 includes multiple feature extraction models.
[0173] The function of each of the multiple feature extraction models is the same as that of the feature extraction model described in Embodiment 1. That is, each feature extraction model is a model that extracts features from an input image.
[0174] For example, the feature extraction unit 340 takes the first image as input to each of the feature extraction models and extracts multiple first features from the input first image. The feature extraction unit 340 also takes a transformed image as input to each of the feature extraction models and extracts multiple second features from the input transformed image. As a result, multiple combinations of first and second features corresponding to each of the multiple feature extraction models are extracted from a set of first and transformed images.
[0175] Each of the feature extraction models may be a machine learning model that has been pre-trained to extract image features, similar to the feature extraction model described in Embodiment 1. In training, the first image may be used, or other images such as the second image, the first image, and the second image may be used.
[0176] Multiple feature extraction models are different models from each other.
[0177] For example, each feature extraction model may be a machine learning model composed of a neural network. In this case, the multiple feature extraction models may differ in the training data used for training, the loss function used for training, at least some of the layer structure that makes up the multiple feature extraction models, at least some of the parameters included in the multiple feature extraction models, etc.
[0178] (Regarding the update unit 350) The update unit 350 updates the image transformation parameters based on a plurality of first and second features. The image transformation parameters are updated based on a plurality of first features extracted from the first image using each of the plurality of feature extraction models, and a plurality of second features extracted from the transformed image using each of the plurality of feature extraction models.
[0179] The update unit 350 may calculate a loss based on a plurality of first and second feature quantities and update the image transformation parameters based on the loss.
[0180] Figure 13 shows an example of the functional configuration of the update unit 350.
[0181] The update unit 350 includes a loss calculation unit 351 and a parameter update unit 352.
[0182] (Regarding the loss calculation unit 351) The loss calculation unit 351 calculates the loss based on a plurality of first and second features.
[0183] For example, the loss calculation unit 351 calculates the loss using a plurality of first and second features, a first label, a second label, and a predetermined loss function.
[0184] The loss function used by the loss calculation unit 351 is a function in which the loss decreases as the similarity between the first and second features originating from the same object increases, and the loss increases as the similarity between the first and second features originating from different objects increases. The loss function used by the loss calculation unit 351 may be the same as the loss function L1 represented by equation (1) above. However, the loss function used by the loss calculation unit 351 is not limited to that represented by equation (1).
[0185] (Regarding the parameter update unit 352) The parameter update unit 352 updates the image transformation parameters based on the calculated loss. This enables the training of the image transformation model.
[0186] For example, the parameter update unit 352 updates the image conversion parameters to reduce the loss.
[0187] As a result, if multiple first and second features used to calculate the loss originate from the same object, the image transformation parameters are updated to increase the similarity between the multiple first and second features originating from the same object. In this case, an image transformation model can be constructed that transforms the second image so that the multiple second features in the feature space become closer to the multiple first features.
[0188] Furthermore, if the multiple first and second features used to calculate the loss originate from different objects, the image transformation parameters are updated so that the similarity between the multiple first and second features originating from different objects decreases. In this case, an image transformation model can be constructed that transforms the second image so that the multiple second features are farther from the multiple first features in the feature space.
[0189] In this way, the image transformation parameters are updated using multiple sets of features derived from the first and second images. The image transformation parameters and feature extraction parameters may be repeatedly updated using multiple combinations of the first and second images.
[0190] As a result, an image transformation model can be constructed that uses various feature extraction models to extract features suitable for matching the first and second images, that is, features that can clearly distinguish whether or not the objects contained in the first and second images are the same. Therefore, it becomes possible to further improve the accuracy of image-based authentication regardless of the combination of light wavelength ranges used to capture each of the images being matched.
[0191] (Example of processing operation of the information processing device 300) The information processing device 300 executes the information processing shown in the flowchart of Figure 14.
[0192] Steps S110, S120, and S130 are performed in the same manner as in Embodiment 1.
[0193] The feature extraction unit 340 extracts multiple features from the image (step S340).
[0194] The feature extraction unit 340 inputs the first image into multiple feature extraction models and extracts multiple first features. The feature extraction unit 340 inputs the transformed image into multiple feature extraction models and extracts multiple second features.
[0195] The update unit 350 updates the image transformation parameters based on a plurality of first and second feature quantities (step S350).
[0196] Figure 15 is a flowchart showing a detailed example of the processing operation of the update unit 350, i.e., the update process (step S350).
[0197] The loss calculation unit 351 calculates the loss based on a plurality of first and second features (step S351).
[0198] The parameter update unit 352 updates the image conversion parameters based on the calculated loss (step S352).
[0199] The parameter update unit 352 returns to the information processing shown in Figure 14. This information processing may be executed repeatedly.
[0200] The information processing device 300 may be physically configured in the same way as the information processing device 100.
[0201] Embodiment 3 has been described above.
[0202] (Function and Effects) According to this embodiment, the feature extraction unit 340 includes a plurality of feature extraction models for extracting feature quantities from an image. The update unit 350 updates the image conversion parameters based on a plurality of first feature quantities extracted from the first image using each of the plurality of feature extraction models, and a plurality of second feature quantities extracted from the converted image using each of the plurality of feature extraction models.
[0203] This updates the image transformation parameters for transforming the second image using multiple sets of features derived from the first and second images. Therefore, it is possible to construct an image transformation model that extracts features suitable for matching the first and second images using various feature extraction models, i.e., features that can clearly distinguish whether the objects contained in the first and second images are the same or not. Consequently, it becomes possible to further improve the accuracy of image-based authentication regardless of the combination of light wavelength ranges used to capture each of the images being matched.
[0204] <Modification 3> Embodiment 3 described an example in which each of the multiple feature extraction models is a pre-trained machine learning model. However, the multiple feature extraction models may include one or more machine learning models that have not yet been trained. Alternatively, the multiple feature extraction models may include one or more machine learning models that are further trained, as described in Modification 2.
[0205] In this case, the update unit 350 updates the feature extraction parameters along with the image transformation parameters based on one or more losses. The feature extraction parameters updated here are those included in a feature extraction model that has not yet been trained.
[0206] The multiple losses may include losses used in common for updating image transformation parameters and for updating feature extraction parameters included in each of the multiple feature extraction models, or they may include losses that are different from each other.
[0207] This updates the image transformation parameters and feature extraction parameters for transforming the second image using multiple sets of features derived from the first and second images. Therefore, it is possible to construct a feature extraction model and an image transformation model so that features suitable for matching the first and second images are extracted using various feature extraction models. Features suitable for matching the first and second images are, for example, features that can clearly distinguish whether the objects contained in the first and second images are the same or not. Consequently, it becomes possible to further improve the accuracy of image-based authentication regardless of the combination of light wavelength ranges used to capture each image being matched.
[0208] <Embodiment 4> Embodiments 1 to 3 described examples of constructing an image transformation model, or examples of constructing an image transformation model and a feature extraction model. In this embodiment, we will describe an example of performing authentication using images with the constructed image transformation model, or an image transformation model and a feature extraction model.
[0209] In this embodiment, in order to simplify the explanation, descriptions that overlap with other embodiments will be omitted as appropriate.
[0210] As shown in Figure 16, the information processing device 400 includes a second image acquisition unit 420, an image conversion unit 430, a feature extraction unit 440, and a matching unit 470.
[0211] The second image acquisition unit 420 acquires a second image captured from light in the second wavelength region.
[0212] The image conversion unit 430 obtains a converted image by converting the second image. The image conversion unit 430 obtains a converted image by converting the second image using image conversion parameters obtained from prior training using the first training feature quantity extracted from the first training image and the second training feature quantity extracted from the converted image of the second training image.
[0213] The feature extraction unit 440 acquires the features of the image.
[0214] The matching unit 470 compares the first feature quantity, which is a feature quantity of the first image captured in the first wavelength region, with the second feature quantity, which is a feature quantity of the converted image.
[0215] The information processing device 400 performs the information processing shown in the flowchart of Figure 17.
[0216] The second image acquisition unit 420 acquires a second image of light in the second wavelength region (step S420).
[0217] The image conversion unit 430 obtains a converted image by converting the second image (step S430). The image conversion unit 430 obtains a converted image by converting the second image using image conversion parameters obtained from prior training using the first training feature quantity extracted from the first training image and the second training feature quantity extracted from the converted image of the second training image.
[0218] The feature extraction unit 440 acquires the features of the image (step S440).
[0219] The matching unit 470 matches the first feature quantity, which is a feature quantity of the first image captured in the first wavelength region, with the second feature quantity, which is a feature quantity of the converted image (step S470).
[0220] According to this information processing device 400, the image transformation parameters for transforming the second image are updated using feature quantities derived from the first image and the second image, respectively. Therefore, the second image can be transformed so that feature quantities suitable for matching the first and second images are extracted. Consequently, it becomes possible to improve the accuracy of image-based authentication regardless of the combination of light wavelength ranges used to capture each of the images being matched.
[0221] The following describes a detailed example of the information processing device 400.
[0222] (Detailed Example) Figure 18 shows an example of the configuration of an information processing system SYS. The information processing system SYS comprises, for example, one or more imaging devices 401 and an information processing device 400.
[0223] Each of the imaging devices 401 is, for example, a device that captures light in the second wavelength region. Each of the imaging devices 401 is, for example, a camera that captures visible light, short-wave infrared light, light of wavelengths that sunlight is shielded by the atmosphere, etc. Note that each of the imaging devices 401 is not limited to the cameras exemplified here.
[0224] In this embodiment, we will explain using an example in which each of the imaging devices 401 is a camera that captures either short-wave infrared light or light of wavelengths that are blocked by the atmosphere when sunlight is present.
[0225] Each of the imaging devices 401 and the information processing device 400 are connected to each other via the communication network NT, and they send and receive information from each other via the communication network NT. The communication network NT is a communication network that is configured as wired, wireless, or a combination thereof.
[0226] The imaging device 401 may include lighting, which is not shown. The lighting emits light corresponding to the wavelength range of the light that the imaging device 401 captures.
[0227] Figure 19 shows a detailed example of the information processing device 400. The information processing device 400 comprises a first image acquisition unit 410, a second image acquisition unit 420, an image conversion unit 430, a feature extraction unit 440, a matching unit 470, and a registered information storage unit 480.
[0228] The first image acquisition unit 410 acquires a first image captured of light in the first wavelength region.
[0229] The registration information storage unit 480 is a storage unit that stores registration information in advance. The registration information includes, for example, a first image.
[0230] The registered information storage unit 480 may be provided in an external device (not shown). In this case, the external device is connected to the information processing device 400 via a communication network NT, and transmits and receives information to and from the information processing device 400 via the communication network NT.
[0231] The information processing device 400 may, in detail, perform the information processing shown in the flowchart of Figure 20.
[0232] Step S420 described above is performed.
[0233] The first image acquisition unit 410 acquires a first image captured of light in the first wavelength region (step S425).
[0234] Steps S430, S440, and S470 described above are performed.
[0235] (Regarding the first image acquisition unit 410) The first image acquisition unit 410 acquires a first image, similar to the first image acquisition unit 110 in Embodiment 1.
[0236] The first image is, for example, a biometric image that is pre-registered for the purpose of authenticating a subject. The first image acquisition unit 410 in this embodiment acquires the first image from the registration information storage unit 480. The first image may be a visible light image.
[0237] (Regarding the second image acquisition unit 420) The second image acquisition unit 420 acquires a second image, similar to the second image acquisition unit 120 in Embodiment 1.
[0238] The second image acquisition unit 420 acquires a second image from each of the one or more imaging devices 401.
[0239] The second image is, for example, a biometric image that is compared with the first image to authenticate the subject. The second image may be an image taken of either short-wave infrared light or light of wavelengths that are blocked by the atmosphere when sunlight is present.
[0240] The first image may be the first image described in Embodiment 1, and is not limited to a visible light image. The second image may be the second image described in Embodiment 1, and is not limited to an image of either short-wave infrared light or light of wavelengths that are blocked by the atmosphere from sunlight. The registration information may include the second image. In this case, the first image acquisition unit 410 may acquire the first image from each of one or more imaging devices 401. The second image acquisition unit 420 may acquire the second image from the registration information storage unit 480.
[0241] (Regarding the image conversion unit 430) As described above, the image conversion unit 430 obtains a converted image by converting the second image.
[0242] The image conversion unit 430 obtains a converted image by converting the second image using a trained image conversion model. Any of the methods described above may be used to train the image conversion model. That is, the image conversion unit 430 obtains a converted image by converting the second image using image conversion parameters obtained from prior training using the first training feature quantity extracted from the first training image and the second training feature quantity extracted from the converted image of the second training image. The first training image and the second training image correspond to the first and second images included in the training data described above, respectively.
[0243] The image conversion unit 430 may include a trained image conversion model. The image conversion unit 430 may input a second image to the trained image conversion model to generate a converted image.
[0244] The image conversion unit 430 does not necessarily have to include a trained image conversion model. In this case, the trained image conversion model may be provided, for example, in an external device (not shown). The image conversion unit 430 may then transmit a second image to the external device and obtain a converted image generated by the external device.
[0245] (Regarding the feature extraction unit 440) The feature extraction unit 440 acquires image features as described above.
[0246] The feature extraction unit 440 uses a trained feature extraction model to obtain the first and second features extracted from the first and second images, respectively. The trained feature extraction model may be a feature extraction model used to construct the image transformation model, a feature extraction model constructed by training together with the image transformation model, or a feature extraction model constructed by other methods.
[0247] The feature extraction unit 440 may include a pre-trained feature extraction model. The feature extraction unit 440 may input the first image and the second image to the pre-trained feature extraction model and extract the first and second features from them, respectively.
[0248] The feature extraction unit 440 does not necessarily have to include a trained feature extraction model. In this case, the trained feature extraction model may be provided, for example, in an external device (not shown). The feature extraction unit 440 may then transmit the first image and the second image to the external device and obtain the first and second features extracted from them, respectively, from the external device.
[0249] The registration information may include the first feature instead of the first image, or together with the first image. In this case, the first image acquisition unit 410 may acquire the first feature, which is a feature of the first image. The feature extraction unit 440 does not need to use a trained feature extraction model to acquire the first feature.
[0250] (Regarding the matching unit 470) As described above, the matching unit 470 matches the first feature quantity with the second feature quantity.
[0251] The matching unit 470 calculates, for example, the similarity between the first feature and the second feature, and outputs the matching result between the first feature and the second feature based on the similarity. The similarity is a value corresponding to the degree to which the first feature and the second feature are similar, and may be, for example, the L1 norm, cosine similarity, etc., as described above.
[0252] The matching unit 470 determines that the first and second features that were matched originate from the same object if the similarity is equal to or greater than a predetermined threshold. In this case, for example, the matching unit 470 outputs a matching result indicating that the objects are the same and that authentication was successful.
[0253] The matching unit 470 determines that the first and second features that were matched do not originate from the same object, i.e., they originate from different objects, if the similarity is below a predetermined threshold. In this case, for example, the matching unit 470 outputs a matching result indicating that the objects are not the same, are from different objects, or that authentication failed.
[0254] The information processing device 400 may be physically configured in the same way as the information processing device 100.
[0255] Embodiment 4 has been described above.
[0256] (Function and Effects) According to this embodiment, the information processing device 400 includes a second image acquisition unit 420, an image conversion unit 430, a feature extraction unit 440, and a matching unit 470.
[0257] The second image acquisition unit 420 acquires a second image captured from light in the second wavelength region.
[0258] The image conversion unit 430 obtains a converted image by converting the second image. The image conversion unit 430 obtains a converted image by converting the second image using image conversion parameters obtained from prior training using the first training feature quantity extracted from the first training image and the second training feature quantity extracted from the converted image of the second training image.
[0259] The feature extraction unit 440 acquires the features of the image.
[0260] The matching unit 470 compares the first feature quantity, which is a feature quantity of the first image captured in the first wavelength region, with the second feature quantity, which is a feature quantity of the converted image.
[0261] As a result, the second image is transformed using image transformation parameters obtained through prior training, which utilizes the first training feature extracted from the first training image and the second training feature extracted from the transformed image of the second training image. Therefore, accurate authentication becomes possible regardless of the combination of light wavelength ranges used to capture each image being matched.
[0262] (Function and Effect) According to this embodiment, the first image is a biometric image that is pre-registered for the purpose of authenticating the subject. The second image is a biometric image that is compared with the first image for the purpose of authenticating the subject.
[0263] As a result, the second image, which is matched with the first image to authenticate the subject, is transformed using image transformation parameters obtained through prior training using the first training feature extracted from the first training image and the second training feature extracted from the transformed image of the second training image. Therefore, accurate biometric authentication becomes possible regardless of the combination of light wavelength ranges used to capture each of the images being matched.
[0264] <Embodiment 5> Embodiment 4 described an example in which the second acquisition unit acquires only the second image. Now, we will describe an example in which the second acquisition unit acquires the first image in addition to the second image.
[0265] In this embodiment, the diagrams showing the functional and physical configuration of the information processing device, the processing operation of the information processing device, and an example of the configuration of the information processing system may be the same as those in Figures 16 to 20. In this embodiment, in order to simplify the explanation, explanations that overlap with other embodiments will be omitted as appropriate.
[0266] The multiple imaging devices 401 provided in the above-described information processing system SYS may include imaging devices that capture light in multiple different wavelength ranges, such as a camera that captures visible light, a camera that captures short-wave infrared light, and a camera that captures light of wavelengths that sunlight is shielded by the atmosphere. In other words, the imaging device 401 may include an imaging device that captures a first image and an imaging device that captures a second image.
[0267] The second image acquisition unit 420 acquires the first image and the second image. The second image acquisition unit 420 may acquire the first image and the second image from multiple imaging devices 401.
[0268] The image conversion unit 430 may determine whether the image acquired by the second image acquisition unit 420 is the second image. If the acquired image is the second image, the image conversion unit 430 may convert the second image to generate a converted image.
[0269] The feature extraction unit 440 acquires the features of a converted image when a converted image is generated. The feature extraction unit 440 acquires the features of a first image when a second image acquisition unit 420 acquires a first image. The feature extraction unit 440 acquires the features of a first image acquired by a first image acquisition unit 410.
[0270] The matching unit 470 compares the first feature quantity of the first image acquired by the first image acquisition unit 410 with the first feature quantity of the first image acquired by the second image acquisition unit 420 or the second feature quantity of the converted image.
[0271] (Function and Effects) According to this embodiment, the second image acquisition unit 420 acquires the first image and the second image. The image conversion unit 430 determines whether the image acquired by the second image acquisition unit 420 is the second image, and if the acquired image is the second image, it converts the second image to generate a converted image.
[0272] This allows for accurate authentication even if only the first or second image is acquired. Therefore, accurate biometric authentication becomes possible regardless of the combination of light wavelength ranges used to capture each image being matched.
[0273] Although this disclosure has been described above with reference to embodiments, this disclosure is not limited to the embodiments described above. Various modifications to the structure and details of this disclosure are possible, which can be understood by those skilled in the art within the scope of this disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0274] Furthermore, while the flowcharts used in the above description show multiple steps (processes) in sequence, the execution order of the steps performed in each embodiment is not limited to the order in which they are described. In each embodiment, the order of the illustrated steps can be changed to the extent that it does not impair the content.
[0275] Some or all of the above embodiments may also be described as follows, but are not limited to the following: 1. An information processing apparatus comprising: a first image acquisition means for acquiring a first image obtained by capturing light in a first wavelength region; a second image acquisition means for acquiring a second image obtained by capturing light in a second wavelength region; an image conversion means for converting the second image to generate a converted image; a feature extraction means for extracting feature quantities from an image; and an update means for updating image conversion parameters for converting the second image based on a first feature quantity which is the feature quantity of the first image and a second feature quantity which is the feature quantity of the converted image. 2. The information processing apparatus according to 1, wherein the update means updates the image conversion parameters using a loss calculated from the first feature quantity, the second feature quantity, a first label for identifying an object included in the first image, and a second label for identifying an object included in the second image. 3. The information processing apparatus according to 1, wherein the update means updates the image conversion parameters using a loss calculated only from the first feature quantity and the second feature quantity. 4. The information processing apparatus according to any one of 1 to 3, wherein the second image is an image of light in the wavelength range where sunlight is blocked by the atmosphere. 5. The light in the first wavelength range is visible light. The information processing apparatus according to any one of 4. 6. The information processing apparatus according to any one of 1 to 5, wherein the feature extraction means includes a plurality of feature extraction models for extracting features from an image, and the update means updates the image transformation parameters based on a plurality of first features extracted from the first image using each of the plurality of feature extraction models and a plurality of second features extracted from the transformed image using each of the plurality of feature extraction models. 7. The information processing apparatus according to any one of 1 to 6, wherein the features are extracted using a machine learning model that has been pre-trained using the first image to extract features from the image. 8. The information processing apparatus according to any one of 1 to 6, wherein the update means further updates the feature extraction parameters for extracting the features from the image based on the first features and the second features.9. The information processing apparatus according to any one of 1 to 8, wherein the converted image has more channels than the second image. 10. The information processing apparatus according to any one of 1 to 9, wherein the first wavelength region and the second wavelength region include at least partially different wavelength regions. 11. The information processing apparatus comprising: a second image acquisition means for acquiring a second image obtained by photographing light in the second wavelength region; an image conversion means for acquiring a converted image obtained by converting the second image; a feature extraction means for acquiring feature quantities of an image; and a matching means for comparing a first feature quantity, which is the feature quantity of the first image obtained by photographing light in the first wavelength region, with a second feature quantity, which is the feature quantity of the converted image, wherein the image conversion means acquires a converted image obtained by converting the second image using image conversion parameters obtained in prior training using a first training feature quantity extracted from a first training image and a second training feature quantity extracted from the converted image of the second training image. 12. 11. The information processing apparatus described in 13. The first image is a biometric image registered in advance for the purpose of authenticating a subject, and the second image is a biometric image that is matched with the first image for the purpose of authenticating a subject. 13. An information processing method comprising one or more computers acquiring a first image obtained by capturing light in a first wavelength region, acquiring a second image obtained by capturing light in a second wavelength region, converting the second image to generate a converted image, extracting feature quantities from the image, and updating image conversion parameters for converting the second image based on a first feature quantity which is the feature quantity of the first image and a second feature quantity which is the feature quantity of the converted image. 14. The information processing method described in 13. The image conversion parameters are updated using a loss calculated from the first feature quantity, the second feature quantity, a first label for identifying an object included in the first image, and a second label for identifying an object included in the second image. 15. The image conversion parameters are updated using a loss calculated only from the first feature quantity and the second feature quantity. The information processing method described above.16. The second image is an image of light in the wavelength range where sunlight is blocked by the atmosphere. The information processing method according to any one of 13 to 15. 17. The light in the first wavelength range is visible light. The information processing method according to 16. 18. The feature quantities are extracted using a plurality of feature extraction models for extracting feature quantities from the image, and the image transformation parameters are updated based on a plurality of first feature quantities extracted from the first image using each of the plurality of feature extraction models and a plurality of second feature quantities extracted from the transformed image using each of the plurality of feature extraction models. The information processing method according to any one of 13 to 17. 19. The feature quantities are extracted using a machine learning model that has been pre-trained using the first image to extract feature quantities from the image. The information processing method according to any one of 13 to 18. 20. The information processing method according to any one of 13 to 18, wherein updating the image conversion parameters further updates the feature extraction parameters for extracting the features from the image based on the first and second features. 21. The information processing method according to any one of 13 to 20, wherein the converted image has more channels than the second image. 22. The information processing method according to any one of 13 to 21, wherein the first wavelength region and the second wavelength region include at least partially different wavelength regions. 23. A program for causing one or more computers to execute the information processing method according to any one of 13 to 21. 24. A recording medium on which the program according to 23 is recorded.25. An information processing method comprising: one or more computers acquiring a second image obtained by capturing light in a second wavelength region; acquiring a converted image obtained by transforming the second image; acquiring image features; and comparing a first feature, which is the feature of a first image obtained by capturing light in a first wavelength region, with a second feature, which is the feature of the converted image, wherein the acquisition of the converted image is performed using image transformation parameters obtained in prior training using a first training feature extracted from a first training image and a second training feature extracted from the converted image of the second training image, thereby acquiring a converted image obtained by transforming the second image. 26. The information processing method according to 25., wherein the first image is a biometric image registered in advance for the purpose of authenticating a subject, and the second image is a biometric image that is compared with the first image for the purpose of authenticating a subject. 27. A program for causing one or more computers to execute the information processing method according to 25. or 26. 28. A recording medium on which the program according to 27. is recorded.
[0276] 100, 200, 300, 400 Information Processing Unit 110, 410 First Image Acquisition Unit 120, 420 Second Image Acquisition Unit 130, 430 Image Conversion Unit 140, 240, 340, 440 Feature Extraction Unit 150, 250, 350 Update Unit 151, 251, 351 Loss Calculation Unit 152, 252, 352 Parameter Update Unit 160 Training Data Storage Unit 470 Matching Unit 480 Registered Information Storage Unit SYS Information Processing System
Claims
1. An information processing apparatus comprising: a first image acquisition means for acquiring a first image obtained by capturing light in a first wavelength region; a second image acquisition means for acquiring a second image obtained by capturing light in a second wavelength region; an image conversion means for converting the second image to generate a converted image; a feature extraction means for extracting feature quantities from an image; and an update means for updating image conversion parameters for converting the second image based on a first feature quantity which is the feature quantity of the first image and a second feature quantity which is the feature quantity of the converted image.
2. The information processing apparatus according to claim 1, wherein the updating means updates the image transformation parameters using a loss calculated from the first feature, the second feature, a first label for identifying an object included in the first image, and a second label for identifying an object included in the second image.
3. The information processing apparatus according to claim 1 or 2, wherein the second image is an image of light in a wavelength range where sunlight is blocked by the atmosphere.
4. The information processing apparatus according to claim 3, wherein the light in the first wavelength region is visible light.
5. The information processing apparatus according to any one of claims 1 to 4, wherein the feature extraction means includes a plurality of feature extraction models for extracting features from an image, and the update means updates the image conversion parameters based on a plurality of first features extracted from the first image using each of the plurality of feature extraction models and a plurality of second features extracted from the converted image using each of the plurality of feature extraction models.
6. The information processing apparatus according to any one of claims 1 to 5, wherein the features are extracted using a machine learning model that has been pre-trained using the first image to extract the features of the image.
7. The information processing apparatus according to any one of claims 1 to 6, wherein the updating means further updates feature extraction parameters for extracting the feature from the image based on the first feature and the second feature.
8. The information processing apparatus according to any one of claims 1 to 7, wherein the converted image has more channels than the second image.
9. The information processing apparatus according to any one of claims 1 to 8, wherein the first wavelength region and the second wavelength region include at least a portion of different wavelength regions.
10. An information processing device comprising: a second image acquisition means for acquiring a second image obtained by capturing light in a second wavelength region; an image conversion means for acquiring a converted image obtained by converting the second image; a feature extraction means for acquiring feature quantities of an image; and a matching means for comparing a first feature quantity, which is the feature quantity of a first image obtained by capturing light in a first wavelength region, with a second feature quantity, which is the feature quantity of the converted image, wherein the image conversion means acquires a converted image obtained by converting the second image using image conversion parameters obtained in prior training using a first training feature quantity extracted from a first training image and a second training feature quantity extracted from the converted image of the second training image.
11. The information processing apparatus according to claim 10, wherein the first image is a biometric image registered in advance for the purpose of authenticating a subject, and the second image is a biometric image that is compared with the first image for the purpose of authenticating a subject.
12. An information processing method comprising: one or more computers acquiring a first image captured of light in a first wavelength region; acquiring a second image captured of light in a second wavelength region; transforming the second image to generate a transformed image; extracting feature quantities from the image; and updating image transformation parameters for transforming the second image based on the first feature quantities, which are the feature quantities of the first image, and the second feature quantities, which are the feature quantities of the transformed image.
13. A recording medium on which a program is stored that causes one or more computers to perform the following actions: acquire a first image by capturing light in a first wavelength region; acquire a second image by capturing light in a second wavelength region; convert the second image to generate a converted image; extract feature quantities from the image; and update image conversion parameters for converting the second image based on the first feature quantities, which are the feature quantities of the first image, and the second feature quantities, which are the feature quantities of the converted image.
14. An information processing method comprising: one or more computers acquiring a second image obtained by capturing light in a second wavelength region; acquiring a transformed image obtained by transforming the second image; acquiring image features; and comparing a first feature, which is the features of a first image obtained by capturing light in a first wavelength region, with a second feature, which is the features of the transformed image, wherein the acquisition of the transformed image is performed using image transformation parameters obtained in prior training using a first training feature extracted from a first training image and a second training feature extracted from the transformed image of the second training image, thereby acquiring a transformed image obtained by transforming the second image.
15. A recording medium on which a program is recorded that causes one or more computers to acquire a second image obtained by capturing light in a second wavelength region, acquire a converted image obtained by transforming the second image, acquire image features, compare a first feature, which is the features of a first image obtained by capturing light in a first wavelength region, with a second feature, which is the features of the converted image, and acquire a converted image, wherein the program acquires a converted image obtained by transforming the second image using image transformation parameters obtained in prior training using a first training feature extracted from a first training image and a second training feature extracted from the converted image of the second training image.