Image processing device, image processing method, and program

JP7927467B2Active Publication Date: 2026-10-01CANON KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022096455
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2026-10-01
Estimated Expiration
2042-06-15

Smart Images

  • Figure 0007927467000004
    Figure 0007927467000004
  • Figure 0007927467000005
    Figure 0007927467000005
  • Figure 0007927467000006
    Figure 0007927467000006
Patent Text Reader

Abstract

To provide an image processing device that suppresses different persons from being recognized as the same person.SOLUTION: Provided is an image processing system that recognizes a person included in an input image. A recognition device (image recognition device) 103 has: a face image acquisition unit that acquires registered images of a plurality of persons; a face feature extraction unit 301 that generates a feature map corresponding to each learned model on the basis of a plurality of learned models for extracting a feature for each person from the image; a feature extractor selection unit 302 that selects the learned model on the basis of each quality value for the learned model determined based on the feature map; and a collation unit 303 that recognizes a person included in the input image on the basis of the selected learned model.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an individual recognition technique using images. [Background Art]

[0002] An image recognition apparatus that determines whether a person captured in an image is the same person as a registered person generally uses a trained model obtained by training on human features. In Patent Document 1, the trained model is used to determine whether the person is the same based on the features extracted from the face image. [Prior Art Documents] [Patent Documents]

[0003] [Patent Document 1] Japanese Patent No. 6003124 [Summary of the Invention] [Problems to be Solved by the Invention]

[0004] However, in the method of Patent Document 1, when different pre-registered persons look similar to each other, there is a possibility that two different persons may be erroneously determined to be the same person. The present invention has been made in view of the above problem, and an object of the present invention is to suppress recognition of different persons as the same person. [Means for Solving the Problems]

[0005] An image processing apparatus according to the present invention for solving the above problem is an image processing apparatus for recognizing a person included in an input image, comprising: an acquisition unit that acquires registered images of a plurality of persons; a generation unit that generates feature maps corresponding to each of a plurality of trained models based on the plurality of trained models for extracting features for each person from an image; a selection unit that selects a trained model based on each quality value for the trained model determined based on the feature maps; and a recognition unit that recognizes a person included in the input image based on the selected trained model. [Effects of the Invention]

[0006] According to the present invention, it is possible to suppress the recognition of different people as the same person. [Brief explanation of the drawing]

[0007] [Figure 1] Block diagram showing an example of an image processing system configuration. [Figure 2] Block diagram showing an example of the hardware configuration of an image processing device. [Figure 3] Block diagram showing an example of the functional configuration of an image processing device. [Figure 4] A flowchart to explain the processes performed by an image processing device. [Figure 5] A flowchart to explain the processes performed by an image processing device. [Figure 6] Block diagram showing an example of the functional configuration of an image processing device. [Figure 7] A flowchart to explain the processes performed by an image processing device. [Figure 8] A flowchart to explain the processes performed by an image processing device. [Figure 9] A diagram illustrating an example of a feature map. [Figure 10] A diagram illustrating an example of a feature map. [Figure 11] A diagram illustrating an example of display by an output device. [Figure 12] A diagram illustrating an example of display by an output device. [Modes for carrying out the invention]

[0008] The present invention will now be described in detail below with reference to the attached drawings, based on preferred embodiments. Note that the configurations shown in the following embodiments are merely examples, and the present invention is not limited to the illustrated configurations.

[0009] <Embodiment 1> In image-based face recognition, variations in accuracy when detecting facial regions from images, as well as factors such as lighting conditions and changes in facial expression, can lead to some variation in features even when the images are of the same person. Therefore, when multiple registered individuals with high similarity exist, the registered individual with the highest similarity may not necessarily be the correct recognition result, leading to a decrease in reliability. For example, if a given input image contains registered individual A with a similarity of 99% and registered individual B with a similarity of 98%, the recognition result might output registered individual A when in reality it was registered individual B.

[0010] This embodiment describes an image processing device for face recognition that takes multiple registered images and multiple feature extractors as input and includes a feature extractor selection unit that selects a feature extractor based on the results of extracting features from the registered images.

[0011] <Configuration of the image processing system> Below, an example of the configuration of the image processing system 1 according to this embodiment is shown with reference to Figure 1.

[0012] Image processing system 1 comprises a learning device 101, an input device 102, a recognition device 103, an output device 104, a learning data storage device 105, a feature extractor storage device 106, and a registered image storage device 107. All of these devices are connected to a network, allowing them to send and receive data from each other. Here, learning data refers to images of people's faces, each assigned a different ID. A feature extractor (learning model) is trained using this learning data. Registered images are individual face images of multiple people. Test images are face images of specific individuals of interest. In other words, test images refer to images taken using the recognition device of a person of particular interest (a person of interest) that is specifically targeted for identification.

[0013] The learning device 101 trains one or more feature extractors (learning models) based on learning data. Here, the trained feature extractor (trained model) is a deep neural network that has been trained to extract features for identifying a person from an image. The learning device receives a feature extractor from the feature extractor storage device 106 and receives learning data from the learning data storage device 105. Furthermore, a plurality of new feature extractors are generated by training a reference feature extractor using additional learning data.

[0014] The input device 102 inputs an image captured by an imaging device to each device. Specifically, the input device 102 images a target person using a camera and transmits the image to the recognition device 103. In the present embodiment, it is assumed that the imaging device is a monocular color camera, and the captured image is a color image. However, the camera may be a monochrome camera in addition to a color camera. For example, the camera may be a grayscale camera, an infrared camera, a wide-angle lens camera, or a panoramic camera. It may also be a camera capable of panning, tilting, and zooming.

[0015] The recognition device 103 recognizes a person included in an input image using a feature extractor. In order to improve the success probability of identification corresponding to the target person in particular, the recognition device according to the present embodiment performs recognition based on a feature extractor suitable for recognition corresponding to the target person among the plurality of feature extractors. The recognition device 103 receives a plurality of feature extractors from the learning device 101 and receives a test image from the input device 102. The recognition device also receives a registered image from the registered image storage device 107. Furthermore, the recognition device performs recognition processing and transmits a recognition result to the output device 104.

[0016] The output device 104 displays, for example, an input image, a registered image, or a feature map of the feature extractor on a display device such as a display. The output device 104 receives a test image from the input device 102 and receives a registered image from the registered image storage device 107. The output device also receives a recognition result from the recognition device 103. Furthermore, the output device displays the received image and the recognition result on the display. The display device may be a projector in addition to a display.

[0017] <Description of Recognition Apparatus> Figure 2 is a diagram showing an example of the hardware configuration of a recognition apparatus (image recognition apparatus). A central processing unit (CPU) 201 uses a RAM 203 as a work memory to read and execute an OS and other programs stored in a ROM 202 or a storage device 204, controls each component connected to a system bus 209, and performs arithmetic operations, logical judgments, and the like for various types of processing. The processing executed by the CPU 201 includes the image recognition processing of the embodiment. The storage device 204 is a hard disk drive, an external storage device, or the like, and stores programs and various data related to the image recognition processing of the embodiment. An input unit 205 is an input device such as a mouse, a keyboard, or a touch panel for inputting user instructions. The storage device 204 is connected to the system bus 209 via an interface such as SATA, and the input unit 205 is connected to the system bus 209 via a serial bus such as USB, respectively, and detailed descriptions thereof are omitted. A communication I / F 206 communicates with external devices via wireless or wired communication.

[0018] Hereinafter, a configuration example of the recognition apparatus according to the present embodiment will be described with reference to FIG. 3. A recognition apparatus (image recognition apparatus) 103 includes a face image acquisition unit 300, a face feature extraction unit 301, a feature extractor selection unit 302, and a matching unit 303. The recognition apparatus (image recognition apparatus) 103 acquires a registered image from a registered image storage device 107. The feature extractor selection unit 302 of the recognition apparatus 103 acquires a feature extractor from a feature extractor storage device 106. The recognition apparatus 103 is connected to an input device 102 and an output device 104. The recognition apparatus 103 acquires an image from the input device 102. The recognition apparatus 103 outputs a matching result to the output device 104.

[0019] The face image acquisition unit 300 acquires face images by detecting human faces from images. It receives registered images and test images input from an external source. Furthermore, it detects the facial regions of human faces from the received registered images and test images and transmits the detected facial regions to the face feature extraction unit 301. The facial region is indicated by its position, size, etc., in the image. Existing face detection methods can be used to detect the facial region. For example, it can be detected using a trained model that estimates a rectangular region containing a face from the input image, or by matching with a template image that mimics the color or shape of a face. If the registered image and test image are already partial images of faces, the processing of the face image acquisition unit 300 may be skipped.

[0020] The face feature extraction unit 301 extracts face features (feature vectors) by inputting face images into a feature extractor (trained model) that extracts features for each person from an image. The feature extractor (trained model) is a Deep Convolutional Neural Network (hereinafter referred to as DNN). Specifically, first, the face feature extraction unit 301 receives multiple feature extractors input from the feature extractor memory device 106. It also receives the facial region of a person transmitted from the face image acquisition unit 300. Next, it extracts features from the facial region of a person using each of the multiple feature extractors. For example, if there are M multiple feature extractors and N facial regions received, M × N features are extracted. Furthermore, the features extracted from the registered image are sent to the feature extractor selection unit 302, and the features extracted from the test image are sent to the matching unit 303. The trained model refers to a network structure based on a neural network that outputs features corresponding to the input image from the input image, and its parameters. Multiple feature extractors each have different parameter sets or network structures. Neural networks are generated using existing deep learning techniques such as convolutional neural networks.

[0021] The feature extractor selection unit 302 determines a quality value for each trained model based on the features extracted by inputting the registered image into each of the multiple feature extractors and the feature map. Based on the quality value, the feature extractor selection unit 302 selects a trained model for recognizing a specific person. The feature extractor selection unit 302 receives multiple feature extractors input from an external source. It also receives features for the face region of the registered image transmitted from the face feature extraction unit 301. Furthermore, based on the received features, it selects one feature extractor from the received multiple feature extractors and transmits the features for the face region extracted by the selected feature extractor to the matching unit 303. For example, if there are M feature extractors, it selects one feature extractor from the M. Furthermore, it transmits 1 × N features extracted by the selected feature extractor to the matching unit 303. Here, the face feature extraction unit 301 generates a feature map corresponding to each feature extractor based on the multiple feature extractors (trained models). A feature map is a map that shows the distribution of each person's features in the feature space corresponding to a feature extractor. In other words, for the Ith feature extractor, a feature map is a map that plots the N features representing each of the N people in the same feature space.

[0022] The matching unit 303 recognizes a person included in the input image based on the selected feature extractor. The matching unit 303 receives features of the registered image transmitted from the feature extractor selection unit 302. It also receives features for the face region of the test image transmitted from the face feature extraction unit 301. Furthermore, it compares the features of the registered image and the test image and calculates the ID of the registered image that is thought to correspond to the test image. Alternatively, if there is no registered image that is thought to correspond, it calculates the information "No matching ID". Finally, it outputs the calculated registered image ID or the information "No matching ID" as the recognition result to the output device 104.

[0023] <Image recognition method using a recognition device> In the following, an example of an image recognition method performed by the recognition device 103 according to this embodiment is shown with reference to Figure 4. The processes shown in the flowcharts in Figure 4 and other figures are executed by the CPU 201 of Figure 2, which is a computer, according to the computer program stored in the storage device 204. In the following description, each process (step) is indicated by prefixing it with S, and the notation of the process (step) is omitted.

[0024] Note that the following flowchart assumes the feature extractor has been pre-trained. The method for training the feature extractor will be described later.

[0025] In S400, the face image acquisition unit 300 acquires a face image by detecting a person's face from the registered image.

[0026] In S401, the face feature extraction unit 301 extracts face features (feature vectors) by inputting the face image into a feature extractor (a trained model) that extracts features for each person from the image. The face feature extraction unit 301 extracts features for each face region obtained in S400 using each of the multiple feature extractors. In this example, there are four feature extractors: DNN_A, DNN_B, DNN_C, and DNN_D, and features for all face regions are extracted using these four feature extractors. In this example, a one-dimensional vector consisting of 512 numerical values ​​is used as the features, but one-dimensional vectors with different numbers of numerical values ​​or multi-dimensional vectors may also be used.

[0027] In S402, the feature extractor selection unit 302 generates feature maps corresponding to each feature extractor using the facial region features extracted from the registered image. In this embodiment, an example of selecting one feature extractor from four DNN_A, DNN_B, DNN_C, and DNN_D is described. The learning methods for each feature extractor will be described later. The feature extractor selection unit 302 generates four feature maps MAP_A, MAP_B, MAP_C, and MAP_D corresponding to these four feature extractors. Here, a feature map is a collection of data in which the facial region features extracted from all registered images extracted by each feature extractor are plotted in the same feature space.

[0028] In S403, the feature extractor selection unit 302 determines the quality values ​​for each of the multiple feature extractors. Specifically, it calculates quality values ​​Q_A, Q_B, Q_C, and Q_D based on the feature distribution for the feature map corresponding to each feature extractor. Here, the quality value Q is calculated using the following equation 1. The value of equation 1 decreases as the similarity between features that represent different people increases. Conversely, if the similarity between features that represent different people is small, the probability of recognizing different people as the same person decreases. In other words, by increasing the quality value of feature extractors that are less likely to recognize different people as the same person, it becomes possible to identify feature extractors with high accuracy.

[0029]

number

[0030] In this example, the similarity between the two feature vectors f_i and f_j is calculated using cosine similarity, and the feature similarity threshold d is set to 0.5. However, other methods may be used. For example, the Euclidean distance between the vectors could be used to calculate the similarity, and a different threshold corresponding to the Euclidean distance could be set.

[0031] In S404, the feature extractor selection unit 302 selects the feature extractor with the highest quality value from among multiple feature extractors. For example, it compares the magnitudes of four quality values ​​Q_A, Q_B, Q_C, and Q_D, and if Q_B is the largest, it selects DNN_B, which is the feature extractor corresponding to Q_B. Alternatively, the output device may display feature maps corresponding to multiple trained models. By displaying the feature maps, the user can verify that the selected feature extractor is suitable for recognition.

[0032] In S405, the matching unit 303 compares the registered image and the test image using the features extracted by the feature extractor selected in S404, and creates a recognition result. The image processing method for recognizing a specific person will be described in detail below using Figure 5. In S500, the face image acquisition unit 300 acquires the face region from the test image. In S501, if the face image acquisition unit 300 detects a face region in S500, it proceeds to S502; if no face region is detected, it terminates the process. In S502, the face feature extraction unit 301 extracts the features of the face region detected in S500. In S503, the matching unit 303 calculates the similarity between the features of the test image and the features of each registered image, and determines the person in the registered image that meets predetermined criteria as the recognition result. First, the ID of the person in the registered image with the highest feature similarity to the features of the face region of the person in the test image is calculated. Next, it is determined whether the feature similarity is greater than a preset threshold, and if it is greater than the threshold, the recognition result ID is output. If the feature similarity is less than the preset threshold, the recognition result is determined to be "No matching ID". In S504, the matching unit 303 outputs the recognition result to the output device 104. The recognition result may display the registered image of the recognized person, or it may simply display that the person was recognized as a registered person. Alternatively, as a subsequent process, an auto-lock unlock signal may be output, or the recognition result may be output as audio to the speaker.

[0033] <Configuration of the learning device> Here, as a preparatory step for the recognition method shown in Figure 4, the training method for the feature extractors will be explained below. Specifically, the training methods for the four feature extractors DNN_A, DNN_B, DNN_C, and DNN_D will be explained.

[0034] Figure 6 shows an example of the configuration of a learning device for training a feature extractor according to this embodiment. The learning device 101 includes a face image acquisition unit 600, an initial learning unit 601, a learning data selection unit 602, and an additional learning unit 603.

[0035] The face image acquisition unit 600 acquires training data. The face image acquisition unit 600 receives training data input from the training data storage device 105. The training data consists of images containing a person's face. The training data may also contain information indicating the person's ID. Next, the face image acquisition unit 600 detects the face region from the images included in the training data. Furthermore, the face image acquisition unit 600 transmits the detected face region to the initial training unit 601.

[0036] The initial learning unit 601 trains an untrained feature extractor based on training data. The initial learning unit 601 receives face regions from the face image acquisition unit 600. It also receives the pre-trained feature extractor (trained model), which is the target of training, from the feature extractor memory device 106. Furthermore, the initial learning unit 601 uses the received face regions as training data to train the feature extractor and generates a feature extractor (trained model) that has completed initial training. As a training method, in this example, an existing machine learning method or deep learning method is used, but other training methods may also be used. The initial learning unit 601 sets an image in the input layer of the untrained trained model, sets the correct value for the image in the output layer, and adjusts the parameters of the neural network so that the output calculated via the neural network approaches the set correct value. Hereafter, the feature extractor that has completed initial training will be referred to as DNN_A. Furthermore, DNN_A is sent to the training data selection unit 602 and the additional learning unit 603.

[0037] The training data selection unit 602 receives training data from the training data storage device 105. It also receives DNN_A from the initial training unit 601. Furthermore, it selects some images from the training data and sends the selected images to the additional training unit 603.

[0038] The additional learning unit 603 trains the pre-trained model based on additional training data. In other words, it generates pre-trained models with different parameter sets for the pre-trained model. Multiple pre-trained models consist of at least one reference pre-trained model and one or more pre-trained models trained by machine learning or deep learning based on the additional training data. The additional learning unit 603 receives selected images from the training data selection unit 602. It also receives DNN_A from the initial learning unit 601. Furthermore, it performs additional training on DNN_A using the selected images (hereafter, the DNN further trained by the additional learning unit 603 will be referred to as DNN_B). In this example, additional training is performed using only the images selected by the training data selection unit 602, but it is not limited to this; all externally input training data can be used, and the weights of the images selected by the training data selection unit 602 may be increased during training. Furthermore, DNN_A and DNN_B are output to the feature extractor memory device 106 as feature extractors.

[0039] In the example above, the training data selection unit 602 selected only one type of data. However, it is not limited to this, and two or more data selection methods may be applied, and different DNNs may be trained and generated for each data selection method. For example, three types of data selection methods may be provided, and in addition to DNN_A, the additional training unit 603 may generate and output three types of DNNs (DNN_B, DNN_C, DNN_D).

[0040] Although the learning device 101 and the recognition device 103 have been described as separate devices, they may be the same device. For example, the recognition device 103 may include a face image acquisition unit 600, an initial learning unit 601, a learning data selection unit 602, and an additional learning unit 603. When learning and recognition are performed with the same device, one example is to perform face recognition while generating a learning model that matches the local environmental conditions using a surveillance camera installed at a gate or other location where face recognition is actually performed.

[0041] <Learning Method Flowchart> Figure 7 is a flowchart illustrating the learning method of a feature extractor executed by the learning device (image processing device) according to this embodiment. The learning device is assumed to have the same hardware configuration as the recognition device, as shown in Figure 2. The flowchart in Figure 7 is realized by the CPU of the learning device executing a program stored in the storage device within the learning device. First, in S700, the face image acquisition unit 600 acquires training data. The training data consists of images of people's faces taken and stored with a general camera, and information on the ID of the person in the image. In S701, the initial learning unit 601 acquires an untrained feature extractor as the initial feature extractor. In this example, a DNN commonly used in deep learning is used as the untrained feature extractor. In S701, the initial learning unit 601 may initialize the parameter set of a trained feature extractor.

[0042] In S702, the initial learning unit 601 trains an untrained feature extractor based on the training data. In this example, first, an untrained DNN is trained using all of the training data. Hereafter, the DNN trained using all of the training data will be referred to as DNN_A.

[0043] In S703, the training data selection unit 602 selects additional training data. The training data selection unit 602 creates a subset of training data from the acquired training data. The subset may be determined randomly, or a subset that satisfies predetermined conditions may be determined. The former method also allows for the generation of a feature extractor that outputs a different feature map from the original feature extractor, and this feature extractor can improve the recognition accuracy of some individuals. On the other hand, the latter method allows for additional training of individuals that were difficult to recognize with the original feature extractor. In this example, DNN_A is used to extract features from the training data and create a subset in which the person IDs are different from each other and the similarity of features between images is high. In this way, the recognition rate can be improved for individuals that were difficult to recognize with the original feature extractor. Here, for the sake of explanation, we assume that the number of subsets is 3.

[0044] In S704, the additional learning unit 603 trains the trained model based on additional training data. The additional learning unit 603 performs additional training on DNN_A using one of the created subsets. In this example, since there are 3 subsets, the additional training generates 3 DNNs (hereinafter referred to as DNN_B, DNN_C, and DNN_D). In other words, the multiple trained models are trained models that have been further trained based on at least one subset of training data generated according to a specific criterion from the training data of the base trained model. For additional training, in this example, the same existing training method used when training DNN_A can be used. In S705, the additional learning unit 603 stores the parameter set of the trained feature extractor in the feature extractor memory.

[0045] <Variations of learning methods> In S703, a subset is created in which the person IDs are different from each other and the similarity of features between the images is high. However, the process is not limited to this, and subsets may be created based on other predetermined criteria. For example, a subset of images with a specific brightness level may be created, assuming the brightness of the environment in which recognition is performed. Alternatively, a subset of images of a specific race or people wearing glasses may be created. In other words, the learning data selection unit 602 determines a subset of registered images that satisfy predetermined criteria, such as the brightness of the environment being within a predetermined standard, a specific race, and people having specific attributes.

[0046] Below, an example of the procedure for selecting some images in the learning data selection unit 602 is shown using Figure 8 and the following equation 2. Figure 8 is a flowchart illustrating the data selection method performed by the learning device (image processing device) according to this embodiment. In S800, the learning data selection unit 602 calculates the similarity of all the learning data. In this example, the similarity of two feature vectors f_i and f_j is calculated using cosine similarity, but other methods may be used. For example, the Euclidean distance between vectors may be used. In S801, the learning data selection unit 602 assigns 0 to the variable j. In S802, the learning data selection unit 602 creates a set Sj of images whose similarity to one registered image i_j is greater than or equal to a threshold d. In S803, the learning data selection unit 602 adds 1 to the variable j. In S804, the learning data selection unit 602 determines whether variable j and variable n match. If they do not match, the process moves to S802; if they match, the process moves to S805. In S805, the training data selection unit 602 calculates the number of images in each of the image sets Sj created in S802. Furthermore, it creates an image set V consisting of sets of images whose number is equal to or greater than the threshold e, and then terminates the process.

[0047]

number

[0048] Generally, the registered images provided by users of a face recognition system are unknown during the system's development phase. Because the number of registered images and the characteristics of the faces depicted vary, using a particular feature extractor to extract features from registered images may result in the existence of image pairs or groups with very similar features. This can lead to a decrease in face recognition accuracy. Therefore, by utilizing the process described above, the distribution of registered image features is calculated for each of multiple different feature extractors. This allows for the selection of feature extractors that minimize the number of image pairs or groups with very similar features, thereby enabling highly accurate face recognition.

[0049] <Example 1: Setting the importance level of registered individuals> If individuals to be registered in the recognition system are assigned importance levels, and there are individuals for whom recognition errors should be minimized, the automatically selected feature extractor may not be the best for the user. In such cases, the user may input the importance levels of the registered individuals in advance, and the feature map may be evaluated taking importance into account. The feature extractor selection unit 302 further acquires registration information with importance levels assigned to each of the registered images corresponding to multiple individuals, and reflects this in the quality value. In the above example, the quality value Q of the feature map is defined by Equation 1. Equation 1 calculates the quality value of the feature distribution when all registered individuals are evaluated equally. However, this is not the only way; the concept of importance can be introduced to the registered individuals, and the quality value can be calculated by a weighted average based on importance. Here, if the quality value calculated by the weighted average based on importance is Q', then Q' can be defined as shown in Equation 3 below.

[0050]

number

[0051] In other words, the quality value is determined based on importance, with a higher value being desirable the smaller the similarity between the features of a highly important person and the features of a different person. A feature extractor suitable for recognizing a specific person is preferable when the features that identify a specific person and the features that identify others are far apart (larger variance) for the specific person the user particularly wants to correctly identify. In other words, by using the quality value of equation 3, it is possible to suppress the recognition of a specific person as the same person as others. By increasing the weight of a specific person, the user can use a feature extractor with a high recognition rate for that person. To input importance, for example, a user interface can be created where the user can input three levels of importance for each registered image, and data input can be done using touch operation on the display or with a mouse and keyboard.

[0052] <Modification 2: Displaying feature maps using the UI> In the method shown in Figure 4, the best feature extractor is automatically selected from among multiple feature extractors. However, the user may also browse the feature maps, select the best feature map, and then select the feature extractor corresponding to the selected map. Figure 9 shows an example of two different feature maps 900 and 901 displayed by the output device. For example, by converting the feature vectors of an image into two-dimensional vectors using principal component analysis, creating and visualizing a two-dimensional feature map, the user can browse the feature maps and select the desired feature map. This allows the user to visually grasp the feature map that is particularly suitable for identifying the person they want to identify, improving convenience.

[0053] <Variation 3: Displaying easily mistaken faces> Even when using the best feature extractor, depending on the characteristics of the registered images, there may be groups of registered images with similar features. In such cases, it may be helpful to allow the user to confirm the existence of groups of registered images with similar features. For example, as shown in Figure 10, it is advisable to use a UI that highlights groups of registered images with high feature similarity using dotted lines in the feature map. In the feature map 1000, similar groups 1001 and 1002 are indicated by dotted line frames. In other words, based on the selected trained model, information about other people similar to a specific person is displayed on the display means. Alternatively, as shown in Figure 11, a table listing groups of registered images with high feature similarity may be displayed. For example, in the message window 1100, registered image group a and registered image group b are displayed as groups of registered people that are similar to each other. This improves convenience because the user can check in advance whether there are people that are likely to be confused.

[0054] <Variation 4: Select registered images to improve recognition rate> Each feature extractor is evaluated using all registered images. However, if there are multiple images corresponding to each registered person, the total number of images becomes large, and the feature similarity between images with different IDs may become high. In such cases, it is acceptable to select the feature extractor that yields the highest quality, including the selection of multiple images. In other words, some images are selected from all registered images to increase the variance of the feature distribution across multiple registered images. For example, if there are 20 registered images for each of 100 people, there are a total of 2000 registered images. In this case, it is acceptable to use one image per person and select 100 registered images and one feature extractor that yields the highest quality. The feature extractor selected in this way can suppress the recognition of different people as the same person.

[0055] <Modification 5: Re-selection of the feature extractor> During the operation of a facial recognition system, individuals may be removed from the registration list or new individuals may be registered. When registered images are deleted or added, the selected feature extractor may no longer be the optimal one. In such cases, for example, when registered images are deleted or added, the system should automatically re-select a feature extractor. Also, if the number of registered images is enormous and selecting a feature extractor takes a long time, the system should automatically re-select a feature extractor only when, for example, more than 10% of the images have been changed. In other words, the feature extractor selection unit re-selects a feature extractor when there are changes in the registered images. This improves usability because, even when there are changes in the registered images, the performance of the feature extractor can be checked before recognition is performed.

[0056] <Modification 6: Error detection> If there are many registered images with high similarity in addition to the registered image with the highest similarity, the reliability of the recognition result may be low. In such cases, it is advisable to define a region where the features of the registered image are densely concentrated within the feature space created by the feature extractor, and then display "No matching ID found" as the recognition result when the features of the test image are located within that region.

[0057] Furthermore, when the features of the test image are located within a region where the features of the registered image are densely concentrated, it is advisable to output the ID of the registered person with the highest similarity as the recognition result, along with a warning that other registered people with high similarity exist. For example, as shown in Figure 12, registered images of people whose similarity to the person in the input image is above a predetermined value may be displayed. Along with the input image 1200 or registered image 1202, it is advisable to display images of other registered people with high similarity 1202a to 1202c, or text 1201 such as "This person may be one of the following." Displaying these items allows the user to verify whether or not the environment is suitable for the face recognition system to function correctly, thereby improving convenience.

[0058] <Modification 7: Recognition method using multiple feature extractors> When features in a test image are located within a region where features in a registered image are densely concentrated, it is highly likely that the registered person with the highest similarity does not match the person in the test image. In such cases, multiple feature extractors may be selected and used in combination for recognition processing. In other words, the matching unit recognizes a specific person based on multiple pre-trained models that have been determined. For example, first, a registration map is created using the first feature extractor to detect regions where features in the registered images are densely concentrated. Next, for the set of registered images located in the region where features are densely concentrated, the feature extractor with the highest quality feature map is selected as the second feature extractor. By selecting the first and second feature extractors in this way, if features extracted by the first feature extractor in the test image are located in the region where features are densely concentrated, it is advisable to perform recognition processing using the second feature extractor. This helps to suppress the recognition of different people as the same person.

[0059] <Other Embodiments> Furthermore, the present invention can also be realized by performing the following process: that is, supplying software (program) that realizes the functions of the above-described embodiment to a system or device via a network or various storage media, and having the computer (or CPU or MPU, etc.) of that system or device read and execute the program.

Claims

1. An image processing device that recognizes a person included in an input image, A means of obtaining registered images of multiple people, A generation means that generates a feature map corresponding to each of the trained models, which shows a feature distribution in which the features of the multiple people extracted from the registered images of the multiple people are placed in the same feature space, based on multiple trained models that extract features of each person from images, A selection means for selecting a trained model based on the respective quality values ​​for the trained model determined based on the feature map, The system includes a recognition means that recognizes a person included in an input image based on the selected trained model, The image processing apparatus is characterized in that the aforementioned quality value is determined such that the smaller the similarity between the characteristics of a particular person and the characteristics of a different person, the larger the value.

2. The acquisition means further acquires registration information in which importance is assigned to each of the registered images corresponding to the plurality of persons, The image processing apparatus according to claim 1, characterized in that the quality value is determined based on the importance such that the value increases as the similarity between the characteristics of a person with high importance and the characteristics of a different person decreases.

3. The image processing apparatus according to claim 1, further comprising control means for causing the feature maps corresponding to the plurality of trained models to be displayed on an output device.

4. The image recognition device according to claim 1, further comprising control means for displaying information on a display means that is similar to the specific person based on the selected trained model.

5. The selection means selects multiple trained models to be selected, The image processing apparatus according to claim 1, wherein the recognition means recognizes a person included in the input image based on the plurality of trained models determined.

6. The image processing apparatus according to claim 1, further comprising a selection means for selecting some images from all of the registered images such that the variance of the distribution of features of the multiple registered images is large.

7. The image processing apparatus according to claim 1, characterized in that the plurality of trained models include at least one reference trained model and one or more trained models trained by machine learning or deep learning based on additional training data.

8. The system further includes a learning method for training multiple learning models that extract individual features from images, The image processing apparatus according to claim 7, characterized in that the plurality of trained models are trained models that have been further trained based on at least one subset of the training data of the reference trained model that was generated according to a predetermined criterion.

9. An image processing method for recognizing a person included in an input image, The acquisition unit performs an acquisition process to acquire registered images of multiple people, The extraction unit generates a feature map corresponding to each of the trained models, which, based on multiple trained models that extract features for each person from an image, displays a feature distribution in which the features of the multiple people extracted from the registered images of the multiple people are placed in the same feature space. The selection unit performs a selection step of selecting a trained model based on the respective quality values ​​for the trained model determined based on the feature map, The matching unit includes a recognition step of recognizing a person included in the input image based on the selected trained model, The image processing method is characterized in that the aforementioned quality value is determined such that the smaller the similarity between the characteristics of a specific person and the characteristics of a different person, the larger the value.

10. A program for causing a computer to function as each of the means of the image processing apparatus described in claim 1.

Citation Information

Patent Citations

  • Annealing method of compound semiconductor

    JP1985003124A

  • Object identification device, object identification method, and program

    JP2017058833A

  • Inspection device and inspection method

    JP2021103125A

  • Face Recognition Apparatus and Methods

    US20120230545A1