Information processing device, information processing method, and program
A dual-learning model approach with a large and small neural network for person recognition ensures accuracy and reduces processing time by selecting the appropriate model based on feature similarity, addressing the trade-off in existing systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2022-06-13
- Publication Date
- 2026-05-08
AI Technical Summary
Existing image recognition systems using large neural networks for person recognition face a trade-off between processing time and accuracy, with high-accuracy models being computationally intensive and low-accuracy models being faster but less expressive.
Employing a dual-learning model approach, using a large-scale complex neural network for high-accuracy and a smaller-scale model obtained through knowledge distillation for reduced computational load, and selecting the appropriate model based on feature comparison for each individual.
Maintains recognition accuracy while significantly reducing processing time, especially in resource-constrained environments, by dynamically choosing the appropriate model for each person based on feature similarity.
Smart Images

Figure 0007855412000002 
Figure 0007855412000003 
Figure 0007855412000004
Abstract
Description
Technical Field
[0001] The present invention relates to information registration used for collation and information processing technology for collation.
Background Art
[0002] In recent years, in person recognition for collating whether a person shown in an image is a registered person, image recognition technology based on deep learning has been increasingly used. In image recognition technology based on deep learning, feature information is extracted from a face image of a person using a learning model composed of a neural network (hereinafter abbreviated as NN). In particular, the larger and more complex the learning model composed of NN is, the higher the possibility of extracting highly expressive and accurate feature information from a face image of a person becomes. However, the larger and more complex the learning model is, the longer the processing time required for image recognition becomes. On the other hand, if a small-scale learning model is used, the processing time of image recognition can be shortened, but the feature information extracted from a face image of a person may have low expressiveness and undesirable recognition accuracy.
[0003] Patent Document 1 discloses a technique for performing multi-stage collation using a high-speed collation means with low accuracy and a low-speed collation means with high accuracy. That is, Patent Document 1 discloses a technique in which narrowing down of a recognition target is performed by a high-speed collation means with low accuracy, and then collation (recognition) is performed by switching to a low-speed collation means with high accuracy.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In the technology disclosed in Patent Document 1, the recognition target is narrowed down using a low-accuracy but high-speed matching means, and then the matching (recognition) is performed by switching to a high-accuracy but low-speed matching means. Although a certain level of matching accuracy can be ensured, the processing time becomes extremely long.
[0006] Therefore, the present invention aims to shorten processing time while maintaining accuracy during matching. [Means for solving the problem]
[0007] The information processing apparatus of the present invention is characterized by comprising: a first extraction means for performing a first extraction process on an image to be registered to extract first feature information representing the characteristics of the registered object; a second extraction means for performing a second extraction process on the image to be registered to extract second feature information comparable to the first feature information; a comparison means for comparing the first feature information and the second feature information; and a registration means for registering registration information that associates with the registered object whether to use the first extraction process or the second extraction process for the matching process, based on the results of the comparison by the comparison means. [Effects of the Invention]
[0008] According to the present invention, it is possible to maintain accuracy during matching and shorten processing time. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows an example of a system configuration including an information processing device. [Figure 2] This is a functional block diagram of the registration device and the recognition device. [Figure 3] This is a flowchart showing the registration process. [Figure 4] This is a diagram illustrating an example of registration information. [Figure 5] This is a flowchart showing the recognition process flow. [Figure 6] This diagram illustrates an example of the registration screen in the second embodiment. [Modes for carrying out the invention]
[0010] Embodiments of the present invention will be described below with reference to the drawings. The embodiments described below are not limiting to the present invention, and not all combinations of features described in these embodiments are essential to the solutions of the present invention. The configuration of the embodiments may be modified or changed as appropriate depending on the specifications of the device to which the present invention is applied and various conditions (usage conditions, usage environment, etc.). Furthermore, parts of each embodiment described later may be combined as appropriate. In each of the following embodiments, the same components will be denoted by the same reference numerals.
[0011] <Example System Configuration> Figure 1 is a diagram showing an example configuration of a recognition system including an information processing device according to the first embodiment. The information processing device 110 includes a system control unit 111, a ROM (Read Only Memory) 112, a RAM (Random Access Memory) 113, an HDD (Hard Disk Drive) 114, and a communication unit 115. The system control unit 111 has, for example, a CPU (Central Processing Unit) and reads control programs stored in the ROM 112 to execute various processes. A GPU (Graphics Processing Unit) may be used instead of the CPU. The RAM 113 is used as the main memory, work area, and other temporary storage areas of the system control unit 111. The HDD 114 stores various data and various programs. It is assumed that the information processing programs related to the information processing device 110 of this embodiment are stored in the ROM 112 or the HDD 114. That is, each function and process according to this embodiment, which will be described later, is realized by the system control unit 111 reading and executing the information processing programs stored in the ROM 112 or the HDD 114. The communication unit 115 is capable of communicating with the server 100 via the network 120.
[0012] Server 100 comprises a system control unit 101, ROM 102, RAM 103, HDD 104, communication unit 105, and a database (DB) 106. The system control unit 101 has a CPU and reads control programs stored in ROM 102 to execute various processes. RAM 103 is used as the main memory and temporary storage area for the system control unit 101, such as a work area. HDD 104 stores various data and programs. DB 106 stores necessary data to be managed. Note that the data for DB 106 may also be stored on the HDD. The communication unit 105 is capable of communicating with the information processing device 110 via the network 120.
[0013] Here, the recognition system of this embodiment, as an example, performs a person matching process to confirm whether the person is a registered person or not by performing a person recognition process using the person's facial image when a person passes through a security gate for access control. For this reason, in addition to the information processing device 110 and server 100 described above, the recognition system of this embodiment further includes a network camera 150, an information reading device 140, a security gate 160, and an administrator terminal 130. These network camera 150, information reading device 140, security gate 160, and administrator terminal 130 are connected to the server 100 and information processing device 110 via a network 120 so as to be able to communicate. Note that the network 120 is not limited to a wired network and may include a wireless network. The network camera 150 is an imaging device that is installed, for example, in front of the security gate 160 and takes images of people. The information reading device 140 obtains identification information that has been pre-registered for the person from, for example, an identification card (ID card) or a mobile information terminal such as a smartphone that the person is carrying. The administrator terminal 130 is an information terminal, such as a personal computer, of the administrator who manages this recognition system. In this embodiment, access control at the security gate 160 is used as an example, but it can also be applied to other situations, such as payment when purchasing goods in a store, in which case the network camera 150 would be placed, for example, in front of the cash register. The administrator terminal 130 may also be placed, for example, near the security gate 160 or the cash register, or it may be located in a management center, or it may be a portable information terminal such as a smartphone or tablet held by the administrator.
[0014] In this embodiment, we provide an example of using facial feature information (hereinafter referred to as "feature quantities") extracted from a person's facial image when verifying whether the person being recognized is a pre-registered registered person. Furthermore, in this embodiment, we provide an example of image recognition processing using a learning model composed of a neural network (NN) based on deep learning when extracting facial feature quantities from a person's facial image. That is, the information processing device used for person recognition extracts facial feature quantities from a person's facial image captured by a camera using the learning model, compares these feature quantities with the facial feature quantities of pre-registered people, and determines whether the person is a registered person. However, generally, learning models of NNs that enable highly accurate feature extraction are large-scale and complex processes that require a lot of computation, so the information processing device needs a lot of memory and computing resources. On the other hand, information processing devices installed in actual sites where person recognition is performed often do not have sufficient memory or computing resources.
[0015] Therefore, in the recognition system of this embodiment, when a person to be registered is registered, the information processing device 110 extracts facial features from the facial image of the person to be registered using at least two NN learning models with different computational loads. In this embodiment, the learning models with different computational loads are assumed to be a large-scale and complex NN learning model with a high computational load, and a learning model with a low computational load (small scale) obtained by knowledge distillation of the NN of the learning model. Details of these learning models will be described later. Furthermore, the information processing device 110 compares the features extracted from each of these learning models and, based on the comparison result, sets which learning model should be used when the matching process is performed later. The information processing device 110 then generates registration information that associates model information indicating which learning model should be used during the matching process with the identification information of the person to be registered and the facial features of that person. This registration information is sent to the server 100, and the server 100 stores the registration information in the DB 106.
[0016] Also, in the recognition system of the present embodiment, when the information processing device 110 performs the collation process of whether the person to be recognized is a registered person, it acquires the registration information from the DB106 of the server 100 based on the identification information of the person to be recognized. Further, the information processing device 110 extracts feature amounts from the face image of the person to be recognized using the learning model indicated by the model information in the registration information. Then, the information processing device compares the feature amounts extracted from the face image of the person to be recognized using the learning model corresponding to the model information in the registration information with the feature amounts in the registration information, thereby performing the collation of whether the person to be recognized is the same as the registered person.
[0017] <Functional Block Configuration> The information processing device 110 of the present embodiment is made capable of executing either one or both of the functions of a registration device when registering a person to be registered and a collation device when collating whether the person to be recognized is the same as the registered person. FIG. 2(A) is a functional block diagram showing an example of the functional configuration when the information processing device 110 of the present embodiment functions as the registration device 200. Also, FIG. 2(B) is a functional block diagram showing an example of the functional configuration when the information processing device 110 of the present embodiment functions as the collation device 210.
[0018] In the case of the present embodiment, an example is given where both functions of the registration device 200 and the collation device 210 are realized by the information processing device 110 of FIG. 1. However, for example, the function of the registration device 200 may be borne by the server 100, and the function of the collation device 210 may be borne by the information processing device 110. Alternatively, the respective functions of the registration device 200 and the collation device 210 may be appropriately shared between the information processing device 110 and the server 100. The respective functions of the registration device 200 and the collation device 210 may be realized by the CPU (such as the system control unit 111, etc.) executing the information processing program according to the present embodiment, or some or all of those functions may be realized by a hardware configuration.
[0019] The registration device 200 is a device that generates registration information for a person to be registered. For example, it includes an image acquisition unit 201, a target detection unit 202, a first extraction unit 203, a second extraction unit 204, a comparison unit 205, and a registration unit 206.
[0020] The verification device 210 is a device that verifies (i.e., identifies) whether a person to be recognized is the same as a registered person. It includes an image acquisition unit 211, a target detection unit 212, an identification information acquisition unit 213, a registration information acquisition unit 214, an extraction unit 215, and a verification unit 216. In this embodiment, a person entering the security gate 160 is regarded as a person to be recognized, and it is identified using registration information whether the person to be recognized is a registered person. For example, the verification device 210 performs verification processing as a security measure in cases such as when an unregistered person attempts to intrude by pretending to be a registered person.
[0021] First, the case where the information processing device 110 of this embodiment operates as the registration device 200 will be described. The image acquisition unit 201 of the registration device 200 acquires an image captured by the network camera 150. Note that the network camera 150 may be included in the image acquisition unit 201. The image acquired by the image acquisition unit 201 is assumed to be an image of a person to be registered. The image acquired by the image acquisition unit 201 may be an image of a frame constituting a video or a still image. The target detection unit 202 detects an image of the face region of the person to be registered (referred to as a face image) from the image acquired by the image acquisition unit 201.
[0022] The first extraction unit 203 extracts a first feature amount from the face image detected by the target detection unit 202 by a first feature amount extraction process. Although details will be described later, in the case of this embodiment, the first feature amount extraction process is a process of extracting a face feature amount (first feature amount) using a learning model of a large-scale and complex NN with a large amount of computation. Hereinafter, in this embodiment, the learning model of the NN used for the first feature amount extraction process of extracting the first feature amount from the face image will be referred to as the first model. Details of the first model will be described later.
[0023] The second extraction unit 204 extracts a second feature from the face image detected by the target detection unit 202 using a smaller-scale second feature extraction process that requires less computation than the first feature extraction process. As will be described in detail later, the second feature is feature information that represents similar features to the first feature, and the second feature is obtained as a feature comparable to the first feature. As will also be described in detail later, in this embodiment, the second feature extraction process is assumed to be an extraction process using a smaller-scale learning model that requires less computation than the first model in the first feature extraction process. Hereinafter, in this embodiment, the learning model of the NN used in the second feature extraction process that extracts the second feature from the face image will be referred to as the second model. Details of the second model will be described later.
[0024] The comparison unit 205 compares the first feature extracted by the first extraction unit 203 with the second feature extracted by the second extraction unit 204, and based on the comparison result, generates model information indicating which of the first or second models to use for feature extraction during the matching process. In other words, the comparison unit 205 generates information indicating which of the first or second feature extraction processes to perform during the matching process, based on the comparison result between the first and second feature quantities. Details of the comparison process between the first and second feature quantities in the comparison unit 205 will be described later.
[0025] The registration unit 206 generates registration information by associating the feature quantities and model information with the identification information of the person to be registered, as detected by the target detection unit 202. The person's identification information is unique to each person to be registered, and in this embodiment, it is assumed that the identification information (ID data, etc.) pre-registered on the ID card or mobile information terminal such as a smartphone that the person possesses is used. The feature quantity information included in the registration information may be both the first and second feature quantities, but in this embodiment, it is assumed that it is the feature quantity corresponding to the model information. The registration information, which associates these feature quantities and model information with the identification information, is then notified from the information processing device 110 to the server 100 via the communication unit 115 and stored in DB 106. In this embodiment, the registration information is said to be generated by the information processing device 110, but it may also be generated by the server 100 and stored in DB 106.
[0026] Next, we will describe the case in which the information processing device 110 of this embodiment operates as a verification device 210. The image acquisition unit 211 of the matching device 210 acquires images captured by the network camera 150. The image acquisition unit 211 of the matching device 210 may include the network camera 150. The images acquired by the image acquisition unit 211 of the matching device 210 are images of the person to be recognized. The images acquired by the image acquisition unit 211 may be frames that make up a video, or still images.
[0027] The target detection unit 212 detects the face image of the person to be recognized from the image acquired by the image acquisition unit 211. The identification information acquisition unit 213 acquires identification information (ID data, etc.) that has been pre-registered for the person being recognized from an ID card or a mobile information terminal such as a smartphone that the person is possessing.
[0028] The registration information acquisition unit 214 receives identification information from the identification information acquisition unit 213 and acquires the registration information associated with the identification information from the DB 106 of the server 100. As described above, the registration information is information in the registration device 200 in which feature quantities and model information are associated with the identification information.
[0029] The extraction unit 215 obtains model information corresponding to the identification information obtained by the identification information acquisition unit 213 from the registration information obtained by the registration information acquisition unit 214. As will be described in detail later, if the model information indicates the first model, the extraction unit 215 performs a feature extraction process using the first model on the face image of the person to be recognized, that is, it extracts the first feature by the first feature extraction process. On the other hand, if the model information indicates the second model, the extraction unit 215 performs a feature extraction process using the second model on the face image of the person to be recognized, that is, it extracts the second feature by the second feature extraction process.
[0030] The matching unit 216 obtains feature quantities from the registration information obtained by the registration information acquisition unit 214, calculates the similarity between these feature quantities and the feature quantities extracted by the extraction unit 215, and performs a matching process to determine whether the person to be recognized is a registered person or not. As will be described in detail later, if the model information indicates the first model, the matching unit 216 performs a matching process based on the similarity between the first feature quantity extracted by the extraction unit 215 from the face image of the person to be recognized using the first model and the first feature quantity of the registration information. On the other hand, if the model information indicates the second model, the matching unit 216 performs a matching process based on the similarity between the second feature quantity extracted by the extraction unit 215 from the face image of the person to be recognized using the second model and the second feature quantity of the registration information.
[0031] Figure 3 is a flowchart showing the flow of the registration process performed in the registration device 200. In the flowcharts that follow, the symbol S represents a processing step (processing step). First, in S301, the image acquisition unit 201 of the registration device 200 acquires an image including the face of the person to be registered from a network camera 150 installed in front of the security gate 160 or the like. The face image of the person acquired by the image acquisition unit 201 of the registration device 200 may be an image that has been previously taken and stored in the memory of the ID card, an image stored in a mobile information terminal such as a smartphone, or an image stored in another image server.
[0032] Next, in S302, the target detection unit 202 detects the face region of the person in question as a face image from the image acquired by the image acquisition unit 201. Known techniques can be used to detect the face region image. Examples include detection methods based on pre-set facial features such as eyes, nose, and mouth, and face region detection methods that utilize learning, such as deep learning.
[0033] Next, in S303, the first extraction unit 203 extracts a first feature from the face image detected by the target detection unit 202, and the second extraction unit 204 extracts a second feature from the face image detected by the target detection unit 202. As described above, the first extraction unit 203 extracts the first feature quantities using a pre-trained first model, which in this embodiment is a first NN model that is computationally intensive, large-scale, and complex to process. Since the first model is a large-scale NN training model that is computationally intensive and complex to process, the first feature quantities extracted by the first extraction unit 203 are likely to be highly expressive and accurate.
[0034] On the other hand, as mentioned above, the second extraction unit 204 extracts second features using a pre-trained second model, that is, in this embodiment, a second NN model with less computation and a smaller scale than the first model. However, since the second model is a NN training model with less computation, the second features extracted by the second extraction unit 204 may have lower expressive power and less accuracy than the first features.
[0035] In this embodiment, the second model is a learning model obtained by so-called knowledge distillation of the first model's neural network (NN), which is computationally intensive and large in scale. Knowledge distillation (hereinafter referred to as distillation) is a known technique for reducing the size and processing load of an NN. Distillation generates a relatively small and lightweight NN learning model (generally referred to as the student model) by learning using the output of a large-scale NN learning model (generally referred to as the teacher model).
[0036] More specifically, in distillation, the output of the teacher model and the output of the student model are compared; in this embodiment, the features are compared, and the parameters of the student model are adjusted through learning so that the difference between the respective features is small. For example, if the feature extracted by the teacher model is v1 and the feature extracted by the student model is v2, the student model is trained so that |v1-v2|<ε. Here, ε is a predetermined value. In other words, since distillation directly compares the output (features) of the teacher model and the output (features) of the student model during training, it can be said that the features output by the teacher model and the features output by the student model are comparable. Thus, since the student model is a model trained using the output of the teacher model, the features extracted from an image using the student model are likely to have roughly the same accuracy as the features extracted from an image using the teacher model. For this reason, NN distillation is a technique often used in environments with limited machine resources, such as when performing NN-based processing on edge devices such as personal mobile devices. However, depending on the image being processed for feature extraction, the features extracted by the student model and the features extracted by the teacher model may not necessarily have the same accuracy.
[0037] Next, in S304, the comparison unit 205 compares the first feature quantity extracted by the first extraction unit 203 using the first model with the second feature quantity extracted by the second extraction unit 204 using the second model. Specifically, the comparison unit 205 calculates the distance between the first feature quantity and the second feature quantity. The distance is calculated as the Euclidean distance between the first feature quantity (feature vector) and the second feature quantity (feature vector) in the same feature space. For example, if the first feature quantity is v1 and the second feature quantity is v2, their distance is expressed as ||V1-V2||. The process of calculating the feature distance between feature quantities in this comparison unit 205 utilizes the fact that the output of the first model before distillation and the output of the second model after distillation are comparable features, as described above.
[0038] Next, in S305, the comparison unit 205 uses the feature distance between the first and second features to evaluate which of the first and second models should be used when matching the face of the person detected in S302. As mentioned above, the second model after distillation is trained to match the output of the first model before distillation, so the first and second features extracted from the same image should match. In other words, the value of the feature distance ||V1-V2|| should be 0. However, depending on the face image from which features are extracted, the first and second features may not match.
[0039] Therefore, the comparison unit 205 evaluates which of the first and second models to use during the matching process by determining whether the feature distance between the first and second features is below a predetermined threshold. The predetermined threshold is assumed to be a value set in advance. Here, if the feature distance is below the threshold, that is, if the first and second features are close features, it is considered that approximately equivalent features can be extracted regardless of whether the first or second model is used. On the other hand, if the feature distance exceeds the threshold, that is, if the first and second features are far apart, it is considered highly likely that the first feature obtained using the first model will be different from the second feature obtained using the second model. In other words, if feature extraction is performed using the lightweight second model, it is considered that the extracted features may be different from the highly accurate first feature obtained using the first model.
[0040] Therefore, in the following S306 and S307, the registration unit 206 determines which of the first and second models to use during the matching process based on the comparison result by the comparison unit 205 in S305, and includes the model information corresponding to that decision in the registration information. In other words, for example, if the comparison unit 205 obtains a comparison result in which the feature distance between the first and second feature quantities is below a threshold, it can be said that this person is one from whom the same feature quantities can be obtained even when using the second model as when using the first model. That is, a person from whom the feature distance between the first and second feature quantities is below a threshold is considered to be a person from whom mismatching is unlikely to occur even when feature extraction processing using the second model is performed during the matching process. Therefore, in S306, the registration unit 206 includes model information indicating the second model in the registration information for a person from whom a comparison result in which the feature distance between the first and second feature quantities is below a threshold is obtained. In other words, this person is registered as a person from whom features will be extracted and matched using the second model during the matching process.
[0041] On the other hand, if the comparison unit 205 obtains a comparison result in which the feature distance between the features exceeds a threshold, using the second model during the matching process may not allow for matching with the registered person, and there is also a possibility of mismatching a different person with the registered person. Therefore, for individuals in whom the comparison unit 205 obtains a comparison result in which the feature distance between the first and second features exceeds a threshold, the registration unit 206 includes model information indicating the first model in the registration information in S307. In other words, individuals in whom the feature distance between the features exceeds a threshold are registered as individuals whose features are extracted and matched using the highly accurate first model during the matching process.
[0042] Figure 4 schematically shows an example of the information registered in DB 106 of server 100 as registration information 400. Registration information 400 is information that associates the person's identification information, the person ID, with the feature quantities (face features) extracted from the face image of the person to be registered, and model information indicating whether to use the first or second learning model for the matching process. In the example in Figure 4, the second model is shown as Model A and the first model as Model B. Furthermore, the person identification information (person ID) is exemplified as 00001, 00002, and 00003. In the case of registration information 400, it is shown that Model A (second model) can be used during the matching process for people with person IDs 00001 and 00002, and Model B (first model) must be used during the matching process for person with person ID 00003.
[0043] Figure 5 is a flowchart showing the flow of the matching process performed in the matching device 210. First, in S501, the identification information acquisition unit 213 acquires the unique identification information of the person to be recognized. The identification information acquisition unit 213 acquires the identification information of the person, for example, which is read by an information reading device 140 installed at an entrance where a security gate 160 is located, from an ID card or smartphone carried by the person to be recognized. In this way, the matching device 210 identifies the person it is about to recognize based on the identification information.
[0044] Next, in S502, the registration information acquisition unit 214 reads the corresponding registration information from the DB 106 of the server 100 based on the identification information acquired by the identification information acquisition unit 213 in S401. As shown in Figure 4, the registration information associates the identification information (person ID) of a registered person, the feature quantities extracted from the face image when the person was registered, and model information indicating whether to use the first or second model during the matching process. The registration information acquisition unit 214 queries the server 100 using the person-specific identification information acquired in S501 as a search key, and acquires the registration information sent from the server 100 in response to the query.
[0045] The process of acquiring registration information in S502 will be explained in detail using the registration information 400 shown in Figure 4 above as an example. If the identification information (person ID) acquired by the identification information acquisition unit 213 is, for example, 00001, the system control unit 101 of the server 100 uses that person ID as a search key to read the registration information from the registration information 400 stored in DB 106. That is, the system control unit 101 of the server 100 reads the registration information corresponding to person ID 00001 from the registration information 400 stored in DB 106 and transmits it to the information processing device 110 via the communication unit 105. If, for example, the person ID is 00003, the server 100 reads the registration information corresponding to that person ID and transmits it to the information processing device 110. The registration information acquisition unit 214, having received the registration information transmitted from the server 100 in this way, stores the registration information in RAM 113, which is the memory of the information processing device 110. This registration information will be used in processing from S405 onwards.
[0046] Next, in S503, the image acquisition unit 211 of the matching device 210 acquires an image including the face of the person to be recognized from a network camera 150 installed in front of the security gate 160 or the like.
[0047] Next, in S504, the target detection unit 212 detects the face region of the person in question as a face image from the image acquired by the image acquisition unit 211. As for the technique for detecting the face region image, known techniques can be used, similar to those used by the target detection unit 202 of the registration device 200.
[0048] Next, in S505, the extraction unit 215 extracts features from the face image detected by the target detection unit 212. At this time, the extraction unit 215 extracts features using a learning model corresponding to the model information included in the registration information acquired in S502. For example, if the registration information for person ID 00001 shown in Figure 4 is obtained, the extraction unit 215 extracts features using the second model, which is the model information (Model A) included in the registration information for person ID, i.e., a learning model with low computational complexity and small scale. Also, for example, if the registration information for person ID 00003 is obtained, the extraction unit 215 extracts features using the first model, which is the model information (Model B) included in the registration information for that person ID, i.e., a learning model that has high computational complexity but can produce high-precision output.
[0049] Next, in S506, the matching unit 216 calculates the similarity between the feature quantity of the registered person's face, which is included in the registration information acquired in S502, and the feature quantity extracted by the extraction unit 215 in S505, and evaluates whether the similarity is above a predetermined threshold. The similarity of the feature quantities can be any similarity as long as the multidimensional feature quantities can be compared, and one example of this is cosine similarity. Cosine similarity can be expressed by the following equation (1). Equation (1) shows the cosine similarity between an n-dimensional feature vector p and a feature vector q, respectively. Cosine similarity is a value that can evaluate the closeness between two feature quantities by the angle between the two feature vectors, and it takes a value range of -1.0 to 1.0. A value of -1.0 indicates the lowest similarity between the feature vectors, and a value of 1.0 indicates the highest similarity between the feature vectors.
[0050]
number
[0051] Then, in S506, if the similarity is evaluated to be above a predetermined threshold, the matching unit 216 recognizes in S507 that the person in the face image detected in S504 is the registered person. On the other hand, if the similarity is evaluated to be below a predetermined threshold in S506, the matching unit 216 recognizes in S508 that the person in the face image detected in S504 is a different person from the registered person. The predetermined threshold used in S506 is a value that is set in advance, for example, a value set by the accuracy when the learning model was evaluated in advance. The accuracy when the learning model was evaluated is set based on the self-rejection rate, which is the rate at which a person who should be recognized as the person is recognized as not being the person, and the other-person acceptance rate, which is the rate at which a different person is recognized as the person, as a result of performing a matching process based on the features extracted in advance by the learning model.
[0052] Next, in S509, the matching unit 216 performs processing according to the recognition result in S507 or S508. For example, the matching unit 216 opens or closes the security gate 160 based on the recognition result. That is, if the matching unit 216 obtains a recognition result that the person is the person in question, it opens the security gate 160 to allow the person to enter, while if the recognition result that the person is not the person in question, it closes the security gate 160 to deny the person entry. The opening and closing of the security gate 160 is performed by another gate opening / closing device (not shown), and the matching unit 216 may send a gate opening / closing control command according to the recognition result to that gate opening / closing device. The recognition result may also be notified to, for example, the administrator terminal 130, and the administrator terminal 130 may display the recognition result on a display or output a sound according to the recognition result.
[0053] As mentioned above, in this embodiment, for individuals who can be recognized by the second model, which is lightweight but not highly accurate, the system registers that the second model should be used during the matching process performed at the security gate. This ensures accuracy of recognition during matching, even in information processing devices with limited memory and computing resources. In this case, the use of a smaller second model with less computational complexity also reduces processing time. On the other hand, for individuals whose accuracy cannot be guaranteed by the second model, the system registers that the first model, which requires more computation but has a higher probability of producing a highly accurate output, should be used, thereby ensuring accuracy of recognition during matching.
[0054] In this embodiment, one specific use case to which the recognition system is applied is a recognition system at a security gate installed in an office building. Here, it is assumed that the majority of users of the office building are between 20 and 50 years old, and that minors and the elderly are less frequent users. In a recognition system installed in such a place, when training the second model after distillation, facial images of the age group most likely to use the office building are prioritized. As a result, for people in the 20s to 50s age group, even using the second model which requires less computation and allows for high-speed processing, recognition can be performed with roughly the same accuracy as when using the first model which covers all age groups from minors to the elderly. On the other hand, even in an office building, there is a possibility that minors and the elderly may also use it, and these people may be misrecognized if the second model is used. Therefore, for these people who may be misrecognized, the first model, which requires more computation and therefore takes longer to process, but has better accuracy, is used. In this way, by using a lightweight second model for a large number of frequently used individuals, and a more accurate first model for individuals who are less frequently used and potentially prone to misrecognition, it is possible to reduce the processing time required for recognition while maintaining overall accuracy.
[0055] Furthermore, this embodiment provides an example of recognition using facial images taken on-site, such as in front of a security gate. When using facial images taken on-site in this way, it is conceivable that there may be differences between the feature quantities of the facial image used to generate the registration information and the feature quantities of the facial image taken when matching with the matching device, for example, due to external factors (such as changes in ambient light).
[0056] Thus, assuming that differences may occur in the feature quantities of the face image between the time of registration and the time of matching, the registration information update process may be performed to re-register the feature quantities, taking into account the differences that occur in the feature quantities held as registration information. In this embodiment, the registration update process that takes into account the possible differences in feature quantities is expected to be performed at a predetermined time, or when a change in the external environment in which the person to be recognized is photographed is detected. After the feature quantities of the registration information have been updated, the matching process is performed using the updated feature quantities.
[0057] Here, a predetermined time for updating registration information could be, for example, a period when the seasons change and the amount of daylight changes, such as in summer or winter. During these periods, not only does ambient light, such as the amount of daylight, change, but so can the items worn by people (such as masks or sunglasses). By updating the features of the registration information in accordance with this predetermined time, it becomes possible to perform matching processing that is more suitable for that period.
[0058] Furthermore, methods for detecting changes in the external environment include, for example, detecting changes in ambient light from images, or detecting changes in ambient light using illuminance sensors. In addition to detecting changes in ambient light, detecting changes in the external environment also includes, for example, changes in the camera's installation location, changes in the field of view, and changes in the arrangement of objects around the camera. Changes in the camera's installation location, field of view, and arrangement of objects around the camera can be detected, for example, based on changes in how a place appears, its composition, its size, etc., within the image captured by the camera, or changes in the position of objects, etc. When such changes in the external environment are detected, updating the feature quantities of the registered information enables matching processing that is adapted to the external environment at that time.
[0059] As a process for updating the features of registered information, for example, when a matching process is performed, the same process as during registration shown in the flowchart of Figure 3 is performed again, and the registered information is updated based on the information obtained from that process. That is, for face images detected from images newly captured by a network camera, the first and second feature extraction processes are performed in the same way as in S304, and then the feature comparison process is performed in the same way as in S305. In the feature comparison process in S305, if the similarity of the first and second features is determined to be above a threshold, the person is registered in S306 as the person to whom the second model is applied. On the other hand, in the feature comparison process, if the similarity of the first and second features is determined to be below a threshold, the person is registered in S307 as the person to whom the first model is applied. When the matching process is performed by the matching device, the registered information is already stored in DB106, so in S306 and S307, the registered information in DB106 is updated. That is, by comparing the features extracted from face images including external factors in the environment in which the matching process is actually performed by the matching device, it becomes possible to use registered information adapted to those times and external environments, and as a result, highly accurate matching processing can be achieved.
[0060] As mentioned above, according to the first embodiment, when a person is registered, it is possible to perform high-speed processing while maintaining recognition accuracy during the matching process by pre-determining and registering whether to use the first model or the second model in the subsequent matching process.
[0061] In addition to the aforementioned example of access control for office buildings, the following use cases are also conceivable for registration and verification processes. For example, when entering a security gate in the morning, a facial image may be captured and registered. Subsequently, when passing through other gates within the office building, a facial image may be captured and verified. Furthermore, when leaving the office building, the registration information may be reset. Also, for example, once registration information is registered when entering a security gate, it may remain on the server unless an update is required later, and the registration process may be omitted when entering the office building the following day. Moreover, this embodiment can also be applied to use cases where registration information is generated in advance from a still image of the face taken correctly from the front and pre-registered on a server, and that registration information is used for verification in various locations, not just office buildings. Furthermore, its application is not limited to office buildings; it can also be applied to facilities such as train stations, retail stores, and banks. Furthermore, the recognition system of this embodiment can also be applied to use cases such as performing the aforementioned registration process when checking in at an airport, and then performing verification when passing through security checkpoints such as baggage inspection or boarding gates. In the case of application at an airport, the unique identification information of a person can be the identification information recorded in an IC passport such as a passport, and the facial image used during registration can also be the facial image recorded in the IC passport. Moreover, the identification information and facial image may be the identification information and facial image registered in the membership registration system of an airline or the like.
[0062] Furthermore, in the embodiment described above, an example was given in which one second model generated by distillation using the first model is used, but there may be multiple second models generated by distillation. For example, multiple learning models with different computational loads generated by distillation using the first model may be used as the second model. In this example, the registration device compares the feature information extracted by the first model and the multiple second models, and based on the result of this comparison, includes model information in the registration information indicating which of the first model or the multiple second models to use for matching. The matching device then determines which of the first model or the multiple second models to use based on the model information included in the registration information. In this example as well, the device may appropriately select which of the multiple lightweight second models to use depending on the memory capacity and processing power of the matching device. In this example as well, it is also possible to appropriately select which of the multiple lightweight second models to use in response to the detection of a predetermined time or a change in the external environment, as described above.
[0063] <Second Embodiment> Next, as a second embodiment, an example of performing matching processing on a person or object to be tracked will be described. In the second embodiment, an example of tracking a specific person using images from a surveillance camera will be given. The system configuration of the second embodiment is generally the same as that of Figure 1, but the network camera 150 is assumed to be a surveillance camera. Also, in the case of the second embodiment, the security gate 160 is not necessarily provided. Furthermore, since the functional block configuration of the information processing device in the second embodiment is the same as that of Figure 2, its illustration is omitted, and the explanation of the same configuration and processing as in the first embodiment is also omitted.
[0064] The registration process performed in the registration device 200 of the second embodiment will be explained using the flowchart in Figure 3. First, in S301, the image acquisition unit 201 of the registration device 200 acquires an image from the network camera 150. In the second embodiment, since the network camera 150 is a surveillance camera, an image of the area to be monitored is acquired. If a person passes through the area to be monitored, an image including that person will be acquired.
[0065] Next, in S302, the target detection unit 202 detects images of people (full body images, upper body images, etc.) from the images captured by the surveillance camera. In the second embodiment, the target detection unit 202 detects the person to be tracked and further detects a full body image of the detected person. Any known method can be used to detect the person to be tracked from the images captured by the surveillance camera. Examples include a method that detects areas with a high probability of being a person as candidate areas based on pre-set characteristics of a person such as the head, torso, and limbs, and a method for detecting person areas by learning, such as deep learning. Furthermore, known techniques such as those described in the first embodiment can be used for image detection. Note that the detected person is not limited to one person but may be multiple people.
[0066] In this method of detecting people from captured images, the detection result is often indicated by surrounding the detected person with a detection frame. Figure 6 shows an example where a person detected in a captured image is indicated by a detection frame 601. In the example in Figure 6, for example, a captured image is displayed on the display screen of the administrator terminal 130, and the detection frame 601 is displayed because a person has been detected in that captured image. In addition, whether or not to make a person detected in a captured image a tracking target may be specified by a user such as an administrator. Figure 6 shows an example where the administrator of the administrator terminal 130 specifies a tracking target. In the example in Figure 6, the administrator selects a person indicated by a detection frame 601 as a tracking target on the screen of the tablet terminal acting as the administrator terminal 130. That is, the administrator selects the person as a tracking target by touching the detection frame 601 on the tablet terminal screen.
[0067] Next, in S303, the first extraction unit 203 performs a first feature extraction process on the image of the person selected as the tracking target as described above, and the second extraction unit 204 performs a second feature extraction process on the image of the person selected as the tracking target, similar to the first embodiment. Furthermore, in the next step S304, the comparison unit 205 performs a comparison process between the first feature quantity and the second feature quantity, similar to the case of the first embodiment. Subsequently, as in the first embodiment, an evaluation is performed in S305 to determine whether the similarity is above a threshold, and based on the evaluation result, the registration information is registered in S306 or S307.
[0068] The tracking process in the second embodiment will be explained below using the flowchart in Figure 5. The tracking process in the second embodiment is assumed to be performed by the matching device 210 shown in Figure 2(B). In the first embodiment described above, in S501, the identification information acquisition unit 213 acquired the unique identification information of the person to be recognized, but in the second embodiment, the identification information acquisition unit 213 acquires the identification information of the person to be tracked. One possible method for acquiring the identification information of the person to be tracked is for the administrator to input the identification information (person ID) via the administrator terminal 130. For example, if the administrator wants to track a person with person ID 00001 from the registration information already stored in DB106, the administrator inputs person ID 00001.
[0069] Next, in S502, the registration information acquisition unit 214 reads the registration information from the DB 106 of the server 100 based on the identification information. The registration information is the information registered by the registration device 200 of the second embodiment as described above, and the registration information acquisition unit 214 acquires the registration information corresponding to the identification information from the server 100.
[0070] Next, in S503, the image acquisition unit 211 acquires an image of the area being monitored by the surveillance camera (network camera 150). In the following step S504, the target detection unit 212 detects the person to be tracked and their full-body image from the captured image acquired by the image acquisition unit 211. Prior to this, however, it detects candidate regions that could be the target of tracking from the captured image. That is, when a person is to be tracked, the target detection unit 212 detects regions from the image that are highly likely to be people as candidate regions. Methods for detecting regions that are highly likely to be people include known methods as described above, such as detection methods based on pre-designed human features such as the head, torso, and limbs, and human region detection methods that utilize learning, such as deep learning. Note that the detected person is not limited to one person but may be multiple people.
[0071] Next, in S504, the extraction unit 215 extracts features from the whole-body image detected by the target detection unit 212. In the second embodiment as well, features can be obtained in the same way as in the first embodiment by replacing the face image with a whole-body image. In other words, when the extraction unit 215 extracts features from the whole-body image, it uses a learning model corresponding to the model information included in the registration information acquired in S502 to extract the features.
[0072] Then, in S506 to S508, the matching unit 216 determines whether or not the person is the target of tracking based on the feature quantities extracted in S505, similar to the case of the first embodiment. Subsequently, in S509, the matching unit 216 performs processing according to the recognition result in S507 or S508. In the second embodiment, based on the recognition result, the matching unit 216 performs tracking display processing, for example, to display a rectangular frame surrounding the recognized tracking target on the screen in accordance with the movement of the tracking target.
[0073] According to the second embodiment, by pre-determining and registering whether to use the first or second model for the person to be tracked, detection and tracking can be performed using the appropriate learning model during person tracking. This makes it possible to achieve high-speed person tracking while maintaining tracking accuracy.
[0074] The present invention can also be realized by supplying a program that implements one or more of the functions of the above embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions. The embodiments described above are merely examples of how the present invention can be implemented, and the technical scope of the invention should not be interpreted as being limited by them. In other words, the present invention can be implemented in various ways without departing from its technical concept or its main features.
[0075] This embodiment includes the following configurations, methods, and programs. (Composition 1) A first extraction means performs a first extraction process on an image to be registered to extract first feature information representing the characteristics of the registered image, A second extraction means performs a second extraction process on the image to be registered to extract second feature information that can be compared with the first feature information, A comparison means for comparing the first characteristic information and the second characteristic information, A registration means that registers registration information associated with the registration target, which of the first extraction process or the second extraction process to be used for the matching process based on the results of the comparison by the comparison means, An information processing device characterized by having the following features. (Configuration 2) The information processing apparatus according to Configuration 1, characterized in that the second extraction process is an extraction process that requires less computational effort to extract the feature information than the first extraction process. (Composition 3) The information processing apparatus according to configuration 1 or 2, characterized in that the second extraction process is an extraction process that requires a smaller configuration to extract the feature information than the first extraction process. (Composition 4) The first extraction means extracts the first feature information by the first extraction process using the first model generated by learning, The information processing apparatus according to any one of configurations 1 to 3, characterized in that the second extraction means extracts the second feature information by a second extraction process using the second model generated by distillation using the first model. (Composition 5) The second extraction means extracts a plurality of second feature information using a plurality of second models generated by distillation using the first model, The comparison means compares the first characteristic information with the plurality of second characteristic information, The information processing apparatus according to configuration 4, characterized in that the registration means generates registration information that associates with the registration target whether to perform a matching process using a first extraction process using the first model or a second extraction process using one of the multiple second models, based on the result of the comparison by the comparison means. (Composition 6) The comparison means calculates the distance between the first feature information and the second feature information, The information processing apparatus according to any one of configurations 1 to 5, characterized in that the registration means uses the second extraction process for matching if the distance calculated by the comparison means is less than or equal to a predetermined threshold. (Composition 7) The information processing apparatus according to any one of configurations 1 to 6, characterized in that the registration means generates registration information which associates information indicating whether the first extraction process or the second extraction process is to be used for the matching process, and feature information extracted by either the first extraction process or the second extraction process which is to be used for the matching process, with identification information unique to the registration target. (Composition 8) Image acquisition means for acquiring a captured image including the aforementioned registration target, The system includes a detection means for detecting the image to be registered from the captured image, The first extraction means extracts the first feature quantity from the detected image to be registered, The information processing apparatus according to any one of configurations 1 to 7, characterized in that the second extraction means extracts the second feature quantity from the detected image to be registered. (Composition 9) The information processing apparatus according to configuration 8, characterized in that, when a predetermined time arrives, the image acquisition means acquires a new image of the target to be registered, and updates the registration information by performing processing from the detection means to the registration means based on the new image of the target to be registered. (Composition 10) The information processing apparatus according to configuration 9, characterized in that the predetermined period is a period when the seasons change or a period when the amount of daylight changes. (Composition 11) It has detection means for detecting changes in the external environment, The information processing apparatus according to configuration 8, characterized in that when a change in the external environment is detected, the image acquisition means acquires a new image of the target to be registered, and updates the registration information by performing processing from the detection means to the registration means based on the new image of the target to be registered. (Composition 12) The information processing device according to configuration 11, characterized in that the detection means detects at least one of the following as a change in the external environment: a change in ambient light, a change in the installation location of the imaging device that acquires the image of the target to be registered, a change in the field of view of the imaging device, or a change in the arrangement of objects around the imaging device. (Composition 13) The information processing device according to any one of configurations 1 to 12, characterized in that the image to be registered is a human face image. (Composition 14) A detection means for detecting an image of the object to be recognized from a captured image that includes the object to be recognized, A means for acquiring registration information that acquires registration information registered by an information processing device described in any one of configurations 1 to 13, An extraction means for extracting feature information from the image of the object to be recognized based on the registration information acquired by the registration information acquisition means, A matching means that uses the feature information extracted by the extraction means to verify whether the recognized target is the registered target, An information processing device characterized by having the following features. (Composition 15) The system has identification information acquisition means for acquiring identification information of the recognized target, The information processing apparatus according to configuration 14, characterized in that the registration information acquisition means acquires registration information of a registration target identified according to the identification information from a database in which the registration information is stored. (Composition 16) The information processing apparatus according to configuration 15, characterized in that the detection means tracks the recognition target identified by the identification information. (Method 1) An information processing method performed by an information processing device, A first extraction step involves performing a first extraction process on an image to be registered to extract first feature information that represents the characteristics of the image to be registered, A second extraction step involves performing a second extraction process on the image to be registered to extract second feature information that can be compared with the first feature information, A comparison step of comparing the first characteristic information and the second characteristic information, A registration step in which, based on the results of the comparison in the comparison step, the registration information is registered that associates with the registration target which of the first extraction process or the second extraction process will be used for the matching process, An information processing method characterized by having the following features. (Method 2) An information processing method performed by an information processing device, A detection step of detecting an image of the object to be recognized from a captured image that includes the object to be recognized, A registration information acquisition step in which registration information registered by an information processing device described in any one of configurations 1 to 13 is acquired, Based on the registration information obtained in the registration information acquisition step, an extraction step is performed to extract feature information from the image of the object to be recognized, A matching step is performed to determine whether the recognition target is the registration target or not, using the feature information extracted in the extraction step. An information processing method characterized by having the following features. (Program 1) A program for causing a computer to function as an information processing device described in any one of configurations 1 to 13. (Program 2) A program for causing a computer to function as an information processing device described in any one of configurations 14 to 16. [Explanation of Symbols]
[0076] 100: Server, 110: Information processing device, 200: Registration device, 203: First extraction unit, 204: Second extraction unit, 205: Comparison unit, 206: Registration unit, 213: Identification information acquisition unit, 214: Registration information acquisition unit, 215: Extraction unit, 216: Verification unit
Claims
1. A first extraction means performs a first extraction process on an image to be registered to extract first feature information representing the characteristics of the registered image, A second extraction means performs a second extraction process on the image to be registered to extract second feature information that can be compared with the first feature information, A comparison means for comparing the first characteristic information and the second characteristic information, A registration means that registers registration information associated with the registration target, which of the first extraction process or the second extraction process to be used for the matching process based on the results of the comparison by the comparison means, An information processing device characterized by having the following features.
2. The information processing apparatus according to claim 1, characterized in that the second extraction process is an extraction process that requires less computational power to extract the feature information than the first extraction process.
3. The information processing apparatus according to claim 1, characterized in that the second extraction process is an extraction process that requires a smaller configuration for extracting the feature information than the first extraction process.
4. The first extraction means extracts the first feature information by the first extraction process using the first model generated by learning, The information processing apparatus according to claim 1, characterized in that the second extraction means extracts the second feature information by a second extraction process using the second model generated by distillation using the first model.
5. The second extraction means extracts a plurality of second feature information using a plurality of second models generated by distillation using the first model, The comparison means compares the first characteristic information with the plurality of second characteristic information, The information processing apparatus according to claim 4, characterized in that the registration means generates registration information that associates with the registration target whether to perform a matching process using a first extraction process using the first model or a second extraction process using one of the multiple second models, based on the result of the comparison by the comparison means.
6. The comparison means calculates the distance between the first feature information and the second feature information, The information processing apparatus according to claim 1, characterized in that the registration means uses the second extraction process for matching when the distance calculated by the comparison means is less than or equal to a predetermined threshold.
7. The information processing apparatus according to claim 1, characterized in that the registration means generates registration information which associates information indicating whether the first extraction process or the second extraction process is to be used for the matching process, and feature information extracted by either the first extraction process or the second extraction process which is to be used for the matching process, with identification information unique to the registration target.
8. Image acquisition means for acquiring a captured image including the aforementioned registration target, The system includes a detection means for detecting the image to be registered from the captured image, The first extraction means extracts the first feature information from the detected image of the target to be registered, The information processing apparatus according to claim 1, wherein the second extraction means extracts the second feature information from the detected image of the target to be registered.
9. The information processing apparatus according to claim 8, characterized in that, when a predetermined time arrives, the image acquisition means acquires a new image of the target to be registered, and updates the registration information by performing processing from the detection means to the registration means based on the new image of the target to be registered.
10. The information processing apparatus according to claim 9, characterized in that the predetermined period is a period when the seasons change or a period when the amount of daylight changes.
11. It has detection means for detecting changes in the external environment, The information processing apparatus according to claim 8, characterized in that when a change in the external environment is detected, the image acquisition means acquires a new image of the target to be registered, and updates the registration information by performing processing from the detection means to the registration means based on the new image of the target to be registered.
12. The information processing apparatus according to claim 11, characterized in that the detection means detects at least one of the following as a change in the external environment: a change in ambient light, a change in the installation location of the imaging device that acquires the image of the target to be registered, a change in the field of view of the imaging device, or a change in the arrangement of objects around the imaging device.
13. The information processing apparatus according to claim 1, characterized in that the image to be registered is a facial image of a person.
14. A detection means for detecting an image of the object to be recognized from a captured image that includes the object to be recognized, A means for acquiring registration information that acquires registration information registered by an information processing device according to any one of claims 1 to 13, An extraction means for extracting feature information from the image of the object to be recognized based on the registration information acquired by the registration information acquisition means, A matching means that uses the feature information extracted by the extraction means to verify whether the recognized target is the registered target, An information processing device characterized by having the following features.
15. The system has identification information acquisition means for acquiring identification information of the recognized target, The information processing apparatus according to claim 14, characterized in that the registration information acquisition means acquires registration information of a registration target identified according to the identification information from a database in which the registration information is stored.
16. The information processing apparatus according to claim 15, characterized in that the detection means tracks the recognition target identified by the identification information.
17. An information processing method performed by an information processing device, A first extraction step involves performing a first extraction process on an image to be registered to extract first feature information that represents the characteristics of the image to be registered, A second extraction step involves performing a second extraction process on the image to be registered to extract second feature information that can be compared with the first feature information, A comparison step of comparing the first characteristic information and the second characteristic information, A registration step in which, based on the results of the comparison in the comparison step, the registration information is registered that associates with the registration target which of the first extraction process or the second extraction process will be used for the matching process, An information processing method characterized by having the following features.
18. An information processing method performed by an information processing device, A detection step of detecting an image of the object to be recognized from a captured image that includes the object to be recognized, A registration information acquisition step of acquiring registration information registered by an information processing device according to any one of claims 1 to 13, Based on the registration information obtained in the registration information acquisition step, an extraction step is performed to extract feature information from the image of the object to be recognized, A matching step is performed to determine whether the recognition target is the registration target or not, using the feature information extracted in the extraction step. An information processing method characterized by having the following features.
19. Computers, A first extraction means performs a first extraction process on an image to be registered to extract first feature information representing the characteristics of the registered image, A second extraction means performs a second extraction process on the image to be registered to extract second feature information that can be compared with the first feature information, A comparison means for comparing the first characteristic information and the second characteristic information, A registration means that registers registration information associated with the registration target, which of the first extraction process or the second extraction process to be used for the matching process based on the results of the comparison by the comparison means, A program that allows the device to function as an information processing device, including the necessary components.
20. Computers, A detection means for detecting an image of the object to be recognized from a captured image that includes the object to be recognized, A means for acquiring registration information that acquires registration information registered by an information processing device according to any one of claims 1 to 13, An extraction means for extracting feature information from the image of the object to be recognized based on the registration information acquired by the registration information acquisition means, A matching means that uses the feature information extracted by the extraction means to verify whether the recognized target is the registered target, A program that allows the device to function as an information processing device, including the necessary components.
Citation Information
Patent Citations
Image processing method and its device and recording medium
JP2002133414A
Parameter controller, parameter control program and multistage collation device
JP2010092119A
Image processing device, method of the same, image input device, image processing system, and program
JP2021005340A
Method and system for updating a user identification system
WO2022084039A1