Image processing device, image processing method, and program
The image processing device detects fraudulent impersonation by comparing faces across multiple locations and times, addressing the inability of existing methods to identify simultaneous appearances, thereby preventing security breaches.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-11-20
- Publication Date
- 2026-06-01
AI Technical Summary
Existing methods fail to detect fraudulent impersonation attempts using artificial facial objects before authentication, as they cannot identify irrational situations where a person's face appears simultaneously in multiple locations.
An image processing device that captures images, compares faces across multiple cameras, and determines if two or more faces belong to the same person based on location and time, issuing an alarm for suspicious activity.
Enables early detection of fraudulent authentication attempts by identifying simultaneous appearances of a person's face in different locations, preventing potential security breaches.
Smart Images

Figure 2026089406000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video processing apparatus, a video processing method, and a program, and particularly to video processing for discovering a situation in which an intention of impersonation or the like is suspected.
Background Art
[0002] In recent years, with regard to an apparatus for performing personal authentication, an attack in which a malicious person attempts to impersonate another person and perform unauthorized authentication has become a problem from the security perspective, and means for detecting such an attack at an early stage and preventing the occurrence of damage have been demanded.
[0003] In Patent Document 1, a method for determining validity based on the position and time difference when there are multiple requests for personal confirmation of the same person from a terminal has been proposed. Further, in Patent Document 2, a method for displaying a series of videos at the time of authentication of the same person from authentication information such as an IC card and making it easy to visually confirm has been proposed.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Non-Patent Documents
[0005]
Non-Patent Document 1
Non-Patent Document 2
[0006] One method of impersonating someone using facial recognition is to use an artificial object that mimics the face of the person being attacked, such as a photograph with a printed image of their face or a disguise mask worn on the head. Using the methods described in Patent Documents 1 and 2, it is possible to detect when authentication is attempted by the same person more than once simultaneously and to reject the authentication attempt.
[0007] On the other hand, when an attacker attempts to carry out such impersonation, it is likely that preparations will be made to use an artificial object that mimics a face during the preparation phase. In such a situation, the face of the same person will exist simultaneously in two or more separate locations: the target person and the person attempting the impersonation attack. This can be detected using video before authentication is attempted. From a security standpoint, it is desirable to detect fraudulent authentication as early as possible, but the methods described in Patent Documents 1 and 2 cannot detect such an irrational situation in advance before authentication is performed.
[0008] This invention has been made in view of the above-mentioned problems and aims to detect fraudulent authentication at an early stage. [Means for solving the problem]
[0009] An image processing device according to one embodiment of the present invention is characterized by comprising: an imaging means for capturing an image; a detection means for comparing a plurality of faces included in the image obtained from the imaging means and detecting that a combination of two or more faces that are presumed to belong to the same person is included; and a determination means for determining, based on at least one of the location and time of the image capture, that at least one of the faces included in the combination of faces is not an actual appearance of the person. [Effects of the Invention]
[0010] According to the present invention, fraudulent authentication can be detected at an early stage. [Brief explanation of the drawing]
[0011] [Figure 1] This is a block diagram showing the hardware configuration of the video processing device 100 according to Embodiment 1 of the present invention. [Figure 2] This is a block diagram showing the functional configuration of the video processing device 100 according to Embodiment 1 of the present invention. [Figure 3] This is a schematic diagram illustrating an example of operation of the video processing device 100 according to Embodiment 1 of the present invention. [Figure 4] This flowchart shows the process performed by the video processing device 100 according to Embodiment 1 of the present invention. [Figure 5] This is a block diagram showing the functional configuration of the video processing device 100 according to Embodiment 2 of the present invention. [Figure 6] This is a schematic diagram showing the overall configuration of system 30 according to Embodiment 3 of the present invention. [Figure 7] This is a block diagram showing the functional configuration and functional groups of system 30 according to Embodiment 3 of the present invention. [Modes for carrying out the invention]
[0012] The present invention will now be described in detail based on its embodiments with reference to the attached drawings. Note that the configurations shown in the following embodiments are merely examples, and the present invention is not limited to the illustrated configurations.
[0013] <Embodiment 1> FIG. 1 is a block diagram showing the hardware configuration of a video processing apparatus 100 according to Embodiment 1 of the present invention. The video processing apparatus 100 includes a control device 11, a storage device 12, an arithmetic device 13, an input device 14, an output device 15, and an I / F device 16. I / F is an abbreviation of Interface. A camera 101 is connected to the I / F device 16. The control device 11, the storage device 12, the arithmetic device 13, the input device 14, the output device 15, and the I / F device 16 are connected by, for example, an internal bus and are configured to be able to communicate with each other. The video processing apparatus 100 may include the camera 101.
[0014] The control device 11 is composed of an MPU (Micro Processing Unit) or the like and controls the entire video processing apparatus 100. The storage device 12 is composed of a recording medium such as a hard disk and an MPU or the like and holds programs and data necessary for the operation of the control device 11. The arithmetic device 13 is composed of an MPU or the like and executes necessary arithmetic processing based on the control from the control device 11.
[0015] The input device 14 is a human interface device or the like and inputs a user's operation to the video processing apparatus 100. The output device 15 is a display or the like and presents the processing result of the video processing apparatus 100 to the user. The I / F device 16 is a wired interface such as a universal serial bus, a local area network, an optical cable, or a wireless interface such as Wi-Fi or Bluetooth.
[0016] A camera 101 is connected to the I / F device 16, and a captured image taken by the camera 101 is input to the video processing device 100. In addition to connecting to the camera 101, the I / F device 16 also has functions such as transmitting the processing results obtained by the video processing device 100 to the outside, and inputting programs, data, etc. necessary for the operation of the video processing device 100 into the video processing device 100. Further, an entrance gate, an electronic lock of a door, etc. may be connected to the I / F device 16, and a signal may be transmitted to perform opening and closing, locking, and unlocking of the gate based on the processing results obtained by the video processing device 100.
[0017] FIG. 2 is a block diagram showing the functional configuration of the video processing device 100 according to Embodiment 1 of the present invention. The video processing device 100 includes a shooting unit 201, a shooting condition management unit 202, a face detection unit 203, a recording unit 204, a same face detection unit 205, a same person determination unit 206, a display unit 207, an operation unit 208, and a recording database 220.
[0018] The shooting unit 201 shoots a video using the camera 101 connected to the I / F device 16 and acquires the video. The shooting unit 201 shoots and acquires a video including a human face using the camera 101. Note that the shooting unit 201 may also acquire video data from the storage device 12 or a communication line connected to the I / F device 16.
[0019] The shooting condition management unit 202 manages parameters such as the installation position and shooting direction of the camera 101 for shooting the video acquired by the shooting unit 201, and controls information on the geographical location corresponding to the shot video.
[0020] The face detection unit 203 receives the video acquired by the shooting unit 201 and detects a human face from the input video using, for example, the method shown in Non-Patent Document 1. The recording unit 204 records the face information detected by the face detection unit 203 in the recording database 220. The recording database 220 is a database configured in the storage device 12, records the recording of the video shot by the shooting unit 201, and also records the information on the position and time on the image of the face detected by the face detection unit 203 together.
[0021] The identical face detection unit 205 compares images of two or more faces recorded in the video database 220 and detects faces that are presumed to belong to the same person. Specifically, the identical face detection unit 205 calculates feature quantities from the face images using a statistical model, calculates the similarity between the feature quantities, and determines that they belong to the same person if the similarity is higher than the matching threshold. For the calculation of feature quantities and the comparison between feature quantities, for example, the method shown in Non-Patent Document 2 is used.
[0022] The same person determination unit 206 determines whether the facial images, which have been estimated to belong to the same person by the same face detection unit 205, belong to the same person. The same person determination unit 206 makes this determination using the shooting parameters managed by the shooting condition management unit 202 and the time information recorded in the recording database 220. Details of the determination method will be described later.
[0023] The display unit 207 controls and displays on the output device 15 alarms, necessary information, and visual information of the user interface (UI) that should be notified to the user when using the video processing device 100. In addition to displaying alarms on the display unit 207, alarms may also be notified to the user through sound from a speaker or by flashing a warning light. The operation unit 208 receives user input from the input device 14 and transmits control information to each part of the video processing device 100.
[0024] Next, an example of the operation of the video processing device 100 according to this embodiment will be explained using Figure 3. Figure 3 is a schematic diagram illustrating an example of the operation of the video processing device 100 according to Embodiment 1 of the present invention. Figure 3(A) is a map of the area near the entrance of the entrance management system where the video processing device 100 is installed. Visitors proceed from area 301, which is an unrestricted area, complete the entrance procedure at terminal 302, and then enter the restricted area 304 through gate 303. Such an entrance management system can be used for passport checks at airports, ticket checks at stadiums, and so on.
[0025] Area 304 is equipped with multiple surveillance cameras (corresponding to camera 101) connected to the I / F device 16 of the video processing device 100. Figure 3(B) is a schematic diagram of video footage captured by a surveillance camera installed in Area 304, acquired by the camera 201, and then subjected to face detection by the face detection unit 203. If the video footage acquired by the camera 201 is footage of a corridor, it is expected that multiple visitors' faces will be detected in the video, as shown in Figure 3(B). Multiple such surveillance cameras are installed, and the footage is monitored in a separate room by a supervisor who is a user of the video processing device 100 for security and congestion control purposes.
[0026] Figure 3(C) is a schematic diagram of the image from a camera (corresponding to camera 101) installed on terminal 302, which is connected to the I / F device 16 of the video processing device 100. Visitors perform entry procedures in front of terminal 302, such as verifying their personal information and entry qualifications, and the procedure is completed through verification by staff using camera footage or facial recognition.
[0027] Here, let's assume that Figure 3(B) and Figure 3(C) are images taken at the same time, and that faces 305 and 306, shown in shaded areas, are so similar that they are presumed to belong to the same person. This would mean that two identical individuals exist simultaneously in separate locations: inside the restricted area 304 and in front of terminal 302. Naturally, this is illogical, as only one identical person can exist simultaneously in the world. However, if such a situation were actually detected, it would be suspected that one or both of the detected faces are fake. For example, it is possible that the images capture someone attempting to bypass authentication by impersonating the target, wearing a mask that resembles the target's face, or using a photograph of the target's face or one displayed on a tablet device. Therefore, such a situation is suspicious of fraudulent activity such as impersonation, and notifying the monitor of an alert helps prevent damage. The present invention aims precisely to automatically notify such alerts.
[0028] On the other hand, such detection results can also be obtained due to simple misidentification, such as misidentifying similar individuals like twins or detecting identical posters. For many applications, measures to reduce such false positives will likely be necessary.
[0029] The above is an example using surveillance cameras and terminal device cameras, but cameras installed for other purposes may be added, and the same applies to three or more cameras. For example, even if another surveillance camera capturing images like those in Figure 3(D) is installed in a different area outside of Area 301 or Figure 3(A), the same person can still be detected. Furthermore, cameras installed in geographically distant locations may be used via the I / F device 16.
[0030] Furthermore, as shown in Figures 3(B) and 3(D), if multiple people are captured in each image, it is possible to detect the same person, such as face 305 and face 307, by comparing the combinations of faces. Similarly, as shown in Figure 3(E), if two or more faces presumed to belong to the same person are captured simultaneously by the same camera, such as face 308 and face 309, it is considered appropriate to issue an alarm as this is an irrational situation.
[0031] To achieve the above-described operation, the processing procedure performed by the video processing device 100 according to this embodiment will be explained below using the flowchart in Figure 4. Figure 4 is a flowchart showing the processing performed by the video processing device 100 according to Embodiment 1 of the present invention. The processing in this flowchart is performed for the purpose of detecting faces that are presumed to belong to the same person from the video acquired by the camera unit 201 and issuing a warning when detected. The video processing device 100 repeats the processing in this flowchart, for example, once every 0.2 seconds. Note that the processing may be limited to certain time periods, such as during business hours. Furthermore, the frequency of execution may be increased during times when intensive monitoring is desired, such as before and after the start of entry.
[0032] In this configuration, N cameras are connected to the video processing device 100, designated as camera C1, camera C2, ..., camera CN. For each camera's video, for example, once every 0.1 seconds, the face detection unit 203 repeatedly detects faces in the video captured by the shooting unit 201, and the recording unit 204 records the position and time of the detected faces in the video recording database 220.
[0033] First, in step S401, the identical person detection unit 206 obtains a combination of camera Ci and camera Cj from the pair of two cameras that has not yet undergone processing in step S402 or later. i and j are integers between 1 and N. If the identical person detection unit 206 determines that there is no combination of camera Ci and camera Cj that has not yet undergone processing in step S402 or later, the processing in the flowchart of Figure 4 ends. If the identical person detection unit 206 is able to obtain a combination of camera Ci and camera Cj that has not yet undergone processing in step S402 or later, the processing in step S402 is executed. In this way, the processing from step S402 onward is executed for all combinations of two cameras. The processing in step S401 is an example of an imaging means for capturing images.
[0034] Next, in step S402, the same person determination unit 206 obtains information on camera Ci and camera Cj from the shooting condition management unit 202 and calculates the matching time T(Ci, Cj). The matching time is the time during which the same person may appear in both camera Ci and camera Cj when moving at a normal speed. The matching time T(Ci, Cj) is defined as T(Ci, Cj) = L(Ci, Cj) / 1.2 [s], where L(Ci, Cj) [m] is the physical distance between camera Ci and camera Cj and the average human walking speed is, for example, 1.2 [m / s]. In other words, if the face of the same person appears in both camera Ci and camera Cj within a time difference of less than the matching time, it is considered that an unnatural situation has occurred, rather than simply the person walking by.
[0035] Furthermore, the matching time T(Ci, Cj) can be determined in various ways, taking into account the installation environment. For example, if there are many children or elderly people among the subjects, a lower average walking speed may be assumed. Also, if the camera installation locations are far apart and it is expected that cars or trains will be used for transportation, the time may be determined based on those speeds.
[0036] Next, in step S403, the same-face detection unit 205 obtains information on faces detected in the current video from camera Ci from the recording database 220. The process in step S403 is an example of a face detection means for detecting faces from the video. Next, in step S404, the same-face detection unit 205 compares two pairs of faces from camera Ci obtained in step S403 and stores the pairs of faces determined to be the same person. If such pairs exist, it means that two or more faces of the same person are simultaneously captured on camera Ci.
[0037] In the next step S405, the identical face detection unit 205 determines whether there was a pair of faces that were determined to belong to the same person in step S404. If the identical face detection unit 205 determines that there was a pair of faces that were determined to belong to the same person, the process in step S409 is executed. If the identical face detection unit 205 determines that there was no pair of faces that were determined to belong to the same person, the process in step S406 is executed. The process in step S405 is an example of a detection means that compares multiple faces included in the video obtained from the imaging means and detects whether a combination of two or more faces that are presumed to belong to the same person is included. The detection means uses the faces detected by the face detection means to detect the appearance of multiple faces that are presumed to belong to the same person from the video.
[0038] Next, in step S406, the same-face detection unit 205 obtains information on faces recently detected between T(Ci, Cj) for camera Cj from the recording database 220. The process in step S406 is an example of a face detection means for detecting faces from the video. Next, in step S407, the same-face detection unit 205 compares the face from camera Ci obtained in step S403 with the face from camera Cj obtained in step S406 and stores pairs of faces determined to be the same person. If such pairs exist, it means that the faces of the same person have been captured on camera Ci and camera Cj more than once in a short period of time that would not be possible with the person's normal movement.
[0039] In the next step S408, the identical face detection unit 205 determines whether there was a pair of faces that were determined to belong to the same person in step S407. If the identical face detection unit 205 determines that there was a pair of faces that were determined to belong to the same person, the process in step S409 is executed. If the identical face detection unit 205 determines that there was no pair of faces that were determined to belong to the same person, the process proceeds to step S401 and the process is repeated. The process in step S408 is an example of a detection means that compares multiple faces included in the video obtained from the imaging means and detects whether there is a combination of two or more faces that are presumed to belong to the same person. The detection means uses the faces detected by the face detection means to detect the appearance of multiple faces that are presumed to belong to the same person from the video.
[0040] In step S409, the identical person determination unit 206 examines the pair of faces that the identical face detection unit 205 determined to be the same person in step S404 or step S407, and determines whether they are actually the same person. If the identical face detection unit 205 is using a statistical model trained on general faces, there is a possibility that people with similar faces but who are different people, such as twins, may be mistakenly detected as the same person. The identical person determination unit 206 performs this check to prevent such false alarms. Note that if the performance of the identical face detection unit 205 is sufficiently high, this step may be omitted. Specifically, the identical person determination unit 206 compares the pair of faces using a statistical model that has better twin discrimination ability than the one used by the identical face detection unit 205, and if it is still determined that they are the same person's faces, it is retained as information to alert; otherwise, it is discarded.
[0041] Furthermore, the detailed analysis performed in step S409 is not limited to twin discrimination and can encompass a variety of other processes. For example, instead of using the statistical model of the identical face detection unit 205, a statistical model with higher performance but higher computational complexity could be used. Alternatively, the same statistical model as that of the identical face detection unit 205 could be used, but with a stricter threshold set.
[0042] Furthermore, the review performed in step S409 may include checking whether consecutive frames of the video are identified as the same person's face to ensure that it is not a sudden false detection. Additionally, to avoid false positives on posters or other displays featuring the same person, location information such as walls may be taken into account during the determination, or certain individuals may be ignored as exceptions even if they are identified as the same person. Finally, the system may check the size and relative positions of faces, assuming situations such as when the person is wearing an ID card with their photo around their neck.
[0043] Furthermore, the scrutiny performed in step S409 may use a method to determine whether the image is a living being, in order to directly determine whether it is an impersonation. If one or both are determined not to be living beings, i.e., to be masks or printed materials, this should be recorded as information that should be used as a warning.
[0044] In the next step S410, the same person determination unit 206 determines whether there were any pairs of faces that were determined to be the same person in the detailed examination in step S409. If the same person determination unit 206 determines that there are no pairs of faces that have been determined to be the same person, the process returns to step S401 and is repeated. If the same person determination unit 206 determines that there is one or more pairs of faces that have been determined to be the same person, the process in step S411 is executed. The process in step S410 is an example of a determination means that determines, based on at least one of the location and time of imaging, that at least one of the faces included in the combination of faces is not an actual appearance of the person. Alternatively, the determination means may also determine based on location information of the location where the appearance of multiple faces presumed to be the same person was detected and the time of detection.
[0045] In step S411, the display unit 207 displays a warning to the user and presents the user with information such as the camera and image location and time of the pair of faces determined to be the same person. This allows the user to become aware of the suspicious situation and take appropriate action. The process in step S411 is an example of a notification means that notifies the monitor of the video when the determination means determines that the person has not actually appeared.
[0046] According to this embodiment, it is possible to warn the user of a suspicious situation in which the face of the same person appears simultaneously in different locations.
[0047] In steps S404 and S407 above, all faces appearing in the video are matched using a brute-force method. However, if the computational load is high, appropriate optimization can be performed. For example, as a preliminary step, a lightweight brute-force method can be used to exclude most pairs of faces that can be easily identified as belonging to different people, and then a highly accurate model can be used to determine if the remaining pairs belong to the same person. One lightweight method used for the initial brute-force is to use a statistical model that is lightweight but has low accuracy. In this case, it is desirable to use a model that is trained to tolerate a certain degree of omissions rather than misidentifying pairs of different people as belonging to the same person, as this is more likely to lead to missed results than misidentifying pairs of people as belonging to the same person. Another method is to estimate attributes instead of identifying people and exclude pairs with different attributes. Attributes can include age and gender, and depending on the location, they can also include, for example, the wearing of a specific hat or uniform.
[0048] Furthermore, in order to change the timing of the processing load, the recording unit 204 may simultaneously calculate and record the above-mentioned features and attributes when recording to the recording database 220, and then compare the results in steps S404 and S407. Alternatively, when the face detection unit 203 performs face detection, it may track faces and manage them according to the tracking trajectory, and for faces belonging to the same trajectory, it may be assumed that they have the same features or attributes, thus omitting sequential calculation. In this case, for each trajectory, a face image suitable for calculation is selected based on size, image quality, etc., and features or attributes are calculated using that face image and recorded in the recording database 220 as representative of that trajectory. Alternatively, the results may be determined by averaging or majority voting based on the results calculated from multiple face images.
[0049] <Embodiment 2> Embodiment 2 of the present invention describes a method that combines authentication with an entrance gate or the like. Note that parts common to Embodiment 1 will be omitted from the explanation, and only the differences will be described.
[0050] Figure 5 is a block diagram showing the functional configuration of the video processing device 100 according to Embodiment 2 of the present invention. In addition to the functional configuration of Embodiment 1 shown in Figure 2, the video processing device 100 according to Embodiment 2 further includes a registration unit 209, a matching unit 210, an entry management unit 211, and a registration database 230. Furthermore, the same face detection unit 205 according to Embodiment 2 has a different detection method compared to the same face detection unit 205 according to Embodiment 1.
[0051] The registration unit 209 receives input of facial images and personal information, the matching unit 210 calculates facial feature quantities to be used for matching, and stores them in the registration database 230. The registration database 230 is a database configured on the storage device 12, where information and feature quantities of multiple people are stored, and each person is assigned an identification person ID. The registration unit 209 and the registration database 230 are an example of a registration means for registering facial information of multiple people.
[0052] The matching unit 210 compares the input face with faces registered in the registration database 230 to determine whether the person is the same as one of the registered faces or not, and then determines the person ID. The matching unit 210 calculates feature quantities from the face image detected by the face detection unit 203 using a statistical model, calculates the similarity with the feature quantities included in the registration database 230, and determines that the person is the same if the similarity is higher than the matching threshold. For the calculation of feature quantities and the comparison of feature quantities, for example, the method shown in Non-Patent Document 2 is used. The matching unit 210 is an example of a matching means that matches the input face image with a person registered in the registration means.
[0053] The entrance control unit 211 is equipped with an entrance gate and controls the opening and closing of the gate based on the verification result of the verification unit 210, so that only authorized persons can pass through. The entrance control unit 211 may also control the locking and unlocking of doors instead of gates, or it may control the illumination of lights or the generation of electronic sounds in conjunction with authentication.
[0054] Furthermore, the identical face detection unit 205 according to Embodiment 2 detects faces of the same person when the person ID determined by the matching unit 210 matches.
[0055] The processing procedure in Embodiment 2 is the same as in Embodiment 1, except for steps S404 and S407, and is executed repeatedly, for example, once every 0.2 seconds, for the video captured by the shooting unit 201. N cameras are connected to the video processing device 100 according to Embodiment 2, and these are designated as camera C1, camera C2, ..., camera CN. As in Embodiment 1, the face detection unit 203 repeatedly detects faces in the video captured by the shooting unit 201 for the video from each camera, and the recording unit 204 records the position and time of the detected faces in the video recording database 220. In Embodiment 2, the matching unit 210 further compares each detected face with faces registered in the registration database 230 and saves the person ID (or "not found") of the person determined to be the same person.
[0056] In steps S404 and S407 of Embodiment 2, the same face detection unit 205 compares the person IDs recorded in the video database 220 and determines that they are the same person's face if they are the same. If there are no matches, it is determined that they are not the same person. In steps S405 and S408 of Embodiment 2, if there are multiple faces appearing in the imaging means that have been matched by the matching means and are of the same person, it is detected as the appearance of multiple faces that are presumed to be the same person.
[0057] The matching unit 210 according to Embodiment 2 can be used, for example, to register and authenticate individuals who are allowed to enter an entrance gate. In Embodiment 2, computational costs can be reduced by using authentication information to determine if multiple individuals are the same person. Furthermore, if an attempt is made to impersonate someone with the aim of bypassing authentication, the target of the impersonation is likely to be a registered individual, so not performing identity determination on unrelated individuals who are not authenticated also contributes to reducing computational costs.
[0058] Furthermore, in Embodiment 2, the authentication history may be used when calculating the matching time T(Ci,Cj) in step S402. For example, if the time of entry and exit of a person into an area with restricted access can be obtained from the authentication history, then that person must have been inside the restricted area. From this, the matching time may be set to the range that is assumed to be within walking distance from the time of entry and exit from the restricted area. Alternatively, an alarm may be triggered if the person is detected by a camera outside the restricted area while they are inside the restricted area.
[0059] <Modified form of Embodiment 2> Here, we show one modification of Embodiment 2. In this modification, the same face detection unit 205 determines that a face belongs to the same person not only when the person ID is the same, but also when the person is similar in appearance. This detection means detects the appearance of multiple faces that are presumed to belong to the same person, even if the result of matching the faces that appear in the video with the matching means is not that of the same person but of similar faces.
[0060] In this modified version, the registration unit 209 searches for similar individuals by comparing them with individuals in the registration database 230 when registering a person. Similar individuals are defined, for example, as faces whose similarity score, when comparing feature quantities, is above a predetermined value. For example, twins or parent and child are different people, but tend to have a high similarity score when compared to each other, making misidentification more likely. If a similar individual is found during registration, the registration database 230 records the person ID of the similar individual in both the record of the newly added person and the record of the similar individual.
[0061] In this modified version, the same-face detection unit 205 refers to the registration database 230 when the person IDs are different, and if it is recorded as the person ID of a similar person, it is treated as the "face of the same person" and further examined by the same-person determination unit 206. This makes it possible to reduce the chances of overlooking a person by examining the possibility that they are the same person, even if misidentification of people with similar faces occurs.
[0062] <Embodiment 3> Embodiment 3 describes a method that combines multiple authentication devices, such as entrance gates. Note that the parts common to Embodiments 1 and 2 will be omitted from the explanation, and only the differences will be described.
[0063] Figure 6 is a schematic diagram showing the overall configuration of system 30 according to Embodiment 3 of the present invention. System 30 according to Embodiment 3 consists of two authentication gate systems 601 and 602 and an integrated server 603. System 30 realizes the functions of the video processing device 100 according to the present invention. Authentication gate system 601 consists of a camera 611 that functions as a shooting unit 201, a gate 612 that functions as an entry management unit 211, and a computer 613. Computer 613 functions as a face detection unit 203, a recording unit 204, a registration unit 209, a matching unit 210, a recording database 220, and a registration database 230. Authentication gate systems 601 and 602 have an I / F device 16 connected to the integrated server 603. Authentication gate system 602 also consists of a camera 621, a gate 622, and a computer 623 that correspond to the same functional configuration as authentication gate system 601.
[0064] Authentication gate systems 601 and 602 are systems that perform entry authentication at entrance gates, and are installed in different locations and operate independently. Authentication gate systems 601 and 602 are each connected to the integrated server 603 (described later) via the I / F device 16. Authentication gate systems 601 and 602 and the integrated server 603 may be connected via a network and may be installed in geographically separated locations.
[0065] Authentication gate systems 601 and 602 each have independent recording databases 220 and registration databases 230, and furthermore, the statistical models used in the verification unit 210 that authenticates entrants are also different. The features calculated from different statistical models cannot be compared with each other.
[0066] The integrated server 603 is a computer connected to the authentication gate systems 601 and 602 via an I / F device 16. The integrated server 603 functions as a shooting condition management unit 202, a same face detection unit 205, a same person determination unit 206, a display unit 207, and an operation unit 208.
[0067] The integrated server 603 is a device intended to issue a warning if a face presumed to belong to the same person appears simultaneously in both the authentication gate system 601 and the authentication gate system 602. The shooting condition management unit 202 in the integrated server 603 queries, acquires, and stores the parameters of cameras 611 and 621 installed in the authentication gate systems 601 and 602, and manages them so that they can be used by the integrated server 603. The statistical model used by the same face detection unit 205 in the integrated server 603 may be different from the one used by the matching unit 210 of the authentication gate systems 601 and 602, and will be described as being different below.
[0068] Figure 7 is a block diagram showing the functional configuration and functional groups of the system 30 according to Embodiment 3 of the present invention. Functional components identical to those in Figure 2 are denoted by the same reference numerals and their descriptions are omitted. The shooting unit 201, face detection unit 203, recording unit 204, registration unit 209, matching unit 210, entry management unit 211, recording database 220, and registration database 230 are functions performed by the authentication gate system 601 and authentication gate system 602. The shooting condition management unit 202, identical face detection unit 205, identical person determination unit 206, display unit 207, and operation unit 208 are functions performed by the integrated server 603.
[0069] In Embodiment 3, the integrated server 603 executes the flow shown in Figure 4. In steps S403 and S406 of Embodiment 3, the identical face detection unit 205 communicates with the authentication gate systems 601 and 602 to obtain the current face image from the recording database 220 held by each authentication gate system. In steps S404 and S407 of Embodiment 3, the identical face detection unit 205 extracts and compares feature quantities from the face images obtained from the authentication gate systems 601 and 602 to determine whether the faces belong to the same person. According to Embodiment 3, in this way, even between the independently operating authentication gate systems 601 and 602, a warning can be issued if faces presumed to belong to the same person appear simultaneously.
[0070] Although Embodiment 3 showed two examples of authentication gate systems, the same detection of the same person can be achieved with three or more systems, or by using a facial recognition system other than an authentication gate.
[0071] Furthermore, although the integrated server 603 is provided as an independent server in Embodiment 3, the authentication gate system 601 or 602 may also function as the integrated server 603. For example, the authentication gate system 601 may also function as the integrated server 603. In this case, the authentication gate system 601 will execute the flow shown in Figure 4 in addition to the control processing of the authentication gate, and the authentication gate system 602 will transmit a face image to the authentication gate system 601. In this case, instead of the identical face detection unit 205 performing feature extraction in steps S404 and S407, the feature quantities extracted by the matching unit 210 of the authentication gate system 601 can be reused. That is, feature quantities are extracted from the face image received from the authentication gate system 602 using the same statistical model used by the matching unit 210 of the authentication gate system 601, and compared with the previously extracted feature quantities from the face image obtained by the authentication gate system 601. This way, the feature extraction process for the face image of the authentication gate system 601 is omitted once, making it more efficient. Furthermore, this configuration remains valid even if the roles of authentication gate systems 601 and 602 are swapped, allowing both authentication gate systems 601 and 602 to function as the integrated server 603, each handling half of the process of detecting the same person. Distributing the processing in this way reduces the load on each system, thus improving throughput. This can also be extended to cases with three or more facial recognition systems, where some or all of the systems are distributed to perform the process of detecting the same person in parallel.
[0072] <Modified form of Embodiment 3> Here, we show one modified example of Embodiment 3. In this modified example, the face detection unit 205 of the integrated server 603 performs matching using feature quantities calculated by the matching units 210 of the authentication gate systems 601 and 602, rather than using face images captured by the imaging unit 201.
[0073] In steps S403 and S406 of this modified example, the identical face detection unit 205 communicates with the authentication gate systems 601 and 602 to obtain the facial feature quantities calculated by the matching unit 210 of each authentication gate system. The processing in steps S403 and S406 for the authentication gate system 601 is an example of a first calculation means for calculating a first feature quantity from the faces included in the video. The processing in steps S403 and S406 for the authentication gate system 602 is an example of a second calculation means for calculating a second feature quantity from the faces included in the video. Then, in steps S404 and S407 of this modified example, the identical face detection unit 205 determines whether the faces belong to the same person by comparing the feature quantities. At this time, if the statistical models used by the matching unit 210 of the authentication gate systems 601 and 602 are different, the feature quantities cannot be directly compared, so the feature quantities are converted into a comparable format. Such a converter can be created by learning features obtained in advance from the same face using two statistical models, and comparison can be made by converting one to the other. The processing in steps S404 and S407 is an example of a detection means for detecting the appearance of multiple faces that are estimated to belong to the same person by converting the first feature and the second feature into a comparable format and comparing them.
[0074] Alternatively, a simpler method is to use only the common parts when the statistical models are partially common. For example, suppose that in authentication gate system 601, the feature vectors obtained from model A and model B are combined in series to form the feature vectors, and in authentication gate system 602, the feature vectors obtained from model A and model C are combined in series to form the feature vectors. In this case, the identical face detection unit 205 of the integrated server 603 can detect identical faces using common features by using only the feature vectors of the dimension derived from model A.
[0075] <Other Embodiments> Although embodiments of the present invention have been described in detail above, the present invention can take the form of, for example, a system, apparatus, method, program, or recording medium (storage medium). Specifically, it may be applied to a system consisting of multiple devices (for example, a host computer, interface devices, imaging devices, web applications, etc.), or to an apparatus consisting of a single device.
[0076] Furthermore, the object of the present invention can also be achieved as follows: a recording medium (or storage medium) containing program code (computer program) of software that realizes the functions of the embodiments described above is supplied to a system or device. The storage medium is a computer-readable storage medium. The computer (or CPU or MPU) of the system or device then reads and executes the program code stored on the recording medium. In this case, the program code read from the recording medium itself realizes the functions of the embodiments described above, and the recording medium containing that program code constitutes the present invention.
[0077] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0078] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its essence.
[0079] This embodiment includes the following configurations and methods. (Composition 1) An imaging means for capturing images, A detection means that compares multiple faces contained in the image obtained from the imaging means and detects that a combination of two or more faces that are presumed to belong to the same person is included. A determination means for determining, based on at least one of the location and time of the image capture, that at least one of the faces included in the combination of faces is not an actual appearance of the person, A video processing device characterized by comprising: (Configuration 2) The determination means further comprises a notification means for notifying the monitor of the video when it determines that the person has not actually appeared. The video processing apparatus according to configuration 1, characterized in that... (Composition 3) The determination means makes a determination based on location information of the place where the appearance of multiple faces presumed to belong to the same person was detected, and the time of detection. The image processing apparatus according to configuration 1 or configuration 2, characterized in that it is a video processing apparatus. (Composition 4) A registration method for registering facial information of multiple people, A matching means for comparing the input facial image with a person registered in the registration means, Furthermore, The detection means detects the appearance of multiple faces presumed to belong to the same person when, among the faces appearing in the imaging means, the matching means determines that multiple faces belong to the same person. An image processing apparatus according to any one of configurations 1 to 3, characterized by the above. (Composition 5) The detection means detects the appearance of multiple faces presumed to belong to the same person, even if, among the faces appearing in the video, the result of matching by the matching means is that the faces are not those of the same person but are similar. The video processing apparatus according to configuration 4, characterized by the above. (Composition 6) The system further comprises face detection means for detecting faces from the aforementioned video, The detection means uses the faces detected by the face detection means to detect the appearance of multiple faces that are presumed to belong to the same person from the video. An image processing apparatus according to any one of configurations 1 to 5, characterized by the above. (Composition 7) A first calculation means for calculating a first feature quantity from the faces contained in the aforementioned video, A second calculation means for calculating a second feature quantity from the face contained in the aforementioned video, Furthermore, The detection means detects the appearance of multiple faces that are presumed to belong to the same person by converting the first feature and the second feature into a comparable format and comparing them. An image processing apparatus according to any one of configurations 1 to 6, characterized by the above. (Method 1) The imaging process involves capturing images, A detection step involves comparing multiple faces included in the image obtained in the imaging step and detecting whether a combination of two or more faces that are presumed to belong to the same person is included. A determination step of determining, based on at least one of the location and time of the image capture, that at least one of the faces included in the combination of faces is not an actual appearance of the person, A video processing method characterized by comprising the following: (Program 1) Computers, An imaging means for capturing images, A detection means that compares multiple faces included in the image obtained from the imaging means and detects that a combination of two or more faces that are presumed to belong to the same person is included, and A determination means for determining, based on at least one of the location and time of the image capture, that at least one of the faces included in the combination of faces is not an actual appearance of the person. A program characterized by being designed to function as such. [Explanation of Symbols]
[0080] 100 Video Processing Devices 201 Photography Department 202 Shooting Conditions Management Department 203 Face detection unit 204 Records Department 205 Identical Face Detection Unit 206 Same person determination section 207 Display section 208 Operation section 220 Recording Database
Claims
1. An imaging means for capturing images, A detection means that compares multiple faces contained in the image obtained from the imaging means and detects that a combination of two or more faces that are presumed to belong to the same person is included. A determination means for determining, based on at least one of the location and time of the image capture, that at least one of the faces included in the combination of faces is not an actual appearance of the person, A video processing device characterized by comprising:
2. The determination means further comprises a notification means for notifying the monitor of the video when it determines that the person has not actually appeared. The image processing apparatus according to feature 1.
3. The determination means makes a determination based on location information of the place where the appearance of multiple faces presumed to belong to the same person was detected, and the time of detection. The image processing apparatus according to feature 1.
4. A registration method for registering facial information of multiple people, A matching means for comparing the input facial image with a person registered in the registration means, Furthermore, The detection means detects the appearance of multiple faces presumed to belong to the same person when, among the faces appearing in the imaging means, the matching means determines that multiple faces belong to the same person. The image processing apparatus according to feature 1.
5. The detection means detects the appearance of multiple faces presumed to belong to the same person, even if, among the faces appearing in the video, the result of matching by the matching means is that the faces are not those of the same person but are similar. The image processing apparatus according to feature 4.
6. The system further comprises face detection means for detecting faces from the aforementioned video, The detection means uses the faces detected by the face detection means to detect the appearance of multiple faces that are presumed to belong to the same person from the video. The image processing apparatus according to feature 1.
7. A first calculation means for calculating a first feature quantity from the face contained in the aforementioned video, A second calculation means for calculating a second feature quantity from the faces contained in the aforementioned video, Furthermore, The detection means detects the appearance of multiple faces that are presumed to belong to the same person by converting the first feature quantity and the second feature quantity into a comparable format and comparing them. The image processing apparatus according to feature 1.
8. The imaging process involves capturing images, A detection step involves comparing multiple faces included in the image obtained in the imaging step and detecting whether a combination of two or more faces that are presumed to belong to the same person is included. A determination step of determining, based on at least one of the location and time of the image capture, that at least one of the faces included in the combination of faces is not an actual appearance of the person, A video processing method characterized by comprising the following:
9. Computers, An imaging means for capturing images, A detection means that compares multiple faces included in the image obtained from the imaging means and detects that a combination of two or more faces that are presumed to belong to the same person is included, and A determination means for determining, based on at least one of the location and time of the image capture, that at least one of the faces included in the combination of faces is not an actual appearance of the person. A program characterized by being designed to function as such.