Information processing device, information processing system, information processing method, and computer program
Patent Information
- Application Number
- JP2024567167
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2042-12-28
AI Technical Summary
Existing technologies face challenges in accurately distinguishing between real and fake videos, particularly those generated using techniques like deepfakes, which can deceive authentication systems by synthesizing images and videos that mimic real events.
An information processing device and system that detects landmarks from input videos and still images, generates a composite video by aligning these landmarks, and compares the input video with the composite video to determine if it is fake or real, utilizing feature extraction models to assess similarity and authenticity.
This approach effectively identifies fake videos by capturing characteristic differences, enabling accurate determination of video authenticity and enhancing the reliability of identity verification processes.
Abstract
Description
Information processing device, information processing system, information processing method, and recording medium
[0001] The present disclosure relates to the technical fields of an information processing device, an information processing system, an information processing method, and a recording medium.
[0002] Patent Document 1 describes a technology in which a first facial image is detected from a first captured image taken with a camera, the first facial image is nonlinearly deformed using a template facial image, the deformed first facial image is recorded as a registered facial image, a second facial image is detected from a second captured image taken with the camera, the second facial image is nonlinearly deformed using the template facial image, and the deformed second facial image is compared with the registered facial image.
[0003] Patent Document 2 describes a technology that detects a first face from an input search target image, calculates and records feature amounts of the detected first face, detects a second face from an input key image for search, calculates the facial angle of the second face, determines a synthesis pattern based on the calculated facial angle, generates a synthetic face image based on the determined synthesis pattern, calculates feature amounts of the second face using the generated synthetic face image, searches a database using the calculated multiple facial feature amounts as a query, and integrates the multiple search results, making it less likely that bias will occur in the search results even if the face is photographed at a different angle.
[0004] Patent document 3 describes a technology that processes multiple video data captured under different shooting conditions of a subject, synthesizes the processed video data to generate a video, and has multiple operating modes that perform at least one of a first subject detection process and a second subject detection process on either the processed video data or the generated video, switching the operating mode according to the state of the subject or the characteristics of the captured image.
[0005] Patent document 4 describes a technology that accepts input of multiple pieces of biometric information of a person to be registered, generates a composite image by combining the multiple pieces of biometric information of the person to be registered using a generated synthesis function, and stores the composite image.When an authentication request including the composite image combined using the multiple pieces of biometric information of the person to be authenticated is received from a terminal device, the composite image is compared with the stored composite image.
[0006] Patent Document 5 describes a technology in which, when facial image data is acquired, facial image data is generated from the facial image data by removing noise using a specific algorithm, differential image data is generated between the acquired facial image data and the generated facial image data, and based on information contained in the differential image data, it is determined whether the acquired facial image data is a composite image, and if it is not determined that the acquired facial image data is a composite image, it is determined whether the acquired facial image data is a composite image based on information contained in frequency data generated from the differential image data.
[0007] International Publication No. 2015 / 128961 International Publication No. 2013 / 176263 Japanese Patent Application Laid-Open No. 2015-012567 Japanese Patent Application Laid-Open No. 2010-044588 International Publication No. 2022 / 162760
[0008] An object of this disclosure is to provide an information processing device, an information processing system, an information processing method, and a recording medium that aim to improve upon the techniques described in prior art documents.
[0009] One aspect of the information processing device includes a detection means for detecting landmarks from an input video, a generation means for generating a composite video using the input still image and the landmarks, and a determination means for determining whether the input video is a composite video based on a comparison between the input video and the composite video.
[0010] One aspect of the information processing system includes an information processing device having a detection means for detecting landmarks from an input video, a generation means for generating a composite video using an input still image and the landmarks, and a determination means for determining whether the input video is a composite video based on a comparison between the input video and the composite video, as well as a comparison means for matching at least one of an object appearing in the still image and an object appearing in the input video, and an authentication means for authenticating the object based on at least one of the determination result by the determination means and the matching result by the comparison means.
[0011] One aspect of the information processing method detects landmarks from an input video, generates a composite video using the input still image and the landmarks, and determines whether the input video is a composite video based on a comparison between the input video and the composite video.
[0012] One aspect of the recording medium has recorded thereon a computer program for causing a computer to execute an information processing method that detects landmarks from an input video, generates a composite video using the input still image and the landmarks, and determines whether the input video is a composite video based on a comparison between the input video and the composite video.
[0013] FIG. 1 is a block diagram showing the configuration of an information processing device in a first embodiment. FIG. 2 is a block diagram showing the configuration of an information processing device in a second embodiment. FIG. 3 is a flowchart showing the flow of information processing operations of the information processing device in the second embodiment. FIG. 4 is a block diagram showing the configuration of an information processing device in a third embodiment. FIG. 5 is a flowchart showing the flow of information processing operations of the information processing device in the third embodiment. FIG. 6 is a block diagram showing the configuration of an information processing device in a fourth embodiment. FIG. 7 is a flowchart showing the flow of information processing operations of the information processing device in the fourth embodiment. FIG. 8 is a block diagram showing the configuration of an information processing system in a fifth embodiment. FIG. 9 is a flowchart showing the flow of identity verification operations of the information processing system in the fifth embodiment.
[0014] Hereinafter, embodiments of an information processing device, an information processing system, an information processing method, and a recording medium will be described with reference to the drawings. [1: First Embodiment]
[0015] A first embodiment of an information processing device, an information processing system, an information processing method, and a recording medium will be described below. Hereinafter, the first embodiment of the information processing device, the information processing system, the information processing method, and the recording medium will be described using an information processing device 1 to which the first embodiment of the information processing device, the information processing system, the information processing method, and the recording medium is applied. [1-1: Configuration of Information Processing Device 1]
[0016] 1 is a block diagram showing the configuration of an information processing device 1 according to the first embodiment. As shown in FIG. 1, the information processing device 1 includes a detection unit 11, a generation unit 12, and a determination unit 13.
[0017] The detection unit 11 detects landmarks from the input video. The generation unit 12 generates a composite video using the input still images and landmarks. The determination unit 13 determines whether the input video is a composite video based on a comparison between the input video and the composite video. [1-2: Technical Effects of the Information Processing Device 1]
[0018] The information processing device 1 in the first embodiment determines whether the input video is a composite video or not based on a comparison between the input video and a composite video generated using still images and landmarks, and is therefore able to accurately determine whether the input video is a composite video or not. [2: Second Embodiment]
[0019] Next, a second embodiment of the information processing device, the information processing system, the information processing method, and the recording medium will be described. Hereinafter, the second embodiment of the information processing device, the information processing system, the information processing method, and the recording medium will be described using an information processing device 2 to which the second embodiment of the information processing device, the information processing system, the information processing method, and the recording medium is applied. [2-1: Fake video]
[0020] There is a technology that synthesizes an image of a person based on information from a single photograph of the person's face. Deepfake, for example, is known as a technology for synthesizing images of a person. Deepfake is known as a technology that synthesizes fake videos that show things that did not actually happen. Hereinafter, videos that show things that did not actually happen may be referred to as fake videos. Furthermore, videos that show things that actually happened may be referred to as real videos. For example, a real video may include a video that shows actions performed by person B in front of the camera as captured by a camera. In contrast, a fake video may include a video in which actions performed by person B in front of the camera as captured by a camera are synthesized to appear as if they were performed by person A, who is different from person B.
[0021] There is a technique called reenactment, which generates a fake video in which the facial expression of a person in an original image is changed to a desired facial expression or the person in the original image is facing a desired direction. For example, a technique is known in which, based on at least one facial image of person A, the facial expression of person A in the facial image is changed to match the facial expression of person B, thereby generating a video that makes it appear as if person A is changing his or her facial expression (hereinafter, this may be referred to as "animating a still image").
[0022] By using an original image and an original video, a still image can be animated. The original video may be, for example, a video showing the actions of person B in front of the camera as captured by the camera. The original image may also be a still image showing person A, who is different from person B. When animating a still image, landmarks are first detected in the original image. Landmarks are also detected from each of the video frames that make up the original video. Next, for each video frame that makes up the original video, the original image is edited so that the landmarks in the original image and the landmarks in the corresponding video frame are aligned to generate a composite frame. The still image can be animated by connecting each of the generated composite frames. Landmarks may be characteristic positions of a subject appearing in the image.
[0023] For example, the facial orientation, expression, etc. of person A appearing in the original image can be changed using landmarks on the face of person B appearing in the original video, thereby synthesizing a video in which the facial orientation, expression, etc. of person A changes. The landmarks that change the facial orientation, expression, etc. of a person may be characteristic parts of the face. The characteristic positions on the face may be specific points on parts such as the eyes, nose, and mouth.
[0024] If the acquired video is similar to a composite video created using still images, the acquired video is likely to be a fake video. In this embodiment, this property is used to determine whether the video is a fake or not. That is, in this embodiment, a composite video is generated using still images, and the acquired video is compared with the composite image to determine whether the acquired video is a fake or not. [2-2: Configuration of Information Processing Device 2]
[0025] 2 is a block diagram showing the configuration of an information processing device 2 in the second embodiment. As shown in FIG. 2, the information processing device 2 includes a calculation device 21 and a storage device 22. The information processing device 2 may further include a communication device 23, an input device 24, and an output device 25. However, the information processing device 2 does not necessarily have to include at least one of the communication device 23, the input device 24, and the output device 25. The calculation device 21, the storage device 22, the communication device 23, the input device 24, and the output device 25 may be connected via a data bus 26.
[0026] The arithmetic device 21 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array). The arithmetic device 21 reads a computer program. For example, the arithmetic device 21 may read a computer program stored in the storage device 22. For example, the arithmetic device 21 may read a computer program stored in a computer-readable, non-transitory recording medium using a recording medium reading device (e.g., an input device 24 described later) not shown in the drawings that is included in the information processing device 2. The arithmetic device 21 may acquire (i.e., download or read) the computer program from a device (not shown) located outside the information processing device 2 via the communication device 23 (or another communication device). The arithmetic device 21 executes the read computer program. As a result, logical functional blocks for executing the operations to be performed by the information processing device 2 are realized within the arithmetic device 21. In other words, the arithmetic device 21 can function as a controller for realizing logical functional blocks for executing the operations (in other words, processing) to be performed by the information processing device 2.
[0027] 2 shows an example of logical functional blocks implemented within the arithmetic device 21 to execute information processing operations. As shown in FIG. 2 , the arithmetic device 21 implements a detection unit 211, which is a specific example of a "detection means" described in the appendix below, a generation unit 212, which is a specific example of a "generation means" described in the appendix below, a determination unit 213, which is a specific example of a "determination means" described in the appendix below, a still image reception unit 214, and a video reception unit 215. However, either the still image reception unit 214 or the video reception unit 215 does not necessarily have to be implemented within the arithmetic device 21. Details of the operations of the detection unit 211, the generation unit 212, the determination unit 213, the still image reception unit 214, and the video reception unit 215 will be described later with reference to FIG. 3 .
[0028] The storage device 22 can store desired data. For example, the storage device 22 may temporarily store a computer program executed by the arithmetic device 21. The storage device 22 may temporarily store data that the arithmetic device 21 temporarily uses when the arithmetic device 21 is executing a computer program. The storage device 22 may store data that the information processing device 2 stores long-term. The storage device 22 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device. In other words, the storage device 22 may include a non-temporary recording medium.
[0029] The communication device 23 is capable of communicating with devices external to the information processing device 2 via a communication network (not shown). The communication device 23 may be a communication interface based on standards such as Ethernet (registered trademark), Wi-Fi (registered trademark), Bluetooth (registered trademark), or USB (Universal Serial Bus).
[0030] The input device 24 is a device that accepts information input to the information processing device 2 from outside the information processing device 2. For example, the input device 24 may include an operation device (e.g., at least one of a keyboard, a mouse, and a touch panel) that can be operated by an operator of the information processing device 2. For example, the input device 24 may include a reading device that can read information recorded as data on a recording medium that can be externally attached to the information processing device 2.
[0031] The output device 25 is a device that outputs information to the outside of the information processing device 2. For example, the output device 25 may output information as an image. That is, the output device 25 may include a display device (a so-called display) capable of displaying an image showing the information to be output. That is, in the second embodiment, the information processing device 2 includes a display D as the output device 25. For example, the output device 25 may output information as sound. That is, the output device 25 may include an audio device (a so-called speaker) capable of outputting sound. For example, the output device 25 may output information on paper. That is, the output device 25 may include a printing device (a so-called printer) capable of printing desired information on paper. [2-3: Information Processing Operation Performed by Information Processing Device 2] The information processing operation performed by the information processing device 2 will be described with reference to FIG. 3. FIG. 3 is a flowchart showing the flow of the information processing operation performed by the information processing device 2.
[0032] 3, the still image receiving unit 214 receives input of a still image of the subject (step S20). The still image receiving unit 214 may acquire a still image including a facial region of the subject. The detection unit 211 may detect landmarks from the still image. The detection unit 211 may detect characteristic positions in the facial region as landmarks from the still image. The detection unit 211 may detect specific points of body parts such as the eyes, nose, and mouth as landmarks from the still image.
[0033] The video accepting unit 215 accepts input of an input video of the subject (step S21). The video accepting unit 215 may acquire an input video that includes a facial region of the subject. Each input frame constituting the input video may include a facial region of the subject. Here, it is unclear whether the input video acquired by the video accepting unit 215 is a real video or a fake video, but the video accepting unit 215 may acquire an input video that shows the same subject as the subject appearing in the still image acquired by the still image accepting unit 214.
[0034] The detection unit 211 detects landmarks from the input video (step S22). The detection unit 211 may detect landmarks from each input frame that constitutes the input video. The detection unit 211 may detect landmarks from the input video at positions equivalent to the landmarks detected from the still image.
[0035] The generation unit 212 generates a composite video using the input still images and landmarks (step S23). First, the generation unit 212 may generate a composite frame for each input frame constituting the input video by editing the landmarks of the still image so as to match the landmarks of the corresponding input frame. Next, the generation unit 212 may animate the still images by connecting each composite frame together to generate the composite video.
[0036] The determination unit 213 compares the input video with the composite video (step S24), and determines whether the input video and the composite video are more similar than a reference level (step S25). If the input video and the composite video are more similar than a reference level (step S25: Yes), the determination unit 213 determines that the input video is a forged fake video (step S26). On the other hand, if the input video and the composite video are not more similar than a reference level (step S25: No), the determination unit 213 determines that the input video is a genuine video that is not forged (step S27). [2-4: Technical Effects of Information Processing Device 2]
[0037] A composite video generated using still images and landmarks can capture the characteristics of a fake video generated using technology such as deepfake. The information processing device 2 in the second embodiment utilizes the property that if the characteristics of a composite video generated using input still images are similar to those of the input video, the input video is likely to be a fake video. The information processing device 2 can accurately determine whether the input video is a genuine video that has not been forged, or a fake video that has been forged. [3: Third Embodiment]
[0038] Next, a third embodiment of an information processing device, an information processing system, an information processing method, and a recording medium will be described. Hereinafter, the third embodiment of the information processing device, the information processing system, the information processing method, and the recording medium will be described using an information processing device 3 to which the third embodiment of the information processing device, the information processing system, the information processing method, and the recording medium is applied. [3-1: Configuration of Information Processing Device 3]
[0039] As shown in FIG. 4 , the information processing device 3 in the third embodiment includes a calculation device 21 and a storage device 22, similar to the information processing device 2 in the second embodiment. Furthermore, the information processing device 3 in the third embodiment may include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 in the second embodiment. However, the information processing device 3 does not necessarily include at least one of the communication device 23, the input device 24, and the output device 25. The information processing device 3 in the third embodiment differs from the information processing device 2 in the second embodiment in that the determination unit 313 includes an extraction unit 3131 and a calculation unit 3132. Other features of the information processing device 3 may be the same as those of the information processing device 2 in the second embodiment. Therefore, the following will describe in detail the differences from the previously described embodiments, and will omit descriptions of other overlapping features as appropriate. [3-2: Information Processing Operation Performed by Information Processing Device 3] The information processing operation performed by the information processing device 3 will be described with reference to FIG. 5 . FIG. 5 is a flowchart showing the flow of the information processing operation performed by the information processing device 3.
[0040] 5 , the still image receiving unit 214 receives input of a still image of the subject (step S20). The video receiving unit 215 receives input of an input video of the subject (step S21). The detection unit 211 detects landmarks from the input video (step S22). The generation unit 212 generates a composite video using the input still image and landmarks (step S23).
[0041] The extraction unit 3131 extracts input features of a facial image included in the input moving image and composite features of a facial image included in the composite moving image (step S30). The extraction unit 3131 may extract the input features and composite features using a feature extraction model. The feature extraction model may be a model that outputs features when a facial image is input. The feature extraction model may be a model used in a face recognition mechanism. The feature extraction model may be a model constructed by machine learning. The extraction unit 3131 may extract input features of a facial image included in each input frame constituting the input moving image and composite features of a facial image included in each composite frame corresponding to the input frames constituting the composite moving image.
[0042] The calculation unit 3132 calculates the similarity between the input feature and the combined feature (step S31). The calculation unit 3132 may calculate the similarity between the input feature and the combined feature for each input frame. The calculation unit 3132 may calculate the cosine similarity between the input feature and the combined feature. Furthermore, the calculation unit 3132 may calculate the Euclidean distance between the input feature and the combined feature.
[0043] The determination unit 313 determines whether the similarity between the input feature and the combined feature is equal to or greater than a predetermined value (step S32). For example, if the calculation unit 3132 calculates the cosine similarity, the determination unit 313 may determine whether the cosine similarity is closer to "1" than a predetermined value. Alternatively, if the calculation unit 3132 calculates the Euclidean distance, the determination unit 313 may determine whether the Euclidean distance is closer than a predetermined value.
[0044] Furthermore, for example, the determination unit 313 may determine whether an amount indicating an average degree of similarity for each input frame constituting the input moving image is equal to or greater than a predetermined value. Furthermore, the determination unit 313 may determine whether the number of input frames constituting the input moving image, each of which has a degree of similarity equal to or greater than a predetermined value, is equal to or greater than a predetermined value.
[0045] If the similarity between the input feature amount and the combined feature amount is equal to or greater than a predetermined value (step S32: Yes), the determination unit 313 determines that the input video is a forged fake video (step S26). On the other hand, if the similarity between the input feature amount and the combined feature amount is less than a predetermined value (step S25: No), the determination unit 313 determines that the input video is a genuine video that is not forged (step S27). [3-3: Technical Effects of the Information Processing Device 3]
[0046] The information processing device 3 in the third embodiment determines whether an input video is a fake video based on the similarity between the feature amounts of the input video and the feature amounts of a composite video synthesized using characteristic positions, thereby enabling highly accurate determination. [4: Fourth Embodiment]
[0047] Next, a fourth embodiment of an information processing device, an information processing system, an information processing method, and a recording medium will be described. Hereinafter, the fourth embodiment of an information processing device, an information processing system, an information processing method, and a recording medium will be described using an information processing device 4 to which the fourth embodiment of the information processing device, the information processing system, the information processing method, and the recording medium is applied. [4-1: Information Processing Operation Performed by Information Processing Device 4] The information processing operation performed by the information processing device 4 will be described with reference to FIG. 7. FIG. 7 is a flowchart showing the flow of the information processing operation performed by the information processing device 4.
[0048] 7 , the still image receiving unit 214 receives input of a still image of the subject (step S20). The video receiving unit 215 receives input of an input video of the subject (step S21). The detection unit 211 detects landmarks from the input video (step S22). The generation unit 212 generates a composite video using the input still image and landmarks (step S23).
[0049] The extraction unit 4131 extracts input features of a face image included in an input moving image and composite features of a face image included in a composite moving image (step S30). The calculation unit 4132 calculates the similarity between the input features and the composite features (step S31).
[0050] The extraction unit 4131 extracts a feature amount of a still image (step S40). The calculation unit 4132 calculates the similarity between the input feature amount and the feature amount of the still image (step S41).
[0051] The determination unit 413 determines whether the similarity between the input feature amount and the composite feature amount is greater than the similarity between the input feature amount and the feature amount of the still image (step S42). If the similarity between the input feature amount and the composite feature amount is greater than the similarity between the input feature amount and the feature amount of the still image (step S42: Yes), the determination unit 413 determines that the input video is a fake video (step S26). That is, if the input video and the composite video are more similar than the input video and the still image, the determination unit 413 determines that the input video is a composited video. On the other hand, if the similarity between the input feature amount and the composite feature amount is not greater than the similarity between the input feature amount and the feature amount of the still image (step S42: No), the determination unit 413 determines that the input video is a genuine video (step S27). [4-2: Technical Effects of the Information Processing Device 4]
[0052] In many cases, the degree of similarity between videos synthesized based on the same still images is higher than the degree of similarity between a still image and a video that captures the same subject. The information processing device 4 in the fourth embodiment can use this property to accurately determine whether an input video is a fake video. [5: Fifth Embodiment]
[0053] Next, a fifth embodiment of an information processing device, an information processing system, an information processing method, and a recording medium will be described. Hereinafter, the fifth embodiment of the information processing device, the information processing system, the information processing method, and the recording medium will be described using an information processing system 5 to which the fifth embodiment of the information processing device, the information processing system, the information processing method, and the recording medium is applied. [5-1: Electronic Identity Verification]
[0054] The use of online identity verification such as electronic Know Your Customer (eKYC) is increasing. Identity verification when opening an account at a financial institution or applying for a credit card is increasingly being carried out online using eKYC or other methods rather than face-to-face.
[0055] For example, eKYC may be performed at an imaging location prepared for the eKYC service. Alternatively, eKYC may be performed at any location using a terminal device that can be used by the target person, such as a smartphone with imaging and communication functions. An example of the eKYC flow will be described.
[0056] (Step 1) Capturing a still image: One side of the face photo on an identification document such as a driver's license or My Number card is captured. Images of the other side of the identification document and the thickness of the identification document may also be captured. (Step 2) Capturing a video: For example, the subject is instructed to move from a forward-facing position to a right-facing position, and the subject is captured as they move. (Step 3) Facial image matching: A match is made to determine whether the subject in the facial photo on the identification document and the subject standing in front of the camera are the same person. In facial image matching, features are extracted from the facial photo, features are extracted from the face area detected from the video, and the extracted features are compared to determine whether the similarity between the features is greater than or equal to a predetermined value. (Step 4) Impersonation determination: Whether the subject is impersonating someone else is determined based on the response to the instruction. (Step 5) Identity authentication: Whether identity verification was successful is determined based on the results of Steps 3 and 4.
[0057] As mentioned above, there is a technology that synthesizes an image of a person based on a single photograph of that person's face, posing a threat to impersonation in eKYC. Accurately determining whether a video is fake is an important issue in increasing the reliability of services such as eKYC. Input to eKYC includes facial images from official documents that serve as information for synthesizing fake videos. In other words, one possible method of impersonating an eKYC user is to synthesize and input fake videos based on limited information, such as facial images from official documents such as a driver's license or My Number card.
[0058] The information processing system 5 in the fifth embodiment may be applied to online identity verification such as eKYC. The information processing system 5 in the fifth embodiment may determine whether an input video is a fake video by comparing the input video with a composite video generated based on a facial photograph on an identity verification document such as a driver's license or My Number card. [5-2: Configuration of Information Processing System 5]
[0059] As shown in FIG. 8 , the information processing system 5 of the fifth embodiment includes a calculation device 21 and a storage device 22, similar to the information processing device 2 of the second embodiment to the information processing device 4 of the fourth embodiment. Furthermore, the information processing system 5 of the fifth embodiment may include a communication device 23, an input device 24, and an output device 25, similar to the information processing device 2 of the second embodiment to the information processing device 4 of the fourth embodiment. However, the information processing system 5 does not necessarily include at least one of the communication device 23, the input device 24, and the output device 25. The information processing system 5 of the fifth embodiment differs from the information processing device 2 of the second embodiment to the information processing device 4 of the fourth embodiment in that a matching unit 516, a spoofing determination unit 517, and an authentication unit 518 are further implemented within the calculation device 21. Other features of the information processing system 5 may be the same as at least one other feature of the information processing device 2 of the second embodiment to the information processing device 4 of the fourth embodiment. Therefore, in the following, only the parts that differ from the embodiments already described will be described in detail, and the description of other overlapping parts will be omitted as appropriate.
[0060] The information processing system 5 may be a system capable of performing biometric authentication of a subject. The information processing system 5 may be a device that performs a matching operation using an image and determines whether the subject is impersonating someone else using the image, thereby authenticating the subject. [5-3: Information Processing Operation Performed by Information Processing System 5] The information processing operation performed by the information processing system 5 will be described with reference to FIG. 9. FIG. 9 is a flowchart showing the flow of the information processing operation performed by the information processing system 5.
[0061] As shown in FIG. 9 , the still image receiving unit 214 receives input of a still image of the subject (step S20). In the fifth embodiment, the still image may be a facial image on an official document such as a driver's license or a My Number card. Step S20 may correspond to step 1 of the eKYC flow example described above. If the official document is a My Number card, the matching unit 616 may acquire a facial image stored in an integrated circuit incorporated in the My Number card. For example, the matching unit 616 may acquire a facial image on the My Number card read using a short-range wireless communication function installed in a smartphone carried by the subject. The video receiving unit 215 receives input of a video of the subject (step S21). Step S21 may correspond to step 2 of the eKYC flow example described above.
[0062] The matching unit 516 matches the facial image of the subject (step S50). If the still image is a facial image on an official document such as a driver's license or My Number card, the matching unit 516 may match the subject appearing in the still image with the subject appearing in the input video. In this case, if the matching between the subject appearing in the still image and the subject appearing in the input video fails, the information processing operation may be terminated. Alternatively, the matching unit 516 may match the received still image with a pre-registered registered facial image. Alternatively, the matching unit 516 may match the received input video with a pre-registered registered facial image. In other words, the matching unit 516 may match at least one of the subject appearing in the still image and the subject appearing in the input video. Step S50 may correspond to step 3 of the example eKYC flow described above.
[0063] Furthermore, since a still image and a composite video synthesized based on that still image are similar, there is a high possibility that the still image and the input video will be successfully matched even if the input video is a fake video.
[0064] The spoofing determination unit 517 performs spoofing determination using the input video (step S51). In the fifth embodiment, the input video may be used to determine whether it is a fake video and to perform spoofing determination. For example, the input video may be a video showing an action performed by a subject in response to an instruction from the information processing system 5. The information processing system 5 may instruct the face direction, gaze direction, and face position. The information processing system 5 may guide the gaze. The information processing system 5 may instruct a gesture. The spoofing determination unit 517 may perform active liveness determination using the input video. Step S51 may correspond to step 4 of the example of the eKYC flow described above.
[0065] The detection unit 211 detects landmarks from the input video (step S22). The generation unit 212 generates a composite video using the input still images and landmarks (step S23). The determination unit 213 compares the input video with the composite video (step S24), and determines whether the input video and the composite video are more similar than a reference level (step S25). If the input video and the composite video are more similar than a reference level (step S25: Yes), the determination unit 213 determines that the input video is a forged fake video (step S26). On the other hand, if the input video and the composite video are not more similar than a reference level (step S25: No), the determination unit 213 determines that the input video is a genuine video that is not forged (step S27).
[0066] If the determination unit 213 determines that the input video is not a fake video, the authentication unit 518 authenticates the subject based on the collation result by the comparison unit 516 and the determination result by the impersonation determination unit 517 (step S52). Alternatively, the authentication unit 518 may authenticate the subject on the condition that the determination unit 213 determines that the input video and the synthetic video are less similar than the reference value and the impersonation determination unit 517 determines that the subject has acted in accordance with the instructions. A case where the authentication unit 518 has successfully authenticated the subject may be a case where the subject's identity has been confirmed. Step S52 may correspond to step 5 of the example eKYC flow described above.
[0067] Note that the spoofing determination operation in step S51 and the fake video determination operations in steps S22 to S27 after the comparison operation in step S50 may be performed in parallel. [5-4: Technical Effects of Information Processing System 5] The information processing system 5 in the fifth embodiment determines whether an input video is a fake video or not, thereby enabling accurate identity verification. [6: Supplementary Notes] The following supplementary notes are further disclosed regarding the above-described embodiment. [Supplementary Note 1] An information processing device comprising: a detection means for detecting landmarks from an input video; a generation means for generating a composite video using an input still image and the landmarks; and a determination means for determining whether the input video is a composite video based on a comparison between the input video and the composite video. [Supplementary Note 2] The information processing device described in Supplementary Note 1, wherein the determination means determines that the input video is a composite video if the input video and the composite video are more similar than a reference value. [Supplementary Note 3] The information processing device according to Supplementary Note 1 or 2, wherein the determination means comprises: extraction means for extracting input features of a facial image included in the input video and composite features of a facial image included in the composite video; and calculation means for calculating a similarity between the input features and the composite features. [Supplementary Note 4] The information processing device according to Supplementary Note 1 or 2, wherein the determination means determines that the input video is a composite video if the input video and the composite video are more similar than the input video and the still image. [Supplementary Note 5] An information processing device comprising: detection means for detecting landmarks from the input video; generation means for generating a composite video using the input still image and the landmarks; and determination means for determining whether the input video is a composite video based on a comparison between the input video and the composite video; and an information processing system including: matching means for matching at least one of an object appearing in the still image and an object appearing in the input video; and authentication means for authenticating the object based on at least one of a determination result by the determination means and a matching result by the matching means.[Supplementary Note 6] An information processing method for detecting landmarks from an input video that has been input, generating a composite video using input still images and the landmarks, and determining whether or not the input video is a composite video based on a comparison between the input video and the composite video. [Supplementary Note 7] A recording medium having recorded thereon a computer program for causing a computer to execute an information processing method for detecting landmarks from an input video that has been input, generating a composite video using input still images and the landmarks, and determining whether or not the input video is a composite video based on a comparison between the input video and the composite video.
[0068] This disclosure may be modified as appropriate within the scope of the claims and the technical idea that can be read from the entire specification. Information processing devices, information processing systems, information processing methods, and recording media that involve such modifications are also included in the technical idea of this disclosure.
[0069] 1, 2, 3, 4 Information processing device 11, 211 Detection unit 12, 212 Generation unit 13, 213, 313, 413 Determination unit 214 Still image reception unit 215 Video image reception unit 3131, 4131 Extraction unit 3132, 4132 Calculation unit 5 Information processing system Matching unit 516 Spoofing determination unit 517 Authentication unit 518
Claims
1. A detection means for detecting landmarks from an input video; a generating means for generating a composite video using the input still image and the landmarks; a determination means for determining whether the input moving image is a composite moving image based on a comparison between the input moving image and the composite moving image; An information processing device comprising:
2. The determining means determines that the input moving image is a combined moving image when the input moving image and the combined moving image are more similar than a reference value. The information processing device according to claim 1 .
3. The determination means extraction means for extracting input features of a face image included in the input moving image and composite features of a face image included in the composite moving image; a calculation means for calculating a similarity between the input feature amount and the combined feature amount; have 3. The information processing device according to claim 1 or 2.
4. The determining means determines that the input moving image is a combined moving image when the input moving image and the combined moving image are more similar to each other than the input moving image and the still image.
3. The information processing device according to claim 1 or 2.
5. A detection means for detecting landmarks from an input video; a generating means for generating a composite video using the input still image and the landmarks; a determination means for determining whether the input moving image is a composite moving image based on a comparison between the input moving image and the composite moving image; An information processing device comprising: a matching means for matching at least one of an object appearing in the still image and an object appearing in the input video; an authentication means for authenticating the target based on at least one of the determination result by the determination means and the matching result by the matching means; An information processing system including:
6. Detect landmarks from the input video, Generate a composite video using the input still image and the landmarks; Based on a comparison between the input video and the composite video, it is determined whether the input video is a composite video. A computer-implemented information processing method.
7. On the computer, Detect landmarks from the input video, Generate a composite video using the input still image and the landmarks; Based on a comparison between the input video and the composite video, it is determined whether the input video is a composite video. A computer program for executing an information processing method.