Determining device, determination method, and computer program
Patent Information
- Application Number
- JP2024087499
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2025-12-11
- Estimated Expiration
- 2044-05-29
AI Technical Summary
Existing technologies fail to accurately distinguish between real and synthetic videos generated using facial synthesis techniques like deepfake.
A determination device and method that analyzes frame images from a video to determine similarity with real facial images and trends in similarity data, distinguishing between videos captured naturally and those generated using synthesis technology.
Enables accurate identification of synthetic videos by analyzing frame-to-frame similarity and trend consistency, differentiating between real and synthetic content.
Smart Images

Figure 2025180287000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a determination device, a determination method, and a computer program. [Background technology]
[0002] It has been a long-standing practice to process moving images by detecting a person's face in the moving image and performing image processing on the detected face image. For example, the technology described in Patent Document 1 makes it possible to generate a composite moving image by detecting a person's face in an image and combining an animation element such as glasses at the position of the face. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-151975 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, technology for generating synthetic video has become more advanced, making it possible to synthesize video that does not actually exist by using a specific person's facial image. For example, by superimposing a facial image of another person onto the face of a person in an existing video, and then deforming and superimposing that person's facial image to match the face's orientation and expression, a synthetic video that does not actually exist can be generated. Furthermore, the video itself can be synthesized using generative AI (artificial intelligence) technology, and by superimposing a specific person's facial image onto the face, a synthetic video that does not actually exist can be generated. A specific example of this technology is an AI-based technology known as deepfake.
[0005] The present invention has been made in view of the above-mentioned circumstances, and provides a technique that makes it possible to determine whether a certain moving image is a composite moving image generated by synthesis. [Means for solving the problem]
[0006] One aspect of the present invention is a determination device that includes a control unit that acquires multiple frame images, which are still images, from a target video that is a video to be determined, performs a first process to acquire the degree of similarity between the frame images and real facial images of facial images appearing in the video, performs a second process to acquire the degree of similarity between the frame images and other frame images obtained from the target video, and determines that the target video is a video in which the real facial images were captured if the degree of similarity acquired in the first process and the degree of similarity acquired in the second process are similar based on a predetermined criterion, and determines that the target video is a video obtained using synthesis technology if they are not similar based on the predetermined criterion.
[0007] One aspect of the present invention is the above-mentioned judgment device, wherein the control unit acquires a group of matching data in which multiple matching degrees are arranged in the order in which the frame images to be processed are played back in the first process and the second process, and makes a judgment based on whether the matching data group is similar according to the specified criteria.
[0008] In one aspect of the present invention, in the determination device, the control unit makes a determination based on whether or not the trend of change in the group of coincidence data is similar based on the predetermined criterion.
[0009] One aspect of the present invention is the above-mentioned determination device, wherein the control unit determines the tool used to generate the target moving image when assuming that the target moving image is a moving image obtained using the synthesis technology.
[0010] One aspect of the present invention is a determination method comprising: a frame image acquisition step in which a computer acquires a plurality of frame images, which are still images, from a target video, which is a video to be determined; a first processing step in which the computer executes a first process in which the computer acquires a degree of similarity between the frame images and a genuine face image of a face image appearing in the video; a second processing step in which the computer executes a second process in which the computer acquires a degree of similarity between the frame images and another frame image obtained from the target video; and a determination step in which the computer determines that the target video is a video in which the genuine face image was captured if the degree of similarity acquired in the first process and the degree of similarity acquired in the second process are similar according to a predetermined criterion, and determines that the target video is a video obtained using a synthesis technique if the degree of similarity acquired in the first process and the degree of similarity acquired in the second process are not similar according to the predetermined criterion.
[0011] One aspect of the present invention is a computer program for causing a computer to function as a determination device, which includes a control unit that acquires multiple frame images, which are still images, from a target video that is a video to be determined, performs a first process to obtain the degree of similarity between the frame images and real facial images of facial images appearing in the video, performs a second process to obtain the degree of similarity between the frame images and other frame images obtained from the target video, and determines that the target video is a video in which the real facial images were captured if the degree of similarity obtained in the first process and the degree of similarity obtained in the second process are similar based on a predetermined criterion, and determines that the target video is a video obtained using synthesis technology if they are not similar based on the predetermined criterion. [Effects of the Invention]
[0012] According to the present invention, it is possible to determine whether a certain video is a composite video generated by synthesis. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a diagram showing an outline of a moving image. [Figure 2]FIG. 10 is a diagram showing a specific example of a first process performed on N frame images. [Figure 3] FIG. 10 is a diagram showing a specific example of a group of coincidence data obtained in the first process. [Figure 4] 10A and 10B are diagrams showing a specific example of a second process performed on N frame images. [Figure 5] FIG. 10 is a diagram showing a specific example of a group of coincidence data (coincidence graph) obtained in the second process. [Figure 6] FIG. 10 is a diagram illustrating a modified example of the second process. [Figure 7] This is a diagram showing two similarity graphs obtained for target moving images obtained using synthesis technology such as deepfake. [Figure 8] FIG. 10 is a diagram showing two graphs of degree of coincidence obtained for a target video sequence obtained by filming a real person (the person himself / herself). [Figure 9] FIG. 10 is a diagram showing two graphs of degree of coincidence obtained for a target video sequence obtained by filming a real person (the person himself / herself). [Figure 10] 1 is a schematic block diagram showing the system configuration of a determination system 100 according to the present invention. [Figure 11] 2 is a schematic block diagram showing a specific example of the functional configuration of the terminal device 10. FIG. [Figure 12] 2 is a schematic block diagram showing a specific example of the functional configuration of a determination device 20 (determination device 20a) in the first embodiment. FIG. [Figure 13] 1 is a sequence chart showing a specific example of the flow of operations of the determination system 100. [Figure 14] FIG. 10 is a schematic block diagram showing a specific example of the functional configuration of a determination device 20 (determination device 20b) in a second embodiment. [Figure 15] FIG. 10 is a schematic block diagram showing a specific example of the functional configuration of a determination device 20 (determination device 20c) in a third embodiment. [Figure 16] FIG. 10 is a diagram showing an outline of the processing of the candidate fake video generation unit 233. [Figure 17]FIG. 2 is a diagram illustrating an example of the hardware configuration of an information processing device 90 applied to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] [principle] First, the principle of the present invention will be explained. FIG. 1 is a diagram showing an outline of a moving image. A moving image can be regarded as a collection of multiple still images. Here, the moving image to be processed (hereinafter referred to as the "target moving image") is decomposed into multiple frame images (still images). In the example of FIG. 1, the target moving image is divided into N frame images (N is an integer equal to or greater than 1). It is desirable that N is equal to or greater than 2. Each frame image has a predetermined position in the target moving image (the playback time for that frame image when the moving image is played from the beginning). For example, each key frame of the target moving image may be used as a multiple frame image.
[0015] FIG. 2 is a diagram showing a specific example of a first process performed on N frame images. In the first process, a degree of match between each of the N frame images and a facial image of a real person (hereinafter referred to as a "real still image") is obtained for each frame image. A "real person" refers to the person whose facial image appears in the target video. The degree of match may be expressed using a value representing similarity or the likelihood that the person is the same person. It is preferable that the degree of match be a value representing that the people appearing in the images are the same person, rather than a value representing that the images are the same. Any technology may be used as an algorithm for obtaining the degree of match. For example, the degree of match may be obtained using image features at positions corresponding to each organ in the facial image (right eye, left eye, nose, etc.). In this embodiment, the closer the degree of match value is to 100%, the more likely it is that the person appearing in the frame image and the person appearing in the real still image are the same person.
[0016] FIG. 3 is a diagram showing a specific example of a matching data group obtained in the first process. The matching data group is a data group in which the matching values obtained for each frame are arranged in frame order. FIG. 3 shows a matching graph as a specific example of a matching data group. The matching graph arranges the matching degrees for each frame image in the order in which the frame images are played, and connects adjacent matching values in frame order with lines. The vertical axis of the matching graph represents the matching degree, and the horizontal axis represents the frame number (corresponding to the playback order). A matching graph is obtained that shows the matching degrees between the target moving image and each frame image of the genuine still image.
[0017] Fig. 4 is a diagram showing a specific example of the second processing performed on N frame images. In the second processing, the degree of match between each of the N frame images and one frame image selected from the N frame images (hereinafter referred to as the "target representative image") is obtained for each frame image. This degree of match is preferably a value obtained using the same algorithm as the degree of match used in the first processing using genuine still images. Fig. 5 is a diagram showing a specific example of a group of match data (match graph) obtained in the second processing.
[0018] FIG. 6 is a diagram showing a modified example of the second process. In the second process shown in FIG. 6, multiple target representative images are selected, and the degree of match with each frame image (first frame image to Nth frame image) of the target video sequence is obtained for each target representative image. However, the degree of match may be obtained for the same frame image as the image selected as the target representative image, or, as shown in FIG. 6, the degree of match may not be obtained. Once the degree of match has been obtained for all target representative images, a statistical value of the degree of match obtained for each target representative image is calculated for each frame image of the target video sequence. This statistical value may be used as the final value of the degree of match (value of the group of match data) in the second process. A group of match data (a match graph) may be obtained based on the degree of match obtained by the second process shown in FIG. 6.
[0019] When the first and second processes were actually performed using multiple target videos to obtain groups of matching data, the following trends were found: In other words, differences in trends were found between the two groups of matching data obtained for target videos obtained by filming real people (themselves) and the two groups of matching data obtained for target videos obtained using synthesis technology such as deepfake.
[0020] FIG. 7 is a diagram schematically illustrating two matching graphs obtained for a target video obtained using synthesis technology such as deepfake. The upper part of FIG. 7 is a diagram schematically illustrating a matching graph obtained by the first process, and the lower part of FIG. 7 is a diagram schematically illustrating a matching graph obtained by the second process. As such, when the target video is a video obtained using synthesis technology such as deepfake, there is a significant difference between the trend of change in the matching data group obtained by the first process (e.g., the shape of the matching graph) and the trend of change in the matching data group obtained by the second process.
[0021] One possible reason is as follows. Synthesis technologies such as deepfakes generate facial images of various orientations and expressions by performing image processing on real still images, but each algorithm may have strengths and weaknesses depending on the type of orientation and expression. As a result, a high degree of match can be obtained in situations where image processing is possible to approximate a real still image (facial orientation and expression), but a low degree of match can be obtained in situations where image processing is difficult to approximate a real still image. This results in frame-to-frame variation in the degree of match, as shown in the upper panel of Figure 7. On the other hand, with the second processing, a degree of match with a facial image obtained using the same image processing algorithm can be obtained for each frame image. Therefore, although some variation occurs, the range of variation in the degree of match is small. This results in less frame-to-frame variation in the degree of match, as shown in the lower panel of Figure 7.
[0022] FIG. 8 is a diagram showing two similarity graphs obtained for target moving images obtained by filming a real person (the person himself / herself). The upper part of FIG. 8 is a diagram showing a similarity graph obtained by the first process, and the lower part of FIG. 8 is a diagram showing a similarity graph obtained by the second process. As such, when the target moving images are real moving images, the tendency of change in the group of similarity data obtained by the first process (e.g., the shape of the similarity graph) is very similar to the tendency of change in the group of similarity data obtained by the second process. In the case of FIG. 8, this is because both similarities are obtained as the degree of similarity between facial images obtained by actually filming the person himself / herself.
[0023] FIG. 9 is a diagram showing two correspondence graphs obtained for target videos obtained by filming a real person (the person himself / herself). In FIG. 8, both the upper and lower correspondence graphs were obtained with nearly linear shapes, whereas in FIG. 9, both the upper and lower correspondence graphs were obtained with uneven shapes. For example, a low degree of correspondence can be obtained for various reasons, such as the subject's face being turned more than 90 degrees to the side, facing directly upward or nearly directly downward, wearing a mask, or making an overly playful expression in the target videos. As a result, unevenness appears in the correspondence graph in FIG. 9. However, even if unevenness appears in the correspondence graph, the change trends (e.g., the shape of the correspondence graph) of the two groups of correspondence data obtained for target videos obtained by filming a real person (the person himself / herself) are very similar.
[0024] Based on this tendency, the technology of this embodiment acquires at least two groups of matching data by performing a first process and a second process on a video to be judged (target video). If the trend of change in the matching data groups (for example, the shape of the matching graph) is similar based on a predetermined standard, it is determined that the video was obtained by filming a real person (the person), and if it is not similar based on the predetermined standard, it is determined that the video was obtained using synthesis technology such as deepfake.
[0025] [First embodiment] FIG. 10 is a schematic block diagram showing the system configuration of a determination system 100 of the present invention. The determination system 100 includes a terminal device 10 and a determination device 20. The terminal device 10 and the determination device 20 are communicably connected via a network 70. The network 70 may be a network using wireless communication or a network using wired communication. The network 70 may be configured using, for example, the Internet or a local area network (LAN). The network 70 may be configured by combining a plurality of networks. In the first embodiment, a determination device 20a is used as the determination device 20. Note that in the second and third embodiments described below, a determination device 20b and a determination device 20c are used as the determination device 20, respectively.
[0026] 11 is a schematic block diagram showing a specific example of the functional configuration of terminal device 10. Terminal device 10 is configured using information equipment such as a smartphone, tablet, personal computer, dedicated device, etc. Terminal device 10 includes a communication unit 11, an input unit 12, an output unit 13, an image input unit 14, a storage unit 15, and a control unit 16.
[0027] The communication unit 11 is a communication device. The communication unit 11 may be configured as, for example, a network interface. The communication unit 11 communicates data with other devices via the network 70 in accordance with the control of the control unit 16. The communication unit 11 may be a device that performs wireless communication or a device that performs wired communication.
[0028] The input unit 12 is configured using existing input devices such as a keyboard, a pointing device (mouse, tablet, etc.), buttons, a touch panel, etc. The input unit 12 is operated by a user when inputting user instructions to the terminal device 10. The input unit 12 may be an interface for connecting the input device to the terminal device 10. In this case, the input unit 12 inputs an input signal generated in the input device in response to a user input to the terminal device 10. The input unit 12 may be configured using a microphone and a voice recognition device. In this case, the input unit 12 acquires an acoustic signal generated by the user's speech, performs voice recognition on the words spoken by the user, and inputs character string information of the recognition result to the terminal device 10. The voice recognition process may be performed by the control unit 16. The input unit 12 may be configured in any way as long as it is capable of inputting user instructions to the terminal device 10.
[0029] The output unit 13 outputs information in a form that can be recognized by the user. The output unit 13 may be, for example, an image display device such as a liquid crystal display or an organic EL (Electro Luminescence) display. The output unit 13 may be an interface for connecting an image display device to the terminal device 10. In this case, the output unit 13 generates a video signal for displaying image data and outputs the video signal to the image display device connected to the output unit 13. The output unit 13 may be a device for outputting sound, such as a speaker. The output unit 13 may be an interface for connecting an audio output device, such as a speaker or headphones, to the terminal device 10. In this case, the output unit 13 generates an audio signal for reproducing audio data and outputs the audio signal to the audio output device connected to the output unit 13. The output unit 13 may be configured as a touch panel integrated with the input unit 12.
[0030] The image input unit 14 accepts image data input to the terminal device 10. The image input unit 14 may read image data recorded on a recording medium such as a CD-ROM or a USB memory (Universal Serial Bus Memory). The image input unit 14 may receive images captured by a still camera or a video camera from the camera. If the terminal device 10 is equipped with a still camera or a video camera, the image input unit 14 may input images captured by the camera via a bus. The image input unit 14 may receive image data from another information processing device via a network. The image input unit 14 may be configured in a different manner as long as it is capable of receiving input of image data. An image whose input is accepted by the image input unit 14 is referred to as an "input image."
[0031] The storage unit 15 is configured using a storage device such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 15 stores data used by the control unit 16. The storage unit 15 stores data required when the control unit 16 performs processing. The storage unit 15 may store, for example, image data input by the image input unit 14. A specific example of such an image is a face image (real still image) obtained by capturing an image of a person appearing in the target moving image. For example, if the terminal device 10 is the person appearing in the target moving image, such a real still image may be obtained by the person taking a selfie (taking a self-photograph) using the terminal device 10 or an imaging device.
[0032] The control unit 16 is configured using a processor such as a CPU (Central Processing Unit) and a memory (main storage device). The control unit 16 functions when the processor executes a program. Note that all or part of the functions of the control unit 16 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and semiconductor storage devices (e.g., SSDs: Solid State Drives), as well as storage devices such as hard disks and semiconductor storage devices built into computer systems. The program may be transmitted via a telecommunications line.
[0033] The control unit 16 may execute, for example, an application installed on its own device (terminal device 10). A specific example of such an application is an application provided to the terminal device 10 as a dedicated application for the determination system 100. Another specific example of such an application is a web browser application. Such an application may be pre-installed on the terminal device 10, or may be downloaded each time a determination process is executed. For example, when implemented as a web browser application, the terminal device 10 may download and execute the application from a device specified by a specific web server (for example, the web server itself or another server) in response to the terminal device 10 connecting to the web server. The control unit 16 operates according to the program of the application being executed.
[0034] The control unit 16 controls the terminal device 10 in response to user operations and information received from the determination device 20a. For example, the control unit 16 transmits information input by the user operating the input unit 12 to the determination device 20a using the communication unit 11. For example, the control unit 16 transmits facial image data input by the user operating the image input unit 14 to the determination device 20a using the communication unit 11. For example, when information transmitted from the determination device 20a is received by the communication unit 11 via the network 70, the control unit 16 generates screen data based on the received information and displays the screen data on the output unit 13. Such screen data includes images and text indicating the information transmitted from the determination device 20a. For example, when information transmitted from the determination device 20a is received by the communication unit 11 via the network 70, the control unit 16 generates audio data based on the received information and displays the audio data from the output unit 13.
[0035] Specific examples of information transmitted from the terminal device 10 to the determination device 20a include information indicating data of a genuine still image and information indicating data of a target moving image. As information indicating data of a genuine still image, the data of the genuine still image itself may be transmitted to the determination device 20a, or information indicating an address (URL, etc.) of a genuine still image that has already been uploaded to the network in an accessible manner may be transmitted to the determination device 20a. As information indicating data of the target moving image, the data of the target moving image itself may be transmitted to the determination device 20a, or information indicating an address (URL, etc.) of the target moving image that has already been uploaded to the network in an accessible manner may be transmitted to the determination device 20a.
[0036] 12 is a schematic block diagram showing a specific example of the functional configuration of the determination device 20 (determination device 20a) in the first embodiment. The determination device 20a is configured using an information processing device such as a personal computer or a server device. The determination device 20a includes a communication unit 21, a storage unit 22a, and a control unit 23a.
[0037] The communication unit 21 is a communication device. The communication unit 21 may be configured as, for example, a network interface. The communication unit 21 communicates data with other devices via the network 70 in accordance with the control of the control unit 23a. The communication unit 21 may be a device that performs wireless communication or a device that performs wired communication.
[0038] The storage unit 22a is configured using a storage device such as a magnetic hard disk drive, a semiconductor storage device, etc. The storage unit 22a stores data used by the control unit 23a.
[0039] The control unit 23a is configured using a processor such as a CPU and a memory. The control unit 23a functions as the determination unit 231 when the processor executes a program. All or part of the functions of the control unit 23a may be realized using hardware such as an ASIC, PLD, or FPGA. The program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and semiconductor storage devices (e.g., SSDs), as well as storage devices such as hard disks and semiconductor storage devices built into computer systems. The program may be transmitted via a telecommunications line.
[0040] The determination unit 231 determines whether the target moving image is a moving image obtained by shooting a real person (the person himself / herself) (hereinafter referred to as a "real moving image") or a moving image obtained using a synthesis technology such as deep fake (hereinafter referred to as a "fake moving image"), using a real still image instructed by the terminal device 10. A specific example of the processing performed by the determination unit 231 will be described below.
[0041] The determination unit 231 divides the target video into multiple (e.g., N) frame images. The determination unit 231 executes a first process to acquire a group of matching data between the genuine image and each frame image. The determination unit 231 executes a second process to acquire a group of matching data between the target representative image and each frame image. At this time, the determination unit 231 may execute the second process using one target representative image (see, for example, FIG. 4 ), or may execute the second process using multiple target representative images (see, for example, FIG. 6 ). If the change trends of the two obtained groups of matching data are similar based on a predetermined criterion, the determination unit 231 determines the target video to be a genuine video. On the other hand, if the change trends of the two obtained groups of matching data are not similar based on the predetermined criterion, the determination unit 231 determines the target video to be a fake video. The determination unit 231 transmits information indicating the determination result to the terminal device 10.
[0042] 13 is a sequence chart showing a specific example of the flow of operations of the determination system 100. First, the user inputs a determination instruction by operating the input unit 12 of the terminal device 10 (step S101). The terminal device 10 accepts the input of the determination instruction. The determination instruction includes at least information indicating data of the genuine still image and information indicating data of the target moving image. Upon accepting the input of the determination instruction, the terminal device 10 generates determination instruction information including the input information and transmits it to the determination device 20a (step S102).
[0043] Upon receiving the determination instruction information from the terminal device 10, the determination device 20a acquires the target video indicated by the received determination instruction information and decomposes the target video into a plurality of frame videos (step S103). The determination device 20a executes a first process using the genuine still image indicated by the received determination instruction information and the frame videos obtained by decomposing the target video, and acquires a group of matching data (step S104). The determination device 20a executes a second process using the frame videos obtained by decomposing the target video and the target representative image, and acquires a group of matching data (step S105). The determination device 20a uses the acquired group of matching data to determine whether the target video is a genuine video or a fake video (step S106).
[0044] The determination device 20a transmits information indicating the determination result (determination result information) to the terminal device 10 (step S107). Upon receiving the determination result information from the determination device 20a, the terminal device 10 outputs the determination result (step S108).
[0045] In the judgment system 100 configured in this manner, it is possible to accurately judge whether the video being judged (target video) is a real video or a fake video based on the difference in the change trends between the two groups of matching data obtained from real video and the two groups of matching data obtained from fake video.
[0046] [Second embodiment] 14 is a schematic block diagram showing a specific example of the functional configuration of the determination device 20 (determination device 20b) in the second embodiment. The determination device 20b is configured using an information processing device such as a personal computer or a server device. The determination device 20b includes a communication unit 21, a storage unit 22b, and a control unit 23b. The communication unit 21 has the same configuration as the communication unit 21 of the determination device 20a in the first embodiment, and therefore a description thereof will be omitted.
[0047] The storage unit 22b is configured using a storage device such as a magnetic hard disk drive, a semiconductor storage device, etc. The storage unit 22b stores data used by the control unit 23b. The storage unit 22b functions as a tool information storage unit 221.
[0048] The tool information storage unit 221 stores information about tools that generate fake videos. For each tool that generates fake videos, the tool information storage unit 221 stores information indicating the characteristics of fake videos generated using that tool. A specific example of such information may be information indicating the tendency of a group of matching data obtained when a first process is performed using fake videos generated using that tool. More specifically, information indicating the attributes of frame images that tend to have low matching values may be used. For example, information indicating the type of facial expression, facial direction, size, and shading of the facial image included in the frame image may be used. For example, a smiling face, an angry face, a narrowed-eyed face, a wide-mouthed face, a face looking up, a face looking right, a face looking diagonally downward and to the right, a face with a vertical width of 200 pixels or more, a face with a horizontal width of less than 300 pixels, a face with an overall average brightness lower than a threshold, a face with a brightness difference between the right and left halves equal to or greater than a threshold, etc. may be used as the attribute values of the frame image.
[0049] The control unit 23b is configured using a processor such as a CPU and a memory. The control unit 23b functions as the determination unit 231 and the tool determination unit 232b when the processor executes a program. All or part of the functions of the control unit 23b may be realized using hardware such as an ASIC, PLD, or FPGA. The above program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and semiconductor storage devices (e.g., SSDs), as well as storage devices such as hard disks and semiconductor storage devices built into computer systems. The above program may be transmitted via a telecommunications line.
[0050] The processing of the determination unit 231 is the same as in the first embodiment, and therefore description thereof will be omitted. The tool determination unit 232b determines the tool (hereinafter referred to as "used tool") used to generate the fake video when it is assumed that the target video is a fake video. The tool determination unit 232b determines, for example, from among the features of each tool stored in the tool information storage unit 221, the tool that is closest to the features of the target video as the used tool.
[0051] For example, when the information indicating the characteristics of a fake video is information indicating a trend in the group of coincidence data, the tool determination unit 232b operates as follows: The tool determination unit 232b acquires the group of coincidence data obtained by the first process of the determination unit 231. From the information indicating the trend in the group of coincidence data of each tool stored in the tool information storage unit 221, the tool determination unit 232b selects, as the tool to be used, the tool whose trend in the group of coincidence data obtained by the first process is closest.
[0052] For example, if the information indicating the characteristics of a fake video is information indicating attributes of frame images that tend to have a low degree of match, the tool determination unit 232b operates as follows: The tool determination unit 232b acquires a group of match data obtained by the first processing of the determination unit 231. From the acquired group of match data, the tool determination unit 232b selects a match lower than a predetermined standard (e.g., a threshold value). The tool determination unit 232b acquires the frame images for which the selected low degree of match was obtained. The tool determination unit 232b determines the attribute by performing image processing on the acquired frame images. The determined attribute is an attribute stored in the tool information storage unit 221 as information indicating the characteristics of the fake video. For example, the type of facial expression of a facial image or the direction of the face is determined as the attribute. The tool determination unit 232b selects a tool having characteristics similar to the acquired attribute as the tool to be used. The specific example of the processing of the tool determination unit 232b described above is merely an example. The tool determination unit 232b transmits information indicating the determination result to the terminal device 10.
[0053] In the second embodiment configured as described above, when it is assumed that the target video is a fake video, it is possible to determine the tool (tool used) that was used to generate the fake video.
[0054] [Third embodiment] 15 is a schematic block diagram showing a specific example of the functional configuration of the determination device 20 (determination device 20c) in the third embodiment. The determination device 20c is configured using an information processing device such as a personal computer or a server device. The determination device 20c includes a communication unit 21, a storage unit 22, and a control unit 23c. The communication unit 21 and the storage unit 22 have the same configurations as the communication unit 21 and the storage unit 22 of the determination device 20a in the first embodiment, respectively, and therefore will not be described here.
[0055] The control unit 23c is configured using a processor such as a CPU and a memory. The control unit 23c functions as the determination unit 231, the tool determination unit 232c, and the candidate fake video generation unit 233 by the processor executing a program. Note that all or part of the functions of the control unit 23c may be realized using hardware such as an ASIC, PLD, or FPGA. The above program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and semiconductor storage devices (e.g., SSDs), as well as storage devices such as hard disks and semiconductor storage devices built into computer systems. The above program may be transmitted via a telecommunications line.
[0056] The processing of the determination unit 231 is the same as in the first embodiment, and therefore a description thereof will be omitted. FIG. 16 is a diagram illustrating an outline of the processing of the candidate fake video generation unit 233. The candidate fake video generation unit 233 uses genuine still images and target videos to generate fake videos (hereinafter referred to as "candidate fake videos") using each of a plurality of tools (hereinafter referred to as "candidate tools") that are candidates for the tool to be used. A candidate tool is a tool that generates a synthetic video by replacing a facial image captured in an input video (in this embodiment, the target video) with a facial image captured in an input still image (in this embodiment, the genuine still image). When replacing a facial image captured in a video with a facial image captured in a still image, the candidate tool performs image processing so that the attributes of the facial image in the still image (e.g., facial expression, facial direction, size, shading, etc.) are similar to or match the attributes of the facial image captured in the video. If there are M types of candidate tools, the candidate fake video generation unit 233 may generate M candidate fake videos.
[0057] The tool determination unit 232c executes a first process for each candidate fake video to obtain a group of matching data for each candidate fake video. The tool determination unit 232c compares the group of matching data obtained by executing the first process on the target video with the group of matching data obtained by executing the first process on each candidate fake video, and selects the candidate fake video for which a group of matching data with a change trend most similar to the group of matching data for the target video has been obtained. The tool determination unit 232c then determines the candidate tool used to generate the selected candidate fake video as the used tool.
[0058] In the third embodiment configured as above, when it is assumed that the target video is a fake video, it is possible to determine the tool (tool used) used to generate the fake video.
[0059] FIG. 17 is a diagram illustrating an outline of an example of the hardware configuration of an information processing device 90 applied to this embodiment. The information processing device 90 includes a processor 91, a main storage device 92, a communication interface 93, an auxiliary storage device 94, an input / output interface 95, and an internal bus 96. The processor 91, the main storage device 92, the communication interface 93, the auxiliary storage device 94, and the input / output interface 95 are communicably connected to each other via the internal bus 96. The information processing device 90 may be applied to, for example, the terminal device 10 and the determination device 20 (20a to 20c). In this case, for example, the communication unit 11 and the communication unit 21 may be configured using the communication interface 93. For example, the memory unit 15, the memory unit 22a, and the memory unit 22b may be configured using the auxiliary storage device 94. Furthermore, the control unit 16 and the control units 23a to 23c may be configured using the processor 91 and the main storage device 92.
[0060] (Variation) In this embodiment, the terminal device 10 and the determination device 20 are configured as separate devices, but they may also be configured as an integrated device. In this case, the determination device 20 includes an input unit, an output unit, and an image input unit. These input unit, output unit, and image input unit function in the same manner as the input unit 12, output unit 13, and image input unit 14 of the terminal device 10, respectively. The control units 23a to 23c operate in response to user operations on the input unit, perform determination processing using the input information, and output the determination result using the output unit.
[0061] The determination device 20 may be implemented using a plurality of information processing devices. For example, the determination device 20 may be implemented using a device such as a cloud. For example, in the determination device 20, the storage unit 22 and the control units (23a to 23c) may be implemented in different information processing devices. For example, the storage unit 22 of the determination device 20 may be distributed and implemented in a plurality of information processing devices. For example, the control units (23a to 23c) of the determination device 20 may be distributed and implemented in a plurality of information processing devices.
[0062] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Explanation of symbols]
[0063] 100...Determination system, 10...Terminal device, 11...Communication unit, 12...Input unit, 13...Output unit, 14...Image input unit, 15...Memory unit, 16...Control unit, 20...Determination device, 21...Communication unit, 22...Memory unit, 221...Tool information memory unit, 23a to 23c...Control unit, 231...Determination unit, 232b to 232c...Tool determination unit, 233...Candidate fake video generation unit
Claims
1. A determination device comprising a control unit that acquires multiple frame images, which are still images, from a target video that is a video to be determined, performs a first process to obtain the degree of similarity between the frame images and real facial images of facial images appearing in the video, performs a second process to obtain the degree of similarity between the frame images and other frame images obtained from the target video, and determines that the target video is a video in which the real facial images were captured if the degree of similarity obtained in the first process and the degree of similarity obtained in the second process are similar based on a predetermined criterion, and determines that the target video is a video obtained using synthesis technology if they are not similar based on the predetermined criterion.
2. The control unit of the determination device of claim 1 acquires a group of similarity data in which multiple similarities are arranged in the order in which the frame images to be processed are played back in the first process and the second process, and determines whether the group of similarity data is similar according to the specified criteria.
3. The determination device according to claim 2 , wherein the control unit makes a determination based on whether a trend of change in the group of coincidence data is similar based on the predetermined criterion.
4. The determination device according to claim 1 , wherein the control unit determines a tool used to generate the target video when it is assumed that the target video is a video obtained using the synthesis technique.
5. a frame image acquisition step in which a computer acquires a plurality of frame images, which are still images, from a target moving image, which is a moving image to be determined; a first processing step in which a computer executes a first process to acquire a degree of coincidence between the frame images and a real face image of a face image captured in the moving image; a second processing step in which the computer executes a second process to obtain a degree of coincidence between the frame image and another frame image obtained from the target video sequence; a determining step in which the computer determines that the target video is a video in which the real face image is captured if the degree of coincidence obtained in the first process and the degree of coincidence obtained in the second process are similar based on a predetermined criterion, and determines that the target video is a video obtained using a synthesis technique if the degree of coincidence obtained in the first process and the degree of coincidence obtained in the second process are not similar based on the predetermined criterion; A determination method having the following.
6. A computer program for causing a computer to function as a determination device comprising a control unit that acquires multiple frame images, which are still images, from a target video that is a video to be determined, executes a first process to obtain the degree of similarity between the frame images and the real face images of the facial images appearing in the video, executes a second process to obtain the degree of similarity between the frame images and other frame images obtained from the target video, and determines that the target video is a video in which the real face images were captured if the degree of similarity obtained in the first process and the degree of similarity obtained in the second process are similar based on a predetermined criterion, and determines that the target video is a video obtained using synthesis technology if they are not similar based on the predetermined criterion.
Citation Information
Patent Citations
Synthetic video data generation system and program
JP2016151975A