Learning device, information processing device, learning method, and computer program
Patent Information
- Application Number
- JP2025527350
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2023-06-22
- Filing Date
- 2023-06-22
- Publication Date
- 2026-03-05
AI Technical Summary
Current technologies face challenges in accurately detecting whether images are composite or synthesized, particularly in distinguishing between real and fake images, which is crucial for authentication and verification processes.
A learning device and method that utilize a feature extraction model trained through distance learning, using image sets labeled as non-composite and composite, to extract and compare feature quantities, determining the similarity between images to identify synthesized or fake images by optimizing the feature extraction operation to differentiate between images with the same and different labels.
The solution effectively enhances the ability to detect fake images and improve image authentication by generating a feature extraction model that accurately identifies composite or synthesized images, outperforming comparative examples in performance.
Abstract
Description
Learning device, information processing device, learning method, and recording medium
[0001] The present disclosure relates to the technical fields of a learning device, an information processing device, a learning method, and a recording medium.
[0002] Non-Patent Document 1 describes a technology in which a feature extraction model for face recognition is constructed by metric learning and the feature detection model is used to detect deep fakes.
[0003] Sreeraj Ramachandran, Aakash Varma Nadimpalli, Ajita Rattani, An Experimental Evaluation on Deepfake Detection using Deep Face Recognition, IEEE International Carnahan Conference on Security Technology(ICCST)2021
[0004] An object of the present disclosure is to provide a learning device, an information processing device, a learning method, and a recording medium that are intended to accurately detect whether images have been combined.
[0005] One aspect of the learning device includes feature extraction means for extracting features from images, and learning means for causing the feature extraction means to perform a first learning of a feature extraction operation using a first image set including images labeled with a non-composite label indicating that the images are not composited, and images labeled with a composite label indicating that the images are composited, such that a first feature extracted from a first image labeled with the non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with the non-composite label than to a second feature extracted from a second image labeled with the composite label.
[0006] A first aspect of the information processing device includes a first feature extraction means that performs first learning of a feature extraction operation using a first image set including images labeled with a non-composite label indicating that the image is not composited and images labeled with a composite label indicating that the image is composited, such that a first feature extracted from a first image labeled with the non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with the non-composite label than a second feature extracted from a second image labeled with the composite label; a first calculation means that calculates a similarity between a target feature extracted by the feature extraction means from a target image to be determined as to whether or not the image is composited, and a reference feature extracted by the feature extraction means from a reference image; a determination means that determines whether or not the target image is composited by determining a threshold value of the similarity; and an output means that produces an output according to the determination result by the determination means.
[0007] A second aspect of the information processing device is a second feature extraction means that performs second learning of a feature extraction operation using a second image set that is labeled to identify individuals, such that feature amounts extracted from images labeled with the same label are more similar than feature amounts extracted from images labeled with different labels, and that uses a first image set that includes images labeled with a non-composite label indicating that the image is not composited, and images labeled with a composite label indicating that the image is composited, and extracts a first feature amount extracted from a first image labeled with the non-composite label more similar than feature amounts extracted from an image labeled with the composite label. the first feature extraction means having performed first learning of the feature extraction operation so that a third feature extracted from a third image different from the first image to which the non-synthetic label is applied is more similar to a second feature extracted from a second image to which the non-synthetic label is applied than to a second feature extracted from a second image to which the non-synthetic label is applied; a second calculation means having calculated a similarity between a target feature extracted by the feature extraction means from a target image to be determined as to whether or not the target is the actual person and a reference feature extracted by the feature extraction means from a reference image; a second determination means having determined whether or not the target is the actual person by threshold-based determination of the similarity; and an output means having output according to the determination result by the determination means.
[0008] One aspect of the learning method involves extracting features from images using a feature extraction model, and using a first image set including images labeled with a non-composite label indicating that the images are not composited, and images labeled with a composite label indicating that the images are composited, causing the feature extraction model to perform a first learning of a feature extraction operation such that a first feature extracted from the first image labeled with the non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with the non-composite label than to a second feature extracted from a second image labeled with the composite label.
[0009] In one aspect of the recording medium, a computer program is recorded on the recording medium to cause a computer to execute a learning method for extracting features from images using a feature extraction model, and using a first image set including images labeled with a non-composite label indicating that the images are not composited and images labeled with a composite label indicating that the images are composited, the feature extraction model is caused to perform a first learning of a feature extraction operation such that a first feature extracted from a first image labeled with the non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with the non-composite label than to a second feature extracted from a second image labeled with the composite label.
[0010] The learning device, information processing device, learning method, and recording medium according to the present disclosure can accurately detect whether images have been combined.
[0011] FIG. 1 is a block diagram showing an example of a configuration of a learning device according to the present disclosure. FIG. 2 is a block diagram showing an example of a configuration of a learning device according to the present disclosure. FIG. 3 is a block diagram showing an example of a configuration of a learning device according to the present disclosure. FIG. 4 is a block diagram showing an example of a configuration of a learning device according to the present disclosure. FIG. 5 is a block diagram showing an example of a configuration of a learning device according to the present disclosure. FIG. 6 is a block diagram showing an example of a configuration of an information processing device according to the present disclosure. FIG. 7 is a block diagram showing an example of a configuration of an information processing device according to the present disclosure. FIG. 8 is a block diagram showing an example of a configuration of an information processing device according to the present disclosure. FIG. 9 is a flowchart showing an example of a processing operation of an information processing device according to the present disclosure.
[0012] Hereinafter, embodiments of a learning device, an information processing device, a learning method, and a recording medium will be described with reference to the drawings. [1: First Embodiment]
[0013] A first embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described. Hereinafter, the first embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described using a learning device 1 according to the present disclosure. [1-1: Configuration of learning device 1]
[0014] 1 is a block diagram showing the configuration of a learning device 1 according to the present disclosure. As shown in FIG. 1, the learning device 1 includes a feature extraction unit 11 and a learning unit 12.
[0015] The feature extraction unit 11 extracts features from images. The learning unit 12 causes the feature extraction unit 11 to perform a first learning of the feature extraction operation. The first learning is distance learning. The learning unit 12 causes the feature extraction unit 11 to perform the first learning of the feature extraction operation using a first image set including images labeled with a non-composite label indicating that the images are not combined, and images labeled with a composite label indicating that the images are combined. The learning unit 12 causes the feature extraction operation to perform the first learning so that a first feature extracted from a first image labeled with a non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with a non-composite label than to a second feature extracted from a second image labeled with a composite label. [1-2: Technical Effects of the Learning Device 1]
[0016] The learning device 1 according to the present disclosure can generate a feature extraction unit that extracts features that can detect whether an image is synthesized or not by performing metric learning using an image set that includes images labeled as non-synthesis and images labeled as synthesis. [2: Second Embodiment]
[0017] A second embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described. Hereinafter, the second embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described using a learning device 2 according to the present disclosure. [2-1: Fake Image]
[0018] There is a technology that synthesizes an image of a target based on information from a single photograph of the target. For example, there is a technology that synthesizes an image of a person based on information from a single photograph of the person's face. Deepfake, for example, is known as a technology for synthesizing images of people. Deepfake is known as a technology that synthesizes fake images that depict things that did not actually happen. Hereinafter, an image that depicts things that did not actually happen may be referred to as a fake image. A synthesized image may also be referred to as a fake image. An image that depicts things that actually happened may also be referred to as a real image.
[0019] If it is possible to extract features that indicate the likelihood of a fake image and features that indicate the likelihood of a real image, it will be possible to determine whether a target image is a fake image or a real image based on the features extracted from the target image. Hereinafter, a feature extraction model f that can extract features appropriate for determining whether an image is fake or not will be referred to as a feature extraction model f. θ [2-2: Configuration of learning device 2]
[0020] 2 is a block diagram showing the configuration of the learning device 2. As shown in FIG. 2, the learning device 2 includes a calculation device 21 and a storage device 22. The learning device 2 may further include a communication device 23, an input device 24, and an output device 25. However, the learning device 2 does not have to include at least one of the communication device 23, the input device 24, and the output device 25. The calculation device 21, the storage device 22, the communication device 23, the input device 24, and the output device 25 may be connected via a data bus 26.
[0021] The arithmetic device 21 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array). The arithmetic device 21 loads a computer program. For example, the arithmetic device 21 may load a computer program stored in the storage device 22. For example, the arithmetic device 21 may load a computer program stored in a computer-readable, non-transitory storage medium using a storage medium reading device (e.g., the input device 24 described later) not shown in the drawing provided in the learning device 2. The arithmetic device 21 may acquire (i.e., download or load) the computer program from a device (not shown) located outside the learning device 2 via the communication device 23 (or another communication device). The arithmetic device 21 executes the loaded computer program. As a result, logical functional blocks for executing the operations to be performed by the learning device 2 are realized within the arithmetic device 21. In other words, the arithmetic device 21 can function as a controller for realizing logical functional blocks for executing the operations (in other words, processing) to be performed by the learning device 2.
[0022] The storage device 22 can store desired data. For example, the storage device 22 may temporarily store a computer program executed by the arithmetic device 21. The storage device 22 may temporarily store data that the arithmetic device 21 temporarily uses when the arithmetic device 21 is executing a computer program. The storage device 22 may store data that the learning device 2 stores long-term. The storage device 22 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device. In other words, the storage device 22 may include a non-temporary recording medium.
[0023] The communication device 23 can communicate with devices external to the learning device 2 via a communication network (not shown). The communication device 23 may be a communication interface based on standards such as Ethernet (registered trademark), Wi-Fi (registered trademark), Bluetooth (registered trademark), or USB (Universal Serial Bus).
[0024] Input device 24 is a device that accepts information input to learning device 2 from outside learning device 2. For example, input device 24 may include an operating device (e.g., at least one of a keyboard, a mouse, and a touch panel) that can be operated by an operator of learning device 2. For example, input device 24 may include a reading device that can read information recorded as data on a recording medium that can be attached externally to learning device 2.
[0025] The output device 25 is a device that outputs information to the outside of the learning device 2. For example, the output device 25 may output information as an image. That is, the output device 25 may include a display device (a so-called display) that can display an image showing the information to be output. For example, the output device 25 may output information as sound. That is, the output device 25 may include an audio device (a so-called speaker) that can output sound. For example, the output device 25 may output information on paper. That is, the output device 25 may include a printing device (a so-called printer) that can print desired information on paper.
[0026] 2 shows an example of logical functional blocks realized in the arithmetic device 21 to execute information processing operations. As shown in FIG. 2, the arithmetic device 21 implements a feature extraction unit 211, which is a specific example of the "feature extraction means" described in the appendix below, and a learning unit 212, which is a specific example of the "learning means" described in the appendix below. [2-3: Learning operations performed by the learning device 2]
[0027] The feature extraction unit 211 executes a feature extraction operation to extract features from an image. The feature extraction unit 211 may extract, from the image, a feature indicating the likelihood of the image being a fake image and a feature indicating the likelihood of the image being a real image. The feature extraction unit 211 uses a feature extraction model f θ The feature extraction operation is performed using the feature extraction model f θ When an image is input, the function outputs the feature amount of the image.
[0028] The learning unit 212 uses the feature extraction model f θ The first learning of the feature extraction operation is learning to operate so that a first feature extracted from a first image labeled with a non-composite label is more similar to a third feature extracted from a third image labeled with a non-composite label than to a second feature extracted from a second image labeled with a composite label. In other words, the first learning of the feature extraction operation is distance learning.
[0029] Feature extraction model f θFor example, the feature extraction model f may be a model including a convolutional neural network (CNN) in its architecture. θ may be generated by deep learning.
[0030] The learning unit 212 uses the feature extraction model f θ , the first image set D df The first learning of the feature extraction operation is performed using the first image set D df includes images labeled with a non-composite label indicating that the image is not composited, and images labeled with a composite label indicating that the image is composited. The learning unit 212 performs first learning of the feature extraction operation so that a first feature extracted from a first image labeled with a non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with a non-composite label than to a second feature extracted from a second image labeled with a composite label.
[0031] Feature extraction model f θ is a model that has not undergone any other learning before the first learning. θ The initial value of the parameter θ of the feature extraction model f may be any value. θ As the initial value of the parameter θ, a value set randomly may be used so that the weight distribution is not biased during learning.
[0032] The learning unit 212 optimizes the learnable parameter θ. For example, if the parameter θ is fixed, the parameter θ cannot be learned. In the second embodiment, the feature extraction model f θ All the parameters θ may be learnable.
[0033] The learning unit 212 calculates a first image set D expressed as the following formula 1. df [Formula 1] may also be used. x real 1 , and x real 2 is the real image, and x fakeare fake images. That is, the first image set D df The learning unit 212 generates a first image set D df may also be used.
[0034] In this case, the first image is x in the above formula 1. real 1 The second image may be expressed as x in the above formula 1. fake The third image may be expressed as x in the above formula 1. real 2 may be.
[0035] The learning unit 212 uses the loss function to generate a feature extraction model f θ The learning unit 212 adjusts the parameter θ of the loss function such that the loss is small when the similarity between the first feature amount and the second feature amount is smaller than the similarity between the first feature amount and the third feature amount.
[0036] Feature extraction model f θ is a feature vector R d In this case, the learning unit 212 may output the first feature vector R d and the second feature vector R d From the distance between d and the third feature vector R d A loss function is used that reduces the loss when the distance between the
[0037] The learning unit 212 may employ, for example, triplet loss as a loss function for distance learning. When triplet loss is employed as the loss function, the loss function can be expressed as in the following Equation 2. [Equation 2]
[0038] The above equation 2 is expressed as x as the first image labeled as non-composite. real 1 and a second image x fake Based on the distance between the second feature vector extracted from the first image and the second feature vector extracted from the first image, x is labeled as a non-synthetic image.real 1 The first feature vector extracted from x and the third image with non-composite labels x real 2 The loss function is expressed such that the loss is small when the distance between the feature vector and the third feature vector extracted from the feature extraction model f is small. θ The feature extraction algorithm learns feature extraction operations so that features extracted from images with the same label are close to each other and features extracted from images with different labels are far from each other. Note that α is a hyperparameter that represents the margin.
[0039] The learning unit 212 learns the feature extraction model f so that the loss of the loss function is minimized. θ The parameter θ of the feature extraction model f θ That is, the learning unit 212 optimizes the feature extraction model f so that the following formula 3 is realized. θ The learning unit 212 adjusts the parameter θ so that the sum of the losses calculated from each of the N sets is minimized, and generates the feature extraction model f θ Optimize [Equation 3].
[0040] Although the case where triplet loss is used as the loss function has been described as an example, the loss function is not limited to triplet loss. For example, any loss function suitable for distance learning, such as ArcFace, can be used.
[0041] The learning operation performed by the learning device 2 will be described with reference to Fig. 3. Fig. 3 is a flowchart showing an example of the flow of the learning operation performed by the learning device 2.
[0042] As shown in FIG. 3, the learning unit 212 df The learning unit 212 acquires the first image set D via the communication device 23 or the input device 24 (step S20). df The learning unit 212 may acquire the first image set D stored in the storage device 22. df may be obtained.
[0043] The feature extraction unit 211 extracts a feature extraction model fθ The learning unit 212 extracts features from each of the three images included in the set using the loss function f θ The parameter θ is adjusted (step S22).
[0044] The learning unit 212 determines whether to end learning (step S23). Learning may be ended, for example, when the above formula 3 is satisfied. Alternatively, learning may be ended, for example, when the learning operation has been executed for all of the N sets. In this case, the N sets may be an amount of learning data sufficient for sufficient learning. If learning is not to be ended (step S23: No), the process returns to step S21. If learning is to be ended (step S23: Yes), the desired feature extraction model f θ [2-4: Technical Effects of Learning Device 2]
[0045] The technology disclosed in Non-Patent Document 1, in which a feature extraction model for face recognition is constructed by metric learning and the feature detection model is used to detect fake images, is referred to as the "Comparative Example." The feature extraction model for face recognition is not trained for the purpose of distinguishing between fake and real images.
[0046] The learning device 2 according to the present disclosure aims to distinguish between fake images and real images by using an image set including images labeled with a non-composite label indicating that the image is not composited and images labeled with a composite label indicating that the image is composited, and constructs a feature extraction model f by metric learning. θ Therefore, the learning device 2 generates a feature extraction model f that has a higher performance in distinguishing between fake images and real images than the comparative example. θ [3: Third embodiment]
[0047] A third embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described below. Hereinafter, the third embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described using a learning device 3 according to the present disclosure. [3-1: Learning operation performed by the learning device 3]
[0048] In the third embodiment, the learning unit 312 performs second learning of the feature extraction operation on the face matching feature extraction model f θFace That is, in the third embodiment, the feature extraction model f θ However, the second embodiment differs from the second embodiment in that the second learning is performed before the first learning is performed.
[0049] The second learning of the feature extraction operation is performed using a second set of images that are labeled to identify individuals. The second learning of the feature extraction operation is performed so that features extracted from images labeled with the same label are more similar to each other than features extracted from images labeled with different labels. The face matching feature extraction model f that has undergone the second learning θFace When images of the same person are input, it outputs similar features, and when images of different people are input, it outputs dissimilar features.
[0050] The third embodiment is a feature extraction model f that performs first learning of the feature extraction operation. θ In other words, the initial value of the feature extraction model f in the third embodiment is different from that in the second embodiment. θ The initial value of the parameter θ of the feature extraction model f may be a value optimally adjusted by the second learning of the feature extraction operation. θ The initial value of the parameter θ of the feature extraction model f may be a value suitable for identifying individuals. θ The initial value of the parameter θ may be a value suitable for face matching.
[0051] The first learning of the feature extraction operation is performed using the first image set D df The first learning of the feature extraction operation is, as in the second embodiment, learning for operating such that the first feature extracted from the first image labeled with a non-composite label is more similar to the third feature extracted from the third image labeled with a non-composite label than to the second feature extracted from the second image labeled with a composite label.
[0052] The learning unit 312 causes the feature extraction model, which has learned the feature extraction operation to extract features suitable for identifying individuals from an image, to perform additional learning of the feature extraction operation to extract features suitable for determining whether an image is a fake or not. [3-2: Modification]
[0053] The learning unit 312 performs the second learning on the face matching feature extraction model f θFace Layer g θ’ Feature extraction model f θ The additional layer g may be subjected to additional learning. θ’ is the face matching feature extraction model f θFace As an input, the feature extraction model f θ may be expressed as the following formula 4. [Formula 4]
[0054] The learning unit 312 generates a face matching feature extraction model f θFace parameter θ Face , and additional layer g θ’ Alternatively, the learning unit 312 may adjust each of the parameters θ′ in the face matching feature extraction model f θFace parameter θ Face is fixed and not adjusted, and the additional layer g θ’ Additional learning may be performed by adjusting θ′. [3-3: Technical Effects of Learning Device 3]
[0055] The learning device 3 according to the present disclosure generates a feature extraction model f that is suitable for extracting features used in face matching and for detecting fake images. θ [4: Fourth embodiment]
[0056] A fourth embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described. Hereinafter, a fourth embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described using a learning device 4 according to the present disclosure.
[0057] In the fourth embodiment, a constraint is imposed to maintain the performance obtained by the second learning. In the fourth embodiment, additional learning is performed on the feature extraction model generated by the second learning while imposing a constraint to maintain the performance obtained by the second learning. [4-1: Learning Operation Performed by the Learning Device 4]
[0058] When the learning unit 412 performs additional learning on the feature extraction model, the learning unit 412 limits the additional learning. Specifically, the learning unit 412 limits the additional learning on the feature extraction model for face matching f θFace parameter θ Face The feature extraction model f is set so that the more important the parameter, the smaller the change in the parameter in the first learning. θ This limits the adjustment of the parameter θ.
[0059] The learning unit 412 adds a constraint term to the loss function to generate a feature extraction model f θ The constraint term may be set to the second trained feature extraction model for face matching f θFace parameter θ Face The learning unit 412 may add a constraint term, expressed by the following formula 5, that assigns importance to each parameter to the loss function. [Formula 5]
[0060] F i is the i-th parameter θ i In other words, the constraint term is F i The larger the parameter θ i is the feature extraction model for face matching f θFace parameter θ Face where λ is a hyperparameter for adjusting the constraint strength.
[0061] The learning unit 412 may perform adjustments so that the loss and the constraint term are simultaneously minimized. That is, the learning unit 412 may operate so that the following equation 6 is satisfied: [Equation 6]
[0062] The weights may be assigned manually or automatically. When the weights are assigned automatically, they may be assigned using, for example, Fisher information (see Reference 1 below). [Reference 1] Overcoming catastrophic forgetting in neural networks, arXiv2016 [4-2: Technical effects of learning device 4]
[0063] The learning device 4 according to the present disclosure constrains important parameters so that they do not change significantly in the task performed for face recognition, and therefore produces a feature extraction model f suitable for extracting features used in face matching and for detecting fake images. θ [5: Fifth embodiment]
[0064] A fifth embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described. Hereinafter, the fifth embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described using a learning device 5 according to the present disclosure. [5-1: Learning operation performed by the learning device 5]
[0065] The learning unit 512 learns the first image set D df and a third image set D with labels identifying individuals. fr Using the above, the feature extraction model f θ The learning unit 512 performs a first learning of the feature extraction operation. df and a third image set D expressed as in the following Equation 7: fr Using the above, the feature extraction model f θ may perform the first learning of the feature extraction operation. [Equation 7]
[0066] x i 1 , and x i 2 is the face image of individual i, and x j is the face image of individual j. That is, the third image set D frThe learning unit 512 generates a first image set D df and a third image set D including N sets of two images of the person and one image of another person. fr and may also be used.
[0067] The learning unit 512 performs first learning of the feature extraction operation so that the first feature extracted from the real image 1 is more similar to the third feature extracted from the real image 2 than to the second feature extracted from the fake image, and so that the fourth feature extracted from the person's image 1 is more similar to the sixth feature extracted from the person's image 2 than to the fifth feature extracted from the other person's image. The learning unit 512 learns the feature extraction model f using a first loss function that reduces the loss when the similarity between the first feature and the second feature is smaller than the similarity between the first feature and the third feature, and a second loss function that reduces the loss when the similarity between the fourth feature and the fifth feature is smaller than the similarity between the fourth feature and the sixth feature. θ That is, the learning unit 512 adjusts the parameter θ of the feature extraction model f by using a loss function that reduces the loss when the similarity between feature amounts extracted from images with different labels is smaller than the similarity between feature amounts extracted from images with the same label. θ The parameter θ is adjusted.
[0068] When triplet loss is used as the loss function, the first loss function can be expressed as in the above formula 2. Furthermore, when triplet loss is used as the loss function, the second loss function can be expressed as in the following formula 8. [Formula 8]
[0069] The learning unit 512 learns the feature extraction model f so that the loss of the loss function is minimized. θ The parameter θ of the feature extraction model f θ That is, the learning unit 512 optimizes the feature extraction model f so that the following formula 9 is realized. θ The parameter θ is adjusted as follows: [Equation 9]
[0070] Although the triplet loss is used as the second loss function in the above example, any loss function suitable for metric learning can be used as the second loss function.
[0071] The first learning (first image set D) described in the fifth embodiment df and a third image set D fr The first learning of the feature extraction operation using the feature extraction model for face matching f θFace That is, the learning unit 512 may perform additional learning on θ Face may be used as an initial parameter, and additional learning may be performed using the loss functions expressed by the above formula 2 and formula 8. In this case, the second image set and the third image set may be the same image set. Alternatively, the second image set and the third image set may be different image sets. Furthermore, the learning unit 512 may use the constraint term expressed by the above formula 5 to generate a feature extraction model f θ The parameter θ may be adjusted. [5-2: Technical Effects of Learning Device 5]
[0072] The learning device 5 according to the present disclosure learns to extract features used for face matching while learning to extract features used for detecting fake images. Therefore, a feature extraction model f is suitable for both extracting features used for face matching and extracting features used for detecting fake images. θ [6: Sixth embodiment]
[0073] A sixth embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described. Hereinafter, the sixth embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described using an information processing device 6 according to the present disclosure. [6-1: Configuration of Information Processing Device 6]
[0074] The configuration of the information processing device 6 will be described with reference to Fig. 7. Fig. 7 is a block diagram showing the configuration of the information processing device 6.
[0075] 7, the information processing device 6 includes a calculation device 21, a storage device 22, and an output device 25, similar to the learning devices 2 to 5. Furthermore, the information processing device 6 may include a communication device 23 and an input device 24, similar to the learning devices 2 to 5. However, the information processing device 6 does not need to include at least one of the communication device 23 and the input device 24. The information processing device 6 implements a feature extraction unit 611, a reception unit 613, a calculation unit 614, a determination unit 615, and an output unit 616 within the calculation device 21. [6-2: Information Processing Operation Performed by the Information Processing Device 6]
[0076] The fake determination operation performed by the information processing device 6 will be described with reference to Fig. 8. Fig. 8 is a flowchart showing an example of the flow of the fake determination operation performed by the information processing device 6.
[0077] 8 , the receiving unit 613 receives input of a target image to be determined as to whether or not it has been combined (step S60). The receiving unit 613 may receive input of a face image including a person's face area as the target image. The target image may be a still image. The target image may be a moving image. The receiving unit 613 may acquire the target image via the communication device 23 or the input device 24.
[0078] The receiving unit 613 receives input of a reference image (step S61). The receiving unit 613 may receive input of a face image including a person's face region as the reference image. The reference image may be a still image. The reference image may be a moving image. The receiving unit 613 may acquire a registered image that is registered in advance in the storage device 22 as the reference image. Note that in the present disclosure, it is assumed that the reference image is not a fake image but a real image.
[0079] The feature extraction unit 611 extracts a feature extraction model f θ The feature extraction model f is used to extract features from the image. θ is a model trained by the learning device according to any one of the first to fifth embodiments. θ is at least the first image set D df This is a model that has undergone at least a first training using
[0080] The feature extraction unit 611 extracts features from the target image (referred to as "target features") and also extracts features from the reference image (referred to as "reference features") (step S62).
[0081] The calculation unit 614 calculates the similarity between the target feature and the reference feature (step S63). The determination unit 615 determines whether the similarity exceeds a threshold value (step S64).
[0082] If the similarity exceeds the threshold (step S64: Yes), the determination unit 615 determines that the target image is a genuine image (step S65). If the similarity does not exceed the threshold (step S64: No), the determination unit 615 determines that the target image is a fake image (step S66).
[0083] The output unit 616 outputs according to the determination result (step S67). The output unit 616 may control the output device 25 to cause the output device 25 to output according to the determination result. [6-3: Technical Effects of the Information Processing Device 6]
[0084] The information processing device 6 according to the present disclosure receives at least a first image set D df A feature extraction model f that has undergone at least the first learning using θ Since the feature amount extracted using the above is used to determine whether an image is a fake or not, it is possible to detect fake images with high accuracy.
[0085] A seventh embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described. Hereinafter, the seventh embodiment of a learning device, an information processing device, a learning method, and a recording medium will be described using an information processing device 7 according to the present disclosure. [7-1: Configuration of Information Processing Device 7]
[0086] The configuration of the information processing device 7 will be described with reference to Fig. 9. Fig. 9 is a block diagram showing the configuration of the information processing device 7.
[0087] 9, the information processing device 7, like the information processing device 6, includes a calculation device 21, a storage device 22, and an output device 25. Furthermore, like the information processing device 6, the information processing device 7 may include a communication device 23 and an input device 24. However, the information processing device 7 does not need to include at least one of the communication device 23 and the input device 24. The information processing device 7 realizes a first feature extraction unit 711_1, a first calculation unit 714_1, a first determination unit 715_1, a second feature extraction unit 711_2, a second calculation unit 714_2, a second determination unit 715_2, a reception unit 713, and an output unit 716 within the calculation device 21. [7-2: Information Processing Operation Performed by Information Processing Device 7]
[0088] The authentication operation performed by the information processing device 7 will be described with reference to Fig. 10. Fig. 10 is a flowchart showing an example of the flow of the authentication operation performed by the information processing device 7.
[0089] 10A, the receiving unit 713 receives an input of a target image to be determined as to whether or not the person is the person in question (step S70). The receiving unit 713 may receive an input of a face image including a person's facial region as the target image. The target image may be a still image. The target image may be a moving image. The receiving unit 713 may acquire the target image via the communication device 23 or the input device 24.
[0090] The receiving unit 713 receives input of a reference image (step S71). The receiving unit 713 may receive input of a face image including a person's face area as the reference image. The reference image may be a still image. The reference image may be a moving image. The receiving unit 713 may acquire a registered image registered in advance in the storage device 22 as the reference image.
[0091] The first feature extraction unit 711_1 extracts a feature extraction model f θ The first feature extraction unit 711_1 extracts a first feature from the image using the feature extraction model f θ is a model learned by the learning device according to any one of the first to fifth embodiments. θ is at least the first image set Ddf This is a model that has undergone at least a first training using
[0092] The first feature extraction unit 711_1 extracts a first target feature from the target image, and also extracts a first reference feature from the reference image (step S62).
[0093] The first calculation unit 714_1 calculates a first similarity between the first target feature and the first reference feature (step S63). The first determination unit 715_1 performs a threshold determination of the first similarity. The first determination unit 715_1 determines whether the first similarity exceeds a first threshold (step S64).
[0094] If the first similarity exceeds the first threshold (step S64: Yes), the first determination unit 715_1 determines that the target image is a genuine image (step S65).If the first similarity does not exceed the first threshold (step S64: No), the first determination unit 715_1 determines that the target image is a fake image (step S66).
[0095] The output unit 716 outputs according to the determination result (step S67). The output unit 716 may control the output device 25 to cause the output device 25 to output according to the determination result. If the first determination unit 715_1 determines that the target image is a genuine image, step S67 may be skipped and the information processing device 7 may perform the operation shown in Fig. 10(b). On the other hand, if the first determination unit 715_1 determines that the target image is a fake image, the operation may be terminated without performing the operation shown in Fig. 10(b).
[0096] The target image and reference image used in the operation shown in Fig. 10(b) are the target image and reference image used in the operation shown in Fig. 10(a). Therefore, when the operation shown in Fig. 10(b) is executed following the operation shown in Fig. 10(a), the operations of steps S70 and S71 may be skipped.
[0097] As shown in FIG. 10B, the second feature extraction unit 711_2 extracts a feature extraction model f θThe second feature extraction unit 711_2 extracts a second feature from the image using the feature extraction model f θ is at least a second image set D fr The feature extraction model f used by the second feature extraction unit 711_2 is a model that has undergone at least second learning using θ The feature extraction model f used by the second feature extraction unit 711_2 may be a model trained by the learning device according to any one of the third to fifth embodiments. θ is the feature extraction model f used by the first feature extraction unit 711_1. θ The feature extraction model used by the second feature extraction unit 711_2 may be the same as the model of the first image set D df The model may be a model that has not undergone the first learning using the above.
[0098] The second feature extraction unit 711_2 extracts a second target feature from the target image, and also extracts a second reference feature from the reference image (step S72).
[0099] The second calculation unit 714_2 calculates a second similarity between the second target feature and the second reference feature (step S73). The second determination unit 715_2 performs a threshold determination of the second similarity. The first determination unit 715_1 determines whether the second similarity exceeds a second threshold (step S74).
[0100] If the second similarity exceeds the second threshold (step S74: Yes), the second determination unit 715_2 determines that the person appearing in the target image and the person appearing in the reference image are the same person and that the person appearing in the target image is the person in question (step S75).If the second similarity does not exceed the second threshold (step S74: No), the second determination unit 715_2 determines that the person appearing in the target image and the person appearing in the reference image are different people (step S76).
[0101] The output unit 716 outputs according to the determination result (step S77). The output unit 716 may control the output device 25 to cause the output device 25 to output according to the determination result.
[0102] The information processing device 7 may first perform the operation shown in Fig. 10(b), and then perform the operation shown in Fig. 10(a) after the operation shown in Fig. 10(b). In this case, if the second determination unit 715_2 determines that the person appearing in the target image and the person appearing in the reference image are different people, the information processing device 7 may end its operation without performing the operation shown in Fig. 10(a).
[0103] 10(a) and 10(b) may be performed in parallel. In this case, the information processing device 7 may determine whether or not a person appearing in a target image can be authenticated based on the relationship between the first similarity and the first threshold value and the relationship between the second similarity and the second threshold value. [7-3: Technical Effects of the Information Processing Device 7]
[0104] The information processing device 7 of the present disclosure can accurately detect whether an input video to be determined is a fake video or not, and therefore can accurately verify the identity of the person. [8: Note]
[0105] The following supplementary notes are further disclosed regarding the above-described embodiment: [Supplementary Note 1] A learning device comprising: feature extraction means for extracting features from images; and learning means for causing the feature extraction means to perform first learning of a feature extraction operation using a first image set including images labeled with a non-composite label indicating that the images are not composited and images labeled with a composite label indicating that the images are composited, such that a first feature extracted from the first images labeled with the non-composite label is more similar to a third feature extracted from a third image different from the first images labeled with the non-composite label than to a second feature extracted from a second image labeled with the composite label. [Supplementary Note 2] The learning device according to Supplementary Note 1, wherein the feature extraction means, when an image is input, performs the feature extraction operation using a feature extraction model that outputs features of the image, and the learning means performs the first learning by adjusting parameters of the feature extraction model using a loss function that reduces loss when a similarity between the first feature and the second feature is smaller than a similarity between the first feature and the third feature. [Supplementary Note 3] The learning device according to Supplementary Note 1, wherein the learning means causes the feature extraction means to perform first learning of the feature extraction operation using the first image set and a third image set to which labels that identify individuals are attached, such that the first feature is more similar to the third feature than the second feature, and such that features extracted from images that are attached with the same label are more similar to each other than features extracted from images that are attached with different labels.[Supplementary Note 4] The feature extraction means performs the feature extraction operation using a feature extraction model that outputs features of an image when an image is input, and the learning means performs the first learning by adjusting parameters of the feature extraction model using a loss function that reduces loss when a similarity between the first feature and the second feature is smaller than a similarity between the first feature and the third feature, and a loss function that reduces loss when a similarity between features extracted from images with different labels is smaller than a similarity between features extracted from images with the same label. [Supplementary Note 5] The learning device described in any one of Supplementary Notes 1 to 4, wherein the learning means causes the feature extraction means that has performed second learning of the feature extraction operation using a second set of images with labels that identify individuals so that features extracted from images with the same label are more similar to each other than features extracted from images with different labels. [Supplementary Note 6] The learning device according to Supplementary Note 2 or 4, wherein the learning means, when causing the feature extraction means, which has performed second learning of a feature extraction operation using a second set of images labeled to identify individuals so that features extracted from images labeled with the same label are more similar than features extracted from images labeled with different labels, to perform the first learning, limits adjustment of parameters of the feature extraction model used in the second learning so that the more important a parameter is, the smaller the change in the parameter in the first learning.[Supplementary Note 7] An information processing apparatus comprising: a first feature extraction means that performs first learning of a feature extraction operation using a first image set including images labeled with a non-composite label indicating that the image is not composited and images labeled with a composite label indicating that the image is composited, such that a first feature extracted from a first image labeled with the non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with the non-composite label than a second feature extracted from a second image labeled with the composite label; a first calculation means that calculates a similarity between a target feature extracted by the feature extraction means from a target image to be determined as to whether or not the image is composited, and a reference feature extracted by the feature extraction means from a reference image; a determination means that determines whether or not the target image is composited by determining a threshold value of the similarity; and an output means that performs output according to the determination result by the determination means. [Supplementary Note 8] A second feature extraction means that has performed second learning of a feature extraction operation using a second image set that is labeled to identify an individual, such that feature amounts extracted from images labeled with the same label are more similar than feature amounts extracted from images labeled with different labels, and that has performed first learning of a feature extraction operation using a first image set that includes images labeled with a non-composite label indicating that the image is not composited and images labeled with a composite label indicating that the image is composited, such that a first feature amount extracted from a first image labeled with the non-composite label is more similar to a third feature amount extracted from a third image that is different from the first image labeled with the non-composite label than to a second feature amount extracted from a second image labeled with the composite label; a second calculation means that calculates a similarity between a target feature extracted by the feature extraction means from a target image that is a target for determining whether or not the target is the actual person, and a reference feature extracted by the feature extraction means from a reference image; a second determination means that determines whether or not the target is the actual person by determining a threshold value of the similarity; and an output unit that outputs an output according to the determination result by the determination unit.[Supplementary Note 9] A learning method comprising: extracting features from images using a feature extraction model; and using a first image set including images labeled with a non-composite label indicating that the images are not composited and images labeled with a composite label indicating that the images are composited, causing the feature extraction model to perform a first learning of a feature extraction operation such that a first feature extracted from the first image labeled with the non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with the non-composite label than to a second feature extracted from a second image labeled with the composite label. [Supplementary Note 10] A recording medium having recorded thereon a computer program for causing a computer to execute a learning method, the learning method comprising: extracting features from images using a feature extraction model; and using a first image set including images labeled with a non-composite label indicating that the images are not composited and images labeled with a composite label indicating that the images are composited, having the feature extraction model perform a first learning of a feature extraction operation such that a first feature extracted from the first image labeled with the non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with the non-composite label than to a second feature extracted from a second image labeled with the composite label.
[0106] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0107] 1, 2, 3, 4, 5 Learning device 11, 211, 311, 411, 511, 611 Feature extraction unit 12, 212, 312, 412, 512 Learning unit 6, 7 Information processing device 613, 713 Reception unit 614 Calculation unit 615 Determination unit 616, 716 Output unit 711_1 First feature extraction unit 711_2 Second feature extraction unit 714_1 First calculation unit 714_2 Second calculation unit 715_1 First determination unit 715_2 Second determination unit
Claims
1. feature extraction means for extracting features from an image; Using a first image set including images labeled as non-composite, which indicates that the images are not composited, and images labeled as composite, which indicates that the images are composited, The first feature extracted from the first image labeled as non-composite is than the second feature extracted from the second image to which the synthetic label is attached, learning means for causing the feature extraction means to perform first learning of a feature extraction operation so as to resemble a third feature extracted from a third image different from the first image to which the non-composite label is attached; and A learning device comprising:
2. When an image is input, the feature extraction means executes the feature extraction operation using a feature extraction model that outputs a feature of the image; The learning means performs the first learning by adjusting parameters of the feature extraction model using a loss function that reduces loss when a similarity between the first feature and the second feature is smaller than a similarity between the first feature and the third feature. The learning device according to claim 1 .
3. The learning means causes the feature extraction means to Using the first image set and a third image set to which labels that identify individuals are attached, the first feature amount is more similar to the third feature amount than the second feature amount, and A first learning of the feature extraction operation is performed so that feature amounts extracted from images with the same label are more similar to each other than feature amounts extracted from images with different labels. The learning device according to claim 1 .
4. When an image is input, the feature extraction means performs the feature extraction operation using a feature extraction model that outputs a feature of the image; The learning means adjusts parameters of the feature extraction model using a loss function that reduces loss when the similarity between the first feature and the second feature is smaller than the similarity between the first feature and the third feature, and a loss function that reduces loss when the similarity between the feature extracted from the images with different labels is smaller than the similarity between the feature extracted from the images with the same label, thereby performing the first learning. The learning device according to claim 3 .
5. The learning means Using a second set of images labeled to identify individuals, the feature extraction means that has performed second learning of the feature extraction operation so that feature amounts extracted from images labeled with the same label are more similar to each other than feature amounts extracted from images labeled with different labels is made to perform the first learning. The learning device according to any one of claims 1 to 4.
6. The learning means when causing the feature extraction means, which has performed second learning of a feature extraction operation using a second image set to which labels that identify individuals are attached, so that feature amounts extracted from images attached with the same label are more similar to each other than feature amounts extracted from images attached with different labels, to perform the first learning, Among the parameters of the feature extraction model that have undergone the second learning, the adjustment of the parameters of the feature extraction model is limited so that the more important a parameter is, the smaller the change in the parameter in the first learning is. The learning device according to claim 2 or 4.
7. a first feature extraction means that performs first learning of a feature extraction operation using a first image set including images labeled with a non-composite label indicating that the images are not composited and images labeled with a composite label indicating that the images are composited, such that a first feature extracted from a first image labeled with the non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with the non-composite label than to a second feature extracted from a second image labeled with the composite label; a first calculation means for calculating a similarity between a target feature extracted by the first feature extraction means from a target image to be determined as to whether or not the target image is combined and a reference feature extracted by the first feature extraction means from a reference image; a determination means for determining whether the target image has been synthesized based on a threshold value determination of the similarity; an output means for outputting an output according to the determination result by the determination means; An information processing device comprising:
8. a second feature extraction means that performs second learning of a feature extraction operation using a second image set that is labeled to identify an individual, such that feature amounts extracted from images labeled with the same label are more similar to feature amounts extracted from images labeled with different labels; and a second feature extraction means that performs first learning of a feature extraction operation using a first image set that includes images labeled with a non-composite label indicating that the image is not composited and images labeled with a composite label indicating that the image is composited, such that a first feature amount extracted from a first image labeled with the non-composite label is more similar to a third feature amount extracted from a third image different from the first image labeled with the non-composite label than to a second feature amount extracted from a second image labeled with the composite label; a second calculation means for calculating a similarity between a target feature extracted by the second feature extraction means from a target image to be determined as to whether the person is the same as the person in question and a reference feature extracted by the second feature extraction means from a reference image; a second determination means for determining whether the subject is the person in question based on a threshold determination of the degree of similarity; an output means for outputting an output according to the determination result by the determination means; An information processing device comprising:
9. Extract features from the image using a feature extraction model, using a first image set including images labeled with a non-composite label indicating that the images are not composited and images labeled with a composite label indicating that the images are composited, the feature extraction model is caused to perform a first learning of a feature extraction operation such that a first feature extracted from the first image labeled with the non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with the non-composite label than to a second feature extracted from a second image labeled with the composite label; Computer-implemented learning methods.
10. On the computer, Extract features from the image using a feature extraction model, Using a first image set including images labeled with a non-composite label indicating that the images are not composited and images labeled with a composite label indicating that the images are composited, the feature extraction model is made to perform a first learning of a feature extraction operation such that a first feature extracted from the first image labeled with the non-composite label is more similar to a third feature extracted from a third image different from the first image labeled with the non-composite label than to a second feature extracted from a second image labeled with the composite label. A computer program for implementing a learning method.