Learning device, inference device, feature extraction device, learning method, and program

The learning device improves the similarity of facial images by applying noise, denoising, and updating parameters to enhance personalization in image generation AI for human faces, addressing the issue of inaccurate similarity in existing methods.

JP2026071064APending Publication Date: 2026-04-28CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing methods for personalizing image generation AI for human faces, such as Stable Diffusion, often generate multiple face images that are not accurately identified as the same person due to learning based on pixel values, leading to reduced similarity among facial images of the same individual.

Method used

A learning device that applies noise to face images, removes noise through denoising, and updates parameters based on feature information to enhance personal similarity among facial images, using techniques like CLIP encoders and U-Nets to improve facial feature matching.

Benefits of technology

The method enhances the similarity between facial images of the same person while reducing similarity between different individuals, resulting in higher-quality and more accurate personalization of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071064000001_ABST
    Figure 2026071064000001_ABST
Patent Text Reader

Abstract

Improve the similarity of multiple facial images of the same person. [Solution] The learning device includes: acquisition means for acquiring multiple face image groups containing multiple face images of the same person; noise application means for generating a noise-applied face image by applying noise to a noise-applied face image, which is one face image selected from multiple face images for each person included in the face image group; face image generation means for generating a restored face image, which is restored by removing at least noise from the noise-applied face image, based on the face image group, the noise-applied face image, and parameters; and parameter update means for updating the parameters based on feature information, which is information about the feature quantities of at least one of the face images included in the face image group and the restored face image, so that the personal similarity, which indicates the degree to which multiple face images generated as the same person are the same person, increases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning device, an inference device, a feature quantity calculation device, a learning method, and a program.

Background Art

[0002] In personalization that enables an image generation AI (Artificial Intelligence) to generate an image of a specific object, there is a method of personalization specialized for a human face. This method is used for generating desired photo materials for advertising and the like, generating data used for learning a machine learning model, and the like. Non-Patent Document 1 discloses a learning method for enabling an image generation AI such as Stable Diffusion specialized for a human face to perform personalization. Non-Patent Document 1 enables the generation of a human face image included in an input face image by performing learning of an image generation AI such as Stable Diffusion using a plurality of face images of the same person in advance.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in Non-Patent Document 1, since learning is performed based on the pixel values of an image, there are cases where a plurality of face images that cannot be regarded as face images of the same person are generated as face images of the same person.

[0005] Therefore, the present invention provides a technique for improving the similarity of multiple facial images of the same person. [Means for solving the problem]

[0006] To solve this problem, for example, the learning device of the present invention has the following configuration. That is, A means for acquiring multiple sets of facial images, each containing multiple facial images of the same person, A noise application means for generating a noise-applied face image by applying noise to a noise-applied face image, which is a single face image selected from multiple face images for each person included in the aforementioned group of face images, A face image generation means that generates a restored face image in which at least the noise has been removed from the noisy face image, based on the aforementioned group of face images, the noise-added face image, and parameters, A parameter update means updates the parameters based on feature information, which is information about the feature quantities of at least one of the face images included in the group of face images and the reconstructed face image, so that the personal similarity, which indicates the degree to which multiple face images generated as representing the same person are of the same person, increases. It is equipped with. [Effects of the Invention]

[0007] According to the present invention, the similarity between multiple facial images of the same person can be improved. [Brief explanation of the drawing]

[0008] [Figure 1] A block diagram showing the hardware configuration of the information processing device of the embodiment. [Figure 2] A functional block diagram to explain the functions of a learning device. [Figure 3] A diagram showing details of the training data, which includes data from multiple sets of facial images. [Figure 4] A diagram illustrating the selection of facial images for noise generation. [Figure 5]A simplified diagram showing the relationships between facial images used for matching to calculate the degree of similarity to the person. [Figure 6] A simplified diagram showing the relationships between facial images used for matching to calculate similarity between individuals. [Figure 7] A functional block diagram explaining the functions of the inference device. [Figure 8] A diagram illustrating the parameter updates in modified example 2. [Figure 9] A diagram illustrating the parameter updates in modified example 3. [Modes for carrying out the invention]

[0009] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, identical or similar configurations are given the same reference numerals, and redundant descriptions are omitted.

[0010] (Embodiment) The configuration of the information processing device 100 in this embodiment will be described with reference to the block diagram in Figure 1. Figure 1 is a block diagram showing the hardware configuration of the information processing device 100 in this embodiment. As shown in Figure 1, the information processing device 100 may be a so-called computer. The information processing device 100 performs, for example, the personalization process of an image generation AI specialized for human faces. The information processing device 100 has a CPU 101, ROM 102, RAM 103, external storage device 104, input device interface 105, output device interface 106, communication interface 107, and system bus 108.

[0011] The CPU 101 is the abbreviation of Central Processing Unit and is a processor. The CPU 101 controls the entire information processing device 100. The information processing device 100 may have other processors such as an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), a NPU (Neural Processing Unit), and a QPU (Quantum Processing Unit) instead of or in addition to the CPU 101.

[0012] A part or all of each function of the information processing device 100 is realized by one or more processors including the CPU 101 reading out a computer program (hereinafter also referred to as a program) stored in the ROM 102 and the external storage device 104, expanding it in the RAM 103, and executing it. Also, a part or all of each function of the information processing device 100 may be realized by one or more circuits such as an ASIC (Application Specific Integrated Circuit) and a PLD (Programmable Logic Device) including an FPGA (Field Programmable Gate Array).

[0013] The ROM 102 is the abbreviation of Read Only Memory. The ROM 102 may be a non-volatile storage device. The ROM 102 stores programs that do not require modification, parameters necessary for program execution, and the like.

[0014] The RAM 103 is the abbreviation of Random Access Memory. The RAM 103 may be a memory that can read and write data at high speed compared to the ROM 102 and the external storage device 104. The RAM 103 temporarily stores programs supplied from external devices and the like, parameters necessary for program execution, and data to be processed by the program.

[0015] The external storage device 104 is a storage device fixedly installed in the information processing device 100. The external storage device 104 is, for example, a storage device such as a hard disk and a memory card with a larger capacity than the ROM 102 and the like. Note that the external storage device 104 may include a flexible disk (FD), an optical disk such as a Compact Disk (CD), a magnetic card, an optical card, an IC card, a memory card, etc. that are detachable from the information processing device 100.

[0016] The input device interface 105 receives operations from the user's input device 109 and transfers them to the CPU 101 and the like. The input device 109 may be a pointing device capable of inputting data, a keyboard, or the like.

[0017] The output device interface 106 is an interface for outputting the data held by the information processing device 100 to an external device. The external device is, for example, a monitor 110 for displaying an image generated by the information processing device 100.

[0018] The communication interface 107 is a communication interface for connecting to a network line 111 such as the Internet. The communication interface 107 transmits and receives data between the information processing device 100 and an information processing device installed externally via the network line 111.

[0019] The system bus 108 is a transmission path that communicably connects each unit from the CPU 101 to the communication interface 107. Each process described later functions as a process when the CPU 101 executes a program stored in a computer-readable storage medium such as the ROM 102.

[0020] (Configuration of the learning device) The learning device of the present embodiment performs learning so that an image generation AI such as Stable Diffusion can perform personalization specialized for a person's face. The learning is executed by the information processing device 100 shown in FIG. 1.

[0021] The configuration of the learning device 200 in this embodiment will be explained using the block diagram in Figure 2. Figure 2 is a functional block diagram for explaining the functions of the learning device 200. The learning device 200 updates the parameters for generating the reconstructed face image based on feature information, which is information about the feature quantities of at least one of the face images and the reconstructed face image included in the face image group, so that the similarity to the person, which indicates that multiple face images of the same person are of the same person, increases. The feature information may be information that affects the feature quantities. Here, a face feature is a feature that determines whether or not the person shown in multiple face images can be considered to be the same person. Face features include, for example, feature quantities used in face recognition matching.

[0022] The learning device 200 includes a learning data holding unit 201, a face image group acquisition unit 202, a prompt acquisition unit 203, a downsampling unit 204, an image encoding unit 205, a text encoding unit 206, a feature integration unit 207, a noise addition unit 208, a denoising unit 209, an MSE loss calculation unit 210, an upsampling unit 211, a face detection unit 212, a face feature calculation unit 213, an edge value calculation unit 214, a self-similarity calculation unit 215, a non-self-similarity calculation unit 216, an edge value difference calculation unit 217, and a parameter update unit 218. The downsampling unit 204, image encoding unit 205, text encoding unit 206, feature integration unit 207, denoising unit 209, and upsampling unit 211 are examples of face image generation means.

[0023] The functions of the learning device 200, from the learning data holding unit 201 to the parameter update unit 218, are realized when the CPU 101 of the information processing device 100 reads a learning program stored in an external storage device 104 or the like and loads it into the RAM 103.

[0024] The learning data holding unit 201 holds the learning data stored in the external storage device 104. Figure 3 shows the details of the learning data LD, which contains data from multiple face image groups IMGs. The learning data LD contains multiple face image groups IMGs. One face image group IMG contains multiple face images of the same person. Therefore, the learning data LD has multiple face image groups IMGs, each containing multiple face images for a single person. The person in one face image group IMG (for example, person A1) may be different from the person in another face image group IMG (for example, person A2). Furthermore, the learning data LD includes prompts associated with each person's face image. That is, the learning data LD has multiple face image groups IMGs, each grouping the face images of the same person and the prompts associated with each face image. The prompts may be text that describes the image that you want the image generation AI, such as Stable Diffusion, to generate. For example, if you want to generate an image of "a photo of a smiling man," the prompt may be text such as "A photo of a smile man." The prompt may be the text output by an LLM (Large Language Model) that generates image captions after inputting a facial image. In this embodiment, the prompt must include common nouns referring to people, such as "man" and "woman".

[0025] Next, we will explain the facial images used for noise generation, referring to Figure 4. Figure 4 is a diagram illustrating the selection of facial images for noise generation. The training data LD in Figure 4 shows the data acquired in one iteration in mini-batch learning.

[0026] The face image acquisition unit 202 acquires multiple predetermined face image groups IMG from the training data LD held by the training data holding unit 201. In this embodiment, since mini-batch learning is performed, the face image acquisition unit 202 acquires face image groups with the same number of people as the batch size determined in advance. Here, each face image group includes at least two or more face images of the same person.

[0027] When the batch size is determined to be n, the face image acquisition unit 202 acquires n face image groups by acquiring two or more face images from each of the n person face image groups (IMG). The face image acquisition unit 202 also acquires the acquired person (for example, person A). b1 Among the multiple face images included in the group of face images of (san), one face image (for example, face image I) Ab1 +1) is randomly selected as a face image for noise generation. The face image acquisition unit 202 randomly selects a face image for noise generation for each face image group IMG, that is, for each person.

[0028] The prompt acquisition unit 203 selects and acquires prompts corresponding to the face images for noise addition and the face images selected by the face image group acquisition unit 202. The face images for noise addition are randomly selected by the face image group acquisition unit 202 as described above.

[0029] Referring to Figure 4, the prompt for the facial image used for noise generation will be explained. When the predetermined batch size is n, the prompt acquisition unit 203 acquires a set of facial images IMG for each of the same number of people N1 as the batch size. Here, N1 is as follows (1).

number

[0030] Next, the details of the group of face images acquired by the prompt acquisition unit 203 for each person will be explained. First, the prompt acquisition unit 203 acquires at least one face image X1 for each person. Here, X1 is as shown in (2) below.

number

[0031] These facial images are mixed with the prompt features in the feature integration unit 207. Furthermore, the prompt acquisition unit 203 acquires one facial image X2 and its corresponding prompt for each person. Here, X2 is as shown in (3) below.

number

[0032] Face image X2 is a face image used for noise addition and denoising.

[0033] The downsampling unit 204 downsamples the image size of the face images selected as noise-adding face images from the face image group acquired by the face image group acquisition unit 202. Downsampling here refers to reducing the image size through a transformation using a neural network. The downsampling size can be arbitrary. As a downsampling method, a pre-trained Variational Autoencoder encoder may be used. The number of channels in the image may be increased during downsampling. This reduces the processing cost of the subsequent denoising process.

[0034] The image encoding unit 205 calculates feature quantities for face images that were not selected as noise-injecting face images from the face image acquisition unit 202, based on the parameters. The image encoding unit 205 may use the CLIP image encoder, which is a multimodal model of images and text and has pre-trained parameters, to calculate the feature quantities of face images. CLIP is a neural network trained to maximize the similarity between the feature quantities of correlated images and text.

[0035] The text encoding unit 206 calculates the feature vectors of the prompt selected by the prompt acquisition unit 203. The text encoding unit 206 may use the CLIP text encoder, which is a multimodal model of images and text, to calculate the feature vectors of the prompt.

[0036] The feature integrator 207 integrates the features of the face image calculated by the image encoding unit 205 and the features of the prompt calculated by the text encoding unit 206, specifically those features that correspond to common nouns referring to people, such as "man" and "woman" (referred to as person common noun features). The feature integrator 207 replaces the original person common noun features with the integrated features to create new prompt features. If there are multiple face images for a single person, the feature integrator 207 duplicates the person common noun features, integrates each one, and replaces the original person common noun features with the multiple integrated features. As a method of integration, the feature integrator 207 may use a Multi-layer Perceptron with pre-trained parameters as input, taking the concatenated features of the face image features and prompt features as input. As a method of concatenation, if the face image features and prompt features are one-dimensional features, the feature integrator 207 may simply use a process to increase the number of dimensions.

[0037] For example, if the prompt is "A photo of a smile man", the feature integration unit 207 integrates the features calculated from "man" with the features calculated by the image encoding unit 205. As a result, the features calculated from "man" are transformed from features representing an unspecified person to features representing a specific person.

[0038] The noise application unit 208 applies noise to the downsampled face image for noise application by the downsampling unit 204 to generate a noise-applied image. The noise application unit 208 may use a process that repeatedly applies minute noise (e.g., Gaussian noise) as a method of applying noise.

[0039] The denoising unit 209 performs denoising on the noised face image, which has been noised by the noise application unit 208, based on the parameters, to generate a denoised restored face image. The denoising unit 209 may use a process that performs noise estimation multiple times assuming a Markov chain as the denoising method. Alternatively, the denoising unit 209 may use a neural network such as U-Net with pre-trained parameters for noise estimation. The learning of the neural network by the noise application and denoising processes represents a learning method for image generation models known as diffusion models.

[0040] The MSE loss calculation unit 210 compares the face image used for noise application before noise application by the noise application unit 208 with the denoised restored face image after denoising by the denoising processing unit 209, and calculates a numerical value related to the difference in pixel values ​​(numerical values ​​for each pixel) (here, the mean squared error (MSE)). The mean squared error is an example of feature information and difference information. The face image used for noise application here may be an image downsampled to the same size as the denoised face image by the downsampling unit 204. The MSE loss calculation unit 210 is an example of the first calculation means.

[0041] The upsampling unit 211 upsamples the denoised restored face image, which has been denoised by the denoising processing unit 209, and reduces its image size to the same size as the face image before it was downsampled by the downsampling unit 204, thereby generating a restored face image of the same size. In this case, the number of channels is set to 3. If the encoder of a pre-trained Variational Autoencoder is used in the downsampling unit 204, the upsampling unit 211 uses the decoder corresponding to that encoder as the upsampling method. As a result, if the denoising process is performed appropriately, the upsampling unit 211 can restore the face image before noise was added.

[0042] The face detection unit 212 takes the face image group acquired by the face image group acquisition unit 202 and the reconstructed face image upsampled by the upsampling unit 211 as input and detects the bounding boxes of the face regions, which are the face regions of the face images in the face image group and the reconstructed face image. However, the face detection unit 212 detects the bounding boxes in such a way that face images that were in a corresponding relationship in the MSE loss calculation unit 210 are detected in the same bounding box region.

[0043] The face feature calculation unit 213 calculates face features using the face region detected by the face detection unit 212 as input. The face feature calculation unit 213 is an example of a fourth calculation means. The face feature calculation unit 213 may use a pre-trained machine learning model used for face recognition to calculate face features. The face feature calculation unit 213 may calculate features used for matching face recognition generated using face images generated by the inference device described later as training data.

[0044] The edge value calculation unit 214 takes the face region detected by the face detection unit 212 as input and calculates an edge value, which is a value indicating the edge of the face image. The edge value calculation unit 214 is an example of a fourth calculation means.

[0045] The personal resemblance calculation unit 215 calculates personal resemblance, which is the similarity between the reconstructed face image upsampled by the upsampling unit 211 and the face images that were not selected as noise-adding face images from the group of face images of the same person selected by the face image group acquisition unit 202. The face images that were not selected as noise-adding face images from the group of face images of the same person selected by the face image group acquisition unit 202 are other face images of the same person as the original face image of the reconstructed face image. The edge value calculation unit 214 is an example of a fourth calculation means. The personal resemblance calculation unit 215 is an example of a second calculation means. Specifically, the personal resemblance calculation unit 215 calculates personal resemblance by comparing face features calculated based on the reconstructed face image upsampled by the upsampling unit 211 with face features calculated based on face images that were not selected as noise-adding face images from the group of face images of the same person selected by the face image group acquisition unit 202. In other words, personal resemblance is an indicator of the degree to which multiple face images generated as the same person are of the same person. Facial features are an example of feature information.

[0046] Figure 5 is a simplified diagram showing the relationship between face images used for matching to calculate personal similarity. The image generation AI in Figure 5 includes a downsampling unit 204, an image encoding unit 205, a text encoding unit 206, a feature integration unit 207, a denoising unit 209, and an upsampling unit 211. As shown in Figure 5, the personal similarity calculation unit 215 uses a certain person (in this case, person A) to perform the matching. b1 A denoised and restored face image of the same person (for example, Person A) that was not selected as a face image for noise application, and a face image of the same person that was not selected as a face image for noise application. b1 The facial image of the person (1) is compared with the image of the person in question to calculate the degree of similarity.

[0047] The other-person similarity calculation unit 216 calculates the similarity between different individuals by comparing the facial features calculated based on the restored facial image upsampled and restored by the upsampling unit 211. The other-person similarity calculation unit 216 is an example of a third calculation means. Figure 6 is a simplified diagram showing the relationship between facial images used for comparison in calculating the other-person similarity. As shown in Figure 6, the other-person similarity calculation unit 216 compares the first person (in this case, person A) b1 After adding noise to the face image used for noise generation (let's call it Person A), a denoising process is performed to remove the noise, resulting in a noise-free face image, and a second person (in this case, Person A) who is different from the first person. b2 After adding noise to the noise-adding face image (as defined above), denoising is performed to remove the noise, and the similarity between the denoised face image and the original face image is calculated as the similarity between the two individuals.

[0048] The edge value difference calculation unit 217 calculates the difference between the edge values ​​calculated by the edge value calculation unit 214 based on the corresponding face images calculated by the MSE loss calculation unit 210.

[0049] The parameter update unit 218 updates the parameters used by the image encoding unit 205, the feature integration unit 207, and the denoising processing unit 209 to the following state. (1) The mean squared error calculated by the MSE loss calculation unit 210 is reduced. (2) To increase the personal similarity calculated by the personal similarity calculation unit 215. (3) The similarity score calculated by the similarity score calculation unit 216 is made smaller. (4) The difference value of the edge values ​​calculated by the edge value difference calculation unit 217 is made smaller.

[0050] The parameters updated here include the following: (1) Parameters of CLIP, which constitutes the image encoding unit 205. (2) Parameters of the MLP that constitute the feature integration unit 207. (3) Parameters of the U-Net that constitute the denoising processing unit 209.

[0051] The parameter update unit 218 performs this process simultaneously with data of the same number of individuals as the predetermined batch size, and executes the process where all training data has been cycled through as one epoch. The parameter update unit 218 may execute the epoch process any number of times. The parameter update unit 218 may use Lora to update the parameters of the U-Net used in the denoising unit 209 to reduce processing costs.

[0052] As described above, the learning device 200 of this embodiment updates the parameters to increase the similarity to the person, so it can retain the facial features of a specific person and increase the similarity of multiple facial images of the same person. As a result, the learning device 200 can improve the similarity of facial images of the same person generated by the upsampling unit 211 using these parameters, making the facial images more similar.

[0053] The learning device 200 of this embodiment updates the parameters to reduce the similarity between individuals, thereby retaining the facial features of a specific person and reducing the similarity between multiple facial images of different people. As a result, the learning device 200 can make the facial images of different people generated by the upsampling unit 211 using these parameters more dissimilar.

[0054] The learning device 200 of this embodiment updates the parameters so that the mean square error of the pixel values ​​between the noise-added face image and the denoised face image is reduced, thereby reducing the pixel values ​​between the two face images. As a result, the learning device 200 can reduce the difference between face images of the same person generated by the upsampling unit 211 using these parameters, thereby making the face images more similar.

[0055] The learning device 200 in this embodiment updates the parameters so that the difference in edge values ​​of the face region becomes smaller. As a result, the learning device 200 can improve the edges of the face image generated by the upsampling unit 211 using these parameters, thereby generating a high-quality face image.

[0056] (Configuration of the inference device) The inference device 300 of this embodiment takes a facial image of a specific person and a prompt as input, based on parameters updated by the learning device 200, and generates a facial image representing that person in accordance with the prompt. Specifically, the inference device 300 personalizes the person corresponding to the input facial image based on an image generation AI such as Stable Diffusion, and generates an image in accordance with the input prompt. The inference device 300 operates on the information processing device 100 shown in Figure 1.

[0057] The configuration of the inference device 300 in this embodiment will be explained using the block diagram in Figure 7. Figure 7 is a functional block diagram illustrating the functions of the inference device 300. The inference device 300 includes a face image group acquisition unit 301, a prompt acquisition unit 302, an image encoding unit 303, a text encoding unit 304, a feature integration unit 305, a noise generation unit 306, a denoising processing unit 307, and an upsampling unit 308. The CPU 101 may implement each of the functions from the face image group acquisition unit 301 to the upsampling unit 308 by reading an inference program from an external storage device 104 or the like and expanding it into the RAM 103.

[0058] The face image acquisition unit 301 acquires a group of face images that includes at least one face image of the same person who is the target of personalization.

[0059] The prompt acquisition unit 302 acquires an arbitrary prompt.

[0060] The image encoding unit 303 has the same function as the image encoding unit 205 of the learning device 200. The image encoding unit 303 holds the parameters learned by the learning device 200.

[0061] The text encoding unit 304 has the same function as the text encoding unit 206 of the learning device 200.

[0062] The feature integration unit 305 has the same function as the feature integration unit 207 of the learning device 200. The feature integration unit 305 holds the parameters learned by the learning device 200.

[0063] The noise generation unit 306 generates random noise images.

[0064] The denoising section 307 has the same function as the denoising section 209 of the learning device 200. The denoising section 307 holds the parameters learned by the learning device 200.

[0065] The upsampling unit 308 has the same function as the upsampling unit 211 of the learning device 200. Since the learning device 200 is trained to generate a face image of a specific person during its denoising and upsampling processes, the output of the upsampling unit 308 is a personalized face image. As for the upsampling method, if the encoder of the Variational Autoencoder pre-trained by the downsampling unit 204 is used, the decoder corresponding to that encoder is used.

[0066] (Effects of this embodiment) This embodiment uses facial feature matching used in facial recognition to enable image generation AI, including Stable Diffusion, to learn how to personalize for human faces.

[0067] Personalization performance can be improved by training the model to generate images where the similarity of features in facial images of the same person is high, and the similarity of features in facial images of different people is low. Furthermore, the quality of the generated facial images can be improved by training the model so that the edge strength of the generated facial images is the same as the edge strength of the facial images used for noise generation.

[0068] (Variation 1) In this embodiment, the edge value calculation unit 214 calculates edge information of the face image generated and uses it as one of the criteria for parameter updating in the parameter update unit 218. However, image features other than edge information, such as SIFT (Scale-invariant feature transform), may also be used.

[0069] (Modification 2) Figure 8 illustrates the parameter update in Modification 2. The parameter update unit 218 updates the parameters so that the personal resemblance calculated by the personal resemblance calculation unit 215 becomes large (for example, maximum). However, such a restrictive learning process may result in the generation of unnatural face images. Therefore, in Modification 2, as shown in Figure 8, the personal resemblance calculation unit 215 performs a pre-match between the face image group IMG acquired by the face image group acquisition unit 202 and the face image that is not used for noise addition of the same person to calculate personal resemblance. The parameter update unit 218 may update the parameters so that they are approximately the same as the personal resemblance calculated by the pre-match.

[0070] (Variation 3) Figure 9 illustrates the parameter update in Modification 3. The parameter update unit 218 updates the parameters so that the similarity to other people calculated by the similarity to other people calculation unit 216 becomes small (for example, minimal). However, in learning with such strict constraints, the upsampling unit 211 may generate unnatural face images. Therefore, as shown in Figure 9, the similarity is calculated by performing a pre-match with a noise-adding face image of a different person from the face image group IMG acquired by the face image group acquisition unit 202. The parameter update unit 218 may also update the parameters so that the similarity calculated by the pre-match and the similarity to other people are approximately the same.

[0071] (Modification 4) In the embodiments described above, the MSE loss calculation unit 210 compares the face image before noise is added by the noise addition unit 208 with the face image after denoising by the denoising processing unit 209 to calculate the mean squared error (MSE) of the pixel values ​​(numerical values ​​for each pixel). However, the comparison targets for the mean squared error are not limited to these. For example, the MSE loss calculation unit 210 may compare the restored face image recovered by the upsampling unit 211 with the face image after denoising by the denoising processing unit 209 to calculate the mean squared error (MSE) of the pixel values ​​(numerical values ​​for each pixel).

[0072] (Variation 5) In the embodiment described above, an example was shown in which the parameter update unit 218 updates the parameters to increase the personal resemblance (hereinafter referred to as the first personal resemblance) calculated by the personal resemblance calculation unit 215. However, parameter updates based on personal resemblance are not limited to this. For example, the personal resemblance calculation unit 215 may calculate the personal resemblance (hereinafter referred to as the second personal resemblance) between the parameter-assigning face image and other face images of the same person as the parameter-assigning face image. In this case, the parameter update unit 218 may update the parameters to increase the first personal resemblance and to reduce the magnitude of the difference between the first personal resemblance and the second personal resemblance.

[0073] (Experimental variation 6) In the above embodiment, the interpersonal similarity calculation unit 216 calculates an interpersonal similarity score (hereinafter referred to as the first interpersonal similarity score), which is the similarity between reconstructed facial images of different people (unrelated individuals), and the parameter update unit 218 updates the parameters so that the first interpersonal similarity score decreases. However, parameter updates based on interpersonal similarity scores are not limited to this. For example, the interpersonal similarity calculation unit 216 may calculate an interpersonal similarity score (hereinafter referred to as the second interpersonal similarity score), which is the similarity between facial images used for noise addition of different people. In this case, the parameter update unit 218 may update the parameters so that the first interpersonal similarity score decreases and the difference between the first interpersonal similarity score and the second interpersonal similarity score decreases.

[0074] (Example 7) In the above-described embodiment, a learning device 200 having a downsampling unit 204 and an upsampling unit 211 was used as an example, but the downsampling unit 204 and the upsampling unit 211 may be omitted. In this case, the noise application unit 208 may apply noise to a face image for noise application that has not been downsampled. The denoising processing unit 209 may also output the denoised face image, obtained by removing noise from the face image to which noise has been applied, as a reconstructed face image to the MSE loss calculation unit 210 and the face detection unit 212.

[0075] (Variation 8) In the above-described embodiment, an example was given in which the edge value calculation unit 214 calculates edge values ​​for the face region, but the edge value calculation unit 214 may calculate edge values ​​for other regions. For example, the edge value calculation unit 214 may calculate edge values ​​for facial features (eyes, nose, mouth, ears), etc.

[0076] (Extreme variation 9) The learning device of the above-described embodiment may calculate feature information by other processes. For example, the learning device may calculate other feature information by applying various filtering processes.

[0077] (Other examples) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. Furthermore, the present invention can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0078] The disclosures herein include the following learning devices, inference devices, feature extraction devices, learning methods, and programs. (Item 1) A means for acquiring multiple sets of facial images, each containing multiple facial images of the same person, A noise application means for generating a noise-applied face image by applying noise to a noise-applied face image, which is a single face image selected from multiple face images for each person included in the aforementioned group of face images, A face image generation means that generates a restored face image in which at least the noise has been removed from the noisy face image, based on the aforementioned group of face images, the noise-added face image, and parameters, A parameter update means updates the parameters based on feature information, which is information about the feature quantities of at least one of the face images included in the group of face images and the reconstructed face image, so that the personal similarity, which indicates the degree to which multiple face images generated as representing the same person are of the same person, increases. A learning device characterized by being equipped with the following features. (Item 2) The system includes a prompt acquisition means for acquiring a prompt corresponding to the aforementioned facial image for noise addition, The face image generation means further generates the restored face image by restoring the noised face image based on the prompt. A learning device as described in item 1, characterized by the features described herein. (Item 3) The system includes a first calculation means that calculates difference information as feature information based on the difference in pixel values ​​between the restored face image and the face image used for noise application before noise is applied, The parameter update means updates the parameter based on the difference information. A learning device according to item 1 or item 2, characterized by the features described herein. (Item 4) The system includes a second calculation means that calculates a first personal resemblance, which indicates the degree of facial similarity between the reconstructed facial image and other facial images of the same person as the original facial image of the reconstructed facial image, as feature information. The parameter updating means updates the parameters so that the first personal resemblance score increases. A learning device according to any one of items 1 to 3, characterized by the features described herein. (Item 5) The second calculation means calculates a second personal similarity, which indicates the degree of facial similarity between the noise-adding face image and other face images of the same person as the noise-adding face image, as the feature information. The parameter updating means updates the parameters such that the difference between the first personal resemblance and the second personal resemblance becomes smaller. A learning device as described in item 4, characterized by the features described herein. (Item 6) The system includes a third calculation means that calculates a first similarity to another person, which is the similarity between reconstructed facial images of different individuals, as the feature information. The parameter updating means updates the parameters so that the first similarity to another person decreases. A learning device according to any one of items 1 to 5, characterized by the features described herein. (Item 7) The third calculation means calculates a second similarity to other people, which is the similarity between different people's facial images used for noise generation, as the feature information. The parameter updating means updates the parameters such that the magnitude of the difference between the first similarity score and the second similarity score decreases. A learning device as described in item 6, characterized by the features described herein. (Item 8) The system includes a fourth calculation unit that calculates facial feature quantities for the facial region, which is the facial region of the noise-adding facial image and the restored facial image. The parameter update means updates the parameters based on the facial features. A learning device according to any one of items 1 to 7, characterized by the features described herein. (Item 9) The system includes a fourth calculation unit that calculates edge values ​​indicating the edges of the face regions, which are the face regions of the face images in the group of face images and the reconstructed face images. The parameter update means updates the parameter based on the edge value. A learning device according to any one of items 1 to 8, characterized by the features described above. (Item 10) The noise-injecting means applies noise to the downsampled face image for noise injection. A learning device according to any one of items 1 to 9, characterized by the features described herein. (Item 11) The face image generation means generates the restored face image by upsampling the image from which the noise has been removed from the noise-added image. A learning device according to any one of items 1 to 10, characterized by the features described herein. (Item 12) An inference device that, based on the parameters updated by the learning device described in item 1, takes a facial image of a specific person and a prompt as input and generates a facial image representing the person in accordance with the prompt. (Item 13) A feature computed device that calculates features used for matching facial recognition generated using facial images generated by the inference device described in item 12 as training data. (Item 14) Obtain multiple sets of facial images, each containing multiple facial images of the same person. For each person included in the aforementioned group of face images, noise is applied to a single face image selected from multiple face images, which is the noise-inducing face image, to generate a noise-inducing face image. Based on the aforementioned group of face images, the noise-treated face image, and parameters, a restored face image is generated from the noise-treated face image, with at least the noise removed. Based on feature information, which is information about the feature quantities of at least one of the facial images included in the facial image group and the reconstructed facial image, the parameters are updated so that the personal resemblance, which indicates to a certain extent that multiple facial images of the same person are of the same person, increases. A learning method characterized by the following: (Item 15) A program to cause a computer to function as one of the means of a learning device described in any one of items 1 through 11.

[0079] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of Symbols]

[0080] 100...Information processing device, 200...Learning device, 202...Face image group acquisition unit, 203...Prompt acquisition unit, 204...Downsampling unit, 205...Image encoding unit, 206...Text encoding unit, 207...Feature integration unit, 208...Noise addition unit, 209...Denoising processing unit, 210...MSE loss calculation unit, 211...Upsampling unit, 213...Face feature calculation unit, 214...Edge value calculation unit, 215...Self-similarity calculation unit, 216...Other-person-similarity calculation unit, 217...Edge value difference calculation unit, 218...Parameter update unit, 300...Inference device, 301...Face image group acquisition unit, 302...Prompt acquisition unit, 303...Image encoding unit, 304...Text encoding unit, 305...Feature integration unit, 306...Noise generation unit, 307...Denoising unit, 308...Upsampling unit.

Claims

1. A means for acquiring multiple sets of facial images, each containing multiple facial images of the same person, A noise application means for generating a noise-applied face image by applying noise to a noise-applied face image, which is a single face image selected from multiple face images for each person included in the aforementioned group of face images, A face image generation means that generates a restored face image in which at least the noise has been removed from the noisy face image, based on the aforementioned group of face images, the noise-added face image, and parameters, A parameter update means updates the parameters based on feature information, which is information about the feature quantities of at least one of the face images included in the group of face images and the reconstructed face image, so that the personal similarity, which indicates the degree to which multiple face images generated as representing the same person are the same person, increases. A learning device characterized by being equipped with the following features.

2. The system includes a prompt acquisition means for acquiring a prompt corresponding to the aforementioned facial image for noise addition, The face image generation means further generates the restored face image by restoring the noise-added face image based on the prompt. The learning device according to feature 1.

3. The system includes a first calculation means that calculates difference information as feature information based on the difference in pixel values ​​between the restored face image and the face image used for noise application before noise is applied, The parameter update means updates the parameter based on the difference information. The learning device according to feature 1.

4. The system includes a second calculation means that calculates a first personal resemblance, which indicates the degree of facial similarity between the reconstructed facial image and other facial images of the same person as the original facial image of the reconstructed facial image, as feature information. The parameter updating means updates the parameters so that the first personal resemblance score increases. The learning device according to feature 1.

5. The second calculation means calculates a second personal similarity, which indicates the degree of facial similarity between the noise-adding face image and other face images of the same person as the noise-adding face image, as the feature information. The parameter update means updates the parameters such that the difference between the first personal resemblance and the second personal resemblance becomes smaller. The learning device according to feature 4.

6. The system includes a third calculation means that calculates a first similarity to another person, which is the similarity between reconstructed facial images of different individuals, as the feature information. The parameter updating means updates the parameters so that the first similarity to another person decreases. The learning device according to feature 1.

7. The third calculation means calculates a second similarity to another person, which is the similarity between different facial images used for noise generation, as the feature information. The parameter updating means updates the parameters such that the difference between the first similarity score and the second similarity score becomes smaller. The learning device according to feature 6.

8. The system includes a fourth calculation unit that calculates facial feature quantities for the facial region, which is the facial region of the noise-adding facial image and the restored facial image. The parameter update means updates the parameters based on the facial features. The learning device according to feature 1.

9. The system includes a fourth calculation unit that calculates edge values ​​indicating the edges of the face regions, which are the face regions of the face images in the group of face images and the reconstructed face images. The parameter update means updates the parameter based on the edge value. The learning device according to feature 1.

10. The noise-injecting means applies noise to the downsampled face image for noise injection. The learning device according to feature 1.

11. The face image generation means generates the restored face image by upsampling the image from which the noise has been removed from the noise-added image. The learning device according to feature 1.

12. An inference device that, based on the parameters updated by the learning device according to claim 1, takes a facial image of a specific person and a prompt as input and generates a facial image representing the person in accordance with the prompt.

13. A feature calculation device that calculates features used for matching facial recognition generated using facial images generated by the inference device described in claim 12 as training data.

14. Obtain multiple sets of facial images, each containing multiple facial images of the same person. For each person included in the aforementioned group of face images, noise is applied to a single face image selected from multiple face images, which is the noise-inducing face image, to generate a noise-inducing face image. Based on the aforementioned group of face images, the noise-treated face image, and parameters, a restored face image is generated from the noise-treated face image, with at least the noise removed. Based on feature information, which is information about the feature quantities of at least one of the facial images included in the facial image group and the reconstructed facial image, the parameters are updated so that the personal resemblance, which indicates to a certain extent that multiple facial images of the same person are of the same person, increases. A learning method characterized by the following:

15. A program for causing a computer to function as one of the means of a learning device according to any one of claims 1 to 11.