Information processing system, information processing method, and program

A two-stage generative model generates diverse, high-quality, and fair synthetic face images, addressing the limitations of existing methods and enhancing face recognition model performance and ethical compliance.

JP2026061387APending Publication Date: 2026-04-09SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing methods for generating synthetic face images lack diversity, consistency, identifiability, inter-class diversity, and fairness, leading to biased and inefficient training of face recognition models, and raise privacy and ethical concerns.

Method used

A two-stage generative model approach is employed to generate composite face images, ensuring identifiability, diversity, and fairness by conditioning on attributes like gender, race, and age, while maintaining high quality and reducing privacy risks.

Benefits of technology

The solution generates a wide variety of synthetic face images that improve the performance and fairness of face recognition models, reduce privacy issues, and address ethical concerns by controlling demographic distributions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026061387000001_ABST
    Figure 2026061387000001_ABST
Patent Text Reader

Abstract

Easily obtain a wide variety of images. [Solution] The generation unit generates a first-stage composite image by providing a first generation model that generates a composite image with an identifiability attribute, which is an attribute that affects the identifiability of individuals in the image, as a condition. The second generation model that generates a composite image generates a second-stage composite image by providing the feature quantities of the first-stage composite image and a diversity attribute, which is an attribute that affects the diversity of the same individual, as conditions. This technology can be applied, for example, to generating images such as facial images with diverse variations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to an information processing system, an information processing method, and a program, and particularly relates to an information processing system, an information processing method, and a program that enable, for example, easy acquisition of various variations of images.

Background Art

[0002] Development of a learning model as a fair and accurate recognition (identification) model requires a large amount of carefully collected learning data. For example, for learning a face recognition model, which is a recognition model for face recognition, for each of a large number of IDs (identities) for identifying an individual (person), various variations of face images in which the individual of that ID (the individual identified by that ID) appears are required. A fair and accurate recognition model is, for example, a recognition model in which there is no bias in the correct answer rate (recognition rate) of the recognition result due to gender, race, etc., and a high correct answer rate can be obtained.

[0003] Also, for example, for learning a fair and accurate face recognition model, as various variations of face images, demographically balanced (fair) face images, for example, uniform and large-scale face images without bias in attributes such as gender, race, age, etc. are required. Furthermore, it is desirable that various variations of face images are high-quality images.

[0004] However, collecting real face images as various variations of face images as described above is mainly difficult from the perspective of privacy. A real face image is a face image in which a real person appears. Real people include not only people who currently exist but also people who existed in the past.

[0005] Therefore, in order to suppress the risk of privacy and bias in attributes, there is an increasing demand for generating synthetic images, particularly synthetic face images, which are images generated by information processing.

[0006] Various methods for generating composite facial images have been proposed using learning models as generative models for generating composite images.

[0007] Currently, synthetic face images generated using generative models are not always appropriate as images representing a diverse range of face variations. Therefore, there can be a significant difference in performance between face recognition models trained using synthetic face images (or datasets) and face recognition models trained using real face images (or datasets).

[0008] Furthermore, it has been proposed to transform an image so that it has characteristics equivalent to the image to be recognized by the recognition model, and then use the transformed image as training data to train the recognition model (see, for example, Patent Document 1). [Prior art documents] [Patent Documents]

[0009] [Patent Document 1] Japanese Patent Publication No. 2021-082068 [Overview of the project] [Problems that the invention aims to solve]

[0010] In recent years, there has been a demand for technologies that can easily obtain a wide variety of images.

[0011] This technology was developed in light of these circumstances, and aims to make it easy to obtain a wide variety of images. [Means for solving the problem]

[0012] The information processing system or program of this technology includes a generation unit that generates a first-stage composite image by providing a first generation model for generating composite images with an identifiability attribute, which is an attribute that affects the identifiability of individuals depicted in an image, as a condition, and generates a second-stage composite image by providing a second generation model for generating composite images with the feature quantities of the first-stage composite image and a diversity attribute, which is an attribute that affects the diversity of identical individuals, as conditions, or a program for causing a computer to function as such an information processing system.

[0013] The information processing method of this technology includes generating a first-stage composite image by providing a first generative model that generates a composite image with an identifiability attribute, which is an attribute that affects the identifiability of individuals depicted in an image, as a condition, and generating a second-stage composite image by providing a second generative model that generates a composite image with the feature quantities of the first-stage composite image and a diversity attribute, which is an attribute that affects the diversity of identical individuals, as conditions.

[0014] In this technology, a first-stage composite image is generated by providing a first generative model that generates composite images with an identifiability attribute, which is an attribute that affects the identifiability of individuals in an image, as a condition. A second-stage composite image is generated by providing a second generative model that generates composite images with the features of the first-stage composite image and a diversity attribute, which is an attribute that affects the diversity of identical individuals, as conditions.

[0015] An information processing system may be an independent device or an internal block constituting a single device. Furthermore, one or more blocks constituting an information processing system may be configured as separate devices.

[0016] The program is provided by transmitting it via a transmission medium or by recording it on a recording medium. It can be provided.

Brief Description of the Drawings

[0017] [Figure 1] It is a block diagram showing a configuration example of an embodiment of a face processing system to which this technology is applied. [Figure 2] It is a diagram for explaining the first problem when generating a synthetic face image using a generation model. [Figure 3] It is a diagram for explaining the second problem when generating a synthetic face image using a generation model. [Figure 4] It is a diagram for explaining the third problem when generating a synthetic face image using a generation model. [Figure 5] It is a diagram for explaining the fourth problem when generating a synthetic face image using a generation model. [Figure 6] It is a diagram for explaining the fifth problem when generating a synthetic face image using a generation model. [Figure 7] It is a diagram for explaining the influence of the bias of the face images used for learning the face recognition model that appears in the face recognition model. [Figure 8] It is a diagram for explaining a face recognition model that has been learned using the synthetic face image generated by the face image generation unit 21. [Figure 9] It is a diagram for explaining an example of an application using the synthetic face image generated by the face image generation unit 21. [Figure 10] It is a diagram for explaining another example of an application using the synthetic face image generated by the face image generation unit 21. [Figure 11] It is a block diagram showing a configuration example of an information processing system as the face image generation unit 21. [Figure 12] It is a diagram schematically showing the face images stored in the face image DB 41. [Figure 13] It is a block diagram showing a configuration example of the learning unit 43. [Figure 14] It is a block diagram showing a configuration example of the preprocessing unit 61. [Figure 15]This block diagram shows an example configuration of the part of the attribute recognition unit 73 that recognizes the similarity of the diversity attributes. [Figure 16] This is a block diagram showing an example of the configuration of the generation unit 51. [Figure 17] This block diagram shows an example configuration of the first stage processing unit 91 and the second stage processing unit 92. [Figure 18] This is a block diagram showing an example of the configuration of the filter unit 113. [Figure 19] This is a block diagram showing an example of the configuration of the filter unit 123. [Figure 20] This is a flowchart illustrating the processing of the face image generation unit 21. [Figure 21] This is a flowchart illustrating the processing of the first stage in step S13. [Figure 22] This is a flowchart illustrating the processing of the second stage, step S14. [Figure 23] This is a block diagram showing an example configuration of one embodiment of a computer to which this technology is applied. [Modes for carrying out the invention]

[0018] <Facial processing system applying this technology>

[0019] Figure 1 is a block diagram showing an example configuration of one embodiment of a facial processing system to which this technology is applied.

[0020] In Figure 1, the face processing system 10 includes a learning device 20 and a face processing device 30. The learning device 20 performs learning for face processing, that is, learning a (machine) learning model for face processing. The face processing device 30 uses the learned learning model for face processing to perform face processing on face images containing faces, such as face recognition and face image correction. Face image correction, for example, corrects a face image in which a face is turned away from the front to a face image in which a face is facing the front.

[0021] The learning device 20 includes a face image generation unit 21, a face image database 22, a learning face image selection unit 23, and a face image learning unit 24.

[0022] The face image generation unit 21 generates composite face images (data) that include the faces of various individuals. For example, the face image generation unit 21 generates realistic, fair, diverse, and identifiable (individualistic) composite face images through an automated pipeline using an AI (artificial intelligence) model as a learning model. Identifiableness of a face image means that the face image (of the person depicted) can be identified (or recognized) by a specific (ID) individual. Maintaining identifiableness means that a face image that should be identified by a specific individual is identified by that specific individual. The face image generation unit 21 supplies each composite face image to the face image DB 22 in a form associated with an ID (label) that identifies the individual person depicted in the face image.

[0023] In this embodiment, a human individual is used as the individual in the face image, but other individuals, such as dogs or cats, can also be used. Furthermore, although this technology will be explained below using the generation of a (synthesized) face image as an example, this technology can be applied to the generation of full-body images, for example, images showing the entire body. Full-body images can be used, for example, to train learning models used in tasks such as person (human) re-identification.

[0024] The face image DB22 stores the synthesized face images from the face image generation unit 21.

[0025] The training face image selection unit 23 selects composite face images from the composite face images stored in the face image DB 22 to be used as training data for the face processing learning model, and supplies them to the face image learning unit 24. For example, depending on user operations, the training face image selection unit 23 selects composite face images containing individuals with specific attributes, such as individuals of a specific gender, age group, or race, from the composite face images stored in the face image DB 22, and supplies them to the face image learning unit 24. The face image generation unit 21 can select all composite face images stored in the face image DB 22 as training face images.

[0026] The face image learning unit 24 uses the face images to be selected from the face image selection unit 23 to train a learning model for face processing (including updating and fine-tuning), and then supplies the trained model (the trained model) to the face processing device 30.

[0027] The face processing device 30 includes a face detection unit 31 and a face image processing unit 32.

[0028] The face detection unit 31 is supplied with a face image to be processed. The face detection unit 31 detects faces in the face image supplied thereto, identifies a predetermined area containing the face, and supplies it to the face image processing unit 32.

[0029] The face image processing unit 32 is supplied with a trained model that has been provided to the face processing device 30 from the face image learning unit 24. The face image processing unit 32 performs face processing on the face image (faces captured in the image) in which the region containing a face has been identified by the face detection unit 31, using the trained model from the face image learning unit 24, and outputs the result of the face processing. For example, if the trained model for face processing is a face recognition model, the face image processing unit 32 uses the face recognition model to perform face recognition, calculating the ID of the individual whose face is captured in the face image from the face detection unit 31, and outputs the ID obtained by face recognition as the result of face recognition. Alternatively, if the trained model for face processing is a face correction model, the face image processing unit 32 uses the face correction model to generate a face image in which the face captured in the face image from the face detection unit 31 is facing a predetermined direction, such as the front, and outputs that face image as the result of face correction.

[0030] <Problems with generating composite facial images>

[0031] Figure 2 illustrates the first problem when generating a synthetic face image using a generative model.

[0032] The first problem concerns the maintenance of identifiability (identity preservation), specifically the fact that composite facial images that do not maintain identifiability may be generated.

[0033] For example, when using a generative model to generate other face images identifiable by a specific individual based on a face image im21 containing the face of that specific individual, in addition to generating a composite face image im22 identifiable by the specific individual, a composite face image im23 identifiable by an individual other than the specific individual may also be generated. In other words, a composite face image im23 may be generated in which the facial features of the specific individual captured in face image im21 are not preserved.

[0034] The face image generation unit 21 generates a composite face image by providing the generation model with the feature quantities of a face image containing the face of a specific individual (which can also be considered as the feature quantities of the specific individual) as conditions. This suppresses the generation of composite face images that can be identified by individuals other than the specific individual, while promoting the generation of composite face images that can be identified by the specific individual. As a result, a composite face image is generated that maintains the identifiability of the specific individual. Therefore, the first problem can be resolved.

[0035] Here, as features of the face image, any value can be used from which similar values ​​(values ​​that can be considered identical) are extracted from various face images of the same person. For example, the embedding (vector) extracted from the face image in the encoder that constitutes the face recognition model, or the ID output by the face recognition model as a result of face recognition of the face image, can be used as features of the face image. Also, for example, the latent variables extracted from the face image in the encoder of a VAE (Variational AutoEncoder) can be used as features of the face image. In the following, we will use embedding as the features of the face image. As the ID of the face image (the person in it), a value corresponding to the embedding (for example, the vector quantized value of the vector as embedding) can be used.

[0036] Figure 3 illustrates the second problem when generating synthetic facial images using a generative model.

[0037] The second problem concerns image degradation (lack of consistency), specifically the fact that images that are distorted as facial images can sometimes be generated.

[0038] For example, when a generative model is used to generate a synthetic face image im31, the resulting image may contain significant artifacts, or it may contain structures that are difficult to identify as a face.

[0039] The face image generation unit 21 generates a high-quality image as a composite face image by performing a filtering process, and also generates a composite face image (a composite face image that captures a structure that can be recognized as a face) in which the gender and race can be recognized without contradiction (without error) (correctly). Therefore, the second problem can be resolved.

[0040] Furthermore, the filtering process guarantees that, for the generative model described in the first problem (Figure 2), there is only one unique face image for each ID, given the embedding as a feature condition.

[0041] Figure 4 illustrates the third problem when generating synthetic facial images using a generative model.

[0042] The third problem concerns intra-class diversity, specifically the tendency for images to be generated with insufficient diversity as variations of facial images of the same person.

[0043] For example, when using a generative model to generate other face images that can be identified as a specific individual based on a face image im41 showing the face of that specific individual, there is a tendency to generate composite face images im42, im43, and im44 that can be identified as the specific individual, but with little change in attributes such as age, facial posture (face orientation), and facial expression, resulting in a lack of variation.

[0044] The face image generation unit 21 generates a composite face image by providing the generation model with embeddings and attributes (values) of the individuals to be depicted in the composite face image to be generated, such as age, facial posture, facial expression, and similarity, as conditions. As a result, the face image generation unit 21 obtains the embeddings (corresponding IDs) given as conditions and generates a composite face image of an individual who possesses the attributes such as age and similarity given as conditions. Consequently, variations of composite face images of the same person with diversity in attributes (values) such as age, facial posture, facial expression, and similarity are generated. Therefore, the third problem can be resolved.

[0045] Here, similarity refers to the similarity between the individual's face in the composite face image and a representative face that represents that individual. For example, if there are multiple face images of the individual that differ in age, facial posture, expression, lighting conditions, and whether or not they are wearing a mask or glasses, a face that is an average of the faces in those multiple face images can be adopted as the individual's representative face. Alternatively, for example, the face in one face image selected from multiple face images can be adopted as the individual's representative face.

[0046] For example, if a higher similarity score indicates greater resemblance to the representative face, then the higher the similarity score given to the generative model, the more similar the generated composite face image will be to the representative face. For instance, the composite face image will contain faces that are similar to the representative face in terms of age, facial posture, and facial expression. Therefore, the similarity score can influence the age, facial posture, and facial expression of the individuals depicted in the composite face image generated by the generative model.

[0047] Figure 5 illustrates the fourth problem when generating synthetic facial images using a generative model.

[0048] The fourth problem concerns inter-class diversity, specifically the difficulty in generating face images that depict faces of completely different people, beyond the range of faces captured in the face images used to train the generative model.

[0049] For example, when generating face images using a generative model trained with face images im51 and im52, similar face images im53 and im54 will be generated, but face images of completely different people will not be generated. For example, if a generative model trained using only face images of a specific race is used, face images of races other than that specific race will not be generated. Therefore, there is a tendency for images to be generated that lack sufficient diversity as variations of face images of different people (different individuals).

[0050] The face image generation unit 21 generates a composite face image for which embeddings are given as conditions to the generation model, by providing attributes (values) such as gender and race as conditions to other generation models. This generates a composite face image in which individuals with different genders, races, etc., i.e., individuals with different IDs (different people), are depicted. As a result, a composite image with human diversity is generated, i.e., a composite face image in which various different people are depicted. Therefore, the fourth problem can be resolved.

[0051] Figure 6 illustrates the fifth problem when generating a synthetic face image using a generative model.

[0052] The fifth problem concerns the fairness of facial images, specifically the bias in the attributes of the facial images used to train the generative model, such as protected characteristics like gender and race, which is reflected (amplified) in the synthetic facial images generated by that model.

[0053] Real-world facial images (or datasets of real-world facial images) may have biases in protected characteristics such as gender, race, and age. For example, if a generative model is trained using real-world facial images that are biased in gender or race, the synthetic facial images (or datasets of real-world facial images) generated using that trained generative model will have the same biases in gender and race as the real-world facial images used for training. Therefore, it is not possible to obtain fair synthetic facial images, that is, synthetic facial images that are unbiased (even) in attributes such as gender and race.

[0054] As explained in the fourth problem (Figure 5), the face image generation unit 21 generates a composite face image for which embeddings are given as conditions to the generation model by giving attributes (values) such as gender and race as conditions to other generation models. In doing so, the face image generation unit 21 provides gender and race, etc., evenly and uniformly to the other generation models as conditions. As a result, a fair composite face image is generated, that is, a composite face image that is evenly distributed without bias in gender, race, etc. (or a nearly uniform composite face image with suppressed bias). Therefore, the fifth problem can be resolved.

[0055] Furthermore, even with learning models other than generative models, if biased face images are used for training, the resulting trained model will be affected by the bias of the face images used for training.

[0056] Figure 7 illustrates the effect of bias in the face images used to train a face recognition model.

[0057] Figure 7 shows a case where a face recognition model was deployed that was trained using real face images with a bias, specifically with nearly twice as many male faces as female faces.

[0058] The accuracy of face recognition by gender in face recognition models is affected by the gender bias inherent in the real-world face images used to train the model. The accuracy of male face recognition is approximately twice that of female face recognition. Therefore, face recognition models are not fair, exhibiting a bias in accuracy between genders; that is, the accuracy of female face recognition is roughly half that of male face recognition.

[0059] Furthermore, if real face images are used to train a face recognition model, the issue of sensitive data (data that must be protected from unauthorized access) may arise. That is, for example, if real face images are used to train a face recognition model, adversarial attacks on the face recognition model could detect the real face images used as training data. Therefore, for example, if the real face images used for training are images collected from the internet or elsewhere without consent, privacy issues and ethical concerns may arise.

[0060] Figure 8 illustrates a face recognition model trained using a composite face image generated by the face image generation unit 21.

[0061] As explained in Figure 6, the face image generation unit 21 generates unbiased composite face images, that is, composite face images that are evenly distributed without bias in terms of gender, race, etc. Therefore, the composite face images include roughly equal numbers of males and females. When a face recognition model is trained using such unbiased composite face images as training data, the face recognition model reflects the unbiasedness of the composite face images used as training data, resulting in an unbiased face recognition model with no gender bias in accuracy. In other words, the face recognition model becomes an unbiased face recognition model in which the accuracy of male face recognition is roughly equal to that of female face recognition. As a result, the performance of the face recognition model can be improved compared to when the face recognition model is trained using biased face images as explained in Figure 7.

[0062] Furthermore, when synthetic facial images are used to train a facial recognition model, the sensitive data issue does not arise. That is, synthetic facial images are non-sensitive data, and only non-sensitive synthetic facial images can be detected (reproduced) by adversarial attacks on the facial recognition model. Therefore, the privacy issues and ethical concerns that arise when real facial images are used to train a facial recognition model can be suppressed.

[0063] The face image generation unit 21 generates a composite face image through an automated pipeline using an AI model as a learning model. Specifically, the face image generation unit 21 generates a composite face image (second-stage composite image) by providing the generation model (second generation model) with embeddings and attributes such as the age and similarity of the individuals to be depicted in the composite face image to be generated as conditions. The face image generation unit 21 also generates a composite face image (first-stage composite image) from which embeddings to be provided as conditions to the generation model are obtained by providing other generation models (first generation model) with attributes such as gender and race as conditions.

[0064] Therefore, the face image generation unit 21 can generate an unlimited number of new synthetic face images (an enormous number that can be considered unlimited) by changing the information given as conditions to the generation model and other generation models.

[0065] Furthermore, the face image generation unit 21 generates high-quality composite face images that correctly recognize gender and race through filtering, and guarantees that there is only one composite face image for each person (one ID) for which embedding is given as a condition in the generation model.

[0066] Furthermore, according to the face image generation unit 21, embedding is given as a condition to the generation model, and a composite face image is generated. This makes it possible to generate a composite face image that maintains the identifiability to be identified (recognized) by the individual with the ID corresponding to the embedding.

[0067] Furthermore, according to the face image generation unit 21, a composite face image is generated by providing attributes (values) such as gender and race as conditions to other generative models, thereby obtaining the embeddings that are given as conditions to the generative models. Therefore, the distribution of demographic characteristics such as gender and race of individuals depicted in the composite face image can be controlled by the gender and race given as conditions to other generative models. Moreover, a composite face image is generated by providing the embeddings of the composite face image generated by other generative models as conditions to the generative model. Therefore, the composite face image generated by the generative model is an image with reduced privacy risks compared to a real face image.

[0068] Furthermore, according to the face image generation unit 21, the number of composite face images generated by the generation model can be controlled for each ID corresponding to one embedding given to the generation model as a condition, for example, by age or similarity given to the generation model along with the embedding as a condition. This also allows the total number of composite face images generated by the face image generation unit 21 to be controlled.

[0069] Here, the collection of real facial images and the use of facial recognition AI (facial recognition models) based on these images may raise ethical, legal, and social issues.

[0070] The collection of real facial images and the use of facial recognition AI may infringe upon the privacy of individuals (subjects) depicted in those images, depending on the method of collection and the form of use of the facial recognition AI. For example, facial recognition AI used by government agencies and police for law enforcement purposes may lead to human rights violations. Facial recognition AI can be used, for example, in criminal investigations to identify individuals from facial images acquired by security cameras, or for airport censorship and the search for terrorists in public spaces. The use of facial recognition AI by government agencies and police may lead to excessive surveillance of people and undermine respect for people's fundamental human rights. The limitations of facial recognition AI's performance may lead to wrongful arrests. Discrimination may arise due to differences in the recognition accuracy of facial recognition AI based on sensitive attributes such as race and gender.

[0071] US-based Clearview AI has built a database of over 20 billion facial images sourced solely from public sources such as social media, websites, news media, and arrest photos. The company develops and provides facial recognition AI using these images without the subjects' consent. This facial recognition AI is used by government agencies and police forces worldwide and has sparked controversy. In the UK, Clearview AI was ordered to pay a fine of $9.4 million (approximately 1.4 billion yen) and to delete facial images of UK residents and cease operations, after being found to have violated several UK laws by collecting and retaining facial images without individual consent. Disputes against Clearview AI have been filed not only in the UK, but also in California and Illinois in the US, Canada, France, Italy, Sweden, and Australia. Many disputes have already erupted due to the GDPR (General Data Protection Regulation) in Europe and local personal data protection laws.

[0072] Clearview AI's facial recognition technology is being used for military purposes. It is said that the database for this facial recognition technology contains over 2 billion photos collected from the Russian social networking service VKontakte.

[0073] Against this backdrop, legal regulations regarding the collection of real facial images and facial recognition AI have been further strengthened. On March 13, 2024, the European Parliament adopted the European AI Act, a comprehensive AI regulation. Article 5(db) of the European AI Act stipulates that "introducing to the market, commencing use for this particular purpose, or using an AI system that scrapes facial images without limit from the Internet or CCTV footage to create or expand a facial recognition database" is prohibited, thus banning AI systems that collect data by crawling in the creation of datasets for facial recognition applications as "unacceptable AI."

[0074] The generation of composite facial images by the facial image generation unit 21 can be useful in addressing the ethical, legal, and social issues mentioned above.

[0075] <Applications that use synthesized facial images>

[0076] Figure 9 illustrates an example of an application that uses the composite face image generated by the face image generation unit 21.

[0077] The face image generation unit 21 generates a composite face image by providing the generation model with embeddings and attributes such as the age, facial posture, expression, and similarity of the person to be captured in the composite face image to be generated, as conditions. As a result, the face image generation unit 21 generates a composite face image for the person with the ID corresponding to the embedding given as a condition, showing the face with the facial posture and age given as conditions. Therefore, by controlling the facial posture, age, etc., given as conditions to the generation model, for example, by providing the generation model with various facial postures and ages as conditions, it is possible to generate variations of composite face images of the same person with different facial postures or variations of facial images with different ages, as shown in Figure 9.

[0078] The composite face images generated by the face image generation unit 21, as described above, can be used, for example, to train a learning model that performs face recognition when tracking the faces of the same person. Furthermore, the composite face images generated by the face image generation unit 21 can be used to train a learning model that generates passport photos by, for example, generating a face image showing a face facing forward from a face image showing a tilted face. In addition, the composite face images generated by the face image generation unit 21 can be used to train a learning model that performs face recognition on passport photos for verification purposes.

[0079] Figure 10 illustrates another example of an application that uses the composite face image generated by the face image generation unit 21.

[0080] For example, still images and videos of crowds may include many people who have not given their permission, as shown in Figure 10. To protect the privacy of the many people who appear in the still images and videos, the faces of those people can be replaced with faces from composite face images generated by the face image generation unit 21.

[0081] By replacing the faces of people in still images and videos of crowds with faces in composite face images generated by the face image generation unit 21, the distribution of demographic characteristics of people in still images and videos can be controlled.

[0082] For example, as explained in Figure 9, the face image generation unit 21 can generate variations of face images of different ages as a composite face image. Furthermore, as explained in the fourth problem (Figure 5), the face image generation unit 21 can generate composite face images containing individuals of different genders, races, etc., by providing attributes (values) such as gender and race as conditions to other generation models to obtain embeddings that are given as conditions to the generation model. As described above, the face image generation unit 21 can generate composite face images with controlled distribution of demographic characteristics such as gender, race, and age. Therefore, by replacing the faces of people in still images or videos with faces in composite face images generated by the face image generation unit 21, the distribution of demographic characteristics of people in still images or videos can be controlled. For example, it is possible to limit the people in still images or videos to only those of a specific race, such as Caucasians, or only those of a specific age group, such as the elderly. Also, for example, still images or videos can be replaced with images containing diverse races with roughly equal male-female ratios.

[0083] <Example of the configuration of the face image generation unit 21>

[0084] Figure 11 is a block diagram showing an example configuration of the information processing system as the face image generation unit 21 in Figure 1.

[0085] In Figure 11, the face image generation unit 21 includes a learning device 40 and a generation device 50.

[0086] The learning device 40 includes a face image database 41, an acquisition unit 42, and a learning unit 43, and performs training on a learning model for generating face images. The learning device 40 supplies the trained learning model to the generation device 50.

[0087] The face image database 41 stores face images used to train the learning model. Each face image is associated with an ID as a label that identifies the individual whose face is in the image. For example, the face image database 41 stores either one face image or multiple different face images associated with each of several IDs. Multiple different face images associated with the same ID are images of the same individual, but differ in facial posture, age, expression, lighting conditions, presence or absence of accessories such as masks or glasses, etc. The face images stored in the face image database 41 may be real face images, composite face images, or both.

[0088] The acquisition unit 42 acquires face images from the face image DB 41 as training data to be used for training the learning model in the learning unit 43, and supplies them to the learning unit 43.

[0089] The learning unit 43 uses face images (predetermined face images) as training data from the acquisition unit 42 to train a learning model, and supplies the trained learning model to the generation unit 51 of the generation device 50. As the learning model, for example, a learning model consisting of a diffusion model, VAE, GAN (Generative Adversarial Network), or transformer can be employed. For training the learning model, for example, deep learning can be employed.

[0090] The generation device 50 has a generation unit 51. The generation unit 51 generates a composite face image using a learned model from the learning unit 43.

[0091] Figure 12 is a schematic diagram showing the facial images stored in the facial image DB 41.

[0092] The face image DB41 can store a dataset of 2D (dimensional) RGB (red, green, blue) images of real faces for individuals with multiple IDs, differing in facial posture, age, facial expression, lighting conditions, and the presence or absence of accessories such as masks or glasses, as shown in Figure 12. In Figure 12, the horizontal axis represents differences in facial posture, age, facial expression, lighting conditions, and the presence or absence of accessories such as masks or glasses, while the vertical axis represents the ID. Therefore, in Figure 12, each row contains real faces of the same individual with the same ID (the same person), but differing in facial posture, age, facial expression, lighting conditions, and the presence or absence of accessories such as masks or glasses.

[0093] For example, the WebFace260M dataset can be used as the dataset of real face images stored in the face image DB41.

[0094] <Example of the configuration of Learning Section 43>

[0095] Figure 13 is a block diagram showing an example configuration of the learning unit 43 in Figure 11.

[0096] In Figure 13, the learning unit 43 includes a preprocessing unit 61, and learning units 62 and 63.

[0097] The preprocessing unit 61 is supplied with actual face images, each associated with an ID, which are stored in the face image DB 41, from the acquisition unit 42 (Figure 11).

[0098] The preprocessing unit 61 performs preprocessing.

[0099] For example, the preprocessing unit 61, as a preprocessing step, uses the actual face images to which the ID supplied there is associated to train a face recognition model, which is one of the learning models supplied from the learning unit 43 (Figure 11) to the generation unit 51. Here, the (actual) face images stored in the face image DB 41 that are used to train the learning model in the learning unit 43 are also called original face images.

[0100] As a face recognition model, a learning model, such as an encoder, can be used to output features of face images, such as embeddings or IDs. Based on the embeddings or IDs of face images output by the encoder as a face recognition model, it is possible to determine whether the individuals in two face images are the same person. If the vectors of the embeddings of the two face images are similar enough to be considered identical, or if the IDs of the two face images are identical, it can be determined that the individuals in the two face images are the same person. The ID of the face image (and the individual in it) is identified by the embeddings of face images output by the encoder as a face recognition model.

[0101] The preprocessing unit 61 generates embeddings for each original face image by recognizing the original face image using an encoder that is a trained face recognition model as a preprocessing step. Furthermore, the preprocessing unit 61 recognizes the identifiability attribute and diversity attribute (attribute values) of individuals depicted in each original face image using a trained model that recognizes the identifiability attribute and diversity attribute of individuals, respectively, as a preprocessing step.

[0102] Distinguishing attributes are attributes that affect the identifiability of an individual, such as race and gender. Diversity attributes are attributes that affect the diversity of the same individual, such as age, facial posture, and similarity. Pre-trained models that recognize both distinguishing attributes and diversity attributes are available.

[0103] The preprocessor 61 outputs an encoder as a face recognition model, along with each original face image, its embedding, identifiability attribute (labels representing the attribute values), and diversity attribute (labels representing the attribute values). The encoder as a face recognition model output by the preprocessor 61 is supplied to the generation unit 51 as one of the learning models supplied from the learning unit 43 (Figure 11) to the generation unit 51. The original face images output by the preprocessor 61 are supplied to the learning units 62 and 63. The identifiability attribute output by the preprocessor 61 is supplied to the learning unit 62. The embedding and diversity attribute output by the preprocessor 61 are supplied to the learning unit 63.

[0104] The learning unit 62 uses each original face image from the preprocessing unit 61 and the individual identifiability attributes depicted in the original face image to train one of the learning models, the identifiability attribute-conditional model. The identifiability attribute-conditional model is a generative model (first generative model) (other generative model) that generates a composite face image (first stage composite image) conditioned by the identifiability attribute given as a condition. For example, a diffusion model can be adopted as the identifiability attribute-conditional model.

[0105] The learning unit 62 trains the identifiability attribute-conditional model of the diffusion model, which is an identifiability attribute-conditional model of the identifiability attribute of an individual in the original face image, so that the original face image is generated when the identifiability attribute of the individual in the original face image is given as a condition. The learning unit 62 supplies the trained (post-training) identifiability attribute-conditional model to the generation unit 51 as one of the other trained models supplied to the generation unit 51 from the learning unit 43 (Figure 11).

[0106] The learning unit 63 uses each original face image from the preprocessing unit 61, the individual diversity attributes and embeddings depicted in the original face images, to train one of the learning models, the embedding / diversity attribute conditional model. The embedding / diversity attribute conditional model is a generative model (second generative model) that generates a composite face image (second stage composite image) conditioned by the diversity attributes and embeddings, given as conditions. For example, a diffusion model can be adopted as the embedding / diversity attribute conditional model.

[0107] The learning unit 63 trains the embedding / diversity attribute conditional model of the diffusion model as an embedding / diversity attribute conditional model so that when the diversity attributes and embedding of individuals in the original face image are given as conditions, the original face image is generated. The learning unit 62 supplies the trained embedding / diversity attribute conditional model to the generation unit 51 as yet another learning model supplied to the generation unit 51 from the learning unit 43 (Figure 11).

[0108] <Example of configuration of pre-processing unit 61>

[0109] Figure 14 is a block diagram showing an example configuration of the pre-processing unit 61 shown in Figure 13.

[0110] The preprocessing unit 61 includes a face recognition model learning unit 71, a face recognition unit 72, and an attribute recognition unit 73.

[0111] Each of the face recognition model learning unit 71 to attribute recognition unit 73 is supplied with a real face image to which an ID has been associated, which is supplied to the preprocessing unit 61.

[0112] The face recognition model learning unit 71 uses the ID and real face image supplied thereto to train the encoder as a face recognition model. For example, the face recognition model learning unit 71 trains the encoder as a face recognition model so that when a real face image is input, an embedding corresponding to the ID of the real face image is output. The face recognition model learning unit 71 supplies the trained encoder as a face recognition model to the face recognition unit 72, and also supplies it to the generation unit 51 as one of the learning models supplied from the learning unit 43 (Figure 11) to the generation unit 51.

[0113] The face recognition unit 72 generates embeddings for each original face image by recognizing the original face image using an encoder, which is a face recognition model, from the face recognition model learning unit 71. The face recognition unit 72 supplies the embeddings for each original face image to the attribute recognition unit 73. Furthermore, the face recognition unit 72 supplies each original face image to the learning units 62 and 63 (Figure 13), and also supplies the embeddings for each original face image to the learning unit 63.

[0114] The attribute recognition unit 73 recognizes the individual's identifiable attributes (or their attribute values) in each original face image using various pre-prepared learning models that recognize individual identifiable attributes such as race and gender, and supplies them to the learning unit 62 (Figure 13).

[0115] The attribute recognition unit 73 recognizes the attributes (and their attribute values) of the individual's diversity attributes in each original face image, excluding similarity, using various pre-prepared learning models that recognize diversity attributes such as an individual's age and facial posture, and supplies them to the learning unit 63 (Figure 13). The attribute recognition unit 73 also recognizes (calculates) the similarity of the diversity attributes for each original face image with the same ID, using embedding from the face recognition unit 72, and supplies it to the learning unit 63.

[0116] Furthermore, to recognize demographic information such as the race, gender, and age of individuals in the original facial image, learning models such as CLIP can be used. In addition, as a facial pose (attribute value) among the diversity attributes, for example, the angles of tilt in the roll, yaw, and pitch directions of the face, based on the state of a face facing forward, can be adopted.

[0117] <Example configuration of attribute recognition unit 73>

[0118] Figure 15 is a block diagram showing an example configuration of the part of the attribute recognition unit 73 in Figure 14 that recognizes the similarity among the diversity attributes.

[0119] In Figure 15, the attribute recognition unit 73 includes an embedding acquisition unit 81, a representative embedding calculation unit 82, and a similarity calculation unit 83.

[0120] The embedding acquisition unit 81 acquires the embeddings emb[1], emb[2],... emb[N] of the original face images with the same ID from the embeddings from the face recognition unit 72 (Figure 14), and supplies them to the representative embedding calculation unit 82 and the similarity calculation unit 83.

[0121] The representative embedding calculation unit 82 uses the embeddings emb[1] to emb[N] of the original face images with the same ID from the embedding acquisition unit 81 to calculate the embedding of a representative face that represents the faces in the original face images with the same ID, and supplies it to the similarity calculation unit 83.

[0122] For example, the representative embedding calculation unit 82 uses the average value of the embeddings emb[1] to emb[N] of original face images with the same ID, or the embedding of one original face image randomly selected from the original face images with the same ID, as the embedding of the representative face. As the average value of the embeddings emb[1] to emb[N] of original face images with the same ID, for example, the mean vector emb{mean} = 1 / N * (emb[1] + emb[2] + ... emb[N]) as the mean of the embeddings emb[1] to emb[N] can be normalized by the norm |emb{mean}| of that mean vector emb{mean} to get emb{mean} / |emb{mean}|.

[0123] If the embedding of the representative face is the average of the embeddings emb[1]~emb[N] of the original face images with the same ID, the representative face will be a face that is an average of the faces captured in each of the original face images with the same ID. If the embedding of the representative face is the embedding of one original face image randomly selected from the original face images with the same ID, the representative face will be the original face image from which the randomly selected embedding was obtained.

[0124] The similarity calculation unit 83 calculates the similarity of each original face image with the same ID to the representative face of that original face image with the same ID, and outputs it as one of the diversity attributes of the original face image (and the person depicted in it). The similarity calculation unit 83 calculates the similarity using the embedding of the original face image with the same ID from the embedding acquisition unit 81 and the embedding of the representative face from the representative embedding calculation unit 82. For example, for each original face image with the same ID, the similarity calculation unit 83 calculates the similarity of the original face image to the representative face as the cosine similarity between the vector as the embedding of the original face image and the embedding of the representative face.

[0125] In the learning unit 63 (Figure 13), diversity attributes, including similarity, are given as conditions for the embedding / diversity attribute conditional model, and the embedding / diversity attribute conditional model is trained. In the generation of a composite face image using such an embedding / diversity attribute conditional model, diversity attributes, including similarity, and embeddings are given as conditions for the embedding / diversity attribute conditional model, and a composite face image conditioned by the similarity and embedding given as conditions is generated. That is, a composite face image is generated in which a face is identified (recognized) by the individual with the ID corresponding to the embedding given as a condition, and the face is similar to the representative face of that individual with that ID by the degree of similarity represented by the similarity given as a condition.

[0126] Therefore, according to an embedding / diversity attribute conditional model in which diversity attributes including similarity and embedding are given as conditions, it is possible to generate a composite face image in which the similarity to a representative face is controlled by the similarity given as a condition. Furthermore, it is possible to generate a composite face image with diverse similarities to a representative face.

[0127] <Example of the configuration of the generation unit 51>

[0128] Figure 16 is a block diagram showing an example configuration of the generation unit 51 in Figure 11.

[0129] In Figure 16, the generation unit 51 includes a first stage processing unit 91 and a second stage processing unit 92.

[0130] The first stage processing unit 91 is supplied with an encoder for the identifiability attribute conditional model and the face recognition model from the learning unit 43 (Figure 11). The second stage processing unit 92 is supplied with an embedding / diversity attribute conditional model from the learning unit 43.

[0131] The first stage processing unit 91 performs the first stage processing to generate an embedded composite face image as a composite image of the first stage, using the identifiable attribute conditional model and the encoder as a face recognition model from the learning unit 43.

[0132] Specifically, the first-stage processing unit 91 provides a identifiable attribute (or a predetermined attribute value thereof) as a condition to the identifiable attribute condition model, thereby generating a composite face image of an individual possessing the identifiable attribute given as a condition, as the first-stage composite image. Furthermore, the first-stage processing unit 91 uses an encoder as a face recognition model to generate embeddings as feature quantities for the composite face image, which is the first-stage composite image. The first-stage processing unit 91 supplies the embeddings for the first-stage composite image to the second-stage processing unit 93.

[0133] The second stage processing unit 92 performs the second stage processing to generate a composite face image as the second stage composite image, using the embedding / diversity attribute conditional model from the learning unit 43 and the embedding of the first stage composite image from the first stage processing unit 91.

[0134] In other words, the second stage processing unit 92 provides the embedding of the first stage composite image from the first stage processing unit 91 and the diversity attribute (a predetermined attribute value) as conditions to the embedding / diversity attribute conditional model, thereby generating a composite face image as the second stage composite image, which shows an individual whose ID corresponds to the embedding given as a condition and who possesses the diversity attribute given as a condition. The second stage processing unit 92 outputs the second stage composite image as the final composite face image generated by the generation unit 51.

[0135] The generation unit 51 can generate fair and diverse composite face images through a pipeline as the first stage processing by the first stage processing unit 91 and a pipeline as the second stage processing by the second stage processing unit 92.

[0136] In other words, in the first stage processing by the first stage processing unit 91, for example, attribute values ​​such as gender and race can be generated with equal proportions as identifiable attributes. For example, the gender attribute values ​​can be generated in equal proportions for males and females, and the race attribute values ​​can be generated in equal proportions for various races such as Caucasians and Black people, thereby generating all possible combinations of these gender and race attribute values. Then, by providing the combination of gender and race (attribute values) as a condition to the identifiable attribute conditional model, a composite face image conditioned by the gender and race given as a condition can be generated as the first stage composite image. In other words, for example, a composite face image in which individuals identified (recognized) by the gender and race given as a condition can be generated as the first stage composite image. Therefore, as the first stage composite image, a fair composite face image in which a diverse range of individuals with a demographic balance can be generated, for example, a composite face image with (almost) no bias in gender or race.

[0137] In the second stage processing by the second stage processing unit 92, for example, random values ​​or values ​​corresponding to user operations can be generated as attribute values ​​such as age, face posture, and similarity as diversity attributes. For example, as an attribute value for age, multiple values ​​randomly selected from a pre-set age range can be generated. Also, as an attribute value for face posture, multiple values ​​randomly selected from a pre-set range of face tilt angles for each of the roll, yaw, and pitch directions can be generated. Furthermore, for example, as an attribute value for similarity, multiple values ​​randomly selected from a pre-set range of similarity can be generated, thereby generating all possible combinations of attribute values ​​for age, face posture, and similarity. Then, by providing the embedding of the first stage composite image, as well as combinations of age, face posture, and similarity (attribute values), as conditions to the embedding of the first stage composite image, as well as age, face posture, and similarity given as conditions, a composite face image conditioned by the first stage composite image embedding, as well as age, face posture, and similarity, can be generated as the second stage composite image. In other words, for example, a composite face image can be generated as a second-stage composite image, which includes individuals identified (recognized) by an ID corresponding to an embedded image given as a condition, and whose age, facial posture, and similarity are also given as conditions. Therefore, as a second-stage composite image, a fair composite face image can be generated, similar to the first-stage composite image. Furthermore, as a second-stage composite image, a diverse composite face image can be generated, in which each individual identified by an ID corresponding to the embedded image of the first-stage composite image has a variety of ages, facial postures, and similarities.

[0138] Therefore, according to the generation unit 51, it is possible to easily obtain (face) images of diverse individuals with a balanced demographic, varying in age, facial posture, and similarity.

[0139] Figure 17 is a block diagram showing an example configuration of the first stage processing unit 91 and the second stage processing unit 92 in Figure 16.

[0140] The first stage processing unit 91 includes an identifiable attribute generation unit 111, a composite face image generation unit 112, a filter unit 113, and a face recognition unit 114. The first stage processing unit 91 can be configured, for example, without the filter unit 113.

[0141] The identifiable attribute generation unit 111 generates attribute values ​​for identifiable attributes such as race and gender, and supplies them to the composite face image generation unit 112. For example, the identifiable attribute generation unit 111 can generate uniform and consistent attribute values ​​for identifiable attributes such as race and gender. Uniform attribute values ​​mean, for example, attribute values ​​in which each possible value appears in the same proportion. For example, if the identifiable attribute is gender, an attribute value in which male and female, which can be possible values ​​for gender, appear in the same proportion is a uniform attribute value. For example, if the identifiable attribute is race, an attribute value in which various races, such as white and black, which can be possible values ​​for race, appear in the same proportion is a uniform attribute value.

[0142] Furthermore, the identifiable attribute generation unit 111 can generate identifiable attributes (or their attribute values) in other ways, such as in response to user operations.

[0143] The composite face image generation unit 112 is supplied with a identifiability attribute conditional model from the learning unit 43 (Figures 11 and 13). The composite face image generation unit 112 generates a composite face image by providing the identifiability attribute (or its attribute value) from the identifiability attribute generation unit 111 as a condition to the identifiability attribute conditional model. For example, the composite face image generation unit 112 generates a composite face image of an individual whose gender and race are given as identifiability attributes as a condition. The composite face image generation unit 112 supplies the composite face image to the filter unit 113. The composite face image supplied by the composite face image generation unit 112 to the filter unit 113 is also called the first composite face image. The first composite face image is also the first stage composite image.

[0144] The filter unit 113 performs a first filtering process on the first composite face image from the composite face image generation unit 112, and supplies the resulting composite face image to the face recognition unit 114. The composite face image supplied by the filter unit 113 to the face recognition unit 114 is also called the second composite face image. The second composite face image is also the composite image from the first stage.

[0145] The face recognition unit 114 is supplied with an encoder as a face recognition model from the learning unit 43 (Figures 11 and 13). The face recognition unit 114 recognizes the second synthesized face image from the filter unit 113, generates an embedding of the second synthesized face image, and supplies it to the synthesized face image generation unit 122 of the second stage processing unit 92.

[0146] The second stage processing unit 92 includes a diversity attribute generation unit 121, a composite face image generation unit 122, and a filter unit 123. The second stage processing unit 92 can be configured, for example, without the filter unit 123.

[0147] The diversity attribute generation unit 121 generates attribute values ​​for diversity attributes such as age, facial posture, and similarity, and supplies them to the composite face image generation unit 122. For example, the diversity attribute generation unit 121 can generate attribute values ​​for diversity attributes such as age, facial posture, and similarity by randomly sampling within a range of possible attribute values.

[0148] Furthermore, the diversity attribute generation unit 121 can generate diversity attributes (or their attribute values) in other ways, such as in response to user operations. For example, it can set the range for randomly sampling the attribute values ​​of the diversity attributes according to user operations, or it can generate values ​​corresponding to user operations as attribute values ​​of the diversity attributes.

[0149] The composite face image generation unit 122 is supplied with an embedding / diversity attribute conditional model from the learning unit 43 (Figures 11 and 13). The composite face image generation unit 122 generates a composite face image by providing the embedding of the second composite face image from the face recognition unit 114 and the diversity attributes (attribute values) from the diversity attribute generation unit 121 as conditions to the embedding / diversity attribute conditional model. For example, the composite face image generation unit 122 generates a composite face image of an individual whose ID corresponds to the embedding of the second composite face image given as a condition, and whose diversity attributes such as age, face posture, and similarity are given as conditions. The composite face image generation unit 122 supplies the composite face image to the filter unit 123. The composite face image supplied by the composite face image generation unit 122 to the filter unit 123 is also called the third composite face image. The third composite face image is also the second stage composite image.

[0150] The filter unit 113 performs a second filtering process on the third composite face image from the composite face image generation unit 122, and outputs the composite face image obtained as a result of the second filtering process as the final composite face image generated by the generation unit 51. The composite face image output by the filter unit 113 is also called the fourth composite face image. The fourth composite face image is also the composite image of the second stage.

[0151] Figure 18 is a block diagram showing an example configuration of the filter section 113 in Figure 17.

[0152] In Figure 18, the filter unit 113 includes an identifiable attribute recognition unit 141, an image selection unit 142, an image quality recognition unit 143, an image selection unit 144, a face recognition unit 145, an image selection unit 146, and a storage unit 147.

[0153] The Identifiable Attribute Recognition Unit 141 is supplied with a first composite face image from the composite face image generation unit 112 (Figure 17). The Identifiable Attribute Recognition Unit 141 uses various pre-prepared learning models that recognize identifiable attributes such as an individual's race and gender to recognize the identifiable attributes (or attribute values) of the individual depicted in each first composite face image from the composite face image generation unit 112. The Identifiable Attribute Recognition Unit 141 supplies the first composite face image, along with the recognition results of the identifiable attributes of the individual depicted in the first composite face image, to the image selection unit 142.

[0154] The image selection unit 142 selects a face image from the first composite face image from the identifiable attribute recognition unit 141 that has correctly recognized identifiable attributes, based on the recognition result of the identifiable attributes of the first composite face image from the identifiable attribute recognition unit 141, and supplies it to the image quality recognition unit 143. For example, the image selection unit 142 selects a first composite face image from the first composite face image from the identifiable attribute recognition unit 141 that matches the recognition result of the identifiable attributes of the identifiable attributes (or their attribute values) given as conditions to the identifiable attribute conditional model when the first composite face image is generated by the composite face image generation unit 112 (Figure 17), as a face image with correctly recognized identifiable attributes.

[0155] Therefore, the identifiable attribute recognition unit 141 and the image selection unit 142 perform a first filtering process of the filter unit 113, which removes face images other than those in which the identifiable attributes have been correctly recognized from the first composite face image, i.e., face images in which the identifiable attributes have been misrecognized. This ensures that the first composite face image is a face image in which the identifiable attributes such as gender and race are not contradictory, i.e., a face image in which the identifiable attributes are correctly identified (recognized).

[0156] Face images in which identifiable attributes such as gender and race are misrecognized may be images that are broken as face images, for example, images that show structures that are difficult to identify as faces. When embedding is generated from such a broken image and given as a condition to the embedding / diversity attribute conditional model, the image output by the embedding / diversity attribute conditional model (probabilistic generative model) may be an erroneous image, for example, an image that is broken as a face image. By selecting a first synthesized face image in which the identifiable attributes are correctly recognized in the image selection unit 142, it is possible to (significantly) suppress the output of the image that is an erroneous image by the embedding / diversity attribute conditional model.

[0157] The image quality recognition unit 143 uses a pre-prepared learning model, for example, to recognize the image quality of an image, to recognize the image quality of each first composite face image from the image selection unit 142, such as the signal-to-noise ratio (S / N) and resolution. The image quality recognition unit 143 supplies the first composite face image, along with the image quality (recognition result) of the first composite face image, to the image selection unit 144.

[0158] The image selection unit 144 selects a high-quality face image from the first composite face image from the image quality recognition unit 143 based on the image quality of the first composite face image from the image quality recognition unit 143 and supplies it to the face recognition unit 145. For example, the image selection unit 144 selects a first composite face image from the first composite face image from the image quality recognition unit 143 whose signal-to-noise ratio (S / N) is above a threshold as a high-quality face image.

[0159] Therefore, the image quality recognition unit 143 and the image selection unit 144 perform a first filtering process of the filter unit 113, which involves deleting face images other than high-quality face images, i.e., images with poor image quality, from the first composite face image. This ensures that the first composite face image is a high-quality face image.

[0160] In the image selection unit 144, the image quality of the first composite face image (after selection) can be managed by adjusting thresholds such as the signal-to-noise ratio (S / N) when selecting (deleting) the first composite face image.

[0161] The face recognition unit 145 is supplied with an encoder, for example, as the same face recognition model that is supplied from the learning unit 43 (Figure 13) to the face recognition unit 114 (Figure 17). The face recognition unit 145 uses the encoder, for example, as the face recognition model, to recognize each of the first composite face images from the image selection unit 144, thereby generating the embedding of the first composite face image. The face recognition unit 145 supplies the first composite face image, along with the embedding of the first composite face image, to the image selection unit 146.

[0162] The image selection unit 146 selects a face image from the first composite face image from the face recognition unit 145 that contains a different person from the person in the first composite face image stored in the storage unit 147, based on the embedding of the first composite face image from the face recognition unit 145, and supplies it to the storage unit 147. For example, the image selection unit 146 selects a first composite face image from the first composite face image from the face recognition unit 145 that has a cosine similarity with the embedding of each first composite face image stored in the storage unit 147 that is, for example, 0.3 or less, as a face image containing a different person.

[0163] Therefore, the face recognition unit 145 and the image selection unit 146 perform a first filtering process in the filter unit 113 to remove face images from the first composite face image that contain individuals identified as the same person as individuals appearing in other face images (the first composite face image stored in the storage unit 147). This ensures that the first composite face image supplied from the image selection unit 146 to the storage unit 147 contains only one face image with the same ID (there are no multiple face images with the same ID).

[0164] The memory unit 147 stores the first composite face image from the image selection unit 146 and supplies the stored first composite face image as the second composite face image to the face recognition unit 114 (Figure 17).

[0165] Therefore, just like the first composite face image supplied from the image selection unit 146 to the storage unit 147, it is guaranteed that the second composite face image will also consist of only one face image with the same ID. In other words, the second composite face image supplied from the filter unit 113 to the face recognition unit 114 will contain only one face image of a given individual, and there will not be multiple such images.

[0166] As described above, the composite face image generation unit 122 generates a composite face image of the individual whose ID corresponds to the embedding of the second composite face image as the third composite face image. Therefore, if there are multiple face images of individual A and only one face image of other individual B as the second composite face image, more third composite face images of individual A will be generated than third composite face images of individual B. Consequently, the final fourth composite face image will also have more fourth composite face images of individual A than fourth composite face images of individual B, resulting in an unfair outcome.

[0167] As mentioned above, by ensuring that there is only one face image with the same ID in the second composite face image, it is possible to suppress any lack of fairness.

[0168] In this case, the image selection unit 146 selects a first composite face image from the first composite face image from the face recognition unit 145, which has an embedding with a cosine similarity of 0.3 or less with the embedding of the first composite face image stored in the storage unit 147, as a face image depicting a different individual. Therefore, if the cosine similarity of the embeddings of the two face images is greater than 0.3, the individuals depicted in the two face images will be identified (recognized) as the same person (same ID).

[0169] Figure 19 is a block diagram showing an example configuration of the filter section 123 in Figure 17.

[0170] The filter unit 123 includes a face recognition unit 151, an image selection unit 152, and a storage unit 153.

[0171] The face recognition unit 151 is supplied with a third composite face image from the composite face image generation unit 122 (Figure 17). Furthermore, the face recognition unit 151 is supplied with an encoder, which is the same as the one supplied to the face recognition unit 114 (Figure 17) from the learning unit 43 (Figure 13). The face recognition unit 151 generates an embedding for each of the third composite face images from the composite face image generation unit 122 by recognizing them, for example, using the encoder as a face recognition model. The face recognition unit 151 supplies the embedding for the third composite face image along with the third composite face image to the image selection unit 152.

[0172] The image selection unit 152 is supplied with the embedding of the second composite image, which is given as a condition to the embedding / diversity attribute conditional model when generating the third composite face image supplied from the face recognition unit 151. Hereinafter, the second composite face image from which the embedding given as a condition to the embedding / diversity attribute conditional model is obtained when generating the third composite face image will also be referred to as the second composite image corresponding to the third composite face image.

[0173] The image selection unit 152 selects a face image from the third composite face image from the face recognition unit 151 that is identified (recognized) as the same individual as the individual in the second composite face image corresponding to the third composite face image, based on the embedding of the third composite face image from the face recognition unit 145 and the embedding of the second composite face image corresponding to the third composite face image, and supplies it to the storage unit 153. For example, the image selection unit 152 selects a third composite face image from the third composite face image from the face recognition unit 151 that has a cosine similarity of greater than 0.3 to the embedding of the second composite face image corresponding to the third composite face image, as a face image of the same individual as the individual in the corresponding second composite face image.

[0174] Therefore, the face recognition unit 151 and the image selection unit 152 perform a second filtering process in the filter unit 123, which removes face images from the third composite face image that depict individuals identified as different from the individuals depicted in the corresponding second composite face image. This ensures that the third composite face image supplied from the image selection unit 152 to the storage unit 153 contains face images of individuals identified as the same person as those depicted in the corresponding second composite face image.

[0175] The memory unit 153 stores the third composite face image from the image selection unit 152 and outputs the stored third composite face image as the fourth composite face image.

[0176] Therefore, the fourth composite face image is also guaranteed to be a face image of an individual identified as the same person as the individual in the corresponding second composite face image, similar to the third composite face image supplied from the image selection unit 152 to the storage unit 153. In other words, the fourth composite face image is a face image of an individual identified as the same person as the individual in the second composite face image from which the embedding given as a condition to the embedding / diversity attribute conditional model was obtained when the fourth composite face image (the third composite face image that becomes the fourth composite face image) was generated by the composite face image generation unit 122.

[0177] <Processing by the facial image generation unit 21>

[0178] Figure 20 is a flowchart illustrating the processing of the face image generation unit 21.

[0179] In step S11, the learning unit 43 (Figure 11) of the face image generation unit 21 generates an encoder (face recognition model), which is a learning model that generates and recognizes face image embeddings, by performing preprocessing using the original face images stored in the face image DB 41. Furthermore, the learning unit 43 generates the embeddings of the original face images, the identifiability attributes of the original face images (race, gender, etc.), and the diversity attributes (age, face posture, similarity, etc.), and the processing proceeds to step S12.

[0180] In step S12, the learning unit 43 performs training using the original face image and identifiability attributes (race, gender, etc.) to generate an identifiability attribute-conditional model, which is a learning model that can generate a composite face image by providing the identifiability attributes as conditions. Furthermore, the learning unit 43 performs training using the original face image, the embedding of the original face image, and diversity attributes (age, face pose, similarity, etc.) to generate an embedding / diversity attribute-conditional model, which is a learning model that can generate a composite face image by providing the embedding and diversity attributes as conditions. The learning unit 43 supplies the encoder, the identifiability attribute-conditional model, and the embedding / diversity attribute-conditional model to the generation unit 51 (Figure 11), and the process proceeds from step S12 to step S13.

[0181] In step S13, the generation unit 51 performs a first-stage process to generate the embedding of the (second) composite face image using the identifiable attribute conditional model and encoder, and the process proceeds to step S14.

[0182] In step S14, the generation unit 51 performs a second-stage process to generate a (fourth) composite face image using the embedding and the embedding / diversity attribute conditional model generated in the first-stage process, and then the process is completed.

[0183] Figure 21 is a flowchart illustrating the processing of the first stage, step S13, in Figure 20.

[0184] In step S21, the identifiability attribute generation unit 111 (Figure 17) of the generation unit 51 generates multiple uniform attribute values ​​as attribute values ​​for identifiability attributes (race, gender, etc.). The identifiability attribute generation unit 111 supplies the multiple uniform attribute values ​​of identifiability attributes to the composite face image generation unit 112 (Figure 17), and the process proceeds from step S21 to step S22.

[0185] In the identifiability attribute generation unit 111, by generating multiple equally distributed attribute values ​​as attribute values ​​for identifiability attributes, the subsequent composite face image generation unit 112 can generate a demographically balanced and fair composite face image without bias in identifiability attributes (attribute values) such as gender and race. Furthermore, the composite face image generation unit 112 can generate diverse composite face images, in this case, composite face images that depict diverse individuals with different genders and races.

[0186] In step S22, the composite face image generation unit 112 generates a first composite face image conditioned on each attribute value of the identifiable attribute by providing each attribute value of the identifiable attribute from the identifiable attribute generation unit 111 as a condition to the identifiable attribute conditional model. The composite face image generation unit 112 supplies the first composite face image conditioned on each attribute value of the identifiable attribute to the filter unit 113 (Figure 17), and the process proceeds from step S22 to step S23.

[0187] In step S23, the filter unit 113 performs a first filtering process on the first composite face image, which is conditioned on each attribute value of the identifiability attribute from the composite face image generation unit 112, and generates a second composite face image, which is conditioned on each attribute value of the identifiability attribute. The filter unit 113 supplies the second composite face image, which is conditioned on each attribute value of the identifiability attribute, to the face recognition unit 114 (Figure 17), and the process proceeds from step S23 to step S24.

[0188] In the filter unit 113, the first filtering process is performed on the first composite face image to generate a second composite face image that is conditioned on the attribute values ​​of each identifiable attribute. In this second composite face image, there is only one face image with the same ID. Therefore, the fourth composite face image finally obtained by the generation unit 51 can be made into an unbiased face image.

[0189] In step S24, the face recognition unit 114 generates an embedding for the second composite face image by performing face recognition on the second composite face image, which is conditioned on the attribute values ​​of the identifiability attributes from the filter unit 113, using an encoder. The face recognition unit 114 supplies the embedding for the second composite face image, which is conditioned on the attribute values ​​of the identifiability attributes, to the composite face image generation unit 122 (Figure 17), and the process returns.

[0190] In the composite face image generation unit 122, the embedding of the second composite face image from the face recognition unit 114 is given as a condition to the embedding / diversity attribute conditional model, thereby generating a third composite face image. The third composite face image is a face image conditioned on the embedding of the second composite face image. Furthermore, the second composite face image is a face image generated by giving each attribute value of the uniformly generated identifiability attribute as a condition to the identifiability attribute conditional model, and is not the original face image, which is the real face image. Therefore, the third composite face image, and consequently the fourth composite face image finally obtained by the generation unit 51, is a face image in which the faces of individuals in the original face image, which is the real face image, are (almost) not visible, thereby reducing privacy risks.

[0191] Figure 22 is a flowchart illustrating the processing of the second stage of step S14 in Figure 20.

[0192] In step S31, the diversity attribute generation unit 121 (Figure 17) of the generation unit 51 generates multiple random values ​​as attribute values ​​for diversity attributes (age, facial posture, similarity, etc.), for example, within a pre-set range or within a range set according to user operations. The diversity attribute generation unit 121 supplies the multiple attribute values ​​of the diversity attributes to the composite face image generation unit 122 (Figure 17), and the process proceeds from step S31 to step S32.

[0193] In the diversity attribute generation unit 121, multiple (different) attribute values ​​are generated as attribute values ​​for diversity attributes, so that the subsequent composite face image generation unit 122 can generate a composite face image for the same person (an individual identified by the same ID) with diverse diversity attributes (attribute values) such as age, facial posture, and similarity.

[0194] In step S32, the composite face image generation unit 122 provides the embedding of each second composite face image obtained in the first stage processing (Figure 21) and the attribute values ​​of the diversity attribute from the diversity attribute generation unit 121 as conditions to the embedding / diversity attribute conditional model, thereby generating a third composite face image conditioned on the embedding of each second composite face image and the attribute values ​​(combinations) of the diversity attribute. The composite face image generation unit 122 supplies the third composite face image conditioned on the embedding of each second composite face image and the attribute values ​​of the diversity attribute to the filter unit 123 (Figure 17), and the processing proceeds from step S32 to step S33.

[0195] In step S33, the filter unit 133 performs a second filtering process on the third composite face image, which is conditioned on the embedding and diversity attribute values ​​of each second composite face image from the composite face image generation unit 122, and generates a fourth composite face image, which is conditioned on the embedding and diversity attribute values ​​of each second composite face image. The filter unit 123 outputs the fourth composite face image as the final composite face image generated by the generation unit 51, and the process returns.

[0196] <Description of a computer using this technology>

[0197] The series of processes described above can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up that software are installed on a computer. Here, a computer includes computers built into dedicated hardware, as well as general-purpose personal computers that can perform various functions by installing various programs.

[0198] Figure 23 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above by a program.

[0199] In a computer, the processing circuit 901, ROM (Read Only Memory) 902, and RAM (Random Access Memory) 903 are interconnected by a bus 904.

[0200] An input / output interface 905 is further connected to the bus 904. An input unit 906, an output unit 907, a storage unit 908, a communication unit 909, and a drive 910 are connected to the input / output interface 905.

[0201] The input unit 906 may include physical or virtual operating means that the user operates to input information, such as a keyboard, mouse, or touch panel, as well as means that the user inputs information through voice, eye gaze, etc. Furthermore, the input unit 906 may include sensors that acquire physical quantities such as light or sound, such as a camera or microphone. The output unit 907 may include means that present information to the user by stimulating the user's perception, such as a display, speaker, or haptic device. The storage unit 908 is composed of a hard disk, non-volatile or volatile memory, etc., and stores various types of information (including programs). The communication unit 909 is a network interface, etc., and performs wired or wireless communication with the outside. The drive 910 drives removable media 911 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0202] The processing circuit 901 includes a processor that executes programs such as a CPU (Central Processing Unit) and a DSP (Digital Signal Processor). The processing circuit 901 (its processor) performs the series of processes described above by loading the program stored in the memory unit 908 into the RAM 903 via the input / output interface 905 and the bus 904 and executing it. The processing circuit 901 can output the processing results of the series of processes from the output unit 907, for example, via the bus 904 and the input / output interface 905, as needed. The processing circuit 901 can also store the processing results in the memory unit 908 or transmit them from the communication unit 909.

[0203] The program executed by the computer (processing circuit 901) can be provided by recording it on a removable medium 911, such as a package medium. The program can also be provided via wired or wireless transmission media, such as a local area network, the internet, or digital satellite broadcasting.

[0204] In a computer, a program can be installed in the storage unit 908 via the input / output interface 905 by inserting a removable media 911 into the drive 910. Alternatively, a program can be received by the communication unit 909 via a wired or wireless transmission medium and installed in the storage unit 908. Furthermore, programs can be pre-installed in the ROM 902 or the storage unit 908.

[0205] The programs executed by the computer may be programs that are processed chronologically in the order described herein, or they may be programs that are processed in parallel or at necessary times, such as when a call is made.

[0206] The processes that a computer performs according to a program do not necessarily have to follow the order described in the flowchart. In other words, the processes that a computer performs according to a program include processes that are executed in parallel or individually (e.g., parallel processing and object-based processing).

[0207] A program may be processed by a single computer (processor), or it may be processed in a distributed manner by multiple computers. Furthermore, a program may be transferred to a remote computer and executed there.

[0208] For example, in the face image generation unit 21 shown in Figure 11, the face image DB 41 corresponds to the storage unit 908, and the learning unit 43 and generation unit 51 correspond to the processing circuit 901 (processor) that executes the program.

[0209] In this specification, a system means one component or a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device containing multiple modules in one enclosure, are both systems. Furthermore, for example, the entire computer described above, or a combination of a computer and other devices such as a server (not shown), are also systems. One or more components of a computer, such as processing circuit 901 alone, or a combination of processing circuit 901 to bus 904, are also systems.

[0210] Furthermore, the embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the spirit of this technology.

[0211] For example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.

[0212] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices.

[0213] Furthermore, if a single step includes multiple processes, those processes can be executed by a single device or shared among multiple devices.

[0214] Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also occur.

[0215] Furthermore, this technology can take the following configuration.

[0216] <1> To the first generative model that generates a composite image, a first-stage composite image is generated by providing an identifiability attribute, which is an attribute that affects the identifiability of individuals in the image, as a condition. The second generative model, which generates the composite image, is given the features of the first stage composite image and the diversity attribute, which is an attribute that affects the diversity of the same individual, as conditions to generate the second stage composite image. Includes the generation unit Information processing system. <2> The system further includes a filter unit that performs a first filtering process, which includes removing images from the composite image of the first stage that contain individuals identified as the same individuals as those in other images. <1> The information processing system described above. <3> The first filtering process further includes a process to remove images from the composite image of the first stage in which the identifiable attribute is misrecognized. <2> The information processing system described above. <4> The first filtering process further includes a process to remove images with poor image quality from the composite image of the first stage. <2> or <3> The information processing system described above. <5> The system further includes another filter unit that performs a second filtering operation from the composite image of the second stage, which removes images containing individuals that are identified as different individuals from those in the first stage composite image, for which the features were given as conditions to the second generation model when generating the second stage composite image. <1> ~ <4> An information processing system as described in any of the following. <6> The composite image in the first stage and the composite image in the second stage are composite facial images that show human faces. <1> ~ <5> An information processing system as described in any of the following. <7> The aforementioned identifiable attribute includes one or more of the following: race and sex. <6> The information processing system described above. <8> The diversity attribute includes a similarity score that represents the similarity between the faces of individuals depicted in the composite image generated by the second generative model and a representative face that represents that individual. <6> or <7> The information processing system described above. <9> The aforementioned diversity attributes include one or more of the following: age and facial posture. <6> ~ <8> An information processing system as described in any of the following. <10> The unit further includes a learning unit that trains the first generative model and the second generative model using predetermined facial images that show a person's face. <6> ~ <9> An information processing system as described in any of the following. <11> The learning unit trains an encoder that generates the embedding as a feature using the predetermined face image. <10> The information processing system described above. <12> The generation unit uses the encoder to generate embeddings as feature quantities for the composite image of the first stage. <11> The information processing system described above. <13> The aforementioned learning unit, Using the encoder, the embedding of the predetermined face image is generated. The identifiability attribute and diversity attribute of the predetermined face image are recognized, The first generative model is trained using the predetermined face image and the identifiability attribute of the predetermined face image, and the second generative model is trained using the predetermined face image, the embedding of the predetermined face image, and the diversity attribute. <11> or <12> The information processing system described above. <14> The aforementioned predetermined facial image is a facial image that shows the face of an actual person. <10> ~ <13> An information processing system as described in any of the following. <15> The generation unit generates attribute values ​​of the identifiable attributes to be given as conditions to the first generation model. <1> ~ <14> An information processing system as described in any of the following. <16> The generation unit generates a plurality of equal attribute values ​​as attribute values ​​for the identifiable attribute. <15> The information processing system described above. <17> The generation unit generates attribute values ​​of the diversity attribute to be given as conditions to the second generation model. <1> ~ <16> An information processing system as described in any of the following. <18> The generation unit generates random values ​​or values ​​corresponding to user operations as attribute values ​​for the diversity attribute. <17> The information processing system described above. <19> To the first generative model that generates a composite image, a first-stage composite image is generated by providing an identifiability attribute, which is an attribute that affects the identifiability of individuals in the image, as a condition. The second generative model, which generates the composite image, is given the features of the first stage composite image and the diversity attribute, which is an attribute that affects the diversity of the same individual, as conditions to generate the second stage composite image. Information processing methods that include the following. <20> To the first generative model that generates a composite image, a first-stage composite image is generated by providing an identifiability attribute, which is an attribute that affects the identifiability of individuals in the image, as a condition. The second generative model, which generates the composite image, is given the features of the first stage composite image and the diversity attribute, which is an attribute that affects the diversity of the same individual, as conditions to generate the second stage composite image. Generation part A program that makes a computer function. [Explanation of Symbols]

[0217] 10 Face processing system, 20 Learning device, 21 Face image generation unit, 22 Face image DB, 23 Face image selection unit for learning, 24 Face image learning unit, 30 Face processing device, 31 Face detection unit, 32 Face processing unit, 40 Learning device, 41 Face image DB, 42 Acquisition unit, 43 Learning unit, 50 Generation device, 51 Generation unit, 61 Preprocessing unit, 62,63 Learning unit, 71 Face recognition model learning unit, 72 Face recognition unit, 73 Attribute recognition unit, 81 Embedding acquisition unit, 82 Representative embedding calculation unit, 83 Similarity calculation unit, 91 First stage processing unit, 92 Second stage processing unit, 111 Discriminability attribute generation unit, 112 Synthetic face image generation unit, 113 Filter unit, 114 Face recognition unit, 121 Diversity attribute generation unit, 122 Synthetic face image generation unit, 123 Filter unit, 141 Discrimination attribute recognition unit, 142 Image selection unit, 143 Image quality recognition unit, 144 Image selection unit, 145 Face recognition unit, 146 Image selection unit, 147 Storage unit, 151 Face recognition unit, 152 Image selection unit, 153 Storage unit, 901 Processing circuit, 902 ROM, 903 RAM, 904 Bus, 905 Input / Output interface, 906 Input unit, 907 Output unit, 208 Storage unit, Communication unit, 210 Drive, 211 Removable media

Claims

1. To the first generative model that generates a composite image, a first-stage composite image is generated by providing an identifiability attribute, which is an attribute that affects the identifiability of individuals in the image, as a condition. The second generative model, which generates the composite image, is given the features of the first stage composite image and the diversity attribute, which is an attribute that affects the diversity of the same individual, as conditions to generate the second stage composite image. Includes the generation unit Information processing system.

2. The system further includes a filter unit that performs a first filtering process, which includes the process of deleting images from the composite image of the first stage that contain individuals identified as the same individuals as those appearing in other images. The information processing system according to claim 1.

3. The first filtering process further includes a process of removing images from the composite image of the first stage in which the identifiable attribute is misrecognized. The information processing system according to claim 2.

4. The first filtering process further includes a process of removing images with poor image quality from the composite image of the first stage. The information processing system according to claim 2.

5. The system further includes another filter unit that performs a second filtering operation from the composite image of the second stage, which removes images containing individuals that are identified as different individuals from those in the composite image of the first stage, where the feature quantities were given as conditions for the second generation model when generating the composite image of the second stage. The information processing system according to claim 1.

6. The composite image in the first stage and the composite image in the second stage are composite facial images that show human faces. The information processing system according to claim 1.

7. The aforementioned identifiable attribute includes one or more of race and sex. The information processing system according to claim 6.

8. The diversity attribute includes a similarity score that represents the similarity between the faces of individuals depicted in the composite image generated by the second generative model and a representative face that represents that individual. The information processing system according to claim 6.

9. The aforementioned diversity attributes include one or more of the following: age and facial posture. The information processing system according to claim 6.

10. The unit further includes a learning unit that trains the first generative model and the second generative model using predetermined facial images that show a person's face. The information processing system according to claim 6.

11. The learning unit trains an encoder that generates the embedding as a feature using the predetermined face image. The information processing system according to claim 10.

12. The generation unit uses the encoder to generate embeddings as feature quantities for the composite image of the first stage. The information processing system according to claim 11.

13. The aforementioned learning unit, Using the encoder, the embedding of the predetermined face image is generated. The identifiability attribute and diversity attribute of the predetermined face image are recognized, The first generative model is trained using the predetermined face image and the identifiability attribute of the predetermined face image, and the second generative model is trained using the predetermined face image, the embedding of the predetermined face image, and the diversity attribute. The information processing system according to claim 11.

14. The aforementioned predetermined facial image is a facial image that shows the face of an actual person. The information processing system according to claim 10.

15. The generation unit generates attribute values ​​of the identifiable attribute to be given as conditions to the first generation model. The information processing system according to claim 1.

16. The generation unit generates a plurality of equal attribute values ​​as attribute values ​​for the identifiable attribute. The information processing system according to claim 15.

17. The generation unit generates attribute values ​​of the diversity attribute to be given as conditions to the second generation model. The information processing system according to claim 1.

18. The generation unit generates random values ​​or values ​​corresponding to user operations as attribute values ​​for the diversity attribute. The information processing system according to claim 17.

19. To the first generative model that generates a composite image, a first-stage composite image is generated by providing an identifiability attribute, which is an attribute that affects the identifiability of individuals in the image, as a condition. The second generative model, which generates the composite image, is given the features of the first stage composite image and the diversity attribute, which is an attribute that affects the diversity of the same individual, as conditions to generate the second stage composite image. Information processing methods that include the following.

20. To the first generative model that generates a composite image, a first-stage composite image is generated by providing an identifiability attribute, which is an attribute that affects the identifiability of individuals in the image, as a condition. The second generative model, which generates the composite image, is given the features of the first stage composite image and the diversity attribute, which is an attribute that affects the diversity of the same individual, as conditions to generate the second stage composite image. Generation part A program that makes a computer function.

Citation Information

Patent Citations

  • Information processing unit, information processing method and program

    JP2021082068A