Information processing system, information processing method, and program

The system generates diverse and fair synthetic face images through a two-stage process, enhancing face recognition model performance and resolving privacy issues by using AI-based learning models to condition on identity and diversity attributes, thus overcoming the limitations of existing generative models.

WO2026070579A1PCT designated stage Publication Date: 2026-04-02SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing generative models struggle to generate diverse and fair synthetic face images that preserve identity, maintain high quality, and avoid biases in attributes such as gender and race, making it difficult to train accurate face recognition models without privacy concerns.

Method used

An information processing system and method that generates synthetic face images using a two-stage process, first conditioning on identity attributes to preserve identity and then on diversity attributes to enhance variation, while ensuring fairness and quality, using AI-based learning models to create an automated pipeline for generating diverse and unbiased synthetic images.

Benefits of technology

The system produces high-quality, diverse, and fair synthetic face images that accurately represent various attributes, improving the performance of face recognition models and addressing privacy and ethical concerns by eliminating the need for real face image collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025032815_02042026_PF_FP_ABST
    Figure JP2025032815_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Diverse image variations can be obtained easily. A generative unit generates first-stage synthetic images by conditioning a first generative model configured to generate synthetic images on an identity attribute that is an attribute that affects identity of an individual to be depicted in an image. The generative unit generates second-stage synthetic images by conditioning a second generative model configured to generate synthetic images on a feature of the first-stage synthetic images and a diversity attribute that is an attribute that affects diversity of a same individual. The present technique can be applied for generating images such as diverse face image variations, for example.
Need to check novelty before this filing date? Find Prior Art

Description

INFORMATION PROCESSING SYSTEM, INFORMATION PROCESSING METHOD, AND PROGRAM

[0001] The present technique relates to an information processing system, an information processing method, and a program, and particularly to an information processing system, an information processing method, and a program that allow easy acquisition of diverse image variations, for example.

[0002] Development of a fair and accurate recognition (identification) learning model requires a vast amount of carefully collected training data. Training a face recognition model that recognizes faces, for example, requires diverse face image variations of each of multiple IDs (identities) that each specify (identify) a human individual (person), each face image depicting the individual with the corresponding ID (individual specified by the ID). A fair and accurate recognition model is one that achieves a high accuracy rate without bias in the accuracy of the results (recognition rate) depending on gender or race, for example.

[0003] Training a fair and accurate face recognition model requires diverse face image variations, i.e., a large set of demographically balanced (fairly balanced) face images that equally represent attributes such as gender, race, and age, without bias. Moreover, the diverse face image variations should preferably be of high quality.

[0004] However, it is difficult to obtain such diverse image variations by collecting real face images, especially from a privacy perspective. Real face images are face images of real people. Real people include people who exist now, or who existed in the past.

[0005] Accordingly, there is a growing demand for synthetic images, in particular synthetic face images, which are images generated by information processing, to prevent privacy risks and attribute bias.

[0006] Various methods have been proposed to generate synthetic face images using a machine learning model as a generative model for generating synthetic images.

[0007] The synthetic face images generated by currently available generative models are not necessarily appropriate as diverse image face variations. Because of this, there can be a significant difference in the performance between a face recognition model trained on (datasets of) synthetic face images and a face recognition model trained on (datasets of) real face images.

[0008] It has been proposed to convert images to have the same characteristics as the target images to be recognized by a recognizer (recognition model), and to use the converted images as training data to train the recognizer (see, for example, PTL 1).

[0009] JP 2021-082068A

[0010] In recent years, there has been a demand to propose techniques that allow easy acquisition of diverse image variations.

[0011] The present technique has been developed in view of the circumstances described above, with an aim to enable easy acquisition of diverse image variations.

[0012] According to an aspect of the present disclosure, there is provided an information processing system comprising: processing circuitry configured to generate one or more first synthetic images using a first model conditioned on values of identity attributes, extract features from the one or more first synthetic images, and generate one or more second synthetic images using a second model conditioned on the features of the one or more first synthetic images and values of diversity attributes, wherein the one or more second synthetic images include the values of the diversity attributes while preserving identity information corresponding to the features.

[0013] According to an aspect of the present disclosure, there is provided an information processing method comprising: generating, by processing circuitry, one or more first synthetic images using a first model conditioned on values of identity attributes; extracting, by the processing circuitry, features from the one or more first synthetic images; and generating, by the processing circuitry, one or more second synthetic images using a second model conditioned on the features of the one or more first synthetic images and values of diversity attributes, wherein the one or more second synthetic images include the values of the diversity attributes while preserving identity information corresponding to the features.

[0014] According to an aspect of the present disclosure, there is provided a non-transitory computer-readable medium storing instructions which when executed by a processor cause the processor to perform operations comprising: generating one or more first synthetic images using a first model conditioned on values of identity attributes; extracting features from the one or more first synthetic images; and generating one or more second synthetic images using a second model conditioned on the features of the one or more first synthetic images and values of diversity attributes, wherein the one or more second synthetic images include the values of the diversity attributes while preserving identity information corresponding to the features.

[0015] The information processing system may be an independent apparatus, or some internal blocks configuring an apparatus. One or more blocks that configure the information processing system can be implemented as separate apparatus(es).

[0016] The program can be provided by transmission over a transmission medium, or by recording it on a recording medium.

[0017] Fig. 1 is a block diagram showing an example configuration of one embodiment of a face processing system to which the present technique is applied.Fig. 2 is a diagram explaining a first issue in generating synthetic face images using a generative model.Fig. 3 is a diagram explaining a second issue in generating synthetic face images using a generative model.Fig. 4 is a diagram explaining a third issue in generating synthetic face images using a generative model.Fig. 5 is a diagram explaining a fourth issue in generating synthetic face images using a generative model.Fig. 6 is a diagram explaining a fifth issue in generating synthetic face images using a generative model.Fig. 7 is a diagram explaining how a bias in the face images used to train a face recognition model affects that model.Fig. 8 is a diagram illustrating a face recognition model that was trained using synthetic face images generated by a face image generator 21.Fig. 9 is a diagram illustrating an example application that uses synthetic face images generated by the face image generator 21.Fig. 10 is a diagram illustrating another example application that uses synthetic face images generated by the face image generator 21.Fig. 11 is a block diagram showing an example configuration of an information processing system as the face image generator 21.Fig. 12 is a diagram schematically showing the face images stored in a face image DB 41.Fig. 13 is a block diagram showing an example configuration of a training unit 43.Fig. 14 is a block diagram showing an example configuration of a preprocessor 61.Fig. 15 is a block diagram showing an example configuration of a part of an attribute recognizer 73 that recognizes similarity, which is one of diversity attributes.Fig. 16 is a block diagram showing an example configuration of a generative unit 51.Fig. 17 is a block diagram showing an example configuration of a first-stage processor 91 and a second-stage processor 92.Fig. 18 is a block diagram showing an example configuration of a filtering unit 113.Fig. 19 is a block diagram showing an example configuration of a filtering unit 123.Fig. 20 is a flowchart explaining the processing performed by the face image generator 21.Fig. 21 is a flowchart explaining the first-stage processing performed in step S13.Fig. 22 is a flowchart explaining the second-stage processing performed in step S14.Fig. 23 is a block diagram showing an example configuration of one embodiment of a computer to which the present technique is applied.

[0018] <Face processing system to which the present technique is applied>

[0019] Fig. 1 is a block diagram showing an example configuration of one embodiment of a face processing system to which the present technique is applied.

[0020] In Fig. 1, the face processing system 10 includes a training apparatus 20 and a face processing apparatus 30. The training apparatus 20 trains a (machine) learning model for face processing, i.e., for the processing of face images. The face processing apparatus 30 performs face processing of face images depicting faces such as face recognition or face image correction, for example, using the trained learning model for face processing. Face image correction includes, for example, correcting an image of a face that is turned away from the front to an image of a face that is turned toward the front.

[0021] The training apparatus 20 includes a face image generator 21, a face image DB 22, a training face image selector 23, and a face image training unit 24.

[0022] The face image generator 21 generates (data of) synthetic face images of the faces of various people (individuals). For example, the face image generator 21 generates realistic, fair, and diverse synthetic face images in which identity (individuality) is preserved through automated pipelines using an AI (artificial intelligence)-based learning model. Identity of a face image means that a face image (individual in the image) is identified (or recognized) as a specific individual (or ID). Identity being preserved means that a face image that should be identified as that of a specific individual is indeed identified as that specific individual. The face image generator 21 delivers synthetic face images to the face image DB 22 with associated (labels of) IDs that identify the individuals who are (whose faces are) depicted in the face images.

[0023] While this embodiment adopts humans as the individuals depicted in face images, other animals such as dogs or cats can be adopted as “individuals.” While the present technique will be described below with respect to generation of (synthetic) images of faces as one example, this technique can also be applied to generation of other images such as, for example, body images depicting the whole body. Whole body images can be used to train a learning model that is to be employed in person (human) re-identification, for example, and other tasks.

[0024] The face image DB 22 stores synthetic face images delivered from the face image generator 21.

[0025] The training face image selector 23 selects some of the synthetic face images stored in the face image DB 22 as a training set of face images to be used as the data to train a learning model for face processing, and delivers the images to the face image training unit 24. The training face image selector 23 selects, for example, according to a user operation, some of the synthetic face images stored in the face image DB 22 that depict individuals of a specific attribute, e.g., specific gender, age, or race, as a training set of face images, and delivers the images to the face image training unit 24. The face image generator 21 can select all the synthetic face images stored in the face image DB 22 as the training set of face images.

[0026] The face image training unit 24 trains the learning model for face processing (including updating and fine-tuning) using the training set of face images delivered from the training face image selector 23, and delivers the trained model (model after training) to the face processing apparatus 30.

[0027] The face processing apparatus 30 includes a face detector 31 and a face image processor 32.

[0028] The face images to be processed are delivered to the face detector 31. The face detector 31 detects faces in the delivered face images, identifies predetermined regions that include the faces, and delivers the images to the face image processor 32.

[0029] The trained learning model output from the face image training unit 24 to the face processing apparatus 30 is delivered to the face image processor 32. The face image processor 32 performs face processing to the face images (faces in the images), in which the face detector 31 has specified regions containing faces, using the learning model delivered from the face image training unit 24, and outputs the face processing results. For example, when the trained learning model for face processing is a face recognition model that is a learning model for face recognition, the face image processor 32 performs processing, using the face recognition model, to produce the ID of an individual who is (whose face is) depicted in the face image from the face detector 31, and outputs the resultant ID as the face recognition result. Or, when the trained learning model for face processing is a face correction model that is a learning model for face image correction, the face image processor 32 performs processing, using the face correction model, to generate an image of a face turned to a predetermined direction such as the front, for example, from a face image delivered from the face detector 31, and outputs this face image as the result of face correction.

[0030] <Issues in generating synthetic face images>

[0031] Fig. 2 is a diagram explaining a first issue in generating synthetic face images using a generative model.

[0032] The first issue concerns the preservation of identity (ID preservation). That is, sometimes, the generated synthetic face images fail to preserve identity.

[0033] For example, when generating face images that are supposed to be recognized as those of a specific individual based on a face image im21 depicting the face of that specific person, the generative model sometimes generates a synthetic face image im23 that is identified as an individual other than the specified individual, along with a synthetic face image im22 that is identified as that of the specific individual. Namely, sometimes, a synthetic face image im23 that does not preserve the face features of the specific individual depicted in the face image im21 is generated.

[0034] The face image generator 21 generates synthetic face images by conditioning the generative model on features of the face image depicting the face of a specific person (in other words, the features of the specific individual). This prevents the face image generator 21 from generating synthetic face images that are identified as other individuals than the specific individual, as well as facilitates the generation of synthetic face images that are identified as the specific person. This results in the generation of synthetic face images preserving the identity so that the images are identified as the specific individual. The first issue can thus be resolved.

[0035] The feature of a face image can be any value that will be similar (can be considered the same) when extracted from each of various face images depicting the same person. For example, an embedding (vector) extracted from a face image by an encoder that implements the face recognition model, or an ID output as the face recognition result of a face image by the face recognition model, can be adopted as the feature of the face image. Also adoptable as the feature of a face image is, for example, a potential variable extracted from the face image by an encoder such as a VAE (Variational Auto Encoder). In the following, embeddings are used as the features of face images. A value corresponding to an embedding (e.g., a quantized value of a vector as the embedding) can be adopted as the ID of a face image (individual depicted in the image).

[0036] Fig. 3 is a diagram explaining a second issue in generating synthetic face images using a generative model.

[0037] The second issue concerns image corruption (lack of consistency). That is, sometimes, the generated face images are corrupted.

[0038] For example, a synthetic face image im31 generated using the generative model may include a significant artifact, or depict a structure that makes face recognition difficult.

[0039] The face image generator 21 performs a filtering process, so that it generates synthetic face images that are not only of high quality but also depict faces that allow consistent (accurate, or correct) recognition of the gender and race (synthetic face images depicting a structure that is recognizable as a face). The second issue can thus be resolved.

[0040] Additionally, the filtering process guarantees that each face image, whose embedding as its feature is to be used to condition the generative model described with respect to the first issue (Fig. 2), is unique to one ID.

[0041] Fig. 4 is a diagram explaining a third issue in generating synthetic face images using a generative model.

[0042] The third issue concerns intra-class diversity. That is, the generated images of the same person tend to have little variation or insufficient diversity.

[0043] For example, when generating face images that are supposed to be recognized as those of a specific individual based on a face image im41 depicting the face of that specific person, the generative model tends to generate synthetic face images im42, im43, and im44 that are identified as the specific individual but lack variety, i.e., with hardly any change in the attributes such as age, face pose (face orientation), and expression.

[0044] The face image generator 21 generates synthetic face images by conditioning the generative model on the embedding, as well as (values of) attributes such as age, face pose, expression, and similarity of the individual depicted in the synthetic face images to be generated. This allows the face image generator 21 to generate synthetic face images with the conditioning embedding (ID corresponding to the embedding) and depicting the individual having the conditioning attributes such as age and similarity. This results in the generation of a variety of synthetic face images of the same person that are diverse in terms of the attributes (attribute values) such as age, face pose, expression, and similarity. The third issue can thus be resolved.

[0045] Similarity here refers to the similarity between an individual’s face depicted in a synthetic face image and a representative face of that individual. An individual’s representative face can be, for example, an average face of multiple face images depicting the individual with different ages, face poses, expressions, lighting conditions, and items such as a mask or glasses (whether they are worn or not). Alternatively, for example, an individual’s representative face can be the face in a face image selected from multiple face images.

[0046] Assuming that the higher the similarity, the more similar the face is to the representative face, for example, giving the generative model a higher similarity will result in the generation of synthetic face images with faces that are similar to the representative face in terms of age, face pose, and expression, for example. Thus the similarity can affect the age, face pose, and expression of the individual depicted in the synthetic face images generated by the generative model.

[0047] Fig. 5 is a diagram explaining a fourth issue in generating synthetic face images using a generative model.

[0048] The fourth issue concerns inter-class diversity. That is, it is hard to generate face images of a completely different person beyond the range of faces in the face images used to train the generative model.

[0049] For example, a generative model that was trained on face images im51 and im52 generates a face image im53 and a face image im54 that are similar to the face images im51 and im52 used in the training, but does not generate face images of a completely different person. For example, a generative model trained on face images of only one specific race will not generate face images of people of other races than that specific race. Therefore, the generated images tend to lack variety or diversity as the face images of different people (individuals).

[0050] The face image generator 21 generates a synthetic face image to obtain an embedding which the generative model is to be conditioned on, by conditioning another generative model on (values of) attributes such as gender and race. This allows the model to generate synthetic face images of individuals of different genders and races, i.e., different IDs (different people). Thus, synthetic images of diverse people, i.e., synthetic face images of various different people, are generated. The fourth issue can thus be resolved.

[0051] Fig. 6 is a diagram explaining a fifth issue in generating synthetic face images using a generative model.

[0052] The fifth issue concerns fairness of face images. Namely, a bias in attributes, for example, in protected characteristics such as gender and race, of the face images used to train the generative model is reflected (amplified) in the synthetic face images generated by that generative model.

[0053] Sometimes the (datasets of) real face images contain a bias in protected characteristics such as gender, race, and age. When a generative model is trained on real face images with a bias in gender or race, for example, the (datasets) of synthetic face images generated using this trained generative model exhibit a bias in gender and race similar to that in the real face images used for the training. Therefore, fair synthetic face images, i.e., unbiased synthetic face images uniformly (equally) representing the attributes such as gender and race are hard to obtain.

[0054] As described in relation to the fourth issue (Fig. 5), the face image generator 21 generates a synthetic face image to obtain an embedding which the generative model is to be conditioned on, by conditioning another generative model on (values of) attributes such as gender and race. The face image generator 21 conditions the other generative model on gender and race uniformly (equally). This results in the generation of synthetic face images that are fair, equally representing gender and race without bias (or less biased images representing gender and race substantially equally). The fifth issue can thus be resolved.

[0055] It should be noted that any learning model other than generative models that were trained on biased face images will exhibit the influence of the bias in the face images used for the training.

[0056] Fig. 7 is a diagram explaining how a bias in the face images used to train a face recognition model affects that model.

[0057] Fig. 7 shows a case where a face recognition model trained on gender-biased real face images consisting nearly twice as many male faces as female faces is deployed.

[0058] The gender bias in the real face images used to train the face recognition model affects the face recognition accuracy of that face recognition model for each gender, i.e., the accuracy of male face recognition will be nearly twice as high as that of female face recognition. Therefore, the face recognition model will not be fair, as its accuracy is gender-biased -- specifically, the accuracy of female face recognition is approximately half that of male face recognition.

[0059] When real face images are used to train a face recognition model, issues can arise regarding sensitive data (data that needs to be protected from unauthorized access). Namely, when real face images are used to train a face recognition model, for example, a hostile attack to the face recognition model can detect these real face images that were used as training data in the training. Therefore, privacy issues and ethical concerns can arise when real face images used in the training are images that were collected from Internet or somewhere else without consent.

[0060] Fig. 8 is a diagram illustrating a face recognition model that was trained using synthetic face images generated by the face image generator 21.

[0061] The face image generator 21 generates fair synthetic face images, i.e., unbiased and equal in terms of gender and race, as described with reference to Fig. 6. Therefore, the synthetic face images include approximately equal numbers of men and women. A face recognition model trained using such fair synthetic face images as the training data generates accurate results that are fair and unbiased in terms of gender, mirroring the fairness of the synthetic face images as the train data. Namely, the face recognition model will generate gender-fair results, with the accuracy of male face recognition being approximately equal to that of female face recognition. This results in improved performance of the face recognition model compared to a face recognition model that was trained using biased face images described with reference to Fig. 7.

[0062] When synthetic face images are used to train a face recognition model, no issues arise regarding sensitive data. Namely, synthetic face images are non-sensitive data. What will be detected (reproduced) by a hostile attack to the face recognition model will only be synthetic face images that are non-sensitive data. Therefore, privacy issues and ethical concerns that can arise when real face images are used to train the face recognition model can be prevented.

[0063] The face image generator 21 generates synthetic face images through an automated pipeline using an AI-based learning model. That is, the face image generator 21 generates synthetic face images (second-stage synthetic images) by conditioning a generative model (second generative model) on embeddings as well as attributes such as age and similarity of individuals to be depicted in the resultant synthetic face images. The face image generator 21 generates synthetic face images (first-stage synthetic images), from which embeddings are obtained to be used to condition the generative model, by conditioning another generative model (first generative model) on attributes such as gender and race.

[0064] Therefore, the face image generator 21 can generate an unlimited (or an almost infinitesimal) number of new synthetic face images by varying the information for conditioning the generative model and the other generative model.

[0065] Moreover, the filtering process enables the face image generator 21 to generate high-quality synthetic face images that are accurately recognized in terms of gender and race, as well as guarantees that each of the synthetic face images whose embeddings are to be used to condition the generative model is unique to one person (one ID).

[0066] Further, since the face image generator 21 generates synthetic face images by conditioning the generative model on embeddings, the resultant synthetic face images preserve the identity, i.e., the synthetic face images are identified (recognized) as the individuals with the IDs corresponding to the embeddings.

[0067] The face image generator 21 conditions another generative model on (values of) attributes such as gender and race to generate synthetic face images from which embeddings are obtained and used to condition the generative model. This allows the distribution of demographic characteristics such as gender and race of individuals in the synthetic face images to be controlled by varying the gender and race for conditioning the other generative model. Moreover, the generative model that outputs the synthetic face images is conditioned on embeddings of the synthetic face images that were generated by the other generative model. Therefore, the privacy risks entailed in the synthetic face images generated by the generative model are reduced compared to real face images.

[0068] Furthermore, the face image generator 21 can control the number of synthetic face images generated by the generative model for each ID corresponding to an embedding for conditioning the generative model by varying age or similarity, for example, which are used to condition the generative model along with the embedding. Thus the total number of synthetic face images generated by the face image generator 21 can be controlled.

[0069] Collection of real face images and utilization of face recognition AI (face recognition models) can lead to ethical, legal, and social issues.

[0070] The collection of real face images and face recognition AI can violate the privacy of the individuals (objects) of the real face images, depending on the method of collecting the real face images and the method of using the face recognition AI. For example, the use of face recognition AI used by government agencies or police for law enforcement purposes can lead to human rights violation. Face recognition AI can be used, for example, in criminal investigations to identify individuals in face images captured by security cameras, in airport security checks, or in the search for terrorists in public spaces. Allowing government agencies and police to use face recognition AI could lead to excessive surveillance of people, undermining their basic human rights. Limitations in the face recognition AI performance could result in false arrests. Variations in the accuracy of face authentication AI based on sensitive attributes such as race and gender could lead to discrimination.

[0071] Clearview AI, a U.S. company, has built a database of more than 20 billion face images supplied solely from public sources, including SNS, websites, news media, and photographs of suspects or criminals. The company has developed and provided face recognition AI using these unconsented face images. This AI has caused controversy, as it has been used by government agencies and police worldwide. In the U.K., Clearview AI was fined $9.4 million (approximately 1.4 billion yen) for violating several U.K. laws by collecting and retaining face images indefinitely without personal consent. Orders were issued to remove and stop using the face images of U.K. residents. Disputes have been filed against Clearview AI not only in the UK, but also in California and Illinois in the U.S., Canada, France, Italy, Sweden, and Australia. Many disputes have already arisen due to the GDPR (General Data Protection Regulation) in Europe and the personal data protection laws of various regions.

[0072] Clearview AI's face recognition technology is offered for military applications. The database of face recognition technology is said to have collected more than 2 billion photographs from the Russian SNS “VKontakte.”

[0073] Against this backdrop, the legal regulations have been further tightened regarding the collection of real face images and face recognition AI. On March 13, 2024, the European Parliament passed the European AI Act, a comprehensive AI regulation. Article 5(db) of the European AI Act prohibits AI practices including “the placing on the market, the putting into service for this specific purpose, or the use of AI systems that create or expand face recognition databases through the untargeted scraping of face images from the internet or CCTV footage.” Any AI systems that collect face images by crawling to create AI datasets for face recognition purposes are prohibited as “unacceptable AI.”

[0074] The generation of synthetic face images by the face image generator 21 can help address the ethical, legal, and social issues described above.

[0075] <Applications that use synthetic face images>

[0076] Fig. 9 is a diagram illustrating an example application that uses synthetic face images generated by the face image generator 21.

[0077] The face image generator 21 generates synthetic face images by conditioning the generative model on embeddings as well as attributes such as age, face pose, expression, and similarity of the individuals to be depicted in the resultant synthetic face images. This allows the face image generator 21 to generate synthetic face images of the individuals with the IDs corresponding to the conditioning embeddings, and with the conditioning face poses and ages. Accordingly, by controlling the face pose and age used to condition the generative model, e.g., by varying the face pose and age for conditioning the generative model, variations of face images with different face poses and different ages of the same person can be generated as the synthetic face images as shown in Fig. 9.

[0078] The synthetic face images generated by the face image generator 21 described above can be used to train a face recognition learning model for tracking the faces of the same person, for example. The synthetic face images generated by the face image generator 21 can also be used to train a face recognition learning model for generating, for example, passport photographs, wherein a face image with the face turned to the front is generated from a face image with the face tilted or rotated. The synthetic face images generated by the face image generator 21 can also be used to train a face recognition learning model for face verification of passport photographs, for example.

[0079] Fig. 10 is a diagram illustrating another example application that uses synthetic face images generated by the face image generator 21.

[0080] For example, a still or moving image of a crowd can depict many people without permission, as shown in Fig. 10. To protect the privacy of many people in still or moving images, the faces of these people can be replaced with synthetic face images generated by the face image generator 21.

[0081] The faces of the people depicted in still or moving images of a crowd can be replaced with synthetic face images generated by the face image generator 21, thereby to control the distribution of the demographic characteristics of the people in the still or moving images.

[0082] For example, the face image generator 21 can generate a variety of synthetic face images with different ages as described with reference to Fig. 9. Moreover, as described in relation to the fourth issue (Fig. 5), the face image generator 21 can generate synthetic face images of individuals of different genders and races, by obtaining embeddings from the synthetic face images generated by another generative model conditioned on (values of) attributes such as gender and race and using the embeddings for conditioning the generative model. As described above, the face image generator 21 can generate synthetic face images in which the distribution of demographic characteristics such as gender, race, and age are controlled. Therefore, replacing the faces of the people in still or moving images with synthetic face images generated by the face image generator 21 allows for control of the distribution of demographic characteristics of the people in the still or moving images. For example, people in a still or moving image can be replaced with people of a particular race, such as white people, or people of a particular age group, such as older adults. Alternatively, for example, the still or moving image can be replaced with an image of approximately equal numbers of men and women of various different races.

[0083] <Example configuration of face image generator 21>

[0084] Fig. 11 is a block diagram showing an example configuration of an information processing system as the face image generator 21 in Fig. 1.

[0085] In Fig. 11, the face image generator 21 includes a training section 40 and a generative apparatus 50.

[0086] The training section 40 includes a face image DB 41, an acquisition unit 42, and a training unit 43, and trains learning models that generate face images. The training section 40 delivers the trained learning models to the generative apparatus 50.

[0087] The face image DB 41 stores face images to be used to train the learning models. Each face image is labeled with an ID that identifies the individual who is (whose face is) depicted in the image. For example, the face image DB 41 stores one or multiple different face images corresponding to each of the plurality of IDs. Multiple different face images corresponding to an ID refer to multiple images of an individual’s face with the same ID but different face poses, ages, expressions, lighting conditions, and items such as a mask or glasses (whether they are worn or not). The face image DB 41 may store either real face images or synthetic face images, or both.

[0088] The acquisition unit 42 acquires synthetic face images from the face image DB 41 as the training data to be used to train the learning models in the training unit 43, and delivers the data to the training unit 43.

[0089] The training unit 43 trains the learning models using the face images (predetermined face images) from the acquisition unit 42 as the training data, and delivers the trained models to a generative unit 51 of the generative apparatus 50. The learning models can be implemented by a diffusion model, VAE, GAN (Generative Adversarial Networks), or a transformer, for example. Deep learning, for example, can be employed for the training of the learning models.

[0090] The generative apparatus 50 includes the generative unit 51. The generative unit 51 generates synthetic face images using a learning model output from the training unit 43.

[0091] Fig. 12 is a diagram schematically showing the face images stored in the face image DB 41.

[0092] The face image DB 41 can store datasets of two-dimensional RGB (red, green, and blue) images, for example, of real faces of multiple individuals (IDs) with different face poses, ages, expressions, lighting conditions, and items such as a mask or glasses (whether they are worn or not), as shown in Fig. 12. In Fig. 12, the horizontal axis represents the face pose, age, expression, lighting conditions, and the presence / absence of an item such as a mask or glasses. The vertical axis represents the ID. Therefore, each line in Fig. 12 shows real images of the same individual with the same ID (same person) with different face poses, ages, expressions, and lighting conditions, and whether items such as a mask or glasses are worn.

[0093] For example, the face image DB 41 can store WebFace260M (datasets of) real face images.

[0094] <Example configuration of training unit 43>

[0095] Fig. 13 is a block diagram showing an example configuration of the training unit 43 in Fig. 11.

[0096] In Fig. 13, the training unit 43 includes a preprocessor 61, and training modules 62 and 63.

[0097] Real face images with corresponding IDs stored in the face image DB 41 are delivered from the acquisition unit 42 (Fig. 11) to the preprocessor 61.

[0098] The preprocessor 61 performs preprocessing.

[0099] In the preprocessing, the preprocessor 61 trains a face recognition model as one of the learning models that are delivered from the training unit 43 (Fig. 11) to the generative unit 51 using the real face images with corresponding IDs. Here, the (real) face images stored in the face image DB 41 to be used to train the learning model in the training unit 43 shall be referred to also as “original face images.”

[0100] The face recognition model can be any of the learning models that output features of face images, e.g., embeddings or IDs, such as encoders. The embeddings or IDs of face images output from an encoder as a face recognition model allow for determination of whether or not the individuals in two face images are the same person. If the vectors representing the embeddings of two face images are similar enough to be recognized as identical, or if the IDs of two face images are identical, the individuals in the two face images can be determined as the same person. Here, the IDs of the face images (of the individuals in the images) are identified by the embeddings of the face images output from the encoder as the face recognition model.

[0101] In the preprocessing, the preprocessor 61 recognizes the original face images and generates embeddings of the respective face images using the encoder as the trained face recognition model. Further, in the preprocessing, the preprocessor 61 recognizes the identity attributes and (the values of) the diversity attributes of the individuals in the respective original face images, using trained learning models that are trained to each recognize identity attributes and diversity attributes.

[0102] Identity attributes are those that affect the identity of an individual (person) and include race and gender, for example. Diversity attributes are those that affect the diversity of the same individual (person) and include age, face pose, and similarity, for example. The learning models trained to recognize identity attributes and diversity attributes are prepared in advance.

[0103] The preprocessor 61 outputs the encoder as the face recognition model, as well as outputs the original face images and their embeddings, identity attributes (labels indicating the attribute values), and diversity attributes (labels indicating the attribute values). The encoder as the face recognition model output from the preprocessor 61 is delivered to the generative unit 51 as one of the learning models delivered there from the training unit 43 (Fig. 11). The original face images output by the preprocessor 61 are delivered to the training modules 62 and 63. The identity attributes output by the preprocessor 61 are delivered to the training module 62. The embeddings and diversity attributes output by the preprocessor 61 are delivered to the training module 63.

[0104] The training module 62 trains an identity attribute-conditioned model that is one of the learning models, using the original face images and the identity attributes of the individuals in the original face images delivered from the preprocessor 61. The identity attribute-conditioned model here is a generative model (first generative model, or “the other generative model”) that is conditioned on identity attributes to generate synthetic face images (first-stage synthetic images) that reflect the identity attributes. The identity attribute-conditioned model can be a diffusion model, for example.

[0105] The training module 62 trains the identity attribute-conditioned model so that when the diffusion model as the identity attribute-conditioned model is given the identity attributes of an individual in an original face image as a condition, the model generates that original face image. The training module 62 delivers the trained identity attribute-conditioned model (after the training) to the generative unit 51 as one of other learning models delivered there from the training unit 43 (Fig. 11).

[0106] The training module 63 trains an embedding / diversity attribute-conditioned model that is one of the learning models, using the original face images and the diversity attributes and embeddings of the individuals in those original face images from the preprocessor 61. The embedding / diversity attribute-conditioned model here is a generative model (second generative model) that is conditioned on diversity attributes and embeddings to generate synthetic face images (second-stage synthetic images) that reflect the diversity attributes and embeddings. The embedding / diversity attribute-conditioned model can be a diffusion model, for example.

[0107] The training module 63 trains the embedding / diversity attribute-conditioned model so that when the diffusion model as the embedding / diversity attribute-conditioned model is given the diversity attributes and embeddings of an individual in an original face image as conditions, the model generates that original face image. The training module 62 delivers the trained embedding / diversity attribute-conditioned model to the generative unit 51 as another one of other learning models delivered there from the training unit 43 (Fig. 11).

[0108] <Example configuration of preprocessor 61>

[0109] Fig. 14 is a block diagram showing an example configuration of the preprocessor 61 in Fig. 13.

[0110] The preprocessor 61 includes a face recognition model training unit 71, a face recognizer 72, and an attribute recognizer 73.

[0111] Real face images with corresponding IDs delivered to the preprocessor 61 are output to each of the face recognition model training unit 71, face recognizer 72, and attribute recognizer 73.

[0112] The face recognition model training unit 71 trains the encoder as the face recognition model using the delivered IDs and real face images. For example, the face recognition model training unit 71 trains the encoder as the face recognition model such that the model outputs the embedding corresponding to the ID of an input real face image. The face recognition model training unit 71 delivers the encoder as the trained face recognition model to the face recognizer 72, as well as to the generative unit 51 as one of the learning models delivered there from the training unit 43 (Fig. 11).

[0113] The face recognizer 72 recognizes the original face images and generates embeddings of the original face images using the encoder as the face recognition model delivered from the face recognition model training unit 71. The face recognizer 72 delivers the embeddings of the original face images to the attribute recognizer 73. The face recognizer 72 also delivers the original face images to the training modules 62 and 63 (Fig. 13), and delivers the embeddings of the original face images to the training module 63.

[0114] The attribute recognizer 73 recognizes the identity attributes (attribute values) of the individuals in the respective original face images using various learning models that have been prepared in advance to recognize the identity attributes of individuals such as race and gender, and delivers the attributes to the training module 62 (Fig. 13).

[0115] The attribute recognizer 73 recognizes the diversity attributes (attribute values) except similarity of the individuals in the respective original face images using various learning models that have been prepared in advance to recognize the diversity attributes of individuals except similarity, such as age and face pose, and delivers the attributes to the training module 63 (Fig. 13). The attribute recognizer 73 recognizes (calculates) similarity between the original face images with the same ID as one of the diversity attributes using the embeddings from the face recognizer 72, and delivers the similarity to the training module 63.

[0116] A learning model such as CLIP, for example, can be used for the recognition of demographic information such as race, gender, and age of individuals in the original face images. The (attribute values of) face pose that is one of the diversity attributes can be represented, for example, by the inclination angles of the face in roll, yaw, and pitch directions relative to the face facing the front.

[0117] <Example configuration of attribute recognizer 73>

[0118] Fig. 15 is a block diagram showing an example configuration of a part of the attribute recognizer 73 in Fig. 14 that recognizes similarity as one of the diversity attributes.

[0119] In Fig. 15, the attribute recognizer 73 includes an embedding acquisition unit 81, a representative embedding calculator 82, and a similarity calculator 83.

[0120] The embedding acquisition unit 81 acquires embeddings emb[1], emb[2], ... emb[N] of original face images with the same ID from the embeddings delivered from the face recognizer 72 (Fig. 14), and delivers them to the representative embedding calculator 82 and the similarity calculator 83.

[0121] The representative embedding calculator 82 calculates an embedding of a representative one of the faces in the original face images with the same ID using the embeddings emb[1] to emb[N] of the original face images with the same ID delivered from the embedding acquisition unit 81, and delivers the calculated embedding to the similarity calculator 83.

[0122] For example, the representative embedding calculator 82 determines an average value of the embeddings emb[1] to emb[N] of the original face images with the same ID, or the embedding of an image randomly selected from the original face images with the same ID, as the embedding of the representative face. The average value of the embeddings emb[1] to emb[N] of the original face images with the same ID can for example be a value emb{mean} / |emb{mean}| obtained by normalizing the mean vectors emb{mean} = 1 / N*(emb[1] + emb[2] + ... emb[N]) as the average values of the embeddings emb[1] to emb[N] with the norm |emb{mean}| of these mean vectors emb{mean}.

[0123] If the average value of the embeddings emb[1] to emb[N] of the original face images with the same ID is used as the embedding of the representative face, the face will look like an averaged version of the faces depicted in the respective original face images with the same ID. If the embedding of an original face image randomly selected from the original face images with the same ID is used as the embedding of the representative face, the face will be the same face as that of the original face image that is the source of the embedding.

[0124] The similarity calculator 83 calculates the similarity between each of the original face images with the same ID and the representative face of the original face images with the same ID, and outputs the similarity as one of the diversity attributes of the original face images (individuals in the images). The similarity calculator 83 calculates the similarity using the embeddings of the original face images with the same ID delivered from the embedding acquisition unit 81 and the embedding of the representative face delivered from the representative embedding calculator 82. For example, the similarity calculator 83 calculates the cosine similarity between the vectors of the embeddings of the original face images and the representative face as the similarity of the original face images relative to the representative face, for each of the original face images with the same ID.

[0125] The training module 63 (Fig. 13) trains the embedding / diversity attribute-conditioned model by conditioning the model on the diversity attributes including the similarity. In the generation of synthetic face images using this embedding / diversity attribute-conditioned model, the embedding / diversity attribute-conditioned model is conditioned on the diversity attributes including the similarity, so that the model generates synthetic face images reflecting the conditioning similarity and the embeddings. Namely, the synthetic face images generated here depict faces that will be identified (recognized) as belonging to the individual whose ID corresponds to the conditioning embedding. These faces resemble the representative face of that individual to the extent determined by the conditioning similarity.

[0126] Thus the embedding / diversity attribute-conditioned model conditioned on diversity attributes including similarity as well as the embeddings can generate synthetic face images with controlled similarity to the representative face, determined by the similarity used to condition the model. It is also possible to generate synthetic face images with varying similarities relative to the representative face.

[0127] <Example configuration of generative unit 51>

[0128] Fig. 16 is a block diagram showing an example configuration of the generative unit 51 in Fig. 11.

[0129] In Fig. 16, the generative unit 51 includes a first-stage processor 91 and a second-stage processor 92.

[0130] Encoders as the identity attribute-conditioned model and the face recognition model are delivered from the training unit 43 (Fig. 11) to the first-stage processor 91. The embedding / diversity attribute-conditioned model is delivered from the training unit 43 to the second-stage processor 92.

[0131] The first-stage processor 91 performs first-stage processing for generating embeddings of synthetic face images as the first-stage synthetic images using the encoders as the identity attribute-conditioned model and the face recognition model delivered from the training unit 43.

[0132] That is, the first-stage processor 91 conditions the identity attribute-conditioned model on (predetermined values of) identity attributes to generate synthetic face images as the first-stage synthetic images depicting the individuals having the conditioning identity attributes. Further, the first-stage processor 91 generates embeddings as the features of the synthetic face images as the first-stage synthetic images using the encoder as the face recognition model. The first-stage processor 91 delivers the embeddings of the first-stage synthetic images to the second-stage processor 93.

[0133] The second-stage processor 92 performs second-stage processing for generating synthetic face images as the second-stage synthetic images using the embedding / diversity attribute-conditioned model delivered from the training unit 43 and the embeddings of the first-stage synthetic images delivered from the first-stage processor 91.

[0134] That is, the second-stage processor 92 conditions the embedding / diversity attribute-conditioned model on the embeddings of the first-stage synthetic images delivered from the first-stage processor 91 as well as the (predetermined values of) diversity attributes, to generate synthetic face images as the second-stage synthetic images depicting the individuals having the IDs corresponding to the conditioning embeddings and the conditioning diversity attributes. The second-stage processor 92 outputs the second-stage synthetic images as the final synthetic face images generated by the generative unit 51.

[0135] The generative unit 51 can thus generate fair and diverse synthetic face images through the first-stage processing pipeline by the first-stage processor 91 and the second-stage processing pipeline by the second-stage processor 92.

[0136] Namely, the first-stage processing by the first-stage processor 91 can generate evenly distributed attribute values as the identity attributes such as gender and race, for example. For example, gender attribute values can be generated to include male and female in equal proportions, and race attribute values can be generated to include different races such as white people and black people in equal proportions. All combinations of these gender and race attribute values can then be generated. By conditioning the identity attribute-conditioned model on a combination of gender and race (attribute values), synthetic face images reflecting the conditioning gender and race can be generated as the first-stage synthetic images. Namely, for example, synthetic face images depicting individuals that will be identified (recognized) as having the conditioning gender and race can be generated as the first-stage synthetic images. Therefore, the first-stage synthetic images thus generated are synthetic face images that are fair or (almost) unbiased with respect to gender and race, for example, depicting a demographically balanced diversity of individuals.

[0137] The second-stage processing by the second-stage processor 92 can generate random values or values according to user operations as the diversity attributes such as age, face pose, and similarity, for example. For example, multiple values can be randomly selected from a preset range of ages as age attribute values. Multiple values can be randomly selected from a preset range of inclination angles of the face in roll, yaw, and pitch directions as face pose attribute values. Moreover, multiple values can be randomly selected from a preset range of similarities as similarity attribute values. All combinations of these age, face pose, and similarity attribute values can then be generated. By conditioning the embedding / diversity attribute-conditioned model on the embeddings of the first-stage synthetic images as well as on a combination of age, face pose, and similarity (attribute values), synthetic face images reflecting the conditioning embeddings of the first-stage synthetic images and the conditioning combination of age, face pose, and similarity can be generated as the second-stage synthetic images. Namely, for example, synthetic face images depicting individuals that will be identified (recognized) as having the IDs corresponding to the conditioning embeddings, and having the conditioning combination of age, face pose, and similarity can be generated as the second-stage synthetic images. Therefore, similarly to the first-stage synthetic images, fair synthetic face images can be generated as the second-stage synthetic images. Moreover, the second-stage synthetic images thus generated are diverse synthetic face images with various ages, face poses, and similarities with respect to each of the individuals recognized as having the IDs corresponding to the embeddings of the first-stage synthetic images.

[0138] Accordingly, the generative unit 51 enables acquisition of diverse (face) image variations of demographically balanced, diverse individuals with different ages, face poses, and similarities.

[0139] Fig. 17 is a block diagram showing an example configuration of the first-stage processor 91 and the second-stage processor 92 in Fig. 16.

[0140] The first-stage processor 91 includes an identity attribute generator 111, a synthetic face image generator 112, a filtering unit 113, and a face recognizer 114. The first-stage processor 91 can be configured without the filtering unit 113, for example.

[0141] The identity attribute generator 111 generates attribute values of identity attributes such as race and gender and delivers the values to the synthetic face image generator 112. For example, the identity attribute generator 111 can generate uniformly distributed attribute values, equally representing the identity attributes such as gender and race. Here, uniformly distributed attribute values mean that possible values of an attribute are represented in equal proportions. For example, if the identity attribute is gender, uniformly distributed attribute values will include male and female in equal proportions. For example, if the identity attribute is race, uniformly distributed attribute values will include different races such as white people and black people in equal proportions.

[0142] The identity attribute generator 111 can also generate other (attribute values of) identity attributes according to a user operation, for example.

[0143] The identity attribute-conditioned model is delivered from the training unit 43 (Fig. 11 and Fig. 13) to the synthetic face image generator 112. The synthetic face image generator 112 generates synthetic face images by conditioning the identity attribute-conditioned model on the (attribute values of) identity attributes delivered from the identity attribute generator 111. The synthetic face image generator 112 generates synthetic face images depicting individuals with the conditioning identity attributes such as gender and race, for example. The synthetic face image generator 112 delivers the synthetic face images to the filtering unit 113. The synthetic face images delivered from the synthetic face image generator 112 to the filtering unit 113 are also referred to as first synthetic face images. The first synthetic images are also the first-stage synthetic images.

[0144] The filtering unit 113 performs a first filtering process to the first synthetic face images from the synthetic face image generator 112, and delivers the synthetic face images resulting from the first filtering process to the face recognizer 114. The synthetic face images delivered from the filtering unit 113 to the face recognizer 114 are also referred to as second synthetic face images. The second synthetic images are also the first-stage synthetic images.

[0145] The encoder as the face recognition model is delivered from the training unit 43 (Fig. 11, Fig. 13) to the face recognizer 114. The face recognizer 114 recognizes the second synthetic face images from the filtering unit 113 to generate embeddings of those second synthetic face images, and delivers the embeddings to a synthetic face image generator 122 of the second-stage processor 92.

[0146] The second-stage processor 92 includes a diversity attribute generator 121, the synthetic face image generator 122, and a filtering unit 123. The second-stage processor 92 can be configured without the filtering unit 123, for example.

[0147] The diversity attribute generator 121 generates attribute values of diversity attributes such as age, face pose, and similarity and delivers the values to the synthetic face image generator 122. For example, the diversity attribute generator 121 can generate attribute values of diversity attributes by randomly sampling from possible ranges of values representing age, face pose, and similarity.

[0148] The diversity attribute generator 121 can also generate other (attribute values of) diversity attributes according to a user operation, for example. For example, the ranges of diversity attributes for random sampling can be set by a user operation, or, values according to user operations can be generated as attribute values of diversity attributes.

[0149] The embedding / diversity attribute-conditioned model is delivered from the training unit 43 (Fig. 11 and Fig. 13) to the synthetic face image generator 122. The synthetic face image generator 122 generates synthetic face images by conditioning the embedding / diversity attribute-conditioned model on the embeddings of the second synthetic face images delivered from the face recognizer 114 as well as the (attribute values of) diversity attributes delivered from the diversity attribute generator 121. The synthetic face image generator 122 generates synthetic face images of, for example, individuals with the IDs corresponding to the conditioning embeddings of the second synthetic face images. These individuals have the conditioning diversity attributes such as age, face pose, and similarity. The synthetic face image generator 122 delivers the synthetic face images to the filtering unit 123. The synthetic face images delivered from the synthetic face image generator 122 to the filtering unit 123 are also referred to as third synthetic face images. The third synthetic images are also the second-stage synthetic images.

[0150] The filtering unit 113 performs a second filtering process to the third synthetic face images from the synthetic face image generator 122, and outputs the resultant synthetic face images obtained by the second filtering process as the final synthetic face images generated by the generative unit 51. The synthetic face images output by the filtering unit 113 are also referred to as fourth synthetic face images. The fourth synthetic images are also the second-stage synthetic images.

[0151] Fig. 18 is a block diagram showing an example configuration of the filtering unit 113 in Fig. 17.

[0152] In Fig. 18, the filtering unit 113 includes an identity attribute recognizer 141, an image selector 142, an image quality recognizer 143, an image selector 144, a face recognizer 145, an image selector 146, and a storage unit 147.

[0153] The first synthetic face images output from the synthetic face image generator 112 (Fig. 17) are delivered to the identity attribute recognizer 141. The identity attribute recognizer 141 recognizes the (attribute values of) identity attributes of the individuals depicted in the first synthetic face images delivered from the synthetic face image generator 112, using various learning models that have been prepared in advance to recognize identity attributes of individuals such as race and gender, for example. The identity attribute recognizer 141 delivers the recognition results of identity attributes of the first synthetic face images (individuals in the images) along with the first synthetic face images to the image selector 142.

[0154] The image selector 142 selects face images identified as having the correct identity attributes from the first synthetic face images delivered from the identity attribute recognizer 141 based on the recognition results of identity attributes of the first synthetic face images delivered from the identity attribute recognizer 141, and outputs the images to the image quality recognizer 143. For example, when the recognition results of identity attributes of the first synthetic face images from the identity attribute recognizer 141 match the (attribute values of) identity attributes that were used to condition the identity attribute-conditioned model when the first synthetic face images were generated by the synthetic face image generator 112 (Fig. 17), the image selector 142 selects these images as the face images identified as having the correct identity attributes.

[0155] Namely, in the first filtering process in the filtering unit 113, the identity attribute recognizer 141 and the image selector 142 delete the face images other than those that are recognized as having the correct identity attributes, i.e., those recognized as having wrong identity attributes, from the first synthetic face images. This guarantees that the first synthetic face images are the face images without any inconsistencies in the identity attributes such as gender and race, i.e., the face images that are identified (recognized) as having the correct identity attributes.

[0156] The face images that are recognized as having wrong identity attributes such as gender and race can be, for example, corrupted face images, such as an image depicting a structure that makes face recognition hard. If embeddings are generated from such corrupted images and used to condition the embedding / diversity attribute-conditioned model, the images output from the embedding / diversity attribute-conditioned model (statistical generative model) can be erroneous, e.g., corrupted, as face images. The image selector 142 selecting the first synthetic face images that are recognized as having the correct identity attributes (significantly) helps to prevent the embedding / diversity attribute-conditioned model from outputting erroneous images.

[0157] The image quality recognizer 143 recognizes the image qualities of each of the first synthetic face images delivered from the image selector 142 such as S / N (signal-to-noise ratio) or resolution, using learning models that have been prepared in advance to recognize image quality, for example. The image quality recognizer 143 delivers the (recognition results of) image qualities of the first synthetic face images along with the first synthetic face images to the image selector 144.

[0158] The image selector 144 selects high-quality face images from the first synthetic face images delivered from the image quality recognizer 143 based on the image qualities of the first synthetic face images delivered from the image quality recognizer 143, and outputs the images to the face recognizer 145. For example, the image selector 144 selects the first synthetic face images with an S / N of more than a certain threshold, as an indicator of image quality, as high-quality face images from the first synthetic face images delivered from the image quality recognizer 143.

[0159] Namely, in the first filtering process in the filtering unit 113, the image quality recognizer 143 and the image selector 144 delete the face images other than the high-quality images i.e., images of poor quality, from the first synthetic face images. This guarantees that the first synthetic face images are high-quality face images.

[0160] The image quality of the (selected) first synthetic face images can be controlled by adjusting the threshold of S / N or other image quality parameters when the first synthetic face images are selected (or deleted) in the image selector 144.

[0161] The encoder as the same face recognition model as that output from the training unit 43 (Fig. 13) to the face recognizer 114 (Fig. 17), for example, is delivered to the face recognizer 145. The face recognizer 145 recognizes the first synthetic face images delivered from the image selector 144 and generates embeddings of the respective first synthetic face images using the encoder as the face recognition model, for example. The face recognizer 145 delivers the embeddings of the first synthetic face images along with the first synthetic face images to the image selector 146.

[0162] The image selector 146 selects face images depicting individuals different from the individuals depicted in the first synthetic face images stored in the storage unit 147, from the first synthetic face images delivered from the face recognizer 145, based on the embeddings of the first synthetic face images delivered from the face recognizer 145, and delivers the images to the storage unit 147. For example, if the first synthetic face images delivered from the face recognizer 145 include images with embeddings having a cosine similarity of 0.3 or less (less than 0.3) to the embeddings of the first synthetic face images stored in the storage unit 147, the image selector 146 selects these images as the face images depicting different individuals.

[0163] Namely, in the first filtering process in the filtering unit 113, the face recognizer 145 and the image selector 146 delete the face images depicting individuals that are identified as the same as the individuals depicted in the other face images (first synthetic face images stored in the storage unit 147) from the first synthetic face images. This guarantees that there is only one face image per one ID (there are no plural face images with the same ID) in the first synthetic face images that are delivered from the image selector 146 to the storage unit 147.

[0164] The storage unit 147 stores the first synthetic face images from the image selector 146, and delivers the stored first synthetic face images to the face recognizer 114 (Fig. 17) as the second synthetic face images.

[0165] Therefore, similarly to the first synthetic face images delivered from the image selector 146 to the storage unit 147, the second synthetic face images are guaranteed to contain one each face image per one ID. Namely, the second synthetic face images delivered from the filtering unit 113 to the face recognizer 114 each depict only one individual, and no multiple face images depict the same person.

[0166] As described above, the synthetic face image generator 122 generates synthetic face images depicting individuals with the IDs corresponding to the embeddings of the second synthetic face images as third synthetic face images. Therefore, when the second synthetic face images contain multiple face images of individual A and only one face image of another individual B, the third synthetic face images will contain more images of individual A than images of individual B. The fourth synthetic face images that are the final results will also contain more face images of individual A than face images of individual B, and lack fairness.

[0167] The lack of fairness can be avoided by guaranteeing that no two face images in the second synthetic face images have the same ID as described above.

[0168] Here, if the first synthetic face images delivered from the face recognizer 145 include images with embeddings having a cosine similarity of 0.3 or less (less than 0.3) to the embeddings of the first synthetic face images stored in the storage unit 147, the image selector 146 selects these images as the face images depicting different individuals. This means that if the embeddings of two face images have a cosine similarity of 0.3 or more (more than 0.3), the individuals in these two face images will be identified (recognized) as the same person (as having the same ID).

[0169] Fig. 19 is a block diagram showing an example configuration of the filtering unit 123 in Fig. 17.

[0170] The filtering unit 123 includes a face recognizer 151, an image selector 152, and a storage unit 153.

[0171] The third synthetic face images output from the synthetic face image generator 122 (Fig. 17) are delivered to the face recognizer 151. Further, the encoder as the same face recognition model as that output from the training unit 43 (Fig. 13) to the face recognizer 114 (Fig. 17), for example, is delivered to the face recognizer 151. The face recognizer 151 recognizes the third synthetic face images delivered from synthetic face image generator 122 and generates embeddings of the third synthetic face images using the encoder as the face recognition model, for example. The face recognizer 151 delivers the embeddings of the third synthetic face images along with the third synthetic face images to the image selector 152.

[0172] The embeddings of the second synthetic face images that were used to condition the embedding / diversity attribute-conditioned model when generating the third synthetic face images to be output from the face recognizer 151 are delivered to the image selector 152. Hereinafter, the second synthetic face images from which embeddings were obtained to be used to condition the embedding / diversity attribute-conditioned model to generate the third synthetic face images shall also be referred to as “second synthetic images corresponding to the third synthetic face images.”

[0173] Based on the embeddings of the third synthetic face images from the face recognizer 145 as well as the embeddings of the second synthetic images corresponding to the third synthetic face images, the image selector 152 selects face images depicting individuals identified (recognized) as the same individuals depicted in the second synthetic face images corresponding to the third synthetic face images from the third synthetic face images delivered from the face recognizer 151, and delivers the images to the storage unit 153. For example, if the third synthetic face images delivered from the face recognizer 151 include images with embeddings having a cosine similarity of 0.3 or more (more than 0.3) to the embeddings of the second synthetic face images corresponding to the third synthetic face images, the image selector 152 selects these images as the face images depicting the same individuals as those of the corresponding second synthetic face images.

[0174] Namely, in the second filtering process in the filtering unit 123, the face recognizer 151 and the image selector 152 delete the face images depicting individuals that are identified as different from the individuals depicted in the corresponding second synthetic face images from the third synthetic face images. This guarantees that the third synthetic face images delivered from the image selector 152 to the storage unit 153 are face images of individuals identified as the same as the individuals in the corresponding second synthetic face images.

[0175] The storage unit 153 stores the third synthetic face images from the image selector 152, and outputs the stored third synthetic face images as the fourth synthetic face images.

[0176] Therefore, the fourth synthetic face images are also guaranteed to be face images of individuals identified as the same as the individuals in the corresponding second synthetic face images, similarly to the third synthetic face images delivered from the image selector 152 to the storage unit 153. Namely, the fourth synthetic face images are the face images of individuals identified as the same as the individuals in the corresponding second synthetic face images, from which the embeddings were obtained to be used to condition the embedding / diversity attribute-conditioned model when generating the fourth synthetic face images (or the third synthetic face images to become the fourth) in the synthetic face image generator 122.

[0177] <Processing in the face image generator 21>

[0178] Fig. 20 is a flowchart explaining the processing performed by the face image generator 21.

[0179] In step S11, the training unit 43 (Fig. 11) of the face image generator 21 performs preprocessing using the original face images stored in the face image DB 41 to generate an encoder (face recognition model) that is a learning model that generates embeddings of face images and recognizes the face images. Further, the training unit 43 generates embeddings of the original face images, as well as identity attributes (such as race and gender) and diversity attributes (such as age, face pose, and similarity) of the original face images. The process then proceeds to step S12.

[0180] In step S12, the training unit 43 trains a model using the original face images and identity attributes (such as race and gender), conditioning the model on identity attributes, to generate an identity attribute-conditioned model, which is a learning model that can generate synthetic face images. Further, the training unit 43 trains a model using the original face images, the embeddings of the original face images, and diversity attributes (such as age, face pose, and similarity), conditioning the model on the embeddings and diversity attributes, to generate an embedding / diversity attribute-conditioned model, which is a learning model that can generate synthetic face images. The training unit 43 delivers the encoder, identity attribute-conditioned model, and embedding / diversity attribute-conditioned model to the generative unit 51 (Fig. 11). The process then proceeds from step S12 to step S13.

[0181] In step S13, the generative unit 51 performs first-stage processing for generating the embeddings of the (second) synthetic face images using the identity attribute-conditioned model and the encoder. The process then proceeds to step S14.

[0182] In step S14, the generative unit 51 performs second-stage processing for generating the (fourth) synthetic face images using the embeddings and the embedding / diversity attribute-conditioned model generated in the first-stage processing, and the process ends.

[0183] Fig. 21 is a flowchart explaining the first-stage processing performed in step S13 of Fig. 20.

[0184] In step S21, the identity attribute generator 111 (Fig. 17) of the generative unit 51 generates a plurality of uniformly distributed attribute values equally representing the identity attributes (such as race and gender). The identity attribute generator 111 delivers the plurality of attribute values equally representing the identity attributes to the synthetic face image generator 112 (Fig. 17). The process then proceeds from step S21 to step S22.

[0185] Generating multiple uniformly distributed attribute values for identity attributes in the identity attribute generator 111 allows the downstream synthetic face image generator 112 to produce demographically balanced and fair synthetic face images, minimizing bias in (attribute values of) identity attributes such as gender and race. Moreover, the synthetic face image generator 112 can generate diverse synthetic face images, i.e., synthetic face images of diverse individuals of different genders and races.

[0186] In step S22, the synthetic face image generator 112 conditions the identity attribute-conditioned model on the attribute values of the identity attributes delivered from the identity attribute generator 111, to generate first synthetic face images reflecting the attribute values of identity attributes. The synthetic face image generator 112 delivers the first synthetic face images reflecting the respective attribute values of identity attributes to the filtering unit 113 (Fig. 17). The process then proceeds from step S22 to step S23.

[0187] In step S23, the filtering unit 113 performs the first filtering process to the first synthetic face images reflecting the respective attribute values of identity attributes from the synthetic face image generator 112 to generate second synthetic face images reflecting the attribute values of the identity attributes. The filtering unit 113 delivers the second synthetic face images reflecting the respective attribute values of identity attributes to the face recognizer 114 (Fig. 17). The process then proceeds from step S23 to step S24.

[0188] The second synthetic face images reflecting the respective attribute values of identity attributes that are generated by the first filtering process performed to the first synthetic face images in the filtering unit 113 are face images one each corresponding to one ID. This allows the fourth synthetic face images that are finally obtained in the generative unit 51 to be fair face images.

[0189] In step S24, the face recognizer 114 recognizes the faces in the second synthetic face images reflecting the attribute values of identity attribute from the filtering unit 113 using the encoder to generate embeddings of the second synthetic face images. The face recognizer 114 delivers the embeddings of the second synthetic face images reflecting the respective attribute values of the identity attributes to the synthetic face image generator 122 (Fig. 17), and the process is returned.

[0190] The synthetic face image generator 122 generates third synthetic face images by conditioning the embedding / diversity attribute-conditioned model on the embeddings of the second synthetic face images delivered from the face recognizer 114. The third synthetic face images are face images reflecting the embeddings of the second synthetic face images. The second synthetic face images are face images generated by conditioning the identity attribute-conditioned model on equally distributed attribute values of identity attributes, and are different from the original real face images. Therefore, the third synthetic face images as well as the fourth synthetic face images that are finally obtained in the generative unit 51 hardly depict the faces of the individuals depicted in the original real face images, so that privacy risks can be reduced.

[0191] Fig. 22 is a flowchart explaining the second-stage processing performed in step S14 of Fig. 20.

[0192] In step S31, the diversity attribute generator 121 (Fig. 17) of the generative unit 51 generates multiple random attribute values within preset ranges, or, ranges set by a user operation, for the diversity attributes (such as age, face pose, and similarity). The diversity attribute generator 121 delivers the multiple attribute values of the diversity attributes to the synthetic face image generator 122 (Fig. 17). The process then proceeds from step S31 to step S32.

[0193] Generating multiple different attribute values for the diversity attributes in the diversity attribute generator 121 allows the downstream synthetic face image generator 122 to produce diverse synthetic face images with different (attribute values of) diversity attributes such as age, face pose, and similarity of the same person (individuals identified as having the same ID).

[0194] In step S32, the synthetic face image generator 122 conditions the embedding / diversity attribute-conditioned model on the embeddings of the second synthetic face images obtained in the first-stage processing (Fig. 21), as well as the attribute values of diversity attributes delivered from the diversity attribute generator 121, to generate third synthetic face images reflecting the embeddings of the second synthetic face images and the (combinations of) the attribute values of diversity attributes. The synthetic face image generator 122 delivers the third synthetic face images reflecting the embeddings of the second synthetic face images and the attribute values of diversity attributes to the filtering unit 123 (Fig. 17). The process then proceeds from step S32 to step S33.

[0195] In step S33, the filtering unit 133 performs the second filtering process to the third synthetic face images reflecting the embeddings of the second synthetic face images and the attribute values of diversity attributes from the synthetic face image generator 122 to generate fourth synthetic face images reflecting the embeddings of the second synthetic face images and the attribute values of diversity attributes. The filtering unit 123 outputs the fourth synthetic face images as final synthetic face images generated by the generative unit 51. The process is returned.

[0196] <Description of computer to which present technique is applied>

[0197] The series of operations described above can be executed by hardware or software. When executing the series of operations with software, the program that implements the software is installed in a computer. Here, the computers include those incorporated in dedicated hardware, and general-purpose personal computers, for example, that can perform various functions by installing various programs.

[0198] Fig. 23 is a block diagram showing an example configuration of the hardware of a computer that executes the series of operations described above with a program.

[0199] The computer includes a processing circuit 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903 interconnected via a bus 904.

[0200] An input / output interface 905 is further connected to the bus 904. To the input / output interface 905 are connected an input unit 906, an output unit 907, a storage unit 908, a communication unit 909, and a drive 910.

[0201] The input unit 906 can include physical or virtual operating means for a user to operate to input information such as a keyboard, a mouse, or a touchscreen, and other means that allow users to input information by voice or gaze. The input unit 906 can further include sensors that acquire physical quantities of light or sound such as a camera or a microphone. The output unit 907 can include means for presenting information to the user by stimulating the user’s perceptions such as a display, a speaker, and a haptic device. The storage unit 908 can be configured by a hard disk, or a non-volatile or volatile memory to store various types of information (including programs). The communication unit 909 is a network interface and performs wired or wireless communication with external devices. The drive 910 drives removable media 911 including a magnetic disc, an optical disc, a magneto-optical disc, or a semiconductor memory.

[0202] The processing circuit 901 includes a processor such as a CPU (Central Processing Unit) or DSP (Digital Signal Processor) that executes a program. The processing circuit 901 (or the processor) loads the program stored in the storage unit 908 to the RAM 903 via the input / output interface 905 and the bus 904, and executes it, to implement the series of operations described above. The processing circuit 901 can output the results of the series of operations from the output unit 907 via the bus 904 and the input / output interface 905, for example, as required. The processing circuit 901 can also store the operation results in the storage unit 908 or transmit the results from the communication unit 909.

[0203] The program to be executed by the computer (processing circuit 901) can be recorded in a removable medium 911 and provided as a package medium. The program can also be provided through a wired or wireless transmission medium such as a local area network, Internet, and digital satellite broadcast.

[0204] The program can be installed to the storage unit 908 of the computer via the input / output interface 905 by mounting the removable medium 911 to the drive 910. The program can also be received by the communication unit 909 through a wired or wireless transmission medium and installed in the storage unit 908. Alternatively, the program can be pre-installed in the ROM 902 or the storage unit 908.

[0205] The program executed by the computer may be one that performs operations in time series in the order described herein, or in parallel, or at necessary times when called.

[0206] The operations the computer performs in accordance with the program need not necessarily be in the order of the flowchart. Namely, the operations performed by the computer in accordance with the program include parallel or individual operations (e.g., parallel operations and operations by objects).

[0207] The program can be implemented by a single computer (processor), or by multiple computers in a distributed manner. Further, the program can be transferred to and executed by a remote computer.

[0208] In the case with the face image generator 21 in Fig. 11, for example, the face image DB 41 corresponds to the storage unit 908. The training unit 43 and the generative unit 51 correspond to the processing circuit 901 (processor) that has executed a program.

[0209] A system herein refers to a single element or a collection of multiple elements (apparatuses or devices, modules, components, etc.), whether or not all of the multiple elements are contained in the same housing. It follows that multiple devices contained in separate housings and interconnected by a network, as well as a device containing multiple modules in one housing, are both systems. The computer described above in its entirety, and a combination of the computer and other devices that are not shown such as a server, are also systems. A single or multiple elements of a computer, e.g., the processing circuit 901 alone, or the combination of processing circuits 901, bus 904, and others, are also systems.

[0210] The embodiments of the present technique are not limited to the embodiments described above, and can be altered variously without departing from the subject matter of the present technique.

[0211] For example, the present technique can be implemented as cloud computing that shares a function among multiple devices over a network and performs cooperative processing.

[0212] Each of the steps in the flowcharts described above can be performed in a single device, or separately by multiple devices.

[0213] If a step contains multiple processing steps, these steps can be performed in a single device, or separately by multiple devices.

[0214] The advantages described herein are only examples and there may be other advantages.

[0215] The following configurations can be adopted in the present technique:

[0216] <1> An information processing system including a generative unit configured: to generate first-stage synthetic images by conditioning a first generative model configured to generate synthetic images on an identity attribute that is an attribute that affects identity of an individual to be depicted in an image; and to generate second-stage synthetic images by conditioning a second generative model configured to generate synthetic images on a feature of the first-stage synthetic images and a diversity attribute that is an attribute that affects diversity of a same individual. <2> The information processing system set forth in <1> further including a filtering unit configured to perform a first filtering process including a process of deleting images from the first-stage synthetic images, the images to be deleted depicting an individual identified as identical to an individual depicted in other images. <3> The information processing system set forth in <2> wherein the first filtering process further includes a process of deleting images identified as having a wrong identity attribute from the first-stage synthetic images. <4> The information processing system set forth in <2> or <3> wherein the first filtering process further includes a process of deleting images of poor quality from the first-stage synthetic images. <5> The information processing system set forth in any one of <1> to <4>, further including another filtering unit configured to perform a second filtering process of deleting images from the second-stage synthetic images, the images to be deleted depicting an individual identified as different from an individual depicted in the first-stage synthetic images with the feature used to condition the second generative model to generate the second-stage synthetic images. <6> The information processing system set forth in any one of <1> to <5>, wherein the first-stage synthetic images and the second-stage synthetic images are synthetic face images depicting human faces. <7> The information processing system set forth in <6>, wherein the identity attribute includes one or more of race and gender. <8> The information processing system set forth in <6> or <7>, wherein the diversity attribute includes an indicator indicating similarity between an individual’s face depicted in a synthetic image generated by the second-stage generative model, and a representative face that represents that the individual. <9> The information processing system set forth in any one of <6> to <8>, wherein the diversity attribute includes one or more of age and face pose. <10> The information processing system set forth in any one of <6> to <9>, further including a training unit configured to train the first generative model and the second generative model using predetermined face images depicting human faces. <11> The information processing system set forth in <10>, wherein the training unit trains an encoder configured to generate embeddings as the feature using the predetermined face images. <12> The information processing system set forth in <11>, wherein the generative unit is configured to generate embeddings as features of the first-stage synthetic images using the encoder. <13> The information processing system set forth in <11> or <12>, wherein the training unit is configured to generate embeddings of the predetermined face images using the encoder, to recognize the identity attribute and the diversity attribute of the predetermined face images, to train the first generative model using the predetermined face images and the identity attribute of the predetermined face images, and to train the second generative model using the predetermined face images, the embeddings of the predetermined face images, and the diversity attribute. <14> The information processing system set forth in any one of <10> to <13>, wherein the predetermined face images are face images depicting actual human faces. <15> The information processing system set forth in any one of <1> to <14>, wherein the generative unit is configured to generate an attribute value of the identity attribute to be used to condition the first generative model. <16> The information processing system set forth in <15>, wherein the generative unit is configured to generate a plurality of attribute values that are uniformly distributed across identity attributes. <17> The information processing system set forth in any one of <1> to <16>, wherein the generative unit is configured to generate an attribute value of the diversity attribute to be used to condition the second generative model. <18> The information processing system set forth in <17>, wherein the generative unit is configured to generate a random value or a value according to a user operation as an attribute value of the diversity attribute. <19> An information processing method including: generating first-stage synthetic images by conditioning a first generative model configured to generate synthetic images on an identity attribute that is an attribute that affects identity of an individual to be depicted in an image; and generating second-stage synthetic images by conditioning a second generative model configured to generate synthetic images on a feature of the first-stage synthetic images and a diversity attribute that is an attribute that affects diversity of a same individual. <20> A program for causing a computer to function as a generative unit configured: to generate first-stage synthetic images by conditioning a first generative model configured to generate synthetic images on an identity attribute that is an attribute that affects identity of an individual to be depicted in an image; and to generate second-stage synthetic images by conditioning a second generative model configured to generate synthetic images on a feature of the first-stage synthetic images and a diversity attribute that is an attribute that affects diversity of a same individual. <21> An information processing system comprising: processing circuitry configured to generate one or more first synthetic images using a first model conditioned on values of identity attributes, extract features from the one or more first synthetic images, and generate one or more second synthetic images using a second model conditioned on the features of the one or more first synthetic images and values of diversity attributes, wherein the one or more second synthetic images include the values of the diversity attributes while preserving identity information corresponding to the features. <22> The information processing system according to <21>, wherein the first model is conditioned on values of identity attributes selected to provide uniform representation of demographic characteristics. <23> The information processing system according to <21>, wherein the identity attributes include one or more of gender and race, and wherein the diversity attributes include one or more of age, face pose, expression, and similarity. <24> The information processing system according to <21>, wherein the features include embeddings generated by a face recognition model for identifying individuals in the one or more first synthetic images. <25> The information processing system according to <21>, wherein the processing circuitry is further configured to perform a first filtering process on the one or more first synthetic images, wherein each of the one or more filtered first synthetic images corresponds to a unique identity. <26> The information processing system according to <25>, wherein the processing circuitry is further configured to extract identity attributes of individuals depicted in the one or more first synthetic images, select images having correct identity attributes, determine image quality of the selected images, and select images having a quality greater than a predetermined threshold from the images having correct identity attributes. <27> The information processing system according to <21>, wherein the processing circuitry is further configured to perform a second filtering process on the one or more second synthetic images, wherein each of the one or more filtered second synthetic images preserves identity information corresponding to the features. <28> The information processing system according to <21>, wherein the processing circuitry is further configured to control a distribution of demographic characteristics in the one or more second synthetic images by varying the values of identity attributes used to condition the first model. <29> The information processing system according to <21>, wherein the processing circuitry is further configured to train the first model using original images and identity attributes of individuals in the original images, and train the second model using the original images, features of the original images, and diversity attributes of individuals in the original images. <30> The information processing system according to <21>, wherein the processing circuitry is further configured to generate uniformly distributed values of identity attributes. <31> The information processing system according to <21>, wherein the processing circuitry is further configured to generate random values of diversity attributes within preset ranges. <32> The information processing system according to <21>, wherein the one or more first synthetic images and the one or more second synthetic images are face images. <33> The information processing system according to <32>, wherein the one or more second synthetic images are used to train a face recognition model. <34> An information processing method comprising: generating, by processing circuitry, one or more first synthetic images using a first model conditioned on values of identity attributes; extracting, by the processing circuitry, features from the one or more first synthetic images; and generating, by the processing circuitry, one or more second synthetic images using a second model conditioned on the features of the one or more first synthetic images and values of diversity attributes, wherein the one or more second synthetic images include the values of the diversity attributes while preserving identity information corresponding to the features. <35> The information processing method according to <34>, further comprising: performing a first filtering process on the one or more first synthetic images, wherein each of the one or more filtered first synthetic images corresponds to a unique identity; and performing a second filtering process on the one or more second synthetic images, wherein each of the one or more filtered second synthetic images preserves identity information corresponding to the features. <36> The information processing method according to <34>, further comprising: training the first model using original images and identity attributes of individuals in the original images; and training the second model using the original images, features of the original images, and diversity attributes of individuals in the original images. <37> The information processing method according to <34>, further comprising: generating uniformly distributed values of identity attributes; and generating random values of diversity attributes within preset ranges. <38> The information processing method according to <34>, wherein the identity attributes include one or more of gender and race, and wherein the diversity attributes include one or more of age, face pose, expression, and similarity. <39> The information processing method according to <34>, wherein the one or more first synthetic images and the one or more second synthetic images are face images, and wherein the one or more second synthetic images are used to train a face recognition model. <40> A non-transitory computer-readable medium storing instructions which when executed by a processor cause the processor to perform operations comprising: generating one or more first synthetic images using a first model conditioned on values of identity attributes; extracting features from the one or more first synthetic images; and generating one or more second synthetic images using a second model conditioned on the features of the one or more first synthetic images and values of diversity attributes, wherein the one or more second synthetic images include the values of the diversity attributes while preserving identity information corresponding to the features.

[0217] 10 Face processing system 20 Training apparatus 21 Face image generator 22 Face image DB 23 Training face image selector 24 Face image training unit 30 Face processing apparatus 31 Face detector 32 Face processor 40 Training section 41 Face image DB 42 Acquisition unit 43 Training unit 50 Generative apparatus 51 Generative unit 61 Preprocessor 62, 63 Training module 71 Face recognition model training unit 72 Face recognizer 73 Attribute recognizer 81 Embedding acquisition unit 82 Representative embedding calculator 83 Similarity calculator 91 First-stage processor 92 Second-stage processor 111 Identity attribute generator 112 Synthetic face image generator 113 Filtering unit 114 Face recognizer 121 Diversity attribute generator 122 Synthetic face image generator 123 Filtering unit 141 Identity attribute recognizer 142 Image selector 143 Image quality recognizer 144 Image selector 145 Face recognizer 146 Image selector 147 Storage unit 151 Face recognizer 152 Image selector 153 Storage unit 901 Processing circuit 902 ROM 903 RAM 904 Bus 905 Input / output interface 906 Input unit 907 Output unit 208 Storage unit, Communication unit 210 Drive 211 Removable medium

Claims

An information processing system comprising:processing circuitry configured togenerate one or more first synthetic images using a first model conditioned on values of identity attributes,extract features from the one or more first synthetic images, andgenerate one or more second synthetic images using a second model conditioned on the features of the one or more first synthetic images and values of diversity attributes, wherein the one or more second synthetic images include the values of the diversity attributes while preserving identity information corresponding to the features.The information processing system according to claim 1, wherein the first model is conditioned on values of identity attributes selected to provide uniform representation of demographic characteristics.The information processing system according to claim 1, wherein the identity attributes include one or more of gender and race, and wherein the diversity attributes include one or more of age, face pose, expression, and similarity.The information processing system according to claim 1, wherein the features include embeddings generated by a face recognition model for identifying individuals in the one or more first synthetic images.The information processing system according to claim 1, wherein the processing circuitry is further configured toperform a first filtering process on the one or more first synthetic images, wherein each of the one or more filtered first synthetic images corresponds to a unique identity.The information processing system according to claim 5, wherein the processing circuitry is further configured toextract identity attributes of individuals depicted in the one or more first synthetic images,select images having correct identity attributes,determine image quality of the selected images, andselect images having a quality greater than a predetermined threshold from the images having correct identity attributes.The information processing system according to claim 1, wherein the processing circuitry is further configured toperform a second filtering process on the one or more second synthetic images, wherein each of the one or more filtered second synthetic images preserves identity information corresponding to the features.The information processing system according to claim 1, wherein the processing circuitry is further configured tocontrol a distribution of demographic characteristics in the one or more second synthetic images by varying the values of identity attributes used to condition the first model.The information processing system according to claim 1, wherein the processing circuitry is further configured totrain the first model using original images and identity attributes of individuals in the original images, andtrain the second model using the original images, features of the original images, and diversity attributes of individuals in the original images.The information processing system according to claim 1, wherein the processing circuitry is further configured togenerate uniformly distributed values of identity attributes.The information processing system according to claim 1, wherein the processing circuitry is further configured togenerate random values of diversity attributes within preset ranges.The information processing system according to claim 1, wherein the one or more first synthetic images and the one or more second synthetic images are face images.The information processing system according to claim 12, wherein the one or more second synthetic images are used to train a face recognition model.An information processing method comprising:generating, by processing circuitry, one or more first synthetic images using a first model conditioned on values of identity attributes;extracting, by the processing circuitry, features from the one or more first synthetic images; andgenerating, by the processing circuitry, one or more second synthetic images using a second model conditioned on the features of the one or more first synthetic images and values of diversity attributes,wherein the one or more second synthetic images include the values of the diversity attributes while preserving identity information corresponding to the features.The information processing method according to claim 14, further comprising:performing a first filtering process on the one or more first synthetic images, wherein each of the one or more filtered first synthetic images corresponds to a unique identity; andperforming a second filtering process on the one or more second synthetic images, wherein each of the one or more filtered second synthetic images preserves identity information corresponding to the features.The information processing method according to claim 14, further comprising:training the first model using original images and identity attributes of individuals in the original images; andtraining the second model using the original images, features of the original images, and diversity attributes of individuals in the original images.The information processing method according to claim 14, further comprising:generating uniformly distributed values of identity attributes; andgenerating random values of diversity attributes within preset ranges.The information processing method according to claim 14, wherein the identity attributes include one or more of gender and race, and wherein the diversity attributes include one or more of age, face pose, expression, and similarity.The information processing method according to claim 14, wherein the one or more first synthetic images and the one or more second synthetic images are face images, and wherein the one or more second synthetic images are used to train a face recognition model.A non-transitory computer-readable medium storing instructions which when executed by a processor cause the processor to perform operations comprising:generating one or more first synthetic images using a first model conditioned on values of identity attributes;extracting features from the one or more first synthetic images; andgenerating one or more second synthetic images using a second model conditioned on the features of the one or more first synthetic images and values of diversity attributes, wherein the one or more second synthetic images include the values of the diversity attributes while preserving identity information corresponding to the features.

Citation Information

Patent Citations

  • Information processing unit, information processing method and program

    JP2021082068A