Information processing device, data generation method, and program
By using the encoder and decoder of a variable autoencoder to generate facial images through latent variable transformation, the problem of low generation efficiency in existing technologies is solved, achieving efficient generation of facial images with specific attribute values while protecting privacy.
Patent Information
- Application Number
- CN202480019816.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-28
- Filing Date
- 2024-03-01
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies struggle to effectively generate large numbers of facial images with specific attribute values, especially when the learning data lacks images with target attribute values, resulting in low generation efficiency for AI models.
By using the encoder and decoder of a variable autoencoder (VAE), latent variables are transformed based on normal vectors orthogonal to the partitioning plane, the latent space is partitioned and output data is generated. The decoder then generates an output image based on the transformed latent variables, thereby achieving control over the latent variables.
It effectively generates a large number of facial images with specific attribute values, reduces statistical bias, lowers generation costs, and protects personal privacy.
Smart Images

Figure CN120883221A_ABST
Abstract
Description
Technical Field
[0001] This technology relates to information processing apparatus, data generation methods and programs, and more specifically to, for example, information processing apparatus, data generation methods and programs capable of efficiently generating large amounts of data. Background Technology
[0002] For example, as learning data for an artificial intelligence (AI) model used to learn a discrimination model (discriminator) for recognizing faces, one can request facial images (data) in which faces with various attribute values such as gender, age, and race appear.
[0003] AI models used for recognition require tens of thousands of images as learning data. However, collecting such a large number of images is difficult to tailor to the desired use cases. Furthermore, privacy laws and regulations have become increasingly stringent in recent years, making it difficult to collect large amounts of images as learning data by eliminating such laws and regulations.
[0004] Therefore, a technique has been proposed to process and synthesize the pixel values of a face image as digital data to manipulate the apparent age of a person appearing in the face image (see, for example, Patent Document 1).
[0005] In the technique described in Patent Document 1, the pixel values of an image are analyzed by principal component analysis, a feature vector contributing to age is specified, and the feature vector is added to a base image to control apparent age.
[0006] The technique described in Patent Document 1 requires numerous prior steps to control the apparent age of an original image, such as performing principal component analysis on each original image to control the apparent age and performing normalization of each part of the face. Furthermore, in the technique described in Patent Document 1, the target of apparent age control is only the face, and no control is performed on other elements outside the face (such as hair).
[0007] Reference List
[0008] Patent documents
[0009] Patent Document 1: Japanese Patent No. 4893968 Summary of the Invention
[0010] The problem to be solved by the present invention
[0011] As a method for generating a large number of facial images, there are approaches that use AI models (such as Variable Autoencoders (VAEs) or StyleGANs) as facial synthesis models (generative models) to generate facial images with various attribute values. Note that within VAEs, for example, a very deep VAE (VDVAE) is described in CHILD, Rewon's "Very deep VAEs generalize autoregressive models and can outperform them on images. arXiv preprint arXiv:2011.10650,2020". StyleGAN is described, for example, in KARRAS, Tero; LAINE, Samuli; AILA, Timo's "A style-based generator architecture for generic adversarial networks. In: Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019. pp.4401-4410".
[0012] The facial image output by the AI model as a facial synthesis model depends on the learning data used to learn the AI model. Therefore, if the (facial) image used as the learning data has little or no image along the target image (with the attribute values requested by the user, which are expected to be generated) with the attribute values (range) to be targeted, that is, if there is little or no image with attribute values similar to the target image, it is difficult to efficiently generate the target image in an AI model learned using such learning data.
[0013] For example, public facial image datasets exhibit age group bias (such as a small number of elderly people aged 60 or older or young people in their teens or younger) and racial bias (such as a small number of Japanese people). In AI models using datasets such as training data, even when generating 1000 or more images, only a few dozen facial images of young people, elderly people, and Japanese people are generated. Furthermore, it is difficult to efficiently generate a large number of target images when the target image is a young person, elderly person, or Japanese person.
[0014] In this respect, similarity applies not only to images, but also to the generation of other data.
[0015] This technology was developed in view of this situation and makes it possible to efficiently generate large amounts of data.
[0016] Solution to the problem
[0017] The information processing apparatus or program according to the present technology is an information processing apparatus comprising: a control unit, a unit for transforming latent variables based on a normal vector orthogonal to a partitioning plane, wherein the partitioning plane partitions a latent space of latent variables based on a probability distribution generated by a decoder of a variable autoencoder (VAE) through input data input to an encoder of a variable autoencoder (VAE), and a unit for causing the decoder to generate said output data based on the transformed latent variables, wherein the transformed latent variables are transformed latent variables, or a program for causing a computer to function as such an information processing apparatus.
[0018] The data generation method of this technology is a data generation method comprising: transforming latent variables based on a normal vector orthogonal to a partitioning plane, wherein the partitioning plane partitions the latent space of latent variables based on a probability distribution generated by a decoder of a variable autoencoder (VAE) through input data input to the encoder of the VAE; and enabling the decoder to generate output data based on the transformed latent variables, wherein the transformed latent variables are the transformed latent variables.
[0019] In the information processing apparatus, data generation method, and program according to the present technology, latent variables are transformed based on normal vectors orthogonal to the partitioning plane. The partitioning plane is partitioned by input data input to the encoder of the variable autoencoder (VAE) to divide the latent space of latent variables based on probability distribution generated by the decoder of the variable autoencoder of output data. In the decoder, output data is generated based on the transformed latent variables, where the transformed latent variables are the latent variables after transformation.
[0020] It should be noted that the information processing device can be a standalone device or an internal block constituting a device.
[0021] In addition, the program can be provided by transmitting it through a transmission medium or by recording it on a recording medium. Attached Figure Description
[0022] Figure 1 This is a block diagram illustrating a configuration example of an implementation of a data generation apparatus that applies the present technology.
[0023] Figure 2 This is a block diagram showing a configuration example of encoder 21 and decoder 24.
[0024] Figure 3 This is a block diagram showing a configuration example of residual block 31.
[0025] Figure 4This is a block diagram showing a configuration instance of top-down block 41.
[0026] Figure 5 This is a diagram showing an overview of the first generation control for the output image.
[0027] Figure 6 This is a diagram showing the generated instances of the generation probability distribution q[new] in the first generation control.
[0028] Figure 7 This is a diagram illustrating an example of setting weight α through qualitative evaluation.
[0029] Figure 8 This is a diagram illustrating an application layer instance that generates a probability distribution q[new] through qualitative evaluation settings.
[0030] Figure 9 This is a diagram illustrating an example of an output image generated by a decoder 24 that has already performed the first generation control on it.
[0031] Figure 10 This is a diagram showing another instance of the output image generated by the decoder 24, which has already performed the first generation control on it.
[0032] Figure 11 This is a diagram illustrating an example of the distribution of feature quantities of an input image and an output image generated using the input image through a first generation control.
[0033] Figure 12 This is a flowchart illustrating an example of the first generation control process performed by the attribute / ID control unit 23.
[0034] Figure 13 This is a diagram showing an overview of the second generation control for the output image.
[0035] Figure 14 This is a diagram illustrating an example of the transformation of the latent variable z in the second generation control.
[0036] Figure 15 This is a diagram illustrating an instance of the attribute of the contribution vector β.nv.
[0037] Figure 16 This is a flowchart illustrating an example of the process by which the attribute / ID control unit 23 generates the normal vector nv.
[0038] Figure 17 This is a diagram illustrating an example of setting a threshold for separation accuracy through qualitative evaluation.
[0039] Figure 18This is a diagram illustrating an example of an output image generated by a decoder 24 on which a second generation control has been performed.
[0040] Figure 19 This is a diagram illustrating another instance of the output image generated by the decoder 24, on which the second generation control has already been performed.
[0041] Figure 20 This is a flowchart illustrating an example of the second generation control process performed by the attribute / ID control unit 23.
[0042] Figure 21 This is a flowchart illustrating an example of the processing of the latent variable zt during the normalization transformation performed in step S34.
[0043] Figure 22 This is a diagram showing an example of the output image with and without standardization of the latent variable zt after the transformation is performed.
[0044] Figure 23 This is a diagram showing an instance of the UI generated by the UI processing unit 26.
[0045] Figure 24 This is a diagram showing another instance of the UI generated by the UI processing unit 26.
[0046] Figure 25 This is a block diagram illustrating a configuration example of an implementation of the present technology in a computer. Detailed Implementation
[0047] <An embodiment of the data generation device applying this technology>
[0048] Figure 1 This is a block diagram illustrating a configuration example of an implementation of a data generation apparatus that applies the present technology.
[0049] The data generation device 11 generates and outputs an output image as output data in response to the input image as input data.
[0050] In the data generation apparatus 11, for example, a real face image obtained by photographing an actual face, or a virtual face image in which a virtual face generated by an arbitrary generation model appears, is input as a seed image (which will serve as the seed for image generation), and the seed image is input as an input image. The data generation apparatus 11 generates a virtual face image as an output image, which has attributes (attribute values) different from those of the face appearing in the seed image that is the input image. Therefore, according to the data generation apparatus 11, an output image can be obtained as a facial image dataset where privacy is protected and personal information is free.
[0051] The data generation apparatus 11 controls (operates) the generation parameters obtained from pre-prepared seed images (input images) and generates a large number of various images, such as images where the attribute values of the desired attribute have expected values or images with various values from a small number of seed images. Therefore, images with a large amount of data can be generated efficiently, and a large number of images can be enlarged. Note that an image where the attribute values of the desired attribute have expected values is, for example, an image where the race (attribute value) is Japanese. An image where the attribute values of the desired attribute have various values is, for example, an image where the age (attribute value) ranges from young to old. Furthermore, the control of the generation parameters includes not only direct control of the generation parameters but also processing that controls the generation parameters as a result.
[0052] Generative parameters potentially include facial features such as the eyes, skin color, and hairstyle, particularly features related to the appearance of the face. By feeding these generative parameters to a generative model, a facial image can be obtained as the output of the generative model. These generative parameters are also known as latent variables.
[0053] The data generation device 11 labels and cleans a large number of facial images obtained by controlling potential variables as generation parameters, thereby generating a facial image dataset that includes all or part of the large number of facial images.
[0054] In the tag, attribute values for each attribute of the face image (the face appearing in the image) and a face ID (identifier) are assigned to the face image. The tag automatically assigns attribute values and face IDs without manual intervention.
[0055] For facial images, attributes are items related to facial features (characteristics of facial appearance) appearing in the facial image, such as age, gender, facial expression, eye shape, hair color, and ethnicity. A face ID is an ID used to identify the individual (person) whose face appears in a facial image. The same face ID is assigned to multiple facial images identified as the same person, and different face IDs are assigned to multiple facial images identified as different people.
[0056] During cleaning, the statistical values (distribution) of the attribute values of the facial images constituting the facial image dataset, the numerical values of the facial IDs, the resolution of the facial images, and the number of facial images constituting the facial image dataset (the number of data entries) are set to desired values. For example, facial image thinning (i.e., local facial image deletion), downsampling or upsampling of facial images, etc., are performed.
[0057] By performing magnification, labeling, and cleanup of facial images (seed images), it is possible to balance the statistical biases of attributes such as those in facial image datasets. Furthermore, facial image datasets can be constructed at low cost according to the quantity, resolution, attributes (values), and statistical values required by clients such as AI developers.
[0058] In facial image magnification, various facial images with diverse attributes (attribute values) and facial IDs can be generated by controlling latent variables rather than editing the facial image itself. Furthermore, by changing (controlling) the latent variables in the direction in which the attribute values of desired attributes of the facial image change within the latent space of the latent variables, it is easy to generate facial images with attribute values that tend to be insufficient when the face is actually captured, or facial images along attribute values or statistical values expected by clients such as AI developers. Therefore, the cost of constructing a desired facial image dataset can be kept relatively low.
[0059] In the data generation device 11, virtual face images, real face images, etc., provided by the customer can be used as seed images (input images) for magnification.
[0060] The requirements for the facial image dataset requested by the client (hereinafter also referred to as dataset requirements), such as the amount of data, the number of people (face ID numbers), resolution, attributes, attribute statistics (distribution of attribute values for each attribute), etc., can be set by operating the data generation device 11. Furthermore, for example, the dataset requirements can be set in the data generation device 11 based on a file storing dataset requirements.
[0061] The final facial image dataset file may include facial images, an index file in which facial IDs and attribute labels are associated with facial images, thumbnails (images) of facial images, and attribute statistics data files indicating attribute statistics for each attribute.
[0062] In the data generation apparatus 11, the generated facial images (output images) are clustered by recognizing human identifiers or attributes, enabling labeling, assigning facial IDs and attribute values, and calculating attribute statistics (statistics of attribute values). Furthermore, cleanup is performed so that the attribute statistics of the output images are equal to the attribute statistics requested by the client as a dataset, and a final facial image dataset is generated.
[0063] The data generation apparatus 11 includes an encoder 21, a scrambling unit 22, an attribute / ID control unit 23, a decoder 24, an annotation unit 25, and a user interface (UI) processing unit 26. Seed images and dataset requirements are input to the data generation apparatus 11, and the data generation apparatus 11 generates a facial image dataset based on the input seed images and dataset requirements. For example, the data generation apparatus 11 can construct a facial image dataset comprising approximately 50,000 facial images from approximately several hundred seed images.
[0064] Multiple seed images, prepared in advance as input images, are provided to encoder 21. Encoder 21, for example, is an autoencoder (AE), which extracts feature values from the seed images and provides these feature values to scrambling unit 22.
[0065] The scrambling unit 22 scrambles the feature values of the seed image from the encoder 21 and provides the scrambled feature values to the attribute / ID control unit 23. Scrambling is a process that makes the feature values of the seed image irrecoverable, and in the scrambling process, for example, arbitrary random noise, feature values extracted from a pre-prepared image independent of the seed image, etc., are added to the feature values of the seed image. By scrambling, the privacy protection of real people can be enhanced when real people appear in the seed image.
[0066] The scrambling of feature quantities in scrambling unit 22 can be performed as needed. Furthermore, the dataset requirements requested by the customer can be provided to scrambling unit 22. As mentioned above, the dataset requirements are, for example, the amount of facial images constituting the facial image dataset specified by the customer, the resolution of the facial images, attributes such as age to be labeled, attribute statistics indicating the distribution of attribute values for each attribute, etc. Scrambling unit 22 can scramble the feature quantities based on the dataset requirements.
[0067] The attribute / ID control unit 23 performs generation control to generate an output image of a face image generated by the decoder 24, based on information obtained from the feature quantity of the seed image from the scrambling unit 22, etc.
[0068] For example, in the generation control, the attribute / ID control unit 23 generates a large number of output images in which faces recognizable by others appear, based on information obtained from features such as those of the seed image (e.g., probability distribution q, described later), while maintaining the attributes (attribute values) of the seed image. The presence of a large number of output images with faces recognizable by others implies images that can be assigned different face IDs.
[0069] Furthermore, for example, in the generation control, the attribute / ID control unit 23 generates a large number of latent variables for the output image based on information obtained from features such as those of the species image (e.g., contribution vectors, which will be described later), in which the attribute values of specific attributes of the species image (e.g., age (appearance age)) are changed.
[0070] Then, in the generation control, the attribute / ID control unit 23 causes the decoder 24 to generate the output image based on the latent variables of each output image.
[0071] Data set requirements can be provided to the attribute / ID control unit 23. The attribute / ID control unit 23 can then perform generation control based on the data set requirements.
[0072] Decoder 24 is a generative model that outputs an output image (output data) in response to the input image (input data) to encoder 21, and is, for example, a decoder for an image of a photoelectrode (AE). Decoder 24 can generate and output the output image based on the latent variables generated by attribute / ID control unit 23, according to the generation control of attribute / ID control unit 23. The output image output from decoder 24 is provided to annotation unit 25. It should be noted that decoder 24, as a generative model, can generate the output image regardless of the input image to encoder 21. In this case, the output image generated by decoder 24 is an image within the range of the learning data used by decoder 24 for learning.
[0073] The dataset requirements are provided to the annotation unit 25. The annotation unit 25 generates (constructs) a facial image dataset that meets the dataset requirements based on the dataset requirements and the output image from the decoder 24, and outputs the facial image dataset to subsequent stages such as the recording unit (not shown).
[0074] That is, annotation unit 25 marks and cleans the facial image dataset (facial images) relative to attributes and facial IDs, and generates a final facial image dataset that meets the dataset requirements.
[0075] In annotation unit 25, for example, by using existing arbitrary clustering methods, similarity calculation methods, discriminators obtained through pre-learning, etc., labeling and cleaning can be performed without the need for user intervention (i.e., without manual intervention). Users are those who use data generation device 11, and also include operators, customers, etc., of data generation device 11.
[0076] Cleaning is a bias removal process used to remove statistical data biases, such as those arising from multiple output images sharing the same attribute (value). For example, in cleaning, some output images that constitute a facial image dataset are deleted (removed), thereby setting the data volume, attribute statistics, and number of people to the dataset requirements, and adjusting these parameters. In adjusting the number of people, output images assigned the same face ID are deleted.
[0077] The UI processing unit 26 serves as a UI generation unit. It generates images (including information to be provided to the user, buttons to be operated by the user, etc.) as a user interface (GUI) as needed, and displays the UI on a display unit (not shown) to present the UI to the user. Furthermore, the UI processing unit 26 executes various settings and controls the necessary blocks constituting the data generation device 11 based on the user's operations on the UI.
[0078] It should be noted that, for example, when the data generation device 11 is configured as a server-client system, the server can be used as the data generation device 11. In this case, the UI processing unit 26 can transmit the UI to the client for display.
[0079] <Configuration Example of Encoder 21 and Decoder 24>
[0080] Figure 2 This is a block diagram showing a configuration example of encoder 21 and decoder 24.
[0081] Right now, Figure 2 An example configuration of encoder 21 and decoder 24 is shown in the case where a very deep VAE (VDVAE) (decoder) is used as the generative model for generating the output image.
[0082] exist Figure 2 In the diagram, encoder 21 and decoder 24 are the encoder and decoder of VDVAE, respectively. Note that... Figure 2 And will be described later Figure 3 and 4 This is a reference from CHILD and Rewon. A very deep VAE generalization encompasses the autoregressive model and, in the case of the encoder in the "graph output image generation model," it can outperform them in the display unit (not shown) in deletion (removal).
[0083] The encoder 21 includes 66 residual blocks (res blocks) 31 as multiple layers, and a pooling layer (pool) 32 inserted for each of the several residual blocks 31. In the encoder 21, the input image is fed into the lowest residual block 31 in the graph, and the data flows along a path extending from bottom to top (bottom-up path).
[0084] Residual block 31 receives the output of the following residual block 31 or pooling layer 32 as input, extracts (generates) feature values (activations) of the input image from this input, and outputs the feature values (activations) to the following residual block 31 or pooling layer 32. Note that the lowest residual block 31 extracts the feature values of the input image from the input image.
[0085] Pooling layer 32 performs pooling on the output (features of the input image) of the residual block 31 immediately below it, and outputs it to the residual block 31 immediately above it.
[0086] Decoder 24 includes 66 top-down blocks 41 identical to encoder 21, and an unpooling layer 42 inserted for each of the top-down blocks 41. The unpooling layer 42 is inserted at positions corresponding to the pooling layers 32 of encoder 21. In decoder 24, data flows along a top-down path.
[0087] In encoder 21, within residual block 31, the residual block 31 immediately below pooling layer 32 and the topmost residual block 31 also output the feature quantity of the input image from the top-down block 41 at the position corresponding to residual block 31 to the top-down block 41 immediately above the next depooling layer 42 in the downward direction to the top-down block 41.
[0088] The top-down block 41 takes the output of the top-down block 41 immediately above it or the depooling layer 42 as input, and uses the feature values of the input image output by the residual block 31 as input as needed to generate the feature values (features) of the output image, and outputs the feature values to the top-down block 41 immediately below it or the depooling layer 42. Note that the bottommost top-down block 41 outputs the output image.
[0089] Depooling layer 42 performs depooling on the output (features of the output image) of the top-down block 41 immediately above it, and outputs the depooled output to the top-down block 41 immediately below it.
[0090] In the top-down block 41, when the output of the top-down block 41 immediately above or the depooling layer 42 and the feature quantity of the input image output by the residual block 31 are used as inputs to generate the feature quantity of the output image, the output image is a reconstructed image obtained by recovering the input image.
[0091] Here, in Figure 2 In the accompanying drawings, descriptions of some layers in the 66 layers of residual block 31 and top-down block 41 are omitted. Furthermore, regarding the layers of residual block 31 and top-down block 41, in the accompanying drawings, the bottom layer is the upper layer, and the top layer is the lower layer.
[0092] Figure 3 This is a block diagram showing a configuration example of residual block 31.
[0093] The residual block 31 includes convolutional layers 51 to 54 and an additive unit 55.
[0094] The input image, the output of the preceding (upper-layer) residual block 31, or the output of the preceding pooling layer 32 is provided to convolutional layer 51 as input to residual block 31. Convolutional layer 51 performs a 1-input convolution (kernel application) on the input of residual block 31 and outputs the result to convolutional layer 52. Convolutional layer 52 performs a 3-output convolution on the output of convolutional layer 51 and outputs the result to convolutional layer 53. Convolutional layer 53 performs a 3-output convolution on the output of convolutional layer 52 and outputs the result to convolutional layer 54. Convolutional layer 54 performs a 1-output convolution on the output of convolutional layer 53 and outputs the result to addition unit 55.
[0095] In addition to the output of convolutional layer 54, the input of residual block 31 is provided to addition unit 55. Addition unit 55 adds the input of residual block 31 and the output of convolutional layer 54, and outputs the sum as a feature (activation) of the input image.
[0096] Figure 4 This is a block diagram showing a configuration instance of block 41 from top to bottom.
[0097] From top to bottom, block 41 includes a concat 61, convolutional layers 62 to 70, addition units 71 and 72, and a residual block 73.
[0098] A fixed value, the output of the preceding (lower) top-down block 41, or the output of the preceding pooling layer 42 is provided to the connection unit 61 as the input to the top-down block 41. Furthermore, the connection unit 61 is provided with feature values (activations) of the input image output from the residual block 31 of the encoder 21. The connection unit 61 connects the input to the feature values of the top-down block 41 and the input image output from the residual block 31, and outputs the result to the convolutional layer 62.
[0099] Convolutional layer 62 performs a single-row out-convolution on the output of connection layer 61 and outputs the result to convolutional layer 63. Convolutional layer 63 performs a triple-row out-convolution on the output of convolutional layer 62 and outputs the result to convolutional layer 64. Convolutional layer 64 performs a triple-row out-convolution on the output of convolutional layer 63 and outputs the result to convolutional layer 65. Convolutional layer 65 performs a single-row out-convolution on the output of convolutional layer 64 and outputs the result.
[0100] Here, in top-down block 41, parameters defining a predetermined probability distribution (first probability distribution) q are generated based on the output of convolutional layer 65.
[0101] A probability distribution q is the probability distribution of the latent variable z used to generate the image (data) obtained by reconstructing the input image (input data) as the output image (output data). Any probability distribution can be used as the probability distribution q. For example, when using a normal distribution as the probability distribution q, the parameters defining the probability distribution q are the mean and standard deviation (or variance).
[0102] The input of top-down block 41 is provided to convolutional layer 66. Convolutional layer 66 performs a 1-input convolution on the input of top-down block 41 and outputs the result to convolutional layer 67. Convolutional layer 67 performs a 3-output convolution on the output of convolutional layer 66 and outputs the result to convolutional layer 68. Convolutional layer 68 performs a 3-output convolution on the output of convolutional layer 67 and outputs the result to convolutional layer 69. Convolutional layer 69 performs a 1-output convolution on the output of convolutional layer 68 and outputs the result to addition unit 71.
[0103] Here, in top-down block 41, parameters defining a predetermined probability distribution (second probability distribution) p are generated based on the output of convolutional layer 69.
[0104] The probability distribution p is the probability distribution along the latent variable z of the learning data used by the decoder 24 (and encoder 21) as a generative model (the probability distribution followed by the latent variable z of the learning data). Any probability distribution can be used as the probability distribution p. For example, in the case of using a normal distribution as the probability distribution p, the parameters defining the probability distribution p are the mean and standard deviation.
[0105] The latent variable z is provided to the convolutional layer 70. In VDVAE, when the image obtained by reconstructing the input image is generated as the output image, the top-down block 41 samples the latent variable z based on the probability distribution q and supplies the latent variable z to the convolutional layer 70. On the other hand, when a random image is generated as the output image along the range of the learning data, the top-down block 41 samples the latent variable z based on the probability distribution p and supplies the latent variable z to the convolutional layer 70.
[0106] Convolutional layer 70 performs a 1-row convolution on the latent variable z and outputs the result to addition unit 72.
[0107] In addition to the output of convolutional layer 69, the input of top-down block 41 is provided to addition unit 71. Addition unit 71 adds the output of convolutional layer 69 to the input of top-down block 41 and outputs the result to addition unit 72. Addition unit 72 adds the output of addition unit 71 to the output of convolutional layer 70 and provides the result to residual block 73.
[0108] Similar to Figure 3 The residual block 73 is configured in the residual block 31. The residual block 73 uses the output of the addition unit 72 as the input to perform a process similar to that of the residual block 31, and outputs the processing result as a feature quantity (feature) of the output image.
[0109] The attribute / ID control unit 23 can execute a first generation control and / or a second generation control, which are configured as described above, to generate the output image. The first generation control and the second generation control will be described below.
[0110] <First Generation Control>
[0111] Figure 5 This is a diagram showing an overview of the first generation control for the output image.
[0112] In the first generation control, the attribute / ID control unit 23 generates a generation probability distribution q[new] as a new probability distribution based on probability distributions q and p. Then, the attribute / ID control unit 23 generates a latent variable z for generating the output image based on the generation probability distribution q[new], and causes the decoder 24 to generate the output image based on the latent variable z (performing facial imaging of the latent variable z).
[0113] Here, as seen Figure 4 As described, probability distribution q is information obtained from the feature quantities (activations) of the input image obtained by encoder 21 and the feature quantities (features) of the output image obtained by decoder 24. Probability distribution p is information obtained from the feature quantities (features) of the output image obtained by decoder 24.
[0114] As described above, in the first generation control, the output image is generated based on the latent variable z generated according to the generation probability distribution q[new] generated based on probability distributions q and p. As a result, a large number of output images can be efficiently generated from a small number of input images (seed images), in which faces that retain the properties of the input images appear and can be identified as other people, and which are difficult to generate using only probability distributions q or p.
[0115] Furthermore, since the output image is generated based on the latent variable z rather than on image processing such as editing and compositing of the input image, an image showing a natural face can be generated as the output image.
[0116] Figure 6 This is a diagram showing the generated instances of the generation probability distribution q[new] in the first generation control.
[0117] The generated probability distribution q[new] can be any probability distribution, but here we use a normal distribution. In this case, the generated probability distribution q[new] is defined by the mean and standard deviation.
[0118] The attribute / ID control unit 23 can generate a normal distribution as the generated probability distribution q[new], wherein the weighted sum of the mean qm of the probability distribution q and the mean pm based on the probability distribution p is the mean qm[new], and the weighted sum of the standard deviation qv of the probability distribution q and the standard deviation pv based on the probability distribution p is the standard deviation qv[new].
[0119] Here, the value of the mean pm based on the probability distribution p is represented by the function f(pm) with the mean pm as a parameter, and the value of the standard deviation pv based on the probability distribution p is represented by the function g(pv) with the standard deviation pv as a parameter.
[0120] In this case, the mean qm[new] of the generated probability distribution q[new] is represented by the expression qm[new] = α·pm + (1-α)·f(pm), and the standard deviation qv[new] is represented by the expression qv[new] = α·pv + (1-α)·g(pv). α is a weight used to adjust the ratio of the mixture between probability distribution q and probability distribution p, and is not limited to the range of 0 to 1, and can be set to any value including negative values. Note that the case of α = 0 is essentially equivalent to the case of using only probability distribution p, and the case of α = 1 is equivalent to the case of using only probability distribution q.
[0121] As a function of the value based on the mean pm, f(pm) can be expressed as, for example, f(pm) = X to N(pm, pmstd), f(pm) = pm, etc. X to N(pm, pmstd) represents the value sampled from a normal distribution with a mean pm and a standard deviation of pmstd. The standard deviation pmstd is a variable used to adjust for the randomness of X to N(pm, pmstd) and can be set to any value.
[0122] As a function of the value based on the standard deviation pv, g(pv) can be expressed as, for example, g(pv) = X to N(pv, pvstd), g(pv) = pv, etc. X to N(pv, pvstd) represents the value sampled from a normally distributed system with a mean of pv and a standard deviation of pvstd. The standard deviation pvstd is a variable used to adjust for the randomness of X to N(pv, pvstd) and can be set to any value.
[0123] In the following text, f(pm) = X to N(pm, pmstd) is used as the function f(pm), and g(pv) = X to N(pv, pvstd) is used as the function g(pv).
[0124] The weights α and the standard deviations pmstd and pvstd can be set, for example, through a qualitative evaluation of the output image.
[0125] Figure 7 This is a diagram illustrating an example of setting weight α through qualitative evaluation.
[0126] Figure 7 Examples of input and output images as seed images are shown for each of the cases where α = 0.5, 0.7, and 0.9.
[0127] exist Figure 7 The input image features a Japanese woman.
[0128] Given an input image, an output image is generated that contains a woman who is likely to be identified as the same woman in the input image. Therefore, setting α = 0.9 is not suitable for magnifying an image of a person (another person) who is different from the input image.
[0129] With α = 0.5, the influence of images containing a large number of Europeans and Americans in the learning data of decoder 24 (and encoder 21) becomes strong, and the output image is generated where the attribute value of race changes from "Japanese" in the input image to "European and American". Therefore, setting α = 0.5 is not suitable for generating an output image in which a face identified as another person appears while preserving the attributes of the input image.
[0130] With α = 0.7, while preserving the racial attribute value as the "Japanese" attribute of the input image, an output image containing a face identified as another person is generated. Therefore, by setting α = 0.7, an output image containing a face identified as another person can be generated effectively while preserving the attributes of the input image.
[0131] According to the inventors of this application, by setting the standard deviations pmstd and pvstd to pmstd = 0.35 and pvstd = 0.75 through qualitative evaluation, it has been confirmed that output images in which faces recognized by another person appear are generated efficiently while preserving the properties of the input image.
[0132] It should be noted that the weights α and standard deviations pmstd and pvstd that are effective in generating output images showing the desired face can vary depending on the learning data used by the decoder 24, the input image, the attributes of the input image retained in the output image, etc. Therefore, it is expected that the weights α and standard deviations pmstd and pvstd should be specified and set to values that are effective for obtaining output images showing the desired face through qualitative evaluation, etc.
[0133] The output image (features) is generated using the generation probability distribution q[new]. That is, the generation of the output image based on the latent variable z generated by the generation probability distribution q[new] can be performed not on all 66 layers of the decoder 24 (top-down block 41), but on some layers.
[0134] For example, output images using the generation probability distribution q[new] can be generated only for some of the 66 layers of decoder 24 (e.g., only the upper layer, middle layer, lower layer, and other arbitrarily selected layers). The layer that generates the output image using the generation probability distribution q[new] is also called the application layer of the generation probability distribution q[new].
[0135] Which of the 66 layers of decoder 24 will be the application layers that generate the probability distribution q[new] can be set, for example, through a qualitative evaluation of the output image.
[0136] Figure 8 This is a diagram illustrating an application layer instance that generates a probability distribution q[new] through qualitative evaluation settings.
[0137] Figure 8 Examples of seed and output images as input images are shown when the application layers of the generation probability distribution q[new] are layers 10, 21, 43, and 57 in the lower layers, respectively.
[0138] It should be noted that the probability distribution p (based on which the latent variable z is generated) is used in the remaining layers of the 66th layer of decoder 24, excluding the application layer that generates the probability distribution q[new].
[0139] exist Figure 8The input image contains a Japanese woman. When all 66 layers (out of the lower layers 43, 57, or 66) are set as application layers to generate probability distribution q[new], the output images generated are not significantly influenced by the input image (probability distribution q) and do not exhibit much randomness (probability distribution p) along the range of the learning data. That is, output images with low diversity are generated, where women who might be identified as the same person as the woman in the input image appear. Therefore, setting multiple layers out of 66 as application layers to generate probability distribution q[new] is not suitable for amplifying images that are identified as people other than those appearing in the input image.
[0140] With the lower 10 layers of the 66 layers set as application layers for generating probability distribution q[new], images of Europeans and Americans, which are heavily included in the learning data of decoder 24, are strongly influenced, and the attribute value of race as an attribute changes from "Japanese" in the input image to "European and American" in the output image. Therefore, it is inappropriate to set only about 10 layers of the lower layers as application layers for generating probability distribution q[new] to generate output images in which faces identified as other people appear, while preserving the attributes of the input image.
[0141] With the lower 21 layers of the 66 layers set as application layers for generating probability distribution q[new], an output image containing a face identified as another person is generated while maintaining the ethnicity attribute value as "Japanese" of the input image. Therefore, by setting the lower approximately 21 layers as application layers for generating probability distribution q[new], output images containing faces identified as another person can be generated efficiently while preserving the attributes of the input image.
[0142] It should be noted that, depending on which layer is the application layer for generating the probability distribution q[new], the effectiveness of generating an output image in which the desired face appears can be changed based on the learning data used to learn the decoder 24, the input image, the attributes of the input image retained in the output image, etc. Therefore, for a layer to be used as the application layer for generating the probability distribution q[new], it is desirable to specify and set the layer that is effective in obtaining an output image in which the desired face appears through qualitative evaluation, etc.
[0143] Figure 9 This is a diagram illustrating an example of an output image generated by a decoder 24 on which the first generation control has been performed.
[0144] Figure 9 The input image (seed image) and output image used to generate the output image are shown.
[0145] according to Figure 9It can be confirmed that facial features such as eyes, nose, and mouth differ from those of faces appearing in the input image, and an output image is generated that includes faces of people other than those appearing in the input image. Furthermore, although Japanese people appear in the input image, it can be confirmed that an output image is generated that maintains the "Japanese" attribute value, which is an attribute of the input image. Additionally, it can be confirmed that an output image containing other people is generated.
[0146] Figure 10 This is a diagram illustrating another instance of the output image generated by the decoder 24, on which the first generation control has already been performed.
[0147] Figure 10 The input image (seed image) and the output image used to generate the output image are shown.
[0148] here, Figure 9 The output images are shown when f(pm) = X to N(pm, pmstd) is used as the function f(pm) and when g(pv) = X to N(pv, pvstd) is used as the function g(pv). On the other hand, Figure 10 The output images are shown when f(pm) = pm is used as the function f(pm) and when g(pv) = pv is used as the function g(pv).
[0149] exist Figure 10 In the case of the diversity of output images, Figure 9 The situation is slightly reduced compared to before, but it can be confirmed that the input image and the output image containing people other than those in other images are generated while maintaining the attributes of the input image (Japanese).
[0150] Figure 11 This is a diagram illustrating an example of the distribution of feature quantities of an input image and an output image generated using the input image through a first generation control.
[0151] exist Figure 11 In the diagram, the horizontal and vertical axes represent the principal components A and B of the feature quantity, respectively. These are the first and second principal components of the feature quantity obtained by performing principal component analysis on the feature quantity of the input or output image. Here, the feature quantity of the input or output image does not necessarily have to be the same as the feature quantity obtained by the encoder 21 or decoder 24.
[0152] Figure 11 A shows the distribution of the principal components A and B of the features of 200 input images (seed images). Figure 11Figure B shows the distribution of the principal components A and B of the feature quantities in the 4140 output images generated using 200 input images through the first generation control. It can be confirmed that the output images can be magnified without bias in the distribution of the principal components A and B.
[0153] Figure 12 This is a flowchart illustrating an example of the first generation control process performed by the attribute / ID control unit 23.
[0154] In step S11, the attribute / ID control unit 23 causes the decoder 24 to generate a probability distribution q based on the feature quantities of the input image from the encoder 21 (the feature quantities of the input image provided from the encoder 21 via the scrambling unit 22), etc. Furthermore, the attribute / ID control unit 23 causes the decoder 24 to generate a probability distribution p. Then, the attribute / ID control unit 23 generates the generated probability distribution q[new] based on the probability distribution q and the probability distribution p, and the process proceeds from step S11 to step S12.
[0155] In step S12, the attribute / ID control unit 23 generates a latent variable z based on the generation probability distribution q[new], and the process proceeds to step S13. That is, the attribute / ID control unit 23 generates the latent variable z by sampling the generation probability distribution q[new] of the application layer based on the generation probability distribution q[new].
[0156] In step S13, the attribute / ID control unit 23 causes the decoder 24, which serves as the generative model, to generate an output image based on the latent variable z. That is, the attribute / ID control unit 23 generates the output image (feature quantity) based on the latent variable z, which is generated based on the generation probability distribution q[new] of the application layer in the 66 layers of the decoder 24. In the decoder 24, the output image is generated based on the latent variable z in layers other than the application layer of the generation probability distribution q[new], which is generated based on the layer's probability distribution p.
[0157] The attribute / ID control unit 23 performs steps S11 to S13 on each input image, and repeats steps S12 and S13 multiple times for an input image as needed. As a result, the necessary number of output images are generated.
[0158] <Second Generation Control>
[0159] Figure 13 This is a diagram showing an overview of the second generation control for the output image.
[0160] In the second generation control, the attribute / ID control unit 23 generates a new latent variable zt based on the probability distribution-based latent variable z generated by the decoder 24 and the contribution vector. That is, the attribute / ID control unit 23 generates a transformed latent variable zt as a new latent variable by transforming the latent variable z generated by the decoder 24 based on the contribution vector. Then, the attribute / ID control unit 23 instructs the decoder 24 to generate an output image (a facial image of the latent variable zt that has undergone the transformation) based on the transformed latent variable zt. The contribution vector is a vector that contributes to a predetermined attribute (the change in the attribute value) of the input image in the latent space. For example, the contribution vector is generated based on the latent variable z, which is generated based on a probability distribution q, and the probability distribution q is generated based on the feature quantity (activation) of the input image obtained by the encoder 21.
[0161] Here, the latent variable z based on the probability distribution generated by decoder 24 refers to a latent variable z directly or indirectly generated based on one or both of the probability distributions q and p generated by decoder 24. Therefore, the latent variable z based on the probability distribution generated by decoder 24 includes not only latent variables directly generated based on probability distributions q or p generated by decoder 24, but also latent variables generated based on generated probability distributions q[new] generated based on probability distributions q and p. In the case of generating latent variable z based on probability distribution q, the probability distribution q is generated using the feature quantities (activations) of the input image obtained by encoder 21. Therefore, the latent variable z based on the probability distribution generated by decoder 24 can also be referred to as a latent variable generated based on the feature quantities of the input image.
[0162] As described above, in the second generation control, the output image is generated based on the transformed latent variable zt, which is generated by transforming the latent variable z based on the contribution vector. As a result, a large number of output images can be generated efficiently, wherein the attributes (attribute values) of the input images contributed by the contribution vector are changed from a small number of input images (seed images), and faces that can be recognized by the same person and are difficult to generate by using only probability distributions q or p appear.
[0163] Furthermore, since the output image is generated based on the latent variable z rather than on image processing such as editing and compositing of the input image, an image showing a natural face can be generated as the output image.
[0164] Figure 14 This is a diagram illustrating an example of the transformation of the latent variable z in the second generation control.
[0165] In the second generation control, the attribute / ID control unit 23 linearly transforms the latent variable z into the transformed latent variable zt, for example, based on the contribution vector β.nv. That is, the attribute / ID control unit 23 linearly transforms the latent variable z into the transformed latent variable zt, for example, according to zt = z + β.nv.
[0166] In the contribution vector β.nv, nv represents the (unit) normal vector orthogonal to the partitioning plane, which partitions the latent space of the latent variable z based on the attribute values of predetermined attributes of the input image. β is a variable representing the fluctuation width of the attribute values contributed by the contribution vector β.nv. The attribute values (predetermined attributes) contributed by the contribution vector β.nv of the output image are controlled by the value of the variable β.
[0167] Figure 15 This is a diagram illustrating an instance of the attribute of the contribution vector β.nv.
[0168] Right now, Figure 15 This shows an instance of an attribute controlled by the contribution vector β.nv.
[0169] Based on the contribution vector β.nv, attribute values such as facial orientation (Pose), age (appearance age), expression, and glasses (present or absent) can be controlled, and an output image with attribute values different from those of the input image (Original) can be generated.
[0170] Figure 16 This is a flowchart illustrating an example of the process by which the attribute / ID control unit 23 generates the normal vector nv.
[0171] In step S21, the attribute / ID control unit 23 controls the encoder 21 and decoder 24 to generate a latent variable z based on the probability distribution q according to the input image (seed image), and the processing proceeds to step S22.
[0172] In step S22, the attribute / ID control unit 23 selects the latent variable z from the latent variable z of each input image generated in step S21 to be used as the learning data and evaluation data of the linear separator, and the processing proceeds to step S23.
[0173] As a linear separator, any other linear separator such as a support vector machine (SVM) can be used. As the learning and evaluation data for the linear separator, the latent variable z of multiple upper input images and multiple lower input images can be selected, provided that the input images are sorted according to a predetermined attribute (i.e., the attribute values of the attribute to be controlled by a second generation control (hereinafter also referred to as the attention attribute)).
[0174] For example, when using age (appearance age) as the attention attribute, the latent variables z of the first 10% of the input images sorted by age (i.e., input images showing young people) and the latent variables z of the last 10% of the input images (i.e., input images showing older people) can be selected as the training and evaluation data for the linear separator. Then, for example, any 8% of the latent variables z of the first 10% of the input images sorted by age, and any 8% of the latent variables z of the last 10% of the input images sorted by age, can be selected as the training data for the linear separator. Furthermore, the remaining 2% of the latent variables z of the first 10% of the input images and the remaining 2% of the latent variables z of the last 10% of the input images can be selected as the evaluation data for the linear separator.
[0175] In step S23, the attribute / ID control unit 23 uses the latent variable z as the learning data for each of the 66 layers of the decoder 24, learns a hyperplane using an SVM as a linear separator, and the process continues to step S24. By learning the hyperplane, a hyperplane is calculated as a partitioning plane for the latent space using the attribute values of the interest attributes of each layer of the decoder 24. Here, since the interest attribute is age, a hyperplane is calculated that divides the latent space into a subspace for older people and a subspace for younger people.
[0176] In step S24, the attribute / ID control unit 23 calculates the separation accuracy of the evaluation data for each layer of the decoder 24 using a hyperplane that serves as the partitioning plane, and the process proceeds to step S25. The separation accuracy can be measured as the ratio of the number of evaluation data entries where the latent variable z of the evaluation data can be correctly separated into one of the subspaces for older adults and younger adults.
[0177] In step S25, the attribute / ID control unit 23 calculates (generates) the (unit) normal vector of the hyperplane as the age control vector for the age control of the layer in which the separation accuracy exceeds the threshold, and the processing ends.
[0178] The attribute / ID control unit 23 enables the generation of output images (features) based on the latent variable zt of the decoder 24, which has already been computed as the normal vector of the age control vector in step S25. For other layers, the output images (features) are generated based on the latent variable z generated by the decoder 24 (i.e., the latent variable z generated based on the probability distribution q, probability distribution p, or generated probability distribution q[new]).
[0179] The threshold for separation accuracy used in step S25 can be set, for example, through a qualitative evaluation of the output image.
[0180] Figure 17 This is a diagram illustrating an example of setting a threshold for separation accuracy through qualitative evaluation.
[0181] Figure 17 Examples of seed and output images as input images are shown when the separation accuracy thresholds are set to 0.5, 0.8, and 0.9, respectively, and the latent variable zt obtained by the transformation based on the latent variable z, which is the normal vector as the age control vector, is used only for layers whose separation accuracy exceeds the threshold.
[0182] exist Figure 17 The input image shows a young Japanese man.
[0183] When generating output images (feature quantities) of the latent variable zt based on the transformation for all 66 layers of decoder 24, an image with noise that appears slightly yellow overall is generated as the output image.
[0184] Therefore, instead of all 66 layers of decoder 24, the generation of the output image based on the latent variable zt of the transformation can be performed on some layers (e.g., only on the layers where the separation accuracy exceeds the threshold).
[0185] When the threshold is set to 0.5 and output images of the transformation-based latent variable zt are generated only for layers whose separation accuracy exceeds the threshold, similar to the case where output images of the transformation-based latent variable zt are generated for all 66 layers, the slightly yellow noise appears noticeable in the output images.
[0186] On the other hand, when the threshold is set above 0.8 and the output image based on the latent variable zt is generated only for layers whose separation accuracy exceeds the threshold, the slightly yellow noise appears to decrease and become less noticeable.
[0187] Therefore, for example, 0.8 can be used as the threshold for separation accuracy. The threshold for separation accuracy at which noise becomes inconspicuous in the output image can vary depending on the learning data used by the decoder 24 (and encoder 21), the input image, interest attributes, etc. Therefore, regarding the threshold for separation accuracy, it is desirable to specify and set a value that is effective for obtaining an output image where noise is inconspicuous through qualitative evaluation, etc.
[0188] Figure 18 This is a diagram illustrating an example of an output image generated by a decoder 24 on which a second generation control has been performed.
[0189] Figure 18This shows an example of the output image when age is used as the attribute of interest. Figure 18 In the image, the first row of images is the output image obtained by changing the variable β from the input image containing Japanese men, and the second row of images is the output image obtained by changing the variable β from the input image containing Japanese women.
[0190] according to Figure 18 It can be confirmed that, regardless of gender, the system generates an output image showing the face of a person whose age has significantly changed since appearing in the input image. Furthermore, in the output image, it can be confirmed that hairstyle, skin quality, and other characteristics also change with age.
[0191] Figure 19 This is a diagram illustrating another instance of the output image generated by the decoder 24, on which a second generation control has been performed.
[0192] Figure 19 This shows an example of the output image when race is used as the attribute of interest. Figure 19 In the image, the first row of images is the output image obtained by changing the variable β from the input image containing a Western woman, and the second row of images is the output image obtained by changing the variable β from the input image containing a Western man.
[0193] according to Figure 19 This confirms that an output image is generated from Western faces appearing in the input image, capturing Japanese (or Asian) faces regardless of gender. Note that in Figure 19 In this process, the facial contours and skin texture appearing in the output image tend to depend on the input image.
[0194] Figure 20 This is a flowchart illustrating an example of the second generation control process performed by the attribute / ID control unit 23.
[0195] In step S31, as Figure 16 As shown, the attribute / ID control unit 23 generates a normal vector nv as a contribution vector β.nv that contributes to the attribute of interest, and proceeds to step S32. Furthermore, as... Figure 16 As shown, in addition to the input image (seed image), any facial image (group) other than the input image can be used to generate the normal vector nv. Furthermore, if the normal vector nv has already been generated for the interest attribute in the past, in the processing of step S32 and subsequent steps, the processing of step S31 can be omitted by using the normal vector nv as a normal vector orthogonal to the partitioning plane (which partitions the latent space by the attribute values of the interest attribute of the input image).
[0196] In step S32, the attribute / ID control unit 23 sets variable β as a data setting requirement based on attribute statistics of the attribute of interest (e.g., age distribution). Furthermore, the attribute / ID control unit 23 generates a contribution vector β.nv based on variable β, and the process proceeds from step S32 to step S33.
[0197] In step S33, the attribute / ID control unit 23 transforms the latent variable z generated by the decoder 24 based on probability distribution q or p, etc., into the transformed latent variable zt according to the expression zt = z + β based on the contribution vector β.nv, and the process proceeds to step S34. The transformation from latent variable z to transformed latent variable zt is performed only for the layer that generates the normal vector nv (i.e., for the layer whose separation accuracy exceeds the threshold).
[0198] In step S34, the attribute / ID control unit 23 standardizes the potential variable zt as needed for transformation, and the process proceeds to step S35.
[0199] In step S35, the attribute / ID control unit 23 causes the decoder 24, which serves as the generative model, to generate an output image based on the latent variable zt of the transformation. That is, the attribute / ID control unit 23 generates the output image (features) based on the latent variable zt of the transformation of the layer that generates the normal vector nv among the 66 layers of the decoder 24. In the decoder 24, in layers other than the layer that generates the normal vector nv, the output image is generated based on the latent variable z, which is generated based on the probability distribution q or p of the layer.
[0200] The attribute / ID control unit 23 repeats steps S32 to S36 multiple times for each input image as needed. As a result, the required number of output images are generated, wherein the attribute values of the attributes of interest have predetermined values or various values.
[0201] Here, when probability distribution q is used as the probability distribution for generating the latent variable z to be transformed into the latent variable zt in step S33 and the probability distribution for generating the latent variable z of layers other than the layer with normal vector nv in step S35, an output image is generated in which a face appears, which can be identified as the same person as the face appearing in the input image used to generate probability distribution q, and in this face, the attribute value of the attention attribute (e.g., age, etc.) changes according to the variable β. When probability distribution p is used as the probability distribution for generating the latent variable z, an output image is generated in which a face appears along the range of the learning data of decoder 24 and has the attribute value of the attention attribute that changes according to the variable β.
[0202] Figure 21 It is shown in Figure 20The flowchart shows an instance of the processing of the latent variable zt in step S34, which is the standardization transformation performed.
[0203] The latent variable z is a vector with multiple features. In step S41, the attribute / ID control unit 23 calculates the mean and variance of the features of the vector as the latent variable z before transformation, and the process proceeds to step S42.
[0204] In step S42, the attribute / ID control unit 23 corrects the elements of the latent variable zt to be transformed based on the mean and variance of the elements of the latent variable z before the transformation, so that the mean and variance of the elements of the vector of the latent variable zt to be transformed match the mean and variance before the transformation, respectively.
[0205] Figure 22 This is a diagram showing an example of the output image with and without standardization of the latent variable zt after the transformation is performed.
[0206] In the case of using the latent variable zt of the transformation, specifically for the lower layer of layer 66 in decoder 24, when the normalization of the latent variable zt of the transformation is not performed, an image with noise that appears to have increased brightness is generated as the output image.
[0207] Therefore, for example, in the case where one or both of the lower two layers (the bottommost layer and the second layer starting from the bottommost layer) become layers that generate the normal vector nv, the latent variable zt of the transformation can be normalized for that layer. By normalizing the latent variable zt of the transformation, an output image in which noise is suppressed can be obtained, such as... Figure 22 As shown. It should be noted that the standardization of the latent variable zt can be performed not only for the lower two layers but also for the lower three or more layers.
[0208] The first generation control and the second generation control described above can be executed simultaneously. By executing the first generation control and the second generation control simultaneously, a large number of output images can be generated that are identified as other people and have various attribute values of the interest attribute. For example, when age is used as the interest attribute, a large number of output images that are identified as other people and have a wide range of ages, from young people to the elderly, can be generated.
[0209] When both the first and second generation controls are executed simultaneously, as described in the first generation control, for the application layer of the generation probability distribution q[new] in the 66 layers of the decoder 24, the output image (feature quantity) is generated based on the latent variable z generated according to the generation probability distribution q[new]. For layers other than the application layer of the generation probability distribution q[new], the output image is generated based on the latent variable z, which is generated based on the layer's probability distribution p. However, for the layer in which the normal vector nv is generated, i.e., the layer in which the separation accuracy exceeds a threshold, as described in the second generation control, the latent variable z is transformed into a transformed latent variable zt, and the output image is generated based on the transformed latent variable zt.
[0210] <ui>
[0211] Figure 23 It is shown by Figure 1 An illustration of an instance of the UI generated by the UI processing unit 26 in the diagram.
[0212] It should be noted that the data generation device 11 performs both the first generation control and the second generation control, and sets age as the attribute of interest.
[0213] Figure 23 An example of a display of the target distribution input screen 110 as a UI is shown.
[0214] The target distribution input screen 110 is a UI used to set the target values of the attribute value distribution of various attributes of the output image. For example, a user can input the target values of the age and gender distribution (age distribution and gender distribution) of faces (people) appearing in the output image (group) as the attribute values of the output image by operating on the target distribution input screen 110.
[0215] The target distribution input screen 110 includes a selection button 111, a preset name field 112, a storage destination field 113, a distribution input field 114, a storage destination field 115, a distribution display area 116, a preview button 117, an execute button 118, etc.
[0216] Use button 111 to select a preset. Presets include, for example, age and gender distributions (target values) prepared as default values, and a dataset of seed images used as input images. Figure 23 In this context, three preset items are prepared based on the assumed usage scenarios.
[0217] When a preset item is selected via the operation of selection button 111, a preset name is entered in the preset name field 112. Furthermore, the storage destination (path) of the input image file for the preset item selected via the operation of selection button 111 (hereinafter also referred to as the selected preset item) is entered in the storage destination field 113. Additionally, the age distribution and gender distribution of the selected preset item are entered into the distribution input field 114 as target values for the age distribution and gender distribution of the output image.
[0218] When the age distribution or gender distribution entered as a preset option in the distribution input field 114 does not match the age distribution or gender distribution required by the customer's dataset, the user can modify the age distribution or gender distribution entered as a preset option in the distribution input field 114 to match the age distribution or gender distribution required by the dataset. Therefore, if the age distribution and gender distribution as preset options match the age distribution and gender distribution required by the dataset respectively, the user does not need to enter (numerically) the age distribution and gender distribution as required by the dataset. Furthermore, if a portion of the age distribution or gender distribution as a preset option does not match the age distribution or gender distribution required by the dataset, the required age distribution and gender distribution can be easily entered by correcting only the mismatched portion.
[0219] The name of the selected preset is displayed in the preset name field 112.
[0220] The storage destination field 113 displays the storage destination of the input image file. Users can input the storage destination of the input image, which is prepared separately from the preset input image, into the storage destination field 113 by operating the storage destination field 113.
[0221] In the distribution input field 114, the target values for the age distribution and gender distribution can be entered by operating the distribution input field 114. As described above, when a preset item is selected by operating the selection button 111, the age distribution or gender distribution selected as the preset item is entered into the distribution input field 114 as the target value for the age distribution and gender distribution.
[0222] exist Figure 23 In the distribution, the percentage of output images (face occurrences) for each age group from teenagers to sixty years old relative to the total number of output images is input into the distribution input field 114, serving as the target value for the age distribution. The percentage of output images for each age group from teenagers to sixty years old is adjusted so that the sum is 100%.
[0223] In addition, Figure 23 In this process, the percentage of each (face) of the output images of men and women relative to the total number of output images is input as the target value for the gender distribution into the distribution input field 114. If the percentage of the number of output images of one of the men and women is changed based on user actions, the percentage of the number of output images of the other is adjusted so that the sum becomes 100%.
[0224] The storage destination of the output image file is displayed in the storage destination field 115. Users can input the storage destination of the output image into the storage destination field 115 by manipulating the storage destination field 115.
[0225] In the distribution display area 116, the target values for age and gender distributions entered in the distribution input field 114 are displayed as frequency histograms representing the number of output images for each age group and each gender. The total number of output images can, for example, be set according to user input. Figure 23 The total number of output images is 500.
[0226] It should be noted that the target values for the age and gender distributions can be changed by manipulating the distribution input field 114 and the histogram as the target values for the age and gender distributions displayed in the distribution display area 116. When the histogram as the target values for the age and gender distributions displayed in the distribution display area 116 is manipulated, the percentages in the distribution input field 114 change according to the manipulation.
[0227] With the output image preview displayed, operate the preview button 117.
[0228] When the data generation device 11 generates an output image based on the target values of age distribution and gender distribution entered in the distribution input field 114 (displayed in the distribution display area 116), the execution button 118 is operated.
[0229] When the preview button 117 or the execution button 118 is pressed, the input images stored in the storage destination displayed in the storage destination field 113 are classified into male and female input images (where each of the male and female input images appears). Then, the attribute / ID control unit 23 performs a first generation control and a second generation control using the male input images, and the decoder 24 generates various male output images (groups) of various ages based on the target values of the age distribution and gender distribution. Similarly, the first generation control and the second generation control are performed using the female input images, and various female output images (groups) of various ages based on the target values of the age distribution and gender distribution are generated. When generating the output images, in the second generation control, the latent variable z is transformed based on the target value of the age distribution in the distribution input field 114 (distribution display area 116). That is, the variable β is set based on the target value of the age distribution, and the latent variable z is transformed into the transformed latent variable zt based on the contribution vector β.nv. As a result, an output image with an age distribution that is as close as possible to the target value of the age distribution is generated.
[0230] In the data generation apparatus 11, the annotation unit 25 cleans the male and female output images generated by the decoder 24, so that the age distribution and gender distribution of the output images (approximately) match the target values of the age distribution and gender distribution in the distribution input field 114. Then, when the execution button 118 is operated, the output images obtained as a result of the cleaning are stored in a file at the storage destination displayed in the storage destination field 115.
[0231] It should be noted that if the execution button 118 is operated after the preview button 117 is operated, the output image generated in response to the operation of the preview button 177 can be stored in the file of the storage destination displayed in the storage destination field 115, without the need to generate a new output image.
[0232] Figure 24 It is shown by Figure 1 A diagram illustrating another instance of the UI generated by the UI processing unit 26 in the diagram.
[0233] Figure 24 This illustrates a display example of preview screen 120, which displays a preview of the output image during operation. Figure 23 The UI generated in the case of preview button 117.
[0234] The preview screen 120 includes a preview area 121, an age field 122, a gender checkbox 123, an update button 124, and a distributed display area 125.
[0235] When the preview button 117 is pressed, the data generation device 11 generates an output image as described above. Furthermore, the data generation device 11 generates, for example, thumbnails of some randomly selected output images. These thumbnails of the output images generated as described above are displayed in the preview area 121.
[0236] Display a number above the thumbnail that serves as the face ID. Figure 24 In the output image, only the number serving as the face ID is displayed above the thumbnail of the male image, while the number serving as the face ID is displayed in a rectangular shape above the thumbnail of the female image. Based on the display of the face ID, the gender of the person appearing in the output image (thumbnail) can be easily determined. It should be noted that, in addition, the face ID can be displayed in different colors; for example, depending on the gender of the output image, the male face ID is displayed in blue, and the female face ID is displayed in red.
[0237] In age field 122, the age range of the output images (faces appearing in the thumbnails) is displayed. Users can change the age displayed in age field 122 by manipulating it. In preview area 121, thumbnails of some output images selected from those showing faces (people) within the age range displayed in age field 122 are displayed.
[0238] When selecting the gender for the output image used to display the thumbnail, operate the gender checkbox 123. Gender checkbox 123 has a male checkbox and a female checkbox. When the male checkbox is checked, select the output image for displaying the thumbnail from the male output images. When the female checkbox is checked, select the output image for displaying the thumbnail from the female output images.
[0239] If the thumbnail displayed in the update preview area 121 is updated, operate the update button 124. If the update button 124 is operated, the output image used to display the thumbnail in the preview area 121 is reselected, and the thumbnail of the output image is displayed in the preview area 121.
[0240] In the distribution display area 125, similar to... Figure 23 In the distribution display area 116, the target values of age and gender distribution (the age and gender distributions set on the target distribution input screen 110) are displayed as histograms of the frequency of the number of output images for each age group and the number of output images for each gender. Furthermore, in the distribution display area 125, the age and gender distributions of the output images actually generated by the data generation device 11 are displayed in a similar histogram format.
[0241] In the distribution display area 125, for the age distribution, the target value (target) of the age distribution and the age distribution (output) of the output images actually generated by the data generation device 11 are displayed in the form of bar charts representing the frequencies of the same age group. Similarly, regarding the gender distribution, the target value (target) of the gender distribution and the gender distribution (output) of the output images actually generated by the data generation device 11 are displayed in the form of bar charts representing the frequencies of the same gender. That is, the number of male output images actually generated by the data generation device 11 and the target value of the number of male output images are displayed, and the number of female output images actually generated by the data generation device 11 and the target value of the number of female output images are displayed in an arranged manner.
[0242] It should be noted that in this embodiment, facial images are used as both input and output data for the data generation apparatus 11. However, as input and output data, images containing objects other than faces, or media data other than images, such as audio (speech), text (sentences), etc., can be used. Furthermore, different media data, such as images, can be used instead of the same media data as input and output data. For example, text can be used as input data, and an image with content represented by text can be used as output data.
[0243] Furthermore, in this embodiment, a VDVAE decoder is used as the generative model, but any generative model other than VDVAE can be used. For example, as the generative model performing the first and second generative controls, a VAE-based generative model other than VDVAE or a generative model other than VAE can be used. For example, as the generative model performing the first generative control, any generative model other than a VAE-based generative model can be used, wherein a first probability distribution for generating data obtained by reconstructing the input data and a second probability distribution along the learning data used for learning the generative model are obtained as output data.
[0244] Examples of techniques for editing synthetic images include interfaceGAN and encoder4editing. InterfaceGAN is described in SHEN, Yujun et al., "InterfaceGAN: Interpreting the disentangled face representation learned by GANs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020". Encoder4editing is described in TOV, Omer et al., "Designing an encoder for styleGAN image manipulation. ACM Transactions on Graphics (TOG), 2021, 40.4: 1-14". Both interfaceGAN and encoder4editing are based on StyleGAN, which is learned under the premise of image transformation. Therefore, they are difficult to directly apply to VAEs (face synthesis models) based on facial image generation, including VDVAE.
[0245] <Description of a computer using this technology>
[0246] Next, the above series of processes can be performed via hardware or software. In the case where the series of processes are performed by software, the program constituting the software is installed in a general-purpose computer or similar device.
[0247] Figure 25 This is a block diagram illustrating a configuration example of a computer equipped with a program for performing the series of processes described above.
[0248] The program can be pre-recorded in the hard disk 905 or ROM 903, which are built into the computer as recording media.
[0249] Alternatively, the program can be stored (recorded) in a removable recording medium 911 driven by a drive 909. This removable recording medium 911 can be configured as so-called encapsulated software. Examples of removable recording media 911 include, for instance, floppy disks, CD-ROMs, magneto-optical (MO) disks, DVDs, magnetic disks, semiconductor memories, etc.
[0250] It should be noted that the program can be installed on the computer from the removable recording medium 911 as described above, or it can be downloaded to the computer via a communication network or broadcast network and installed on the internal hard disk 905. That is, for example, the program can be wirelessly transmitted to the computer from the download site via a satellite used for digital satellite broadcasting, or it can be wired to the computer via a network such as a local area network (LAN) and the Internet.
[0251] The computer includes a central processing unit (CPU) 902, and an input / output interface 910 is connected to the CPU 902 via a bus 901.
[0252] When a command is input via the input / output interface 910 through the user operation input unit 907, the CPU 902 executes the program stored in the read-only memory (ROM) 903 according to the command. Optionally, the CPU 902 loads the program stored in the hard disk 905 into the random access memory (RAM) 904 and executes the program.
[0253] Therefore, CPU 902 performs processing according to the flowchart above or according to the configuration described in the block diagram above. Then, for example, as needed, CPU 902 outputs the processing result from output unit 906 through input / output interface 910, transmits the processing result from communication unit 908, or records the processing result in hard disk 905.
[0254] Note that the input unit 907 includes a keyboard, mouse, microphone, etc. Furthermore, the output unit 906 includes a liquid crystal display (LCD), speaker, etc.
[0255] Here, in this specification, the processes executed by the computer according to the program do not necessarily follow the time sequence shown in the flowchart. That is, the processes executed by the computer according to the program also include parallel or individual processes (e.g., parallel processing or processing performed by an object).
[0256] Furthermore, a program can correspond to a process executed by a single computer (processor) or a process executed in a distributed manner by multiple computers. Additionally, a program can be transmitted to a remote computer for execution.
[0257] Furthermore, in this specification, "system" refers to a group of multiple configuration elements (devices, modules (components), etc.), and it is irrelevant whether all configuration elements are housed in the same housing. Therefore, multiple devices and multiple modules housed in separate housings and interconnected via a network, and a single device housed in one housing, are all considered systems.
[0258] It should be noted that the implementation of this technology is not limited to the above-described implementation, and various changes can be made without departing from the spirit of this technology.
[0259] For example, this technology can have a cloud computing configuration, in which a function is shared and processed collaboratively by multiple devices via a network.
[0260] Furthermore, each step described in the above flowchart can be performed by a single device or can be performed by multiple devices in a shared manner.
[0261] Furthermore, in cases where a step includes multiple processes, the multiple processes included in that step can be executed by one device or by multiple devices in a shared manner.
[0262] Furthermore, the effects described in this specification are merely illustrative and not limiting, and some other effects may be achieved.
[0263] It should be noted that this technology can have the following configurations.
[0264] <1> An information processing apparatus, comprising:
[0265] Control unit, the control unit:
[0266] The latent variables are transformed based on normal vectors orthogonal to the partitioning plane, wherein the partitioning plane partitions the latent space of probability-distribution-based latent variables generated by the decoder of the variable autoencoder of the output data by the attribute values of predetermined attributes of the input data to the encoder of the variable autoencoder (VAE), and
[0267] The decoder generates output data based on the latent variables of the transformation, where the latent variables are the latent variables after the transformation.
[0268] <2> according to <1> Information processing device, wherein
[0269] A VAE is a very deep VAE (VDVAE), in which the encoder and decoder are configured by multiple layers.
[0270] <3> according to <2> Information processing device, wherein
[0271] The control unit transforms the potential variables of some of the multiple layers.
[0272] <4> according to <3> Information processing device, wherein
[0273] The control unit transforms the latent variables of the layers in which the separation accuracy of the latent variables generated by the partitioning plane exceeds a threshold.
[0274] <5> according to <4> Information processing device, wherein
[0275] The control unit standardizes the latent variables of the transformation, such that the mean and variance of the elements of the vector of each of the latent variables and the transformation latent variables are matched between the latent variables and the transformation latent variables.
[0276] <6> according to <5> Information processing device, wherein
[0277] The control unit standardizes the potential variables of the transformation of the lower layers in multiple layers.
[0278] <7> according to <1> to <6> Information processing device, wherein
[0279] The control unit transforms the latent variables based on the probability distribution of the data obtained by recovering the input data, which is used to generate the output data.
[0280] <8> according to <1> to <7> The information processing device of any one of the following, wherein,
[0281] The control unit linearly transforms the latent variables based on the normal vector.
[0282] <9> according to <1> to <8> The information processing device of any one of the following, wherein,
[0283] The control unit transforms latent variables based on the normal vectors orthogonal to the hyperplane that serves as the partitioning plane calculated by the linear separator.
[0284] <10> according to <9> Information processing device, wherein
[0285] A hyperplane is a hyperplane calculated by a linear separator using multiple higher-order and lower-order input data when the input data is classified according to the order of attribute values.
[0286] <11> according to <9> or <10> Information processing device, wherein
[0287] The linear separator is a support vector machine (SVM).
[0288] <12> according to <1> to <11> The information processing device of any one of the following, wherein,
[0289] The input and output data are images.
[0290] <13> according to <12> Information processing device, wherein
[0291] The image is a face image that appears in the image.
[0292] <14> according to <13> Information processing device, wherein
[0293] The predefined attribute is the age of the face appearing in the facial image.
[0294] The information processing apparatus further includes a user interface (UI) generation unit, which generates a UI for setting the age distribution of faces appearing in the output image as output data.
[0295] The control unit transforms latent variables based on the age distribution set in the UI.
[0296] <15> according to <14> Information processing device, wherein
[0297] The UI generation unit generates a UI that includes the age distribution set in the UI and the age distribution of faces appearing in the actual generated output image.
[0298] <16> according to <14> or <15> Information processing device, wherein
[0299] A VAE is a very deep VAE (VDVAE), in which the encoder and decoder are configured in multiple layers, and
[0300] The control unit generates the generation probability distribution for the latent variables based on the following:
[0301] A first probability distribution is used to generate the image obtained by recovering the input image as the input data, which is the output image, and...
[0302] Along the second probability distribution of the learning data used by the encoder and decoder,
[0303] Latent variables for some layers are generated based on a generation probability distribution, and latent variables for the remaining layers are generated based on a second probability distribution.
[0304] Transform the latent variables of multiple layers whose separation accuracy exceeds a threshold, based on the latent variables generated by the partition plane.
[0305] <17> A data generation method, comprising:
[0306] The latent variables are transformed based on normal vectors orthogonal to the partitioning plane, wherein the partitioning plane partitions the latent space of probability-distribution-based latent variables generated by the decoder of the variable autoencoder of the output data by attribute values of predetermined attributes of the input data to the encoder of the variable autoencoder (VAE); and
[0307] The decoder generates output data based on the latent variables of the transformation, where the latent variables are the latent variables after the transformation.
[0308] <18> A program for making a computer function as:
[0309] Control unit, the control unit:
[0310] The latent space of latent variables, based on probability distributions, is partitioned by the attribute values of predetermined attributes of the input data to the encoder of a variable autoencoder (VAE).
[0311] The decoder generates output data based on the latent variables of the transformation, where the latent variables are the latent variables after the transformation.
[0312] Reference Symbol List
[0313] 11 Data generation device
[0314] 21 Encoders
[0315] 22 scrambling units
[0316] 23 Attribute / ID Control Unit
[0317] 24 Decoders
[0318] 25 Annotation Units
[0319] 26 UI processing units
[0320] 31 Residual Block
[0321] 32 Pooling Layer
[0322] 41 Top-down blocks
[0323] 42. Remove pooling layer
[0324] Convolutional layers 51 to 54
[0325] 55 addition units
[0326] 61 to 70 convolutional layers
[0327] 71, 72 Addition Units
[0328] 73 Residual Block
[0329] 110 Target Distribution Input Screen
[0330] 111 Select button
[0331] 112 Preset Name Field
[0332] 113 Storage Destination Field
[0333] 114 Assign Input Fields
[0334] 115 Storage Destination Field
[0335] 116 Distribution Display Area
[0336] 117 Preview button
[0337] 118 Execute button
[0338] 120 Preview Screen
[0339] 121 Preview Field
[0340] Age group 122
[0341] 123 Gender checkbox
[0342] 124 Update button
[0343] 125 Distribution Display Area
[0344] 901 bus
[0345] 902CPU
[0346] 903ROM
[0347] 904 RAM
[0348] 905 hard drive
[0349] 906 Output Unit
[0350] 907 Input Unit
[0351] 908 Communication Unit
[0352] 909 drive
[0353] 910 Input / Output Interface
[0354] 911 Removable recording media.< / ui>
Claims
1. An information processing apparatus, comprising: Control unit, the control unit: The latent variables are transformed based on normal vectors orthogonal to the partitioning plane, wherein the partitioning plane partitions the latent space of probability-distribution-based latent variables generated by the decoder of the variable autoencoder of the output data by attribute values of predetermined attributes of the input data to the encoder of the variable autoencoder (VAE). The decoder generates the output data based on the latent variables of the transformation, where the latent variables are the transformed latent variables.
2. The information processing apparatus according to claim 1, wherein... The variable autoencoder is a very deep variable autoencoder (VDVAE) in which the encoder and the decoder are configured by multiple layers.
3. The information processing apparatus according to claim 2, wherein... The control unit transforms the potential variables of some of the multiple layers.
4. The information processing apparatus according to claim 3, wherein The control unit transforms the latent variables of the layers in which the separation accuracy of the latent variables generated by the partitioning plane exceeds a threshold.
5. The information processing apparatus according to claim 4, wherein The control unit standardizes the latent variables of the transformation such that the mean and variance of the elements of the vectors of each of the latent variables and the latent variables of the transformation are matched between the latent variables and the latent variables of the transformation.
6. The information processing apparatus according to claim 5, wherein The control unit standardizes the potential variables of the transformation in the lower layer of the plurality of layers.
7. The information processing apparatus according to claim 1, wherein... The control unit conversion is based on the latent variable generated by the probability distribution of the data obtained by recovering the input data, which is used to generate the output data.
8. The information processing apparatus according to claim 1, wherein The control unit linearly transforms the latent variables based on the normal vector.
9. The information processing apparatus according to claim 1, wherein The control unit transforms the latent variables based on a normal vector orthogonal to the hyperplane that is the dividing plane calculated by the linear separator.
10. The information processing apparatus according to claim 9, wherein The hyperplane is a hyperplane calculated by the linear separator using multiple higher-order input data and multiple lower-order input data, when the input data is classified according to the order of the attribute values.
11. The information processing apparatus according to claim 9, wherein The linear separator is a support vector machine (SVM).
12. The information processing apparatus according to claim 1, wherein The input data and the output data are images.
13. The information processing apparatus according to claim 12, wherein The image is a facial image in which a face appears.
14. The information processing apparatus according to claim 13, wherein The predetermined attribute is the age of the face appearing in the facial image. The information processing apparatus further includes a user interface (UI) generation unit that generates a user interface for setting the age distribution of faces appearing in the output image as output data. The control unit transforms the latent variables based on the age distribution set in the user interface.
15. The information processing apparatus according to claim 14, wherein The user interface generation unit generates a user interface that includes the age distribution set in the user interface and the age distribution of faces appearing in the actually generated output image.
16. The information processing apparatus according to claim 14, wherein The variable autoencoder is a very deep variable autoencoder (VDVAE) in which the encoder and decoder are configured in multiple layers, and The control unit generates the generation probability distribution for the latent variables based on the following: A first probability distribution is used to generate the image obtained by recovering the input image as the input data, which is the output image, and... Along the second probability distribution of the learning data learned by the encoder and the decoder, The latent variables of some of the plurality of layers are generated based on the generation probability distribution, and the latent variables of the remaining layers are generated based on the second probability distribution. The latent variables of the layers in which the separation accuracy of the latent variables generated by the partitioning plane exceeds a threshold are transformed.
17. A data generation method, comprising: The latent variables are transformed based on normal vectors orthogonal to the partitioning plane, wherein the partitioning plane partitions the latent space of probability-distribution-based latent variables generated by the decoder of the variable autoencoder of the output data by attribute values of predetermined attributes of the input data to the encoder of the variable autoencoder (VAE); and The decoder generates the output data based on the latent variables of the transformation, where the latent variables are the transformed latent variables.
18. A program for enabling a computer to function as: Control unit, the control unit: The latent variables are transformed based on the normal vector orthogonal to the partitioning plane, where, The partitioning plane partitions the latent space of latent variables based on probability distributions generated by the decoder of the variable autoencoder (VAE) based on the attribute values of predetermined attributes of the input data to the encoder of the VAE. The decoder generates the output data based on the latent variables of the transformation, where the latent variables are the transformed latent variables.
Citation Information
Patent Citations
JP1973093968A