Information processing apparatus, information processing method and program

By dividing the discriminator in a GAN into a feature extraction network and a final layer and learning them separately, the apparatus generates a distance-convertible discriminator, addressing the issue of poor data diversity in GANs and improving the accuracy and diversity of generated data.

JP2025083258APending Publication Date: 2025-05-30SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024011331
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-01-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing Generative Adversarial Networks (GANs) often suffer from biased data generation towards specific patterns or mode collapse, resulting in poor diversity of generated data.

Method used

The proposed solution involves an information processing apparatus that generates an adversarial generation network with a discriminator divided into a feature extraction network and a final layer. The model learns these components separately to evaluate the distance between the probability distributions of generated and actual feature vectors, resulting in a distance-convertible discriminator.

Benefits of technology

This approach enhances the diversity of data generated by the adversarial generation network by effectively evaluating and reducing the distance between the model and target probability distributions, thereby improving the accuracy and diversity of generated data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025083258000001_ABST
    Figure 2025083258000001_ABST
Patent Text Reader

Abstract

To improve the diversity of data generated by a generator of an adversarial generating network.SOLUTION: An information processing apparatus is provided with a model generation unit for generating an adversarial generation network including a discriminator and a generator. The model generation unit divides the discriminator into a feature extraction network for generating feature vectors of data from the data inputted to the discriminator and a final layer for dropping feature vectors distributed in a feature vector space into a one-dimensional space, and causes the feature extraction network and the final layer to learn, respectively, thereby generating a distance determinable discriminator that is a discriminator capable of evaluating the distance between the probability distribution of a generated feature vector that is a feature vector of generated data generated by the generator and the probability distribution of a real feature vector that is the feature vector of real data included in a training data set.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] Conventionally, technology related to generative adversarial networks (GANs) is known. The purpose of a generative adversarial network is to learn a target probability distribution that follows data generated by a neural network called a generator. To achieve this goal, a neural network called a discriminator is introduced, and the generator and discriminator are optimized using the minimax method. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio, "Generative adversarial nets", Neural Information Processing Systems (NeurIPS), 2014 / 6, vol.27, pp.2672‐2680. Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the above-mentioned conventional techniques, the data generated by the generator of the generative adversarial network may be biased toward a specific pattern, or the generator may generate a lot of similar data (also known as mode collapse). In other words, in the above-mentioned conventional techniques, the diversity of the data generated by the generator of the generative adversarial network may be poor.

[0005] Therefore, the present disclosure proposes an information processing device, an information processing method, and a program that can improve the diversity of data generated by a generator of a generative adversarial network. [Means for solving the problem]

[0006] The information processing device of the present disclosure is an information processing device equipped with a model generation unit that generates a generative adversarial network including a classifier and a generator, in which the model generation unit divides the classifier into a feature extraction network that generates feature vectors of data from data input to the classifier, and a final layer that reduces the feature vectors distributed in a feature vector space to a one-dimensional space, and by training the feature extraction network and the final layer, respectively, generates a distance-determinable classifier that is the classifier that can evaluate the distance between the probability distribution of generated feature vectors, which are feature vectors of generated data generated by the generator, and the probability distribution of actual feature vectors, which are feature vectors of actual data included in a training dataset. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a diagram illustrating a generative adversarial network according to a conventional technique. [Figure 2] FIG. 1 is a diagram illustrating a generative adversarial network according to a conventional technique. [Figure 3] FIG. 1 is a diagram for explaining the difference between a generative adversarial network according to a conventional technique and a generative adversarial network according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is a diagram illustrating a configuration example of an information processing device according to an embodiment of the present disclosure. [Figure 5] FIG. 1 is a diagram for explaining a generative adversarial network according to an embodiment of the present disclosure. [Figure 6] FIG. 1 is a diagram for explaining a generative adversarial network according to an embodiment of the present disclosure. [Figure 7]FIG. 2 is a diagram for explaining three conditions that a distance-discriminator according to an embodiment of the present disclosure must satisfy. [Figure 8] FIG. 2 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of an information processing device according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0009] (Embodiment) (1. Introduction) In an embodiment of the present disclosure, a case will be described in which a generative adversarial network generates image data. Note that the generative adversarial network according to an embodiment of the present disclosure is not limited to image data, and may generate data other than image data, such as text data or audio data. Furthermore, the generative adversarial network according to an embodiment of the present disclosure is not limited to the original GAN ​​presented in the paper "Generative Adversarial Nets" by Ian Goodfellow et al. in 2014, but may also be a derivative of various GANs, such as DCGAN (Deep Convolutional GAN), CycleGAN, or StyleGAN.

[0010] In general, a generative adversarial network is a machine learning model that generates pseudo-data similar to actual data (hereinafter sometimes referred to as "actual data") provided as training data. Specifically, a generative adversarial network is composed of two neural networks called a generator and a classifier. A generative adversarial network is generated by training the generator and the classifier to compete with each other (adversarial training). In this way, training the generator and the classifier to compete with each other corresponds to solving a minimax problem of a loss function common to the generator and the classifier. Below, a training method for a generative adversarial network according to the conventional technology will be specifically described using Figures 1 and 2.

[0011] FIG. 1 is a diagram illustrating a generative adversarial network according to the prior art. In FIG. 1, a generator 100 of the generative adversarial network generates image data 4 of a person from random noise 2 sampled from a normal distribution 3 (which may be a uniform distribution). FIG. 1 shows image data of nine different people generated by the generator 100. FIG. 1 also shows image data of nine different people as image data 5 included in a training dataset 200. In practice, the generator 100 learns the features of a large amount of image data, for example, 100,000 images.

[0012] In FIG. 1, the generator 100 assumes that the image data 5 included in the training dataset 200 is generated based on some probability distribution. Hereinafter, the probability distribution that is the source of generating the image data 5 included in the training dataset 200 will be referred to as a target probability distribution μ. In other words, the generator 100 assumes that the image data 5 included in the training dataset 200 is data sampled from the target probability distribution μ. In other words, when the image data 5 included in the training dataset 200 is represented by a variable x, the generator 100 assumes that x follows a target probability distribution function μ(x).

[0013] The generator 100 also assumes that the image data 4 generated by the generator 100 itself is generated based on some probability distribution. Hereinafter, the probability distribution from which the image data 4 generated by the generator 100 is generated is referred to as the model probability distribution μ θ That is, the generator 100 generates image data 4 that is based on the model probability distribution μ θ In other words, when the image data 4 generated by the generator 100 itself is represented by a variable x, the generator 100 generates a model probability distribution function μ θ Assume that (x) is followed.

[0014] The generator 100 also generates a model probability distribution μ θ In other words, the generator 100 trains the model probability distribution μ θ The generator 100 attempts to learn the model probability distribution μ θ As a result of learning to reduce the distance between the target probability distribution μ0 and the model probability distribution μ θ The generator 100 also obtains the obtained model probability distribution μ θ Here, the image data 4 can be generated based on the model probability distribution μ θ is similar to the target probability distribution μ0, so the model probability distribution μ θ The image data sampled from the model probability distribution μ is similar to the image data sampled from the target probability distribution μ. θ By generating image data 4 based on the above, it is possible to generate image data 4 similar to the image data 5 included in the training dataset 200.

[0015] FIG. 2 is a diagram illustrating a generative adversarial network according to the prior art. First, a training method for a classifier 300 of a generative adversarial network will be described. The classifier 300 is trained to distinguish whether data input thereto is real data (hereinafter, sometimes referred to as "real data") or fake data generated by a generator 100 of the generative adversarial network. The classifier 300 is trained while the parameter values ​​of the generator 100 are fixed. Specifically, the generator 100 generates image data from a random vector. When image data generated by the generator 100 (hereinafter, sometimes referred to as "generated data") is input, the classifier 300 is trained to output "0," a value indicating that the generated data is fake data. Furthermore, a training dataset 200 includes a large amount of real data. When real data included in the training dataset 200 is input, the classifier 300 is trained to output "1," a value indicating that the real data is real data.

[0016] More specifically, the classifier 300 calculates the value of the GAN loss function based on its own output result (the "scalar output value") and updates the parameter values ​​of the classifier 300 using backpropagation. Here, the value of the GAN loss function takes a large value when the classifier 300 determines that the real data is real data (when the output value is close to "1"), and a small value when the classifier 300 determines that the real data is fake data (when the output value is close to "0"). Furthermore, the value of the GAN loss function takes a small value when the classifier 300 determines that the generated data is real data (when the output value is close to "1"), and a large value when the classifier 300 determines that the generated data is fake data (when the output value is close to "0"). The classifier 300 learns the parameter values ​​of the classifier 300 so as to maximize the value of the GAN loss function. In this way, the classifier 300 learns to distinguish whether the data input to it is real data (real data) or fake data (generated data). In other words, the classifier 300 determines whether the data input to itself is data (actual data) sampled from the target probability distribution μ 0 or data from the model probability distribution μ θ That is, the classifier 300 according to the prior art learns to distinguish whether the data is sampled (generated data) from the model probability distribution μ θ This can be interpreted as having the role of evaluating the distance between the target probability distribution μ0 and the

[0017] Next, a learning method of the generator 100 will be described. The generator 100 learns to generate data that can be identified as genuine data by the classifier 300. The generator 100 learns while the parameter values ​​of the classifier 300 are fixed. Specifically, the generator 100 generates image data (generated data) from a random vector. An ideal classifier 300 receives the generated data as an input value, and outputs "1" if it determines that the generated data is genuine data (real data), and outputs "0" if it determines that the generated data is fake data. The generator 100 learns to generate data that will result in the output result of the classifier 300 being "1".

[0018] More specifically, the generator 100 calculates the value of the GAN loss function based on the output result (scalar output value) from the discriminator 300, and updates the parameter values ​​of the generator 100 using backpropagation. Here, the value of the GAN loss function takes a small value when the discriminator 300 determines that the generated data is real data (when the output value is close to "1"), and takes a large value when the discriminator 300 determines that the generated data is fake data (when the output value is close to "0"). The generator 100 learns the parameter values ​​of the generator 100 so as to minimize the value of the GAN loss function. In this way, the generator 100 learns to generate generated data that can be identified as real data by the discriminator 300. In other words, the generator 100 calculates a model probability distribution μ that generates generated data that can be identified as data sampled from the target probability distribution μ by the discriminator 300. θ In other words, the generator 100 learns a model probability distribution μ θ To make the generated data sampled from the target probability distribution μ0 closer to the real data sampled from the target probability distribution μ θ The generator 100 trains the model probability distribution μ θ and the target probability distribution μ0, an index D (for example, JS (Jensen-Shannon) divergence) is used. The generator 100 then learns to minimize the value of the index D based on the output result from the classifier 300.

[0019] The training method of the generative adversarial network according to the above-mentioned conventional technology can be expressed as a minimax problem represented by the following formula (1). In the following formula (1), the function corresponding to the generator 100 is θ, the function corresponding to the discriminator 300 is f, and the loss function of the GAN when training the generator 100 is J. GAN , the loss function of the GAN when training the classifier 300 is V GANThe first half of the following equation (1) represents a minimization problem for the function θ. The second half of the following equation (1) represents a maximization problem for the function f. Note that the loss function J GAN and the loss function V GAN is a common function.

[0020]

number

[0021] In the first half of the above equation (1), we fix the parameter values ​​for the function f and then calculate the loss function J GAN Calculate the parameter value of the function θ that reduces the value of . In addition, in the latter part of the above equation (1), after fixing the parameter value of the function θ, calculate the loss function V GAN Calculate the values ​​of the parameters for the function f that increase the value of

[0022] FIG. 3 is a diagram for explaining the difference between a generative adversarial network according to the prior art and a generative adversarial network according to an embodiment of the present disclosure. The generative adversarial network according to an embodiment of the present disclosure is called a Slicing Adversarial Network (SAN). For details of SAN, see the reference (Yuhta Takida, Masaaki Imaizumi, Takashi Shibuya, Chieh-Hsin Lai, Toshimitsu Uesaka, Naoki Murata, Yuki Mitsufuji, "SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear Layer", [online], September 6, 2023, Internet)<URL:https: / / arxiv.org / pdf / 2301.12811v3.pdf> ) for more information.

[0023] The left side of Figure 3 shows a conventional generative adversarial network (GAN). As described in Figures 1 and 2, the conventional generator 100 generates a model probability distribution μ based on the output of the classifier 300. θ The learning is performed to minimize the value of the index D for evaluating the distance between the target probability distribution μ0 and the starting point O. On the left side of Figure 3, the generator 100 learns from the probability distribution (starting point O) at the start of learning, and then learns based on the output results of the classifier 300 to obtain the optimized model probability distribution μ θ 3 shows the path to reach the optimization point B. Also, on the left side of FIG. 3, a classifier 300 according to the prior art optimizes the model probability distribution μ θ Therefore, it is not possible to properly evaluate the distance between the target probability distribution μ0 (destination A) and the model probability distribution μ θ The figure shows that the distance between the optimization points (optimization point B) is far. That is, the generator 100 according to the prior art generates a model probability distribution μ close to the target probability distribution μ θ It is not possible to obtain the above. For details of the problems with the discriminator 300 and the generator 100 according to the prior art described on the left side of FIG. 3, please refer to the above references.

[0024] As illustrated on the left side of FIG. 3, the prior art classifier 300 uses the model probability distribution μ θ Furthermore, since the generator 100 according to the prior art uses the classifier 300 according to the prior art, it is not possible to properly evaluate the distance between the model probability distribution μ θ It is difficult to successfully learn to reduce the distance between the target probability distribution μ0 and the target probability distribution μ0. As a result, it has been difficult for conventional generative adversarial networks to generate highly accurate data. Specifically, in conventional generative adversarial networks, the data generated by the generator is biased toward specific patterns, or the generator may generate a large amount of similar data (also known as mode collapse). In other words, conventional generative adversarial networks have a poor diversity in the data generated by the generator.

[0025] The right side of FIG. 3 shows a generative adversarial network (SAN, hereinafter sometimes referred to as an "adversarial slicing network") according to an embodiment of the present disclosure. On the right side of FIG. 3, a classifier (classifier of the adversarial slicing network) according to an embodiment of the present disclosure generates a model probability distribution μ θ We show that the distance between the model probability distribution μ0 and the target probability distribution μ0 can be properly evaluated. θ and the target probability distribution μ0. More specifically, the classifier of the adversarial slicing network is a classifier that can evaluate the distance between the probability distribution of the generated feature vectors distributed in the feature vector space of the classifier and the probability distribution of the actual feature vectors. In the following, we will define the model probability distribution μ based on the probability distribution of the generated feature vectors distributed in the feature vector space of the classifier and the probability distribution of the actual feature vectors. θ A classifier that can evaluate the distance between the target probability distribution μ0 and the model probability distribution μ0 is sometimes referred to as a "distance-discriminator." Here, the probability distribution of the generated feature vector of a distance-discriminator is theoretically determined by the model probability distribution μ0 of the generator. θ In addition, the probability distribution of the actual feature vectors of the distance-discriminator can theoretically be identified with the target probability distribution μ of the generator. For theoretical details, please refer to the references mentioned above.

[0026] As mentioned above, the classifier of the adversarial slicing network (i.e., the distance-distributable classifier) ​​estimates the distance between the probability distribution of the generated feature vectors distributed in the feature vector space of the classifier and the probability distribution of the real feature vectors, and calculates the model probability distribution μ θ and the target probability distribution μ 0 can be evaluated. As a result, the generator (the generator of the adversarial slicing network) according to the embodiment of the present disclosure can estimate the distance between the model probability distribution μ θ and the target probability distribution μ0. In this way, the distance-discriminator can estimate the distance between the generator's model probability distribution μ θSince the distance between the target probability distribution μ0 and the model probability distribution μ0 can be evaluated, the adversarial slicing network generator can estimate the distance between the target probability distribution μ0 (destination A) and the model probability distribution μ0. θ (optimization point B). In other words, the generator of the adversarial slicing network can learn to appropriately approximate the distance between the model probability distribution μ and the target probability distribution μ. θ For details on the classifier and generator of the adversarial slicing network described on the right side of Figure 3, please refer to the references mentioned above.

[0027] As illustrated on the right side of FIG. 3, the classifier according to the embodiment of the present disclosure uses the model probability distribution μ θ and the target probability distribution μ 0 can be appropriately evaluated. In addition, since the generator according to the embodiment of the present disclosure uses the classifier according to the embodiment of the present disclosure, the distance between the model probability distribution μ θ and the target probability distribution μ0 can be successfully trained to reduce the distance between μ0 and the target probability distribution μ0. Therefore, the generative adversarial network according to the embodiment of the present disclosure can generate data with higher accuracy than the generative adversarial network according to the conventional technology. Specifically, the generative adversarial network according to the embodiment of the present disclosure can reduce the bias of data generated by the generator toward a specific pattern or the generator generating a large amount of similar data (also known as mode collapse). In other words, the generative adversarial network according to the embodiment of the present disclosure can improve the diversity of the data generated by the generator. Please refer to the above references for comparison results between image data generated by the generative adversarial network according to the conventional technology and image data generated by the adversarial slicing network.

[0028] (2. Configuration of Information Processing Device) 4 is a diagram illustrating an example of the configuration of an information processing device according to an embodiment of the present disclosure. As illustrated in FIG. 4, the information processing device 1 includes a communication unit 10, a storage unit 20, and a control unit 30.

[0029] The communication unit 10 is realized by, for example, a network interface card (NIC), etc. The communication unit 10 may be connected to a network via a wired or wireless connection, and may transmit and receive information to and from, for example, other information processing devices.

[0030] The storage unit 20 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. For example, the storage unit 20 stores information related to various programs (for example, the program according to the embodiment).

[0031] The control unit 50 is a controller, and is realized by, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or the like, executing various programs stored in a storage device inside the information processing device 1 using a storage area such as a RAM as a working area. In the example shown in FIG. 4, the control unit 30 has a model generation unit 31 and a data generation unit 32.

[0032] The model generation unit 31 generates a generative adversarial network including a classifier and a generator. Specifically, the model generation unit 31 acquires a GAN classifier (hereinafter may be referred to as a "classifier"). Next, the model generation unit 31 divides the classifier into a feature extraction network that generates a feature vector of data from data input to the classifier, and a final layer that reduces the feature vectors distributed in a feature vector space to a one-dimensional space, and trains the feature extraction network and the final layer, respectively, to generate a distance-determinable classifier that is a classifier that can evaluate the distance between the probability distribution of generated feature vectors, which are feature vectors of generated data generated by the generator, and the probability distribution of actual feature vectors, which are feature vectors of actual data included in a training dataset.

[0033] FIG. 5 is a diagram illustrating a generative adversarial network according to an embodiment of the present disclosure. Hereinafter, the generative adversarial network according to an embodiment of the present disclosure will be referred to as an adversarial slicing network (or SAN). In FIG. 5, a classifier 300 according to the conventional technology is configured with multiple layers. On the left side of FIG. 5, the classifier 300 according to the conventional technology is represented by a function f(x). x represents data (e.g., image data) input to the classifier 300. For example, x may be generated data generated by a generator or actual data included in a training dataset.

[0034] The center of FIG. 5 shows how the model generation unit 31 decomposes the network of the classifier 300 into a feature extraction network that generates feature vectors of data from data input to the classifier 300, and a final layer that reduces feature vectors distributed in a feature vector space to one-dimensional space. Here, the feature extraction network is a network (or function) that generates feature vectors of data from data (x) input to the classifier 300. In the center of FIG. 5, the feature extraction network is represented by the function h(x). The feature extraction network is composed of multiple layers other than the final layer of the classifier 300. The final layer is a layer (or function) that converts the feature vectors generated by the feature extraction network into scalars and reduces the feature vectors to one-dimensional space. Specifically, the final layer receives the feature vectors output from the feature extraction network as input and outputs a scalar value. In the center of FIG. 5, the final layer is represented by w. Specifically, the model generation unit 31 expresses the function f(x) representing the classifier 300 in the form of an inner product of the function h(x) representing the feature extraction network and w representing the final layer. This is expressed mathematically as the following formula (2).

[0035]

number

[0036] Next, the model generation unit 31 expresses the function f(x) representing the classifier 300 in the form of an inner product of the function h(x) representing the feature extraction network and ω, which is w representing the final layer normalized by the norm of w. This is expressed mathematically as the following formula (3).

[0037]

number

[0038] The model generation unit 31 separates the classifier 300 into a feature extraction network and a final layer, and trains the feature extraction network and the final layer, respectively. More specifically, the model generation unit 31 separates the loss function of the classifier 300 into a loss function of the feature extraction network and a loss function of the final layer, and trains the parameter values ​​of the feature extraction network and the parameter values ​​of the final layer, respectively, thereby generating a distance-measuring classifier. More specifically, the model generation unit 31 separates the loss function of the classifier into a loss function of the feature extraction network and a loss function of the final layer, and trains the parameter values ​​of the feature extraction network and the parameter values ​​of the final layer, respectively, thereby generating a distance-measuring classifier.

[0039] Specifically, the training method for the adversarial slicing network can be expressed as a minimax problem represented by the following formula (4). The following formula (4) corresponds to the above formula (1). The model generation unit 31 trains each of the classifier and the generator so as to optimize the minimax problem represented by the following formula (4). In the following formula (4), as in the above formula (1), the function corresponding to the generator is θ, and the loss function of the generator is J GAN On the other hand, the following formula (4) differs from the above formula (1) in that the classifier is expressed in the form of the inner product of ω and h shown in the above formula (3). Also, in the following formula (4), the loss function of the classifier is expressed as V h SAN and the loss function V of the final layer ω SANThe difference from the above equation (1) is that the first half of the following equation (4) represents a minimization problem for the function θ. The second half of the following equation (1) represents a maximization problem for each of the functions h and ω.

[0040]

number

[0041] In the first half of the above equation (4), the model generation unit 31 fixes the parameter values ​​of the functions h and ω, and then calculates the loss function J GAN In addition, in the latter term of the above equation (4), the model generation unit 31 learns the parameter values ​​of the function θ so as to reduce the value of the loss function V h SAN and the loss function V ω SAN The parameter values ​​of the functions h and ω are learned so as to increase the sum of the values ​​of the loss function V h SAN The term is calculated by the model generation unit 31 by fixing the parameter values ​​of the functions θ and ω and then calculating the loss function V h SAN We show that the parameter value of the function h is learned so that the value of V is reduced. h SAN The term is calculated by the model generation unit 31 by fixing the parameter values ​​of the functions θ and h and then calculating the loss function V ω SAN We show that the parameter values ​​of the function ω that reduce the value of

[0042] FIG. 6 is a diagram illustrating a generative adversarial network according to an embodiment of the present disclosure. The feature extraction network (h) described in FIG. 5 receives data (x) input to the classifier as input and outputs a feature vector. That is, the feature extraction network (h) converts the data (x) input to the classifier into a feature vector. The model generation unit 31 uses the feature extraction network (h) to generate a feature vector of data from the data (x) input to the classifier. In FIG. 6, the model generation unit 31 uses the feature extraction network (h) to generate an actual feature vector 5A, which is the feature vector of actual data 5, from actual data 5 included in a training dataset. Furthermore, the model generation unit 31 uses the feature extraction network (h) to generate a generated feature vector 4A, which is the feature vector of generated data 4, from generated data 4 generated by the generator.

[0043] In addition, feature vectors generated by a feature extraction network are generally high-dimensional (e.g., 100-dimensional) vectors. That is, the feature extraction network (h) projects (maps) data (x) input to the classifier into a high-dimensional feature vector space. As a result, feature vectors are distributed in the high-dimensional feature vector space. In FIG. 6, the model generation unit 31 uses the feature extraction network (h) to convert real data 5 input to the classifier into real feature vectors 5A and maps them into a feature vector space 400. As a result, the real feature vectors 5A are distributed in the feature vector space 400. In FIG. 6, 5B indicates the probability distribution of the real feature vectors 5A mapped into the feature vector space 400 by the model generation unit 31. In addition, the model generation unit 31 uses the feature extraction network (h) to convert generated data 4 input to the classifier into generated feature vectors 4A and maps them into the feature vector space 400. As a result, the generated feature vectors 4A are distributed in the feature vector space 400. In FIG. 6, the probability distribution of generated feature vectors 4A mapped to feature vector space 400 by model generation unit 31 is shown as 4B.

[0044] Furthermore, the final layer (ω) of the classifier 300 reduces (high-dimensional) feature vectors distributed in a (generally high-dimensional) feature vector space to a one-dimensional space. In other words, the final layer converts (generally high-dimensional) feature vectors into scalars, thereby projecting (mapping) the feature vectors into a one-dimensional space. In FIG. 6, the model generation unit 31 uses the final layer (ω) to reduce the real feature vector 5A to a one-dimensional space 500. In other words, the model generation unit 31 uses the final layer (ω) to convert the real feature vector 5A to a scalar, thereby projecting (mapping) the real feature vector 5A to the one-dimensional space 500. In this way, the model generation unit 31 uses the final layer (ω) to reduce the probability distribution 5B of the real feature vector 5A distributed in the feature vector space 400 to the one-dimensional space 500. In FIG. 6, the probability distribution 5C of the real feature vector 5A reduced to the one-dimensional space 500 by the model generation unit 31 is indicated by the symbol 5C. Furthermore, the model generation unit 31 uses the final layer (ω) to reduce the generated feature vector 4A to one-dimensional space 500. In other words, the model generation unit 31 uses the final layer (ω) to convert the generated feature vector 4A to a scalar, thereby projecting (mapping) the generated feature vector 4A to the one-dimensional space 500. In this way, the model generation unit 31 uses the final layer (ω) to reduce the probability distribution 4B of the generated feature vector 4A distributed in the feature vector space 400 to the one-dimensional space 500. In FIG. 6, the probability distribution of the generated feature vector 4A reduced to the one-dimensional space 500 by the model generation unit 31 is indicated by 4C.

[0045] Furthermore, the parameters of the final layer (ω) are parameters related to a direction in the feature vector space 400. Specifically, the parameters of the final layer (ω) are parameters related to a direction 6 that separates a probability distribution 4B of generated feature vectors and a probability distribution 5B of actual feature vectors, both of which are distributed in the feature vector space 400. The model generation unit 31 trains the final layer (ω) to take parameter values ​​corresponding to directions that increase the distance between the probability distribution 4B of generated feature vectors and the probability distribution 5B of actual feature vectors, thereby generating a distance-determinable classifier.

[0046] 7 is a diagram for explaining three conditions that a distance-calculated classifier according to an embodiment of the present disclosure must satisfy. The model generation unit 31 generates a distance-calculated classifier by training the classifier so as to satisfy the three conditions of directional optimality, separability, and injectivity.

[0047] First, directional optimality will be described. As shown on the left side of FIG. 7 , directional optimality refers to slicing the feature vector space 400 in a direction 6 that maximizes the distance between the probability distribution 4B of the generated feature vector 4A and the probability distribution 5B of the actual feature vector 5A, both of which are distributed in the feature vector space 400. For example, the direction 6A or 6B shown on the left side of FIG. 7 is not a direction 6 that maximizes the distance between the probability distribution 4B of the generated feature vector 4A and the probability distribution 5B of the actual feature vector 5A, and therefore does not satisfy directional optimality. In other words, directional optimality refers to slicing the feature vector space 400 in a direction 6 that minimizes the overlap between the probability distribution 4B of the generated feature vector 4A and the probability distribution 5B of the actual feature vector 5A, both of which are distributed in the feature vector space 400. In other words, satisfying directional optimality in a classifier corresponds to generating a final layer (ω) that has learned parameter values ​​corresponding to the direction 6 that maximizes the distance between the probability distribution 4B of the generated feature vector 4A and the probability distribution 5B of the actual feature vector 5A, both of which are distributed in the feature vector space 400. In other words, a classifier satisfying directional optimality corresponds to generating a final layer (ω) that has learned parameter values ​​corresponding to a direction 6 that reduces the overlap between the probability distribution 4B of generated feature vectors 4A distributed in feature vector space 400 and the probability distribution 5B of actual feature vectors 5A.

[0048] The model generation unit 31 generates a distance-measuring classifier by training the classifier to satisfy directional optimality. Specifically, the parameters of the final layer are parameters related to the direction separating the probability distribution of generated feature vectors distributed in the feature vector space from the probability distribution of actual feature vectors. The model generation unit 31 generates a distance-measuring classifier by training the final layer to take parameter values ​​corresponding to the direction increasing the distance between the probability distribution of generated feature vectors and the probability distribution of actual feature vectors. In other words, the model generation unit 31 generates a distance-measuring classifier by training the final layer to achieve a transformation to one-dimensional space 500 that reduces the overlap between the probability distribution of generated feature vectors distributed in the feature vector space and the probability distribution of actual feature vectors. More specifically, when a generated feature vector and an actual feature vector are input, the model generation unit 31 generates a distance-measuring classifier by training the final layer to match a directional vector representing the direction of the mean vector of the actual feature vector relative to the mean vector of the generated feature vectors.

[0049] Next, separability will be described. As shown on the right side of FIG. 6, separability refers to a distribution in which the probability distribution 4B of generated feature vector 4A and the probability distribution 5B of real feature vector 5A, which are distributed in feature vector space 400, are separable. In other words, separability refers to a distribution in which the probability distribution 4B of generated feature vector 4A and the probability distribution 5B of real feature vector 5A overlap when the probability distribution 4B of generated feature vector 4A is moved in a specific direction. For example, as an example of satisfying separability, the lower right of FIG. 6 shows how the probability distribution 4C of generated feature vector 4A and the probability distribution 5C of real feature vector 5A overlap when the probability distribution 4C of generated feature vector 4A, which has been reduced to one-dimensional space 500 by model generation unit 31, is moved in a specific direction 7. For example, in the example shown in the center of Figure 7, actual feature vectors 5A are distributed so as to surround probability distribution 4D of generated feature vector 4A. Therefore, even if probability distribution 4B of generated feature vector 4A is moved in a specific direction, probability distribution 4B of generated feature vector 4A and probability distribution 5B of actual feature vector 5A do not overlap, and therefore separability is not satisfied.

[0050] That is, a classifier satisfying separability corresponds to training a feature extraction network (h) to generate a distribution such that the probability distribution 4B of the generated feature vector 4A overlaps with the probability distribution 5B of the actual feature vector 5A when the probability distribution 4B of the generated feature vector 4A is moved in a specific direction. The model generation unit 31 generates a distance-measuring classifier by training the classifier to satisfy separability. Specifically, the model generation unit 31 generates a distance-measuring classifier by training the feature extraction network such that the probability distribution of the generated feature vector overlaps with the probability distribution of the actual feature vector when the probability distribution of the generated feature vector is moved in a specific direction. More specifically, when generated data and actual data are input, the model generation unit 31 generates a distance-measuring classifier by training the feature extraction network to output a probability distribution of the generated feature vector and a probability distribution of the actual feature vector such that the probability distribution of the generated feature vector overlaps with the probability distribution of the actual feature vector when the probability distribution of the generated feature vector is moved in a specific direction.

[0051] Finally, we will explain injectivity. Injectivity means that there is a one-to-one correspondence between the data input to the feature extraction network (h) and the feature vectors distributed in the feature vector space 400. In other words, injectivity means that there is a one-to-one correspondence between the data input to the feature extraction network (h) and the feature vectors distributed in the feature vector space 400. -1 ) exists. The right side of FIG. 7 does not satisfy injectivity because two different real data 51 and 52 correspond to one real feature vector 5A. The model generation unit 31 generates a distance-determinable classifier by training a classifier to satisfy injectivity. Specifically, the model generation unit 31 generates a distance-determinable classifier by training a feature extraction network so that there is a one-to-one correspondence between data and feature vectors. More specifically, the model generation unit 31 generates a distance-determinable classifier by training a feature extraction network so that, when data is input, it outputs a feature vector that corresponds one-to-one with the data.

[0052] As a result of the above learning, the distance-calculated classifier generated by the model generation unit 31 satisfies the three conditions of directional optimality, separability, and injectivity. In this case, the distance-calculated classifier can evaluate the distance between the probability distribution of the generated feature vectors reduced to a one-dimensional space and the probability distribution of the actual feature vectors by using a one-dimensional distance, as shown in the lower right of Fig. 6.

[0053] Furthermore, the model generation unit 31 acquires a generator of the GAN (hereinafter, may be referred to as "generator"). Next, the model generation unit 31 generates a generator that is trained using a distance-calculated classifier to reduce the distance between the probability distribution of the generated feature vector and the probability distribution of the actual feature vector. Specifically, the model generation unit 31 trains the generator to reduce the distance between the probability distribution of the generated feature vector generated using the distance-calculated classifier and the probability distribution of the actual feature vector. Here, the probability distribution of the generated feature vector is the model probability distribution μ θ 3. The probability distribution of the actual feature vector corresponds to the target probability distribution μ0 described in FIGS. 1 to 3. Specifically, the model generation unit 31 trains the generator so that, when a random vector is input, it outputs image data (generated data). More specifically, the model generation unit 31 trains the generator so that it generates data such that the output result of the distance-variable classifier is "1". As a result, the information processing device 1 generates the model probability distribution μ0 as shown on the right side of FIG. θ The distance between the target probability distribution μ (optimization point B) and the target probability distribution μ (destination A) can be reduced. θ It may be possible to obtain the following.

[0054] For details regarding the loss functions used in the training of the classifier and generator by the model generation unit 31, please refer to the above references.

[0055] Data generation unit 32 generates generated data from a random vector using the generator generated by model generation unit 31. Specifically, when a random vector is input, model generation unit 31 outputs image data (generated data).

[0056] (3. Effects) As described above, the information processing device 1 according to an embodiment of the present disclosure includes a model generation unit 31. The model generation unit 31 generates a generative adversarial network including a classifier and a generator. The model generation unit 31 separates the classifier into a feature extraction network that generates a feature vector of data from data input to the classifier, and a final layer that reduces the feature vectors distributed in a feature vector space to a one-dimensional space, and trains the feature extraction network and the final layer, respectively, to generate a distance-determinable classifier that is a classifier capable of evaluating the distance between the probability distribution of generated feature vectors, which are feature vectors of generated data generated by the generator, and the probability distribution of actual feature vectors, which are feature vectors of actual data included in a training dataset.

[0057] In this way, the information processing device 1 generates a distance-determinable classifier, which is a classifier capable of evaluating the distance between the probability distribution of the generated feature vector, which is a feature vector of the generated data generated by the generator, and the probability distribution of the actual feature vector, which is a feature vector of the actual data included in the training dataset. As described above, the probability distribution of the generated feature vector of the distance-determinable classifier is theoretically the model probability distribution μ θ In addition, the probability distribution of the actual feature vector of the distance-variable classifier can theoretically be regarded as the target probability distribution μ of the generator. Thus, the information processing device 1 uses the distance-variable classifier to calculate the model probability distribution μ of the generator. θ The information processing device 1 can appropriately evaluate the distance between the model probability distribution μ θ Since the distance between μ and the target probability distribution μ can be properly evaluated, the generator can generate a model probability distribution μ close to the target probability distribution μ θTherefore, the information processing device 1 can improve the diversity of data generated by the generator of the generative adversarial network.

[0058] The parameters of the final layer are parameters related to the direction separating the probability distribution of the generated feature vectors from the probability distribution of the actual feature vectors, both of which are distributed in the feature vector space. The model generation unit 31 trains the final layer to take parameter values ​​corresponding to the direction increasing the distance between the probability distribution of the generated feature vectors and the probability distribution of the actual feature vectors, thereby generating a distance-determinable classifier.

[0059] This allows the information processing device 1 to generate a distance-determinable classifier that satisfies directional optimality.

[0060] In addition, the model generation unit 31 generates a distance-determinable classifier by training the feature extraction network so that the probability distribution of the generated feature vector overlaps with the probability distribution of the actual feature vector when the probability distribution of the generated feature vector is moved in a specific direction.

[0061] This allows the information processing device 1 to generate a distance-determinable classifier that satisfies separability.

[0062] Furthermore, the model generation unit 31 generates a distance-determinable classifier by training a feature extraction network so that data and feature vectors correspond one-to-one.

[0063] This allows the information processing device 1 to generate a distance-determinable classifier that satisfies the injectivity.

[0064] In addition, the model generation unit 31 separates the loss function of the classifier into a loss function of the feature extraction network and a loss function of the final layer, and generates a distance-determinable classifier by learning the parameter values ​​of the feature extraction network and the parameter values ​​of the final layer, respectively.

[0065] This allows the information processing device 1 to separate the feature extraction network and the final layer and train the feature extraction network and the final layer separately, thereby training the final layer to satisfy directional optimality and training the feature extraction network to satisfy separability and injectivity.

[0066] Furthermore, the model generating unit 31 uses a distance-variable classifier to generate a generator that is trained to reduce the distance between the probability distribution of the generated feature vector and the probability distribution of the actual feature vector.

[0067] As a result, the information processing device 1 uses a distance-determinable classifier to generate the model probability distribution μ θ Since the distance between μ and the target probability distribution μ can be properly evaluated, the generator can generate a model probability distribution μ close to the target probability distribution μ θ It may be possible to obtain the following.

[0068] The information processing device 1 further includes a data generation unit 32. The data generation unit 32 uses the generator generated by the model generation unit 31 to generate generated data from random vectors.

[0069] As a result, the information processing device 1 obtains a model probability distribution μ close to the target probability distribution μ θ Since the generated data can be generated using a generator that has acquired the above, it is possible to improve the diversity of the data generated by the generator.

[0070] (4. Hardware Configuration) The information processing device 1 according to the embodiment described above is realized by a computer 1000 having a configuration as shown in Fig. 8, for example. Fig. 8 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device according to the present disclosure. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, an HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0071] The CPU 1100 operates and controls each unit based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0072] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .

[0073] HDD 1400 is a computer-readable, non-transitory recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a non-transitory recording medium that records a program according to the present disclosure, which is an example of program data 1450.

[0074] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0075] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disk), magneto-optical recording media such as an MO (Magneto-Optical disk), tape media, magnetic recording media, and semiconductor memories.

[0076] For example, when the computer 1000 functions as the information processing device 1 according to the embodiment, the CPU 1100 of the computer 1000 executes a program loaded onto the RAM 1200 to reproduce the functions of the control unit 30, etc. The HDD 1400 stores the program according to the present disclosure and various data. The CPU 1100 reads and executes the program data 1450 from the HDD 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.

[0077] Although the embodiments of the present disclosure have been described above, the technical scope of the present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure. Furthermore, components of different embodiments and modifications may be combined as appropriate.

[0078] Furthermore, the effects of each embodiment described in this specification are merely examples and are not intended to be limiting, and other effects may also be obtained.

[0079] The present technology can also be configured as follows. (1) An information processing device including a model generation unit that generates a generative adversarial network including a classifier and a generator, The model generation unit The classifier is divided into a feature extraction network that generates a feature vector of the data from the data input to the classifier, and a final layer that reduces the feature vectors distributed in a feature vector space to a one-dimensional space, and the feature extraction network and the final layer are trained separately to generate a distance-determinable classifier that is a classifier that can evaluate the distance between the probability distribution of a generated feature vector, which is the feature vector of generated data generated by the generator, and the probability distribution of an actual feature vector, which is the feature vector of actual data included in a training dataset. Information processing device. (2) the parameters of the final layer are parameters relating to a direction separating a probability distribution of the generated feature vector and a probability distribution of the actual feature vector, both of which are distributed in the feature vector space; The model generation unit generating the distance-discriminable classifier by training the final layer to take parameter values ​​corresponding to the direction that increases the distance between the probability distribution of the generated feature vector and the probability distribution of the actual feature vector; The information processing device according to (1) above. (3) The model generation unit generating the distance-determinable classifier by training the feature extraction network so that the probability distribution of the generated feature vector overlaps with the probability distribution of the actual feature vector when the probability distribution of the generated feature vector is moved in a specific direction; The information processing device according to (1) or (2). (4) The model generation unit generating the distance-discriminator by training the feature extraction network so that there is a one-to-one correspondence between the data and the feature vector; The information processing device according to any one of (1) to (3). (5) The model generation unit generating the distance-determinable classifier by dividing the loss function of the classifier into a loss function of the feature extraction network and a loss function of the final layer, and training the parameter values ​​of the feature extraction network and the parameter values ​​of the final layer, respectively; The information processing device according to any one of (1) to (4). (6) The model generation unit generating the generator trained to reduce the distance between the probability distribution of the generated feature vector and the probability distribution of the actual feature vector using the distance-discriminator; The information processing device according to any one of (1) to (5). (7) The method further includes a data generation unit that generates the generated data from a random vector using the generator generated by the model generation unit. The information processing device according to any one of (1) to (6). (8) The computer generating a generative adversarial network including a classifier and a generator; the classifier is divided into a feature extraction network that generates a feature vector of the data from the data input to the classifier, and a final layer that reduces the feature vectors distributed in a feature vector space to a one-dimensional space, and the feature extraction network and the final layer are trained separately to generate a distance-determinable classifier that is a classifier that can evaluate the distance between a probability distribution of a generated feature vector, which is a feature vector of generated data generated by the generator, and a probability distribution of an actual feature vector, which is a feature vector of actual data included in a training dataset; An information processing method, including: (9) generating a generative adversarial network including a classifier and a generator; the classifier is divided into a feature extraction network that generates a feature vector of the data from the data input to the classifier, and a final layer that reduces the feature vectors distributed in a feature vector space to a one-dimensional space, and the feature extraction network and the final layer are trained separately to generate a distance-determinable classifier that is a classifier that can evaluate the distance between a probability distribution of a generated feature vector, which is a feature vector of generated data generated by the generator, and a probability distribution of an actual feature vector, which is a feature vector of actual data included in a training dataset; A program that causes a computer to execute the following. [Explanation of symbols]

[0080] 1. Information processing equipment 10. Communications Department 20 Memory section 30 Control Unit 31 Model Generation Unit 32 Data Generation Unit

Claims

1. An information processing device including a model generation unit that generates a generative adversarial network including a classifier and a generator, The model generation unit The classifier is divided into a feature extraction network that generates a feature vector of the data from the data input to the classifier, and a final layer that reduces the feature vector distributed in a feature vector space to a one-dimensional space, and the feature extraction network and the final layer are trained separately to generate a distance-determinable classifier that is a classifier capable of evaluating the distance between a probability distribution of a generated feature vector, which is a feature vector of generated data generated by the generator, and a probability distribution of an actual feature vector, which is a feature vector of actual data included in a training dataset. Information processing device.

2. the parameters of the final layer are parameters relating to a direction for separating a probability distribution of the generated feature vector and a probability distribution of the actual feature vector, both of which are distributed in the feature vector space; The model generation unit generating the distance-distributable classifier by training the final layer to take parameter values ​​corresponding to the direction that increases the distance between the probability distribution of the generated feature vector and the probability distribution of the actual feature vector; The information processing device according to claim 1 .

3. The model generation unit generating the distance-distributable classifier by training the feature extraction network so that the probability distribution of the generated feature vector overlaps with the probability distribution of the actual feature vector when the probability distribution of the generated feature vector is moved in a specific direction; The information processing device according to claim 1 .

4. The model generation unit generating the distance-discriminator by training the feature extraction network so that there is a one-to-one correspondence between the data and the feature vector; The information processing device according to claim 1 .

5. The model generation unit generating the distance-discriminator by dividing a loss function of the classifier into a loss function of the feature extraction network and a loss function of the final layer, and learning the parameter values ​​of the feature extraction network and the parameter values ​​of the final layer, respectively; The information processing device according to claim 1 .

6. The model generation unit generating the generator, which is trained to reduce the distance between the probability distribution of the generated feature vector and the probability distribution of the actual feature vector, using the distance-discriminator; The information processing device according to claim 1 .

7. The method further includes a data generation unit that generates the generated data from a random vector using the generator generated by the model generation unit. The information processing device according to claim 1 .

8. The computer generating a generative adversarial network including a classifier and a generator; the classifier is divided into a feature extraction network that generates a feature vector of the data from the data input to the classifier, and a final layer that reduces the feature vector distributed in a feature vector space to a one-dimensional space, and the feature extraction network and the final layer are trained separately to generate a distance-determinable classifier that is a classifier capable of evaluating the distance between a probability distribution of a generated feature vector, which is a feature vector of generated data generated by the generator, and a probability distribution of an actual feature vector, which is a feature vector of actual data included in a training dataset; An information processing method comprising:

9. generating a generative adversarial network including a classifier and a generator; the classifier is divided into a feature extraction network that generates a feature vector of the data from the data input to the classifier, and a final layer that reduces the feature vector distributed in a feature vector space to a one-dimensional space, and the feature extraction network and the final layer are trained separately to generate a distance-determinable classifier that is a classifier capable of evaluating the distance between a probability distribution of a generated feature vector, which is a feature vector of generated data generated by the generator, and a probability distribution of an actual feature vector, which is a feature vector of actual data included in a training dataset; A program that causes a computer to execute the following.