Training device and training method

By exponentially tilting the Gaussian distribution and approximating KL divergence, the proposed methods facilitate the learning of high-dimensional and complex data in variational autoencoders, enhancing reconstruction and probability estimation accuracy.

WO2025146722A1PCT designated stage expired Publication Date: 2025-07-10NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/000079
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing variational autoencoders face difficulties in learning high-dimensional and complex data due to the small volume of the high-density region in standard Gaussian distributions, leading to poor reconstruction and probability estimation accuracy, and Tilted VAEs struggle with calculating KL divergence.

Method used

The proposed methods introduce a function to exponentially tilt the standard Gaussian distribution, allowing for a flexible prior distribution that approximates KL divergence accurately without limiting the encoder, using either a simple quadratic approximation or a relative density ratio to stabilize learning.

Benefits of technology

This approach enables efficient learning of high-dimensional and complex data by variational autoencoders, improving reconstruction and probability estimation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024000079_10072025_PF_FP_ABST
    Figure JP2024000079_10072025_PF_FP_ABST
Patent Text Reader

Abstract

A training device according to one aspect of the present invention includes: an input unit for inputting a data set for training a variational autoencoder; and a training unit for training the variational autoencoder from the data set so as to maximize a variational lower bound which includes, as a prior distribution of latent variables with a distribution obtained by exponentially slanting a standard Gaussian distribution using a prescribed function, an approximation of a KL divergence between an encoder of the variational autoencoder and the prior distribution. The approximation of the KL divergence is composed of the KL divergence between the encoder and the standard Gaussian distribution, an expected value of the function related to the encoder, and a prescribed constant.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device and learning method

[0001] The present disclosure relates to a learning device and a learning method.

[0002] One known deep generative model is a model called a variational autoencoder (VAE) (Non-Patent Document 1). A variational autoencoder is a model that can estimate the probability of given data using latent variables, and is trained to maximize the variational lower bound (ELBO).

[0003] Generally, a standard Gaussian distribution is often used as a prior distribution for latent variables, but the standard Gaussian distribution has a problem in that it cannot learn high-dimensional and complex data because the volume of the high-density region is small. In response to this problem, a method called tilted VAE has been proposed as a method for increasing the volume of the high-density region of the prior distribution (Non-Patent Document 2). In tilted VAE, the volume of the high-density region of the prior distribution is increased by exponentially tilting the prior distribution using the L2 norm of the latent variables.

[0004] Diederik P Kingma, Max Welling, "Auto-Encoding Variational Bayes", arXiv:1312.6114 [stat.ML].Griffin Floto, Stefan Kremer, Mihai Nica, "The Tilted Variational Autoencoder: Improving Out-of-Distribution Detection", Published as a conference paper at ICLR 2023.

[0005] However, the Tilted VAE has a problem in that it is difficult to calculate the KL (Kullback-Leibler) divergence included in the variational lower bound.

[0006] The present disclosure has been made in consideration of the above points, and provides a technology that enables high-dimensional and complex data to be easily learned using a variational autoencoder.

[0007] A learning device according to one aspect of the present disclosure includes an input unit that inputs a dataset for training a variational autoencoder, and a learning unit that uses a distribution that is exponentially tilted from a standard Gaussian distribution using a predetermined function as a prior distribution for latent variables to learn the variational autoencoder from the dataset so as to maximize a variational lower bound that includes an approximation of the KL divergence between the encoder of the variational autoencoder and the prior distribution, and the approximation of the KL divergence is composed of the KL divergence between the encoder and the standard Gaussian distribution, the expected value of the function for the encoder, and a predetermined constant.

[0008] Variational autoencoders provide a technique that can easily learn high-dimensional and complex data.

[0009] FIG. 1 is a diagram showing an example of estimating a probability distribution from given data. FIG. 2 is a diagram showing an example of a variational autoencoder. FIG. 3 is a diagram showing an example of a comparison between a standard Gaussian distribution and an exponentially tilted Gaussian distribution. FIG. 4 is a diagram showing an example of the hardware configuration of a learning device according to this embodiment. FIG. 5 is a diagram showing an example of the functional configuration of a learning device according to this embodiment. FIG. 6 is a flowchart showing an example of a learning process in Example 1. FIG. 7 is a flowchart showing an example of a learning process in Example 2. FIG. 8 is a diagram showing the distribution of latent variables when toy data is used. FIG. 9 is a diagram showing the prior distribution in proposed method 2. FIG. 10 is a diagram showing the approximation accuracy of KL divergence. FIG. 11 is a diagram showing the sensitivity of hyperparameters.

[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.

[0011] <Variational Autoencoder and Tilted VAE> First, the existing methods of variational autoencoder (Non-Patent Document 1) and tilted VAE (Non-Patent Document 2) will be described.

[0012] Variational Autoencoder (VAE) A variational autoencoder (hereinafter referred to as VAE) is one of the representative models of deep generative models, a field of machine learning and artificial intelligence. A generative model is a general term for a method of estimating the probability distribution that given data follows, and can be applied to, for example, image generation and anomaly detection. A deep generative model is a generative model that uses deep learning and is capable of learning more complex and high-dimensional data. For example, when given data with a distribution like the one shown in the left diagram of Figure 1, a generative model can estimate a probability distribution like the one shown in the right diagram of Figure 1. As a result, if data exists in the low probability range, as shown in the right diagram of Figure 1, it is possible to treat that data as anomalous data and achieve anomaly detection.

[0013] Note that complex and high-dimensional data refers to, for example, data that has real-valued elements and is expressed as a high-dimensional vector with low sparsity. An example of such complex and high-dimensional data is image data.

[0014] VAE uses a D-dimensional latent variable z to calculate the probability p θ (x) is estimated by the following equation (1): where D is a predetermined number of dimensions.

[0015] Here, p θ (x|z) is the decoder, θ is its parameter, q φ (z|x) is the encoder, φ is its parameter, and p(z) is the prior distribution of the latent variable z.

[0016] Furthermore, the VAE is trained (that is, the parameters θ and φ are optimized) so as to maximize the variational lower bound (ELBO) shown in the following equation (2).

[0017] Here, the first term on the right side of the above equation (2) represents the reconstruction error. Also, the second term on the right side of the above equation (2) represents the encoder q φ (z|x) and the prior distribution p(z).

[0018] For example, encoder q φ(z|x) is a Gaussian distribution, and the decoder p θ When (x|z) is a Bernoulli distribution, the VAE is modeled as shown in Figure 2. In this model, when data x is given to the VAE, it is distributed as a Gaussian distribution N(z; μ φ (x), σ φ 2 The latent variable z is sampled according to the Bernoulli distribution B(x; μ θ In the VAE shown in Figure 2, the prior distribution p(z) is a standard Gaussian distribution N(z; 0, I).

[0019] Encoder Q φ Since the KL divergence between (z|x) and the prior distribution p(z) can be calculated in a closed form, the standard Gaussian distribution is generally used as the prior distribution p(z). However, when the standard Gaussian distribution is used as the prior distribution p(z), it is difficult to learn high-dimensional and complex data for the following reasons.

[0020] Reason: The latent variable z needs to be distributed far enough apart to reconstruct the data x. On the other hand, to reduce the KL divergence, the latent variable z needs to be distributed in the high-density region of the prior distribution p(z). However, since the volume of the high-density region of the standard Gaussian distribution is very small, it is difficult to distribute the latent variable z far enough apart in the high-density region.

[0021] As a result, VAE encounters problems such as the latent variable z overlapping significantly in the high-density region of the prior distribution p(z), making it impossible to reconstruct the data, or the latent variable z being distributed in the low-density region of the prior distribution p(z), resulting in a decrease in the accuracy of probability estimation. Note that, hereinafter, p(z) is assumed to be a standard Gaussian distribution.

[0022] To address the above problem, a method called tilted VAE has been proposed. In tilted VAE, in order to increase the high-density region of the prior distribution, the prior distribution p(z), which is a standard Gaussian distribution, is exponentially tilted using the L2 norm of the latent variable z according to the following equation (3):

[0023] where τ is a hyperparameter and Z τ is expressed by the following equation (4).

[0024] where Γ(·) is the gamma function, the confluent hypergeometric function.

[0025] The exponentially tilted Gaussian distribution p τ (z) has a maximum density over the entire hypersphere of radius τ, and has a maximum density over a wider area than the standard Gaussian distribution. As an example, the standard Gaussian distribution and the exponentially tilted Gaussian distribution p τ A comparison example with (z) is shown in Figure 3. As shown in Figure 3, the exponentially tilted Gaussian distribution p τ (z) has the maximum density over the entire hypersphere with radius τ = 3, and it can be seen that it has the maximum density over a wider area than the standard Gaussian distribution.

[0026] <Problems with Tilted VAE> Tilted VAE exponentially tilts the standard Gaussian distribution to increase the volume of the high-density region of the prior distribution, making it possible to learn high-dimensional and complex data. However, the exponentially tilted Gaussian distribution p τ The problem is that the KL divergence with (z) is difficult to calculate.

[0027] For simplicity, as shown in the following equation (5), the encoder q φ We restrict the variance of the Gaussian distribution, which is (z|x), to the identity matrix I.

[0028] At this time, the encoder q φ (z | x) and the exponentially tilted Gaussian distribution p τ The KL divergence with (z) can be calculated using the generalized Laguerre polynomial L according to the following equation (6).

[0029] However, it is difficult to actually calculate the above equation (6), so a simple quadratic approximation shown in the following equation (7) is used during VAE training.

[0030] where μ τ ★ is expressed by the following equation (8).

[0031] Therefore, the exponentially tilted Gaussian distribution p τ When (z) is used, the ELBO has a lower bound shown in the following equation (9).

[0032] μ in the above formula (9) τ ★ is calculated using numerical differentiation and the gradient method, but when the hyperparameter τ is small or the number of dimensions D of the latent variable z is large, it collapses to 0, resulting in a large error in the quadratic approximation.

[0033] As mentioned above, Tilted VAE is a KL divergence D KL (q φ (z|x)||p τ (z)) is difficult to calculate, so we use a Gaussian distribution N(z; μ φ (x), I) to encoder q φ (z|x) needs to be limited, and even if the encoder is limited, D KL (q φ (z|x)||p τ However, there is a problem in that the calculation of (z) is difficult.

[0034] <Proposed Methods> Proposed Method 1 and Proposed Method 2 are proposed as methods for solving the above-mentioned problems of the tilted VAE.

[0035] Proposed method 1: Proposed method 1 uses encoder q φ This is a highly accurate approximation method of the KL divergence that does not require restrictions on (z|x). This approximation method allows for more flexible exponential tilting of the prior distribution p(z) using functions other than the L2 norm.

[0036] First, we consider the exponentially tilted Gaussian distribution p f (z) is introduced by the following equation (10).

[0037] Here, Z fis expressed by the following equation (11).

[0038] The above function f is Z f Any function can be used as long as it satisfies <∞. Note that when f(z)=τ∥z∥, it coincides with the Tilted VAE.

[0039] At this time, the encoder q φ (z|x) and a prior distribution p, which is an exponentially tilted Gaussian distribution with a function f. f The KL divergence with (z) can be approximated by the following equation (12) using a standard Gaussian distribution p(z).

[0040] The first term in equation (12) above can be calculated in closed form, the second term can be approximated with high accuracy, and the third term is a constant and can be ignored during VAE training.

[0041] In the proposed method 1, in the above equation (10), f(z) = τ||z||. In this case, Z f =Z τ The KL divergence can be calculated by the following equation (13).

[0042] At this time, the encoder q φ (z|x) does not need to be limited (i.e., the Gaussian distribution N(z; μ φ Note that it is not necessary to be limited to (x), I).

[0043] Therefore, the ELBO can be calculated by simply adding the norm to the reconstruction error, that is, by the following equation (14).

[0044] In this way, in proposed method 1, the standard Gaussian distribution is exponentially tilted by f(z = τ||z||, and the Gaussian distribution p f Even if (z) is used as the prior distribution, the encoder q φ Without restricting (z|x), a highly accurate approximation of the KL divergence can be easily computed.

[0045] <<Proposed Method 2>> In the above formula (10), a function f that provides a more flexible prior distribution than that in proposed method 1 is considered.

[0046] The prior distribution that maximizes ELBO is the aggregated posterior probability q φ It is known that (z).

[0047] Here, p D (x) is the distribution (probability distribution) of data x.

[0048] The function f is the standard Gaussian distribution p(z) and the aggregate posterior probability q φ (z), the exponentially tilted Gaussian distribution p f (z) is the aggregate posterior probability q φ (z). That is, the following equation (16) holds true.

[0049] However, if the supports of the two distributions do not match, the density ratio will diverge to infinity, making learning unstable. Therefore, as shown in the following equation (17), the relative density ratio r α We model the function f using

[0050] Here, 0≦α≦1 is a hyperparameter, and α=0 corresponds to the density ratio. α has an upper bound of 1 / α and is always smoother than the density ratio, so it can be seen as a generalization of the stable density ratio.

[0051] Therefore, in the proposed method 2, the function f shown in the above formula (17) is used in the above formula (10). α is the neural network ψ can be learned by minimizing the objective function J(ψ) shown in the following equation (18).

[0052] Here, p α (z) = αq φ (z) + (1-α)p(z).

[0053] From the above, the objective function L shown in the following equation (19) is obtained. f By alternately maximizing (x; θ, φ) and minimizing the objective function J(ψ) shown in the following equation (20), the relative density ratio r α The prior distribution p f It is possible to learn a VAE with (z).

[0054] Here, α is selected using validation data, and empirically, a value around 0.1 is preferable.

[0055] In this way, in proposed method 2, the standard Gaussian distribution is exponentially tilted using function f, which provides a more flexible prior distribution than proposed method 1. This is expected to improve the performance of the VAE compared to proposed method 1.

[0056] A learning device 10 that learns a VAE using the above-described proposed method 1 or 2 will be described below.

[0057] <Example of Hardware Configuration of Learning Device 10> An example of the hardware configuration of the learning device 10 according to this embodiment will be described with reference to Fig. 4. Fig. 4 is a diagram showing an example of the hardware configuration of the learning device 10 according to this embodiment.

[0058] 4, the learning device 10 according to this embodiment includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a random access memory (RAM) 105, a read-only memory (ROM) 106, an auxiliary storage device 107, and a processor 108. Each of these pieces of hardware is connected to each other via a bus 109 so as to be able to communicate with each other.

[0059] The input device 101 is, for example, a keyboard, a mouse, a touch panel, physical buttons, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the learning device 10 does not necessarily have to have at least one of the input device 101 and the display device 102, for example.

[0060] The external I / F 103 is an interface with an external device such as a recording medium 103a. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.

[0061] The communication I / F 104 is an interface for connecting to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory. The processor 108 is a variety of arithmetic devices such as a central processing unit (CPU) or a graphic processing unit (GPU).

[0062] 4 is an example, and the hardware configuration of the learning device 10 is not limited to this. For example, the learning device 10 may have multiple auxiliary storage devices 107 or multiple processors 108, may not have some of the hardware shown in the figure, or may have various hardware other than the hardware shown in the figure.

[0063] <Example of Functional Configuration of Learning Device 10> An example of the functional configuration of the learning device 10 according to this embodiment will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the functional configuration of the learning device 10 according to this embodiment.

[0064] As shown in Figure 5, the learning device 10 according to this embodiment has an input unit 201 and a learning unit 202. These units are realized, for example, by a process in which one or more programs installed in the learning device 10 are executed by the processor 108 or the like. The learning device 10 according to this embodiment also has a learning dataset storage unit 203 and a parameter storage unit 204. These storage units are realized, for example, by a storage area such as the auxiliary storage device 107. Note that at least one of the storage units of the learning dataset storage unit 203 and the parameter storage unit 204 may be realized by a storage area such as a storage device connected to the learning device 10 so as to be able to communicate with it.

[0065] The input unit 201 inputs a dataset of data x used for training the VAE (hereinafter, the data used for training the VAE will also be referred to as "training data" and the dataset will also be referred to as "training dataset") from the training dataset storage unit 203. It is assumed that the training data x is high-dimensional and complex data.

[0066] The learning unit 202 uses the learning data set input by the input unit 201 to learn the VAE by the above-mentioned proposed method 1 or proposed method 2. That is, when the proposed method 1 is used, the learning unit 202 learns the VAE by the objective function L τ On the other hand, when using the proposed method 2, the learning unit 202 updates the parameters θ and φ of the VAE so as to maximize the objective function L f The parameters θ and φ of the VAE are updated so as to maximize (x; θ, φ), and the neural network r is updated so as to minimize the objective function J(ψ) shown in the above equation (20). ψ This is done alternately with updating the parameter ψ of

[0067] The training data set storage unit 203 stores a training data set. The parameter storage unit 204 stores parameters to be trained (i.e., parameters θ and φ when Proposed Method 1 is used, and parameters θ, φ, and ψ when Proposed Method 2 is used).

[0068] <Learning Process> Hereinafter, the case where VAE is learned by proposed method 1 will be referred to as Example 1, and the case where VAE is learned by proposed method 2 will be referred to as Example 2. The learning process in Example 1 and the learning process in Example 2 will be described.

[0069] First Embodiment The learning process in the first embodiment will be described with reference to Fig. 6. Fig. 6 is a flowchart showing an example of the learning process in the first embodiment.

[0070] The input unit 201 inputs a training data set from the training data set storage unit 203 (step S101).

[0071] The learning unit 202 uses the learning data set input in step S101 to calculate the objective function L τ The VAE parameters θ and φ are updated by a known optimization method so as to maximize (x; θ, φ) (step S102). As a result, the Gaussian distribution p, which is an exponentially tilted Gaussian distribution obtained by setting f(z)=τ∥z∥ using the above equation (10), is obtained. f A VAE is trained with (z) as the prior distribution of the latent variable z.

[0072] Second Embodiment A learning process in a second embodiment will be described with reference to Fig. 7. Fig. 7 is a flowchart showing an example of the learning process in the second embodiment.

[0073] The input unit 201 inputs a training data set from the training data set storage unit 203 (step S201).

[0074] The learning unit 202 alternately performs the following (a) and (b) using the learning data set input in step S201 (step S202).

[0075] (a) The objective function L shown in the above equation (19) f The VAE parameters θ and φ are updated by a known optimization method so as to maximize (x; θ, φ).

[0076] (b) The neural network r is optimized by a known optimization method so as to minimize the objective function J(ψ) shown in the above equation (20). ψUpdate the parameter ψ of

[0077] As a result, the Gaussian distribution p, which is an exponentially inclined standard Gaussian distribution obtained by using f(z) shown in equation (17) and equation (10) above, is obtained. f A VAE is trained with (z) as the prior distribution of the latent variable z.

[0078] <Experimental Example> An experimental example of the above-mentioned proposed methods 1 and 2 will be described below. In this experimental example, a normal VAE learning method and a tilted VAE are adopted as comparative methods. Hereinafter, a VAE learned by a normal learning method will be referred to simply as "VAE" or "normal VAE," and a VAE learned by a tilted VAE will be referred to as "tilted VEA." Meanwhile, a VAE learned by proposed method 1 will be referred to as "Proposed 1," and a VAE learned by proposed method 2 will be referred to as "Proposed 2."

[0079] In this experimental example, toy data was first used to visualize the distribution of latent variables for VAE, Tilted VAE, Proposed1, and Proposed2. The results are shown in Figure 8. The upper left diagram in Figure 8 shows the distribution of latent variables using VAE, the upper right diagram shows the distribution of latent variables using Tilted VAE, the lower left diagram shows the distribution of latent variables using Proposed1, and the lower right diagram shows the distribution of latent variables using Proposed2. The toy data was also composed of four types of four-dimensional OneHot vectors (first to fourth OneHot vectors). Specifically, the first OneHot vector is a four-dimensional vector in which only the first element is 1 and the rest is 0; the second OneHot vector is a four-dimensional vector in which only the second element is 1 and the rest is 0; the third OneHot vector is a four-dimensional vector in which only the third element is 1 and the rest is 0; and the fourth OneHot vector is a four-dimensional vector in which only the fourth element is 1 and the rest is 0.

[0080] Next, we compared the test likelihoods of VAE, Tilted VAE, Proposed 1, and Proposed 2 using real-world image data: MNIST, Fashion MNIST, and SVHN. The results are shown in Table 1 below.

[0081] The larger the test likelihood, the better the performance of the generative model.

[0082] As shown in the upper left diagram of Figure 8, the overlap between latent variables is large in the normal VAE, which causes reconstruction failure and performance degradation. Also, as shown in Table 1 above, the performance of the normal VAE is lower than that of Proposed 2.

[0083] As shown in the upper right diagram of Figure 8, the latent variables of the Tilted VAE are completely separated and easy to reconstruct. On the other hand, the KL divergence with the prior distribution becomes large, resulting in poor test likelihood, as shown in Table 1 above.

[0084] As shown in the lower left diagram of Figure 8, Proposed1 has smaller overlap between latent variables than regular VAE. Furthermore, since the KL divergence with the prior distribution is smaller than that of Tilted VAE, as shown in Table 1 above, Proposed1 has a better test likelihood than Tilted VAE. On the other hand, the test likelihood is lower than that of regular VAE and Proposed2. This suggests that Tilted VAE cannot achieve high performance for high-dimensional data such as real-world image data.

[0085] As shown in the lower right diagram of Figure 8, Proposed2 shows that the latent variables are separated. Furthermore, the prior distribution for Proposed2 is shown in Figure 9. Comparing the lower right diagram of Figure 8 with Figure 9, it can be seen that the prior distribution and the distribution of the latent variables almost perfectly match. Therefore, the KL divergence is a small value, and as shown in Table 1 above, Proposed2 achieves higher performance on all datasets than other methods.

[0086] Finally, we conducted more detailed experiments on Proposed 1, Proposed 2, and their comparison methods. First, Figure 10 shows the approximate values ​​of the KL divergence for Proposed 1 and its comparison method, Tilted VAE, when τ and D are set to various values. As shown in Figure 10, Tilted VAE performs well in the low-dimensional case (when D = 2), but fails in the high-dimensional case (when D = 40). In contrast, Proposed 1 achieves highly accurate approximation in all cases. Note that "True" in Figure 10 indicates the correct answer. Next, Figure 11 shows the sensitivity of the hyperparameter α to Proposed 2. As shown in Figure 11, a smaller hyperparameter α is better, but performance deteriorates when the hyperparameter α is less than 0.1. It is also empirically clear that a value of 0.1 is best. In addition, in order to compare with Proposed 2, FIG. 11 also illustrates the sensitivity of VAE, which is a method that does not use the hyperparameter α.

[0087] <Summary> As described above, the learning device 10 according to this embodiment uses the encoder q φ The KL divergence between (z|x) and an exponentially tilted prior distribution using the L2 norm is approximated with high accuracy, and this approximation is used to train a variational autoencoder. Because this approximation can be calculated relatively easily, the learning device 10 according to this embodiment makes it possible to easily train high-dimensional and complex data using a variational autoencoder.

[0088] Moreover, in the learning device 10 according to this embodiment, it is possible to use the relative density ratio as a function other than the L2 norm, and by using this relative density ratio to exponentially tilt the prior distribution, it is possible to improve the performance of the variational autoencoder.

[0089] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.

[0090] 10 Learning device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage device 108 Processor 109 Bus 201 Input unit 202 Learning unit 203 Learning data set storage unit 204 Parameter storage unit

Claims

1. An input unit that inputs a dataset for learning a variational autoencoder, and a learning unit that learns the variational autoencoder from the dataset so as to maximize a variational lower bound including an approximation of the KL divergence between the encoder of the variational autoencoder and the prior distribution, where the prior distribution of the latent variable is a distribution obtained by exponentially tilting a standard Gaussian distribution using a predetermined function. The approximation of the KL divergence is composed of the KL divergence between the encoder and the standard Gaussian distribution, the expected value of the function with respect to the encoder, and a predetermined constant. A learning device.

2. The function according to claim 1, wherein the function is a relative density ratio between the integrated posterior probability of the encoder with respect to the distribution of the data constituting the dataset and the prior distribution.

3. The learning unit according to claim 2, wherein the learning unit learns the variational autoencoder by alternately performing maximization of the variational lower bound and optimization of an objective function for learning a neural network that approximates the relative density ratio.

4. An input procedure in which a computer inputs a dataset for learning a variational autoencoder, and a learning procedure in which the computer learns the variational autoencoder from the dataset so as to maximize a variational lower bound including an approximation of the KL divergence between the encoder of the variational autoencoder and the prior distribution, where the prior distribution of the latent variable is a distribution obtained by exponentially tilting a standard Gaussian distribution using a predetermined function. The approximation of the KL divergence is composed of the KL divergence between the encoder and the standard Gaussian distribution, the expected value of the function with respect to the encoder, and a predetermined constant. A learning method.