Methods and apparatus for random inference among multiple commonly represented random variables

By combining latent variable models and variational posterior training methods, the computational complexity of inference among multiple random variables in deep neural networks is solved, achieving efficient information decomposition and inference, which is suitable for applications such as image generation and image translation.

CN111105012BActive Publication Date: 2025-10-31SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201911010647.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-03
Filing Date
2019-10-23
Publication Date
2025-10-31
Estimated Expiration
2039-10-23

AI Technical Summary

Technical Problem

Existing deep neural network generative models struggle to effectively perform stochastic inferences among multiple random variables when simulating potential sources by learning distributions, especially when dealing with joint and conditional distributions, which suffer from computational complexity and difficulties in information decomposition.

Method used

We employ a joint latent variable model and variational posterior training method. By introducing a joint latent variable model and variational posterior, we utilize deep neural networks for variational learning, and combine reparameterization techniques and regularization terms to achieve information decomposition and inference among multiple random variables.

Benefits of technology

It achieves efficient information decomposition and inference among multiple random variables, and can generate more accurate joint distributions and conditional distributions, making it suitable for practical applications such as image generation, image captioning, and image translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111105012B_ABST
    Figure CN111105012B_ABST
Patent Text Reader

Abstract

This paper discloses a method and system. The method includes developing a joint latent variable model having a first variable, a second variable, and joint latent variables representing common information between the first and second variables; generating a variational posterior of the joint latent variable model; training the variational posterior; and performing inference of the first variable from the second variable based on the variational posterior.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is based on and claims priority to U.S. Provisional Patent Application No. 62 / 751,108, filed October 26, 2018, with the United States Patent and Trademark Office, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to neural networks. In particular, this disclosure relates to methods and apparatus for making stochastic inferences among multiple random variables through common representations. Background Technology

[0004] Inference through learning random representations is one of the most promising research areas in machine learning. The goal of this research is to model potential sources by learning distributions from observed data, a problem known as the "generative problem." In recent years, generative models based on deep neural networks have been proposed, including approximate probabilistic inference based on variational methods and generative adversarial networks. Summary of the Invention

[0005] According to one embodiment, a method is provided. The method includes developing a joint latent variable model having a first variable, a second variable, and a joint latent variable representing common information between the first and second variables; generating a variational posterior of the joint latent variable model; training the variational posterior; and performing inference of the first variable from the second variable based on the variational posterior.

[0006] According to one embodiment, a system is provided. The system includes at least one decoder, at least one encoder, and a processor configured to develop a joint latent variable model having a first variable, a second variable, and joint latent variables representing common information between the first variable and the second variable; generate a variational posterior of the joint latent variable model; train the variational posterior; and perform inference of the first variable from the second variable based on the variational posterior. Attached Figure Description

[0007] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0008] Figure 1 , Figure 2 and Figure 3 This is a diagram of the implicit model according to one embodiment;

[0009] Figure 4 This is a flowchart of a method for performing inference according to one embodiment;

[0010] Figure 5 and Figure 6 It is a diagram of an implicit model with local stochasticity according to one embodiment;

[0011] Figure 7 This is a diagram of a system according to one embodiment;

[0012] Figure 8 It is a graph of a dataset and data pairs based on one embodiment;

[0013] Figure 9 This is a diagram showing the combined generation result based on one embodiment;

[0014] Figure 10 , Figure 11 , Figure 12 and Figure 13 It is a diagram showing the results of conditional generation and style transfer based on one embodiment; and

[0015] Figure 14 This is a block diagram of an electronic device in a network environment according to one embodiment. Detailed Implementation

[0016] In the following description, embodiments of the present disclosure are described in detail with reference to the accompanying drawings. It should be noted that the same elements will be designated by the same reference numerals, although they are shown in different drawings. Specific details such as detailed configurations and components provided in the following description are only to aid in a comprehensive understanding of the embodiments of the present disclosure. Therefore, it will be apparent to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Additionally, for clarity and conciseness, descriptions of well-known functions and structures have been omitted. The terminology described below is defined in consideration of the functions in this disclosure and may vary depending on the user, the user's intent, or habits. Therefore, the definitions of the terms should be determined based on the content throughout the specification.

[0017] This disclosure can have various modifications and embodiments, which are described in detail below with reference to the accompanying drawings. However, it should be understood that this disclosure is not limited to these embodiments, but includes all modifications, equivalents, and alternatives within the scope of this disclosure.

[0018] Although terms including ordinal numbers such as first, second, etc., may be used to describe various elements, structural elements are not limited by these terms. These terms are used only to distinguish one element from another. For example, a first structural element may be referred to as a second structural element without departing from the scope of this disclosure. Similarly, a second structural element may also be referred to as a first structural element. As used herein, the term "and / or" includes any and all combinations of one or more related terms.

[0019] The terminology used herein is for describing various embodiments of this disclosure only and is not intended to limit this disclosure. Unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In this disclosure, it should be understood that the terms “comprising” or “having” indicate the presence of features, numbers, steps, operations, structural elements, components, or combinations thereof, and do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, structural elements, components, or combinations thereof.

[0020] Unless otherwise defined, all terms used herein have the same meaning as understood by one of ordinary skill in the art to which this disclosure pertains. Terms such as those defined in commonly used dictionaries shall be interpreted as having the same meaning as in the context of the relevant field, and shall not be construed as having an ideal or overly formal meaning unless expressly defined in this disclosure.

[0021] An electronic device according to one embodiment can be one of various types of electronic devices. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computers, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. According to one embodiment of this disclosure, the electronic device is not limited to those described above.

[0022] The terminology used in this disclosure is not intended to limit the disclosure, but rather to include various changes, equivalents, or substitutions of corresponding embodiments. Regarding the description of the drawings, similar reference numerals may be used to refer to similar or related elements. Unless otherwise expressly indicated by the relevant context, nouns corresponding to the singular form of an item may include one or more things. As used herein, each phrase such as “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B, or C,” “at least one of A, B, and C,” and “at least one of A, B, or C” may include all possible combinations of items listed together in the corresponding phrase. As used herein, terms such as “first (1)” and “2)” may also include all possible combinations of items listed together in the corresponding phrase. st "Second (2)" nd The terms “first” and “second” can be used to distinguish one component from another, but are not intended to limit other aspects of the components (e.g., importance or order). The intention is that if an element (e.g., the first element) is referred to, with or without the terms “operably” or “communically”, “connected to,” “linked to,” “connected to,” or “connected to” another element (e.g., the second element), it means that the element can be connected to the other element directly (e.g., wired), wirelessly, or via a third element.

[0023] As used herein, the term "module" can include a unit implemented in hardware, software, or firmware, and can be used interchangeably with other terms such as "logic," "logic block," "part," and "circuit." A module can be a single integrated component or minimum unit suitable for performing one or more functions, or a portion thereof. For example, according to one embodiment, a module can be implemented in the form of an application-specific integrated circuit (ASIC).

[0024] Figure 1 This is a diagram of the latent variable model 100 according to the embodiment. Z 102 represents p θ (z) and X104 represents p data (x). The goal of probabilistic inference is to infer from its samples Learning data distribution p data (x). In a given parametric distribution family {p θ In the case of (x)∶θ∈θ}, it may be difficult to directly find the parameter θ such that p θ (x) approximates p well. data (x). Therefore, one approach is to introduce, for example... Figure 1 The latent variable φ is shown as a relaxed problem. Table 1 shows various terms and definitions for reference, although these terms are not limited to those given in Table 1.

[0025] Table 1

[0026]

[0027] Typically, p θ (z) is a fixed prior distribution, such as the standard normal distribution. Assuming this model, the joint distribution can be represented by the latent variable model p. θ (z)p θ (x|z) is used to characterize, to satisfy p data (x)≈p θ (x)∶=∫p θ (z)p θ (x|z)dz. This can be formulated as in equation (1).

[0028]

[0029] Equation (1) is equivalent to solving the maximum likelihood estimate as in equation (2).

[0030]

[0031] However, due to integration, the edge density p θ (x) is often difficult to find. To address this problem, variational methods introduce a variational distribution q. φ (z|x) as the true posterior pθ The approximate value of (z|x) can be obtained. From this, the upper bound of the log-loss can be derived, as shown in equation (3).

[0032]

[0033] In equation (3), the term D(q) φ (z|x)‖p θ (z)) acts as a regular term, while the term This can be interpreted as reconstruction log loss. The true risk of the loss function is shown in equation (4).

[0034]

[0035]

[0036] In equation (4), h(p) represents the differential entropy of density p. Furthermore, note equation (5).

[0037]

[0038] Therefore, for variational learning (e.g., minimizing risk R(θ,φ)), according to equation (5), for the maximum likelihood estimation as in equation (2), it is a relaxed optimization problem; and according to equation (4), it is equivalent to solving the relaxed joint distribution matching problem in equation (6), rather than directly solving the marginal distribution matching problem as in equation (1):

[0039]

[0040] As described in this paper, θ represents the parameters of the latent model, and φ represents the parameters of the variational posterior. If multiple variational posteriors exist in the same model, the corresponding condition variables (e.g., φ) are represented by the subscripts of the parameters φ. x ,φ x,y wait).

[0041] Figure 2 and 3 This is a diagram of a variable model based on one embodiment. Model 200 depicts the entire latent variable model p. θ (z)p θ (x|z)p θ (y|z), where Z 202 represents p θ (z),X 204 represents p data (x) and Y 206 represents p data (y). Model 302 is a partial latent variable model p of Model 200. θ (z)p θ(x|z), Model 304 is a partially latent variable model p of Model 200. θ (z)p θ (y|z).

[0042] For most parameter distributions, the regularization term D(q) φ (z|x)‖p θ (z) can only be analyzed and calculated through the distribution parameters, which are determined by the differentiable mapping of θ and φ. A challenge arises when attempting to obtain the derivative of the reconstruction error. This derivative can be approximated by a simple Monte Carlo approximation as equation (7):

[0043]

[0044] Where L is the number of samples used for approximation. The non-differentiable sampling process z... l ~q φ (z|x) prohibits the optimization process. To avoid this problem, a reparameterization technique is introduced, the key idea of ​​which is z l ~q φ (z|x) can be obtained through z l =g φ (x | ∈) are sampled, where g φ It is a differentiable deterministic mapping, and ∈~p(∈) are considered to be auxiliary random variables that are easy to sample.

[0045] Variational learning that utilizes tractable distributed parameterization based on deep neural networks using reparameterization techniques is also known as Autoencoder Variational Bayes (AEVB).

[0046] Figure 4 This is a flowchart 400 according to an embodiment for performing variational inference through common representation. Given paired data The aim is to learn the statistical relationship between X and Y. Although the embodiment described herein uses two variables, the model can be extended to multiple variables, as will be apparent to those skilled in the art based on the disclosure herein. By learning p data (x|y) and p data (y|x) is used to perform bidirectional or multidirectional inference from X to Y and Y to X. Therefore, the system trains the joint latent variable model p. θ (z)p θ (z|x)p θ (z|y) and then inference is made through the joint representation Z. Z represents the joint representation of X and Y, which contains enough information about X and Y to ensure conditional independence X╨Y∣Z. Therefore, given any observation, a correct guess of Z should be sufficient to infer any variable X or Y.

[0047] At position 402, the system develops a joint latent variable model and a variational posterior for the joint variable model. This is done to infer Y based on the observations {X = x}, and to access the model posterior p. θ (z|x) can be obtained by sampling z. (0) ~p θ (z|x) and sampling y (0) ~p θ (y∣∣z (0) ) to approximate from p data The sampling process of Y for (y|x). The conditional distribution can also be approximated by equation (8).

[0048]

[0049] In equation (8), S is the number of samples used for approximation. Therefore, for the discrete random variable Y, an approximate maximum a posteriori (MAP) test can be performed as in equation (9).

[0050]

[0051] Given the model posterior p θ (z|x), the approximate inference from X to Y can be obtained via the joint representation Z through p θ (y|z)p θ (z|x) is used to perform random inference, and is called random inference through joint representation.

[0052] Model posterior p θ (z|x) is difficult to handle. Therefore, a variational posterior (or a partial approximate posterior) is generated and used. if Approaching p θ (z∣x) can be replaced by... Inference is made, which is called variational inference through joint representation.

[0053] At position 404, the system trains the variational posterior. The training of the variational posterior utilizes the concept that variational learning is equivalent to a matched joint distribution, as shown in equation (10).

[0054]

[0055] If p is given θ (z)p θ (x|z) makes p θ (x)≈p data (x), then minimizing the objective will ensure Therefore, variational posterior can be obtained Training a fully latent variable model p with the help of θ (z)p θ (x|z)p θ(y|z) ensures p θ (x|z) and p θ (y|z) is suitable for the joint distribution p data (x,y). Then, variational posterior. and / or It can be used for inference.

[0056] The joint model and variational posterior can be trained based on various algorithms. The first algorithm is a two-step training algorithm. The latent model p can be trained by solving equation (11). θ (z)p θ (x|z)p θ (y|z) and variational encoder

[0057]

[0058] Then, the variational encoder can be trained by solving equation (12).

[0059]

[0060] In equation (12), it is assumed that equation (11) finds a good model likelihood p. θ (x|z) and p θ (y|z), and the decoder parameters θ are frozen. The encoder based on Y can also be trained in equation (12).

[0061] Alternatively, the joint model and variational posterior can be trained simultaneously using the hyperparameter α in the second algorithm. Given that the hyperparameter α > 0, the latent model p θ (z)p θ (x|z)p θ (y|z) and variational encoder and (That is, the marginal variational posterior) is combined by solving equation (13).

[0062]

[0063] In order to (i.e., marginal variational posterior) can be trained together, and the objective can be set as equation (14).

[0064]

[0065] hyperparameter α x and α y It is specified and is greater than 0.

[0066] The first algorithm has no hyperparameters to be tuned, so it can be easily generalized to multivariate models; however, for multivariate models, the number of hyperparameters in the second algorithm becomes much larger. Nevertheless, the second algorithm becomes advantageous in semi-supervised learning settings because it has only a few paired samples and a relatively large number of unpaired samples, thus it can naturally incorporate semi-supervised datasets.

[0067] At position 406, the system performs inference. Typically, training a latent variable model p can be difficult. θ (z)p θ (x|z)p θ (y|z), because decoder p θ (x|z) and p θ (y|z) needs to extract relevant information for generating X or Y from the joint representation Z. Therefore, the system can use the local random variable U for each variable. x and U y The form introduces randomness into the model to allow for some structural relaxation, thus fitting the desired information decomposition. By introducing random variables, the latent variable model becomes p θ (u x ,u y ,w)p θ (x|u x ,w)p θ (y|u y ,w), where X and Y are only from (U x ,W) and (U y (,W) is generated. To simplify the notation, z = (u x ,u y ,w)z x =(u x w) and z y =(u y ,w).

[0068] W plays the role of a joint representation of X and Y, while local randomness U x and U y Capture the remaining randomness. The model can be trained in the same way as a model without randomness. The variational loss function is now given as in equation (15) and the corresponding risk is given as in equation (16).

[0069]

[0070]

[0071] Inference can be performed through conditional generation, such that given x (0) The system will conditionally generate Y. If If it has already been trained, then the sample It can be acquired, and Y can be generated as in

[0072] Pattern generation can also be performed. Given paired data of (X, Y), where X is a digital image and Y is the image's label. Y is almost entirely determined by the image, and local randomness U can be generated. y Given a reference image x (0) It can generate the same style x with different tags. (0) A set of images. Style generation can be performed in three steps. First, this system samples... And store Secondly, for each label y, this system samples... And store w(y) = w. Third, this system generates...

[0073] This process assumes that W represents only common information (e.g., label information in this example) and that local random variables adopt all other properties (randomness) except for common information (e.g., the pattern of numbers). However, the native optimization equation (16) does not guarantee this information decomposition because the objective only encourages matching the joint distribution. For example, the extreme case where W contains all information about X and Y is still a possible solution.

[0074] Figure 5 and Figure 6 This is a graph of an implicit model with local stochasticity according to one embodiment. Model 500 shows the base model, and model 600 shows the degradation when H(Y|X)≈0.

[0075] At 408, the system extracts common information. The extraction of common information can be performed during the inference step at 406. Common information extraction can be performed by adding a mutual information regularization term and optimizing the cost function, as described in equation (20) below (i.e., optimizing the training joint model of equation (20) to produce a model capable of extracting common information).

[0076] Wyner's common information represents the minimum descriptive rate for simulating the distributions of two random variables. Given two random variables X and Y, the common information is given by equation (17).

[0077]

[0078] XWY forms a Markov chain. This can be interpreted as the minimum number of bits of information sent to each location to approximate X and Y, while allowing each location to use an arbitrary amount of local randomness.

[0079] Furthermore, equation (18) is not easy to handle during training.

[0080] I(X,Y;W)=D(p θ (w)p θ (x,y|w)‖p θ (x,y)p θ (w)) (18)

[0081] Therefore, variational approximation as in equation (19) is used.

[0082] I(X,Y;W)≈D(p data (x,y)q φ (w|x,y)‖p data (x,y)p θ (w)) (19)

[0083] Therefore, by adding the regularization term λD(p) data (x,y)q φ (w|x,y)‖p data (x,y)p θ The loss function for common information extraction (w) is given by equation (20).

[0084]

[0085] The systems and methods disclosed herein include processes / algorithms that can be integrated into electronic devices, networks, etc. These systems and methods can perform inference using image capture devices of electronic devices and other practical applications that will be apparent to those skilled in the art, such as image captioning, image-in-painting, image-to-image translation (e.g., horse to zebra), style transfer, text-to-speech, and missing / future frame prediction in video.

[0086] Figure 7 This is a diagram of system 700 according to an embodiment, wherein the above embodiment is integrated and implemented. In system 700, input 702 is sent to encoder 704, which processes input 702 as in equation (21) for sampling 706 as in equation (22).

[0087]

[0088]

[0089] The output of encoder 704 is also sent to calculate regularization loss 708 as in equation (23) and regularization loss 710 for common information extraction as in equation (24).

[0090]

[0091]

[0092] The regularization loss 708 and the regularization loss 710 for common information extraction are calculated using the prior distribution p(z) 711. Samples 712 are sent to the decoder 714, while the common representation 716, as in equation (25), is post-processed 718 and also sent to the decoder 714.

[0093]

[0094] Decoder 714 outputs prediction 720. Reconstruction loss 722 is calculated based on the output of decoder 714, post-processing of common representation 716, sample 712 and input 702, as shown in equation (26).

[0095]

[0096] The reconstruction loss 722 is combined with the regularization loss 708 and the regularization loss 710 for common information extraction to determine the loss function 724 as in equation (27).

[0097]

[0098]

[0099] in and In addition, each parameter θ1,…,θ K ,φ [K] Representing different networks, and

[0100] Figure 8 This is a graph of the datasets and dataset pairs used in the add-one experiment according to the embodiment. Dataset 800 is the MNIST (Modified National Institute of Standards and Technology) dataset of handwritten digits, and add-one dataset pair 802 is extracted from dataset 800.

[0101] In the following Figure 9-13 In this context, |Z| represents the dimension of the common variables, while |U| and |V| represent the dimensions of the local variables. Figure 9 The graph shows the combined results generated according to the embodiments. The result of 900 comes from the MNIST-MNIST+1 experiment, where |Z| = 16 and |U| = |V| = 4. The result of 902 comes from MNIST-SVHN (street view room number)+1, where |Z| = 8 and |U| = |V| = 2.

[0102] Figure 10 These are the results of conditional generation and style transfer according to the embodiments. Result 1000 comes from the MNIST-MNIST+1 experiment used for conditional generation, where |Z| = 8 and |U| = |V| = 4. Result 1002 comes from the MNIST-MNIST+1 experiment used for style transfer, where |Z| = 8 and |U| = |V| = 4.

[0103] Figure 11 The results are diagrams of conditional generation and style transfer according to the embodiments. The result of 1100 comes from the MNIST-MNIST+1 experiment used for conditional generation, where |Z|=16 and |U|=|V|=4. The result of 1102 comes from the MNIST-MNIST+1 experiment used for style transfer, where |Z|=16 and |U|=|V|=4.

[0104] Figure 12 These are the results of conditional generation and style transfer according to the embodiments. The result of 1200 comes from the MNIST-SVHN+1 experiment used for conditional generation, where |Z| = 8 and |U| = |V| = 2. The result of 1202 comes from the MNIST-SVHN+1 experiment used for style transfer, where |Z| = 8 and |U| = |V| = 2.

[0105] Figure 13 These are the results of conditional generation and style transfer according to the embodiments. The result of 1300 comes from the MNIST-SVHN+1 experiment used for conditional generation, where |Z| = 8 and |U| = |V| = 16. The result of 1302 comes from the MNIST-SVHN+1 experiment used for style transfer, where |Z| = 8 and |U| = |V| = 16.

[0106] Figure 14 This is a block diagram of an electronic device 1401 in a network environment 1400 according to one embodiment. (See reference...) Figure 14In network environment 1400, electronic device 1401 can communicate with electronic device 1402 via a first network 1498 (e.g., a short-range wireless communication network); or communicate with electronic device 1404 or server 1408 via a second network 1499 (e.g., a long-range wireless communication network). Electronic device 1401 can communicate with electronic device 1404 via server 1408. Electronic device 1401 may include processor 1420, memory 1430, input device 1450, sound output device 1455, display device 1460, audio module 1470, sensor module 1476, interface 1477, haptic module 1479, camera module 1480, power management module 1488, battery 1489, communication module 1490, subscriber identification module (SIM) 1496, or antenna module 1497. In one embodiment, at least one component (e.g., display device 1460 or camera module 1480) may be omitted from electronic device 1401; or one or more other components may be added to electronic device 1401. In one embodiment, some components may be implemented as a single integrated circuit (IC). For example, sensor module 1476 (e.g., fingerprint sensor, aperture sensor, or illuminance sensor) may be embedded in display device 1460 (e.g., display).

[0107] Processor 1420 can execute software (e.g., program 1440) to control at least one other component (e.g., hardware or software component) of electronic device 1401 coupled to processor 1420, and can perform various data processing or calculations. As at least part of data processing or calculation, processor 1420 can load commands or data received from another component (e.g., sensor module 1476 or communication module 1490) into volatile memory 1432; process the commands or data stored in volatile memory 1432; and store the resulting data in non-volatile memory 1434. Processor 1420 may include a main processor 1421 (e.g., a central processing unit (CPU) or application processor (AP)) and an auxiliary processor 1423 (e.g., a graphics processing unit (GPU), image signal processor (ISP), sensor hub processor, or communication processor (CP)), which may operate independently of or in conjunction with the main processor 1421. Alternatively or alternatively, the auxiliary processor 1423 may be adapted to consume less power than the main processor 1421, or to perform specific functions. The auxiliary processor 1423 may be implemented as separate from or part of the main processor 1421.

[0108] The auxiliary processor 1423 can control at least some functions or states associated with at least one component of the electronic device 1401 (e.g., display device 1460, sensor module 1476, or communication module 1490) to replace the main processor 1421 when the main processor 1421 is inactive (e.g., in sleep) or to work with the main processor 1421 when the main processor 1421 is active (e.g., executing an application). According to one embodiment, the auxiliary processor 1423 (e.g., an image signal processor or a communication processor) can be implemented as part of another component (e.g., camera module 1480 or communication module 1490) associated with the functionality of the auxiliary processor 1423.

[0109] Memory 1430 may store various data used by at least one component of electronic device 1401 (e.g., processor 1420 or sensor module 1476). The various data may include, for example, software (e.g., program 1440) and input or output data for associated commands. Memory 1430 may include volatile memory 1432 or non-volatile memory 1434.

[0110] Program 1440 may be stored as software in memory 1430 and may include, for example, an operating system (OS) 1442, middleware 1444, or application program 1446.

[0111] Input device 1450 can receive commands or data from outside electronic device 1401 (e.g., a user) to be used by other components of electronic device 1401 (e.g., processor 1420). Input device 1450 may include, for example, a microphone, mouse, or keyboard.

[0112] The sound output device 1455 can output sound signals to the outside of the electronic device 1401. The sound output device 1455 may include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as playing multimedia or recording; and the receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0113] Display device 1460 can visually provide information to the outside of electronic device 1401 (e.g., to a user). Display device 1460 may include, for example, a display, a holographic device, or a projector, and control circuitry for controlling one of the respective display, holographic device, and projector. According to one embodiment, display device 1460 may include touch circuitry adapted to detect touch, or sensor circuitry (e.g., a pressure sensor) adapted to measure the intensity of the force caused by touch.

[0114] The audio module 1470 can convert sound into electrical signals and vice versa. According to one embodiment, the audio module 1470 can obtain sound via an input device 1450; or output sound via a sound output device 1455 or headphones of an external electronic device 1402 that are directly (e.g., wired) or wirelessly connected to the electronic device 1401.

[0115] Sensor module 1476 can detect the operating state of electronic device 1401 (e.g., power or temperature) or the environmental state outside electronic device 1401 (e.g., user state), and then generate an electrical signal or data value corresponding to the detected state. Sensor module 1476 may include, for example, a gesture sensor, gyroscope sensor, atmospheric pressure sensor, magnetic sensor, accelerometer, handle sensor, proximity sensor, color sensor, infrared (IR) sensor, biometric sensor, temperature sensor, humidity sensor, or illuminance sensor.

[0116] Interface 1477 may support one or more specified protocols for electronic device 1401 to connect directly (e.g., wired) or wirelessly to external electronic device 1402. According to one embodiment, interface 1477 may include, for example, a High Definition Multimedia Interface (HDMI), a Universal Serial Bus (USB) interface, a Secure Digital Card (SD) interface, or an audio interface.

[0117] Connection terminal 1478 may include a connector through which electronic device 1401 can be physically connected to external electronic device 1402. According to one embodiment, connection terminal 1478 may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0118] The tactile module 1479 can convert electrical signals into mechanical stimulation (e.g., vibration or motion) or electrical stimulation, which can be recognized by a user through tactile sensation or muscle sensation. According to one embodiment, the tactile module 1479 may include, for example, a motor, a piezoelectric element, or an electrical stimulator.

[0119] Camera module 1480 can capture still or moving images. According to one embodiment, camera module 1480 may include one or more lenses, an image sensor, an image signal processor, or a flash.

[0120] The power management module 1488 can manage the power supplied to the electronic device 1401. The power management module 1488 can be implemented as at least part of, for example, a power management integrated circuit (PMIC).

[0121] Battery 1489 can supply power to at least one component of electronic device 1401. According to one embodiment, battery 1489 may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0122] Communication module 1490 can support the establishment of a direct (e.g., wired) or wireless communication channel between electronic device 1401 and external electronic devices (e.g., electronic device 1402, electronic device 1404, or server 1408) and communication via the established communication channel. Communication module 1490 may include one or more communication processors that can operate independently of processor 1420 (e.g., AP) and support direct (e.g., wired) or wireless communication. According to one embodiment, communication module 1490 may include wireless communication module 1492 (e.g., cellular communication module, short-range wireless communication module, or Global Navigation Satellite System (GNSS) communication module) or wired communication module 1494 (e.g., local area network (LAN) communication module or power line communication (PLC) module). A corresponding one of these communication modules can communicate via a first network 1498 (e.g., a short-range communication network such as Bluetooth). TM The communication module 1492 can communicate with external electronic devices via a standard such as Wi-Fi Direct or Infrared Data Association (IrDA) or a second network 1499 (e.g., a remote communication network such as a cellular network, the Internet, or a computer network (e.g., a LAN or a wide area network (WAN)). These various types of communication modules can be implemented as a single component (e.g., a single IC) or as multiple components that are separate from each other (e.g., multiple ICs). The wireless communication module 1492 can use subscriber information (e.g., International Mobile Subscriber Identity (IMSI)) stored in the subscriber identification module 1496 to identify and authenticate electronic devices 1401 in communication networks such as the first network 1498 or the second network 1499.

[0123] Antenna module 1497 can transmit or receive signals or power to or from the exterior of electronic device 1401 (e.g., an external electronic device). According to one embodiment, antenna module 1497 may include one or more antennas, and thus, at least one antenna suitable for a communication scheme used in a communication network (e.g., a first network 1498 or a second network 1499) can be selected, for example, via communication module 1490 (e.g., wireless communication module 1492). Signals or power can then be transmitted or received between communication module 1490 and the external electronic device via the selected at least one antenna.

[0124] At least some of the aforementioned components can be interconnected and transmit signals (e.g., commands or data) between them via peripheral communication schemes (e.g., bus, general purpose input and output (GPIO), serial peripheral interface (SPI), or mobile industry processor interface (MIPI)).

[0125] According to one embodiment, commands or data can be sent or received between electronic device 1401 and external electronic device 1404 via server 1408 connected to a second network 1499. Each of electronic devices 1402 and 1404 can be a device of the same or different type as electronic device 1401. All or some operations performed in electronic device 1401 can be performed at one or more of the external electronic devices 1402, 1404, or 1408. For example, if electronic device 1401 is required to automatically perform a function or service or in response to a request from a user or another device, instead of performing a function or service, or in addition to performing a function or service, electronic device 1401 can request one or more external electronic devices to perform at least a portion of a function or service. The one or more external electronic devices receiving the request can perform at least a portion of the requested function or service, or additional functions or services related to the request, and transmit the result of the execution to electronic device 1401. Electronic device 1401 can provide the result, with or without further processing, as at least part of a response to the request. For this purpose, cloud computing, distributed computing, or client-server computing technologies can be used, for example.

[0126] One embodiment may be implemented as software (e.g., program 1440) comprising one or more instructions stored in a storage medium (e.g., internal memory 1436 or external memory 1438) readable by a machine (e.g., electronic device 1401). For example, a processor of electronic device 1401 may invoke at least one of the one or more instructions stored in the storage medium and execute it with or without one or more other components under the processor's control. Thus, the machine can be operated to perform at least one function according to the invoked at least one instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. The term "non-transitory" means that the storage medium is a tangible device and does not include signals (e.g., electromagnetic waves), but the term does not distinguish between locations where data is stored semi-permanently in the storage medium and locations where data is temporarily stored in the storage medium.

[0127] According to one embodiment, the methods of this disclosure may be included and provided in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., an optical disc read-only memory (CD-ROM)); or via an app store (e.g., the Play Store). TM The computer program product may be distributed online (e.g., downloaded or uploaded); or directly between two user devices (e.g., smartphones). If distributed online, at least a portion of the computer program product may be temporarily generated or at least temporarily stored in a machine-readable storage medium, such as the memory of a manufacturer's server, an app store server, or a relay server.

[0128] According to one embodiment, each of the above-described components (e.g., a module or program) may include a single entity or multiple entities. One or more of the above-described components may be omitted, or one or more other components may be added. Alternatively or additionally, multiple components (e.g., modules or programs) may be integrated into a single component. In this case, the integrated component can still perform one or more functions of each of the multiple components in the same or similar manner as they were performed by a corresponding component of the multiple components prior to integration. Operations performed by modules, programs, or other components may be performed sequentially, in parallel, repeatedly, or heuristically; or one or more operations may be performed in a different order or omitted; or one or more other operations may be added.

[0129] Although certain embodiments of this disclosure have been described in detail herein, modifications may be made in various forms without departing from the scope of this disclosure. Therefore, the scope of this disclosure is to be determined not only based on the described embodiments, but also on the appended claims and their equivalents.

Claims

1. A method for neural networks, comprising: Develop a joint latent variable model with a first variable, a second variable, and a joint latent variable representing common information between the first variable and the second variable, wherein the first variable and the second variable are data used for image captioning, image in picture, image-to-image translation, style transfer, text-to-speech, or prediction of missing frames or future frames in video; Generate the variational posterior of the joint latent variable model; Training the variational posterior; Based on the variational posterior, inference of the first variable is performed from the second variable, wherein performing the inference includes conditionally generating the first variable from the second variable; and Extracting the common information between the first variable and the second variable, wherein extracting the common information includes adding a regularization term to the loss function.

2. The method of claim 1, further comprising adding local randomness to the joint latent variable model.

3. The method as described in claim 2, wherein, Adding the local randomness includes separating the joint latent variables into common latent variables and local latent variables.

4. The method of claim 2, wherein, Performing the inference includes generating at least one of the first and second variables.

5. The method of claim 1, wherein training the variational posterior comprises training the decoder in the joint latent variable model using a fully approximate posterior of the joint latent variable model.

6. The method of claim 5, wherein training the variational posterior further comprises fixing the parameters of the decoder and training the marginal variational posterior using the trained decoder.

7. The method of claim 1, wherein training the variational posterior comprises jointly training the joint latent variable model, the fully approximate posterior, and the marginal variational posterior using hyperparameters.

8. A system for a neural network, comprising: At least one decoder; At least one encoder; as well as The processor is configured as follows: Develop a joint latent variable model, which has a first variable, a second variable, and joint latent variables representing common information between the first variable and the second variable, wherein the first variable and the second variable are data used for image captioning, image in picture, image-to-image translation, style transfer, text-to-speech, or prediction of missing frames in video, or prediction of future frames in video; Generate the variational posterior of the joint latent variable model; Training the variational posterior; Inference of the first variable is performed from the second variable based on the variational posterior by conditionally generating the first variable from the second variable; and By adding a regularization term to the loss function, the common information between the first variable and the second variable is extracted.

9. The system of claim 8, wherein, The processor is also configured to add local randomness to the joint latent variable model.

10. The system of claim 9, wherein, The processor is also configured to add the local randomness by separating the joint latent variables into common latent variables and local latent variables.

11. The system of claim 9, wherein, The processor is also configured to perform the inference by generating patterns for the first variable and / or the second variable.

12. The system of claim 8, wherein, The processor is also configured to train the variational posterior by training at least one decoder in the joint latent variable model using a fully approximate posterior of the joint latent variable model.

13. The system of claim 12, wherein, The processor is also configured to train the variational posterior by fixing the parameters of the at least one decoder, and to train the marginal variational posterior using the trained at least one decoder.

14. The system of claim 8, wherein, The processor is also configured to train the marginal variational posterior by jointly training the joint latent variable model, the fully approximate posterior, and the variational posterior using hyperparameters.

Citation Information

Patent Citations

  • Method and system used for establishing data model for relational data

    CN106156067A

  • Convergence test device, method and program

    JP2015162233A