Neural network learning device, neural network learning method, and program
By designing a neural network containing encoder and decoder, and using specific loss functions and monotonic relationship learning methods, the problem of high-dimensional data analysis in the existing technology requires professional knowledge and inability to deal with missing information is solved, and efficient and automated data analysis is achieved.
Patent Information
- Application Number
- JP2024140933
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-05-17
AI Technical Summary
Existing analytical methods such as NMF and IRM require the professional knowledge of data analysts when processing high-dimensional data, and cannot effectively handle the missing information, resulting in poor progress in the analysis task.
A method using a neural network including an encoder and a decoder is designed to process the missing situation of input information through a specific loss function, and to make the latent variable have a monotonic relationship with the input vector by learning.
It enables efficient analysis of high-dimensional data without the need for a data analyst and can handle the problem of missing input information, simplifying the analysis process.
Smart Images

Figure 0007673866000008 
Figure 0007673866000009 
Figure 0007673866000010
Abstract
Description
[Technical field]
[0001] The present invention relates to techniques for training neural networks. [Background technology]
[0002] Various methods have been proposed as techniques for analyzing large amounts of high-dimensional data. For example, there are methods that use Non-negative Matrix Factorization (NMF) in Non-Patent Document 1 and Infinite Relational Model (IRM) in Non-Patent Document 2. Using these methods, it is possible to find characteristic properties of data and to group data with common properties into clusters. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Lee, DD and Seung, HS, “Learning the parts of objects by non-negative matrix factorization,” Nature, 401, pp.788-791, 1999. [Non-Patent Document 2] Kemp, C., Tenenbaum, JB, Griffiths, TL, Yamada, T. and Ueda, N., “Learning systems of concepts with an infinite relational model,” AAAI06(Proceedings of the 21st national conference on Artificial intelligence, pp.381-388, 2006. Summary of the Invention [Problem to be solved by the invention]
[0004] Analytical methods using NMF and IRM often require advanced analytical skills such as those possessed by data analysts. However, data analysts are often not familiar with the high-dimensional data to be analyzed (hereafter referred to as the target data), and in such cases, collaboration with an expert on the target data is necessary, but this work sometimes does not go smoothly. Therefore, a method is needed that allows analysis by the target data expert alone, without the need for a data analyst.
[0005] Consider analysis using a neural network including an encoder and a decoder, such as the Variational AutoEncoder (VAE) in Reference Non-Patent Document 1. Here, the encoder is a neural network that converts an input vector into a latent variable vector, and the decoder is a neural network that converts the latent variable vector into an output vector. The latent variable vector is a vector with a lower dimension than the input vector and the output vector, and is a vector with latent variables as elements. If high-dimensional data to be analyzed is converted using an encoder that has been trained to make the input vector and the output vector almost identical, it can be compressed into low-dimensional secondary data, but since the relationship between the data to be analyzed and the secondary data is unknown, it cannot be applied to analysis work as it is. Here, training to make the data almost identical means that ideally it is preferable to train to make them completely identical, but in reality, it is necessary to train to make them almost identical due to restrictions on training time, etc., so that training is performed in a manner that the data are considered to be identical and the processing is terminated when a certain condition is met.
[0006] (Reference non-patent document 1: Kingma, DP and Welling, M., “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114, 2013.) In addition, there may be cases where some of the input vectors representing the input information are missing, and no information is present. In this case, unless the system can properly handle the absence of input information, the system cannot be applied to analysis work.
[0007] Therefore, an object of the present invention is to provide a technique for training a neural network including an encoder and a decoder using a loss function that does not treat the absence of input information as a loss, even when the input information does not exist. [Means for solving the problem]
[0008] One aspect of the present invention is a neural network learning device that trains a neural network including an encoder that converts an input vector into a latent variable vector having latent variables as elements and a decoder that converts the latent variable vector into an output vector so that the input vector and the output vector become substantially identical, and includes a learning unit that performs learning by repeating a parameter update process that updates parameters included in the neural network, and the encoder, when each piece of input information included in a predetermined input information group corresponds to one of three cases, that is, positive information, negative information, or no information exists, represents each piece of input information as a positive information bit that is 1 when the input information corresponds to positive information and 0 when no information exists or the input information corresponds to negative information, and a negative information bit that is 1 when the input information corresponds to negative information and 0 when no information exists or the input information corresponds to positive information. The encoder is configured with a plurality of layers, and the layer that receives the input vector obtains a plurality of output values from the input vector, and each of the output values is obtained by adding together a value of each positive information bit included in the input vector to which a weight parameter has been assigned and a value of each negative information bit included in the input vector to which a weight parameter has been assigned. The parameter update process is performed so as to reduce the value of a loss function including a sum of all input information in a group of input information for learning losses, the sum of which is larger when the input information corresponds to positive information, the smaller the probability that the input information obtained by the decoder corresponds to positive information, and is larger when the input information corresponds to negative information, the smaller the probability that the input information obtained by the decoder corresponds to negative information, and is approximately 0 when the input information does not exist. Effect of the Invention
[0009] According to the present invention, even when input information does not exist, it is possible to train a neural network including an encoder and a decoder using a loss function that does not treat the absence of input information as a loss. [Brief description of the drawings]
[0010] [Figure 1] FIG. 2 is a diagram illustrating an example of analysis target data. [Diagram 2] FIG. 1 is a block diagram showing a configuration of a neural network learning device 100. [Diagram 3] 3 is a flowchart showing the operation of the neural network learning device 100. [Figure 4] FIG. 2 is a diagram illustrating an example of a functional configuration of a computer that realizes each device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Hereinafter, an embodiment of the present invention will be described in detail. Note that components having the same functions are given the same reference numbers and duplicated explanations will be omitted.
[0012] Before describing each embodiment, the notation used in this specification will be described.
[0013] ^ (caret) represents a superscript. For example, x y^z Yes z is a superscript to x, and x y^z Yes z is a subscript to x. Also, _ (underscore) represents a subscript. For example, x y_z Yes z is a superscript to x, and x y_z Yes z is a subscript on x.
[0014] In addition, the superscripts "^" and "~" such as ^x and ~x for a certain letter x should actually be written directly above the "x", but due to the constraints of the description in the specification, they are written as ^x and ~x.
[0015] <Technical background> Here, a learning method of a neural network including an encoder and a decoder used in an embodiment of the present invention will be described. The neural network used in an embodiment of the present invention is a neural network including an encoder that converts an input vector into a latent variable vector and a decoder that converts the latent variable vector into an output vector. In an embodiment of the present invention, this neural network is trained so that the input vector and the output vector become substantially identical. In an embodiment of the present invention, in order to make a certain latent variable included in the latent variable vector larger as the magnitude of a certain property included in the input vector increases, or to make a certain latent variable included in the latent variable vector smaller, the latent variable is trained to have the following characteristic (hereinafter referred to as characteristic 1).
[0016] [Feature 1] The latent variable is trained to have monotonicity with respect to the input vector. Here, the latent variable having monotonicity with respect to the input vector means that the latent variable vector has either a monotonically increasing relationship, that is, the larger the input vector is, the larger the latent variable vector becomes, or a monotonically decreasing relationship, that is, the larger the input vector is, the smaller the latent variable vector becomes. Note that the magnitude of the input vector or the latent variable vector is based on the order relationship with respect to the vector (i.e., a relationship defined using the order relationship with respect to each element of the vector), and for example, the following order relationship can be used.
[0017] Vector v = (v1, …, v n ), v'=(v'1, …, v' n ), v≦v' is true for all elements of vectors v and v', i.e., the i-th element v of vector v i , the i-th element v' of vector v' i (where i = 1, …, n), v i ≦v' i This means that the following holds true.
[0018] Specifically, learning to have monotonicity with respect to an input vector means learning to have either the first or second relationship with the input vector:
[0019] The first relationship is a relationship in which, when two input vectors are designated as a first input vector and a second input vector, the element value of the first input vector is greater than the element value of the second input vector for at least one element of the input vector and the element value of the first input vector is equal to or greater than the element value of the second input vector for all remaining elements of the input vector, then, when a latent variable vector obtained by converting the first input vector is designated as a first latent variable vector and a latent variable vector obtained by converting the second input vector is designated as a second latent variable vector, the element value of the first latent variable vector for at least one element of the latent variable vector is greater than the element value of the second latent variable vector and the element value of the first latent variable vector is equal to or greater than the element value of the second latent variable vector for all remaining elements of the latent variable vector.
[0020] The second relationship is a relationship in which, when two input vectors are designated as a first input vector and a second input vector, the element value of the first input vector is greater than the element value of the second input vector for at least one element of the input vector and the element value of the first input vector is greater than or equal to the element value of the second input vector for all remaining elements of the input vector, then, when a latent variable vector obtained by converting the first input vector is designated as a first latent variable vector and a latent variable vector obtained by converting the second input vector is designated as a second latent variable vector, the element value of the first latent variable vector for at least one element of the latent variable vector is smaller than the element value of the second latent variable vector and the element value of the first latent variable vector is less than or equal to the element value of the second latent variable vector for all remaining elements of the latent variable vector.
[0021] For convenience, the first relationship may be expressed as a relationship in which the latent variable and the input vector have a monotonically increasing relationship, and the second relationship may be expressed as a relationship in which the latent variable and the input vector have a monotonically decreasing relationship. Therefore, the expression that the latent variable has monotonicity with respect to the input vector can be said to be a convenient expression that indicates that the latent variable has either the first relationship or the second relationship.
[0022] In an embodiment of the present invention, since the input vector and the output vector are learned to be substantially the same, instead of learning using the relationship between the input vector and the latent variable vector, learning may be performed using the relationship between the latent variable vector and the output vector. Specifically, learning may be performed so that the output vector has either the third or fourth relationship with the latent variable vector. Note that the third relationship below is equivalent to the first relationship described above, and the fourth relationship below is equivalent to the second relationship described above.
[0023] The third relationship is a relationship in which, when two latent variable vectors are designated as a first latent variable vector and a second latent variable vector, the element value of the first latent variable vector is greater than the element value of the second latent variable vector for at least one element of the latent variable vector and the element value of the first latent variable vector is equal to or greater than the element value of the second latent variable vector for all remaining elements of the latent variable vector, the output vector obtained by converting the first latent variable vector is designated as a first output vector, and the output vector obtained by converting the second latent variable vector is designated as a second output vector, the element value of the first output vector is greater than the element value of the second output vector for at least one element of the output vector and the element value of the first output vector is equal to or greater than the element value of the second output vector for all remaining elements of the output vector.
[0024] A fourth relationship is a relationship in which, when two latent variable vectors are designated as a first latent variable vector and a second latent variable vector, the element value of the first latent variable vector is greater than the element value of the second latent variable vector for at least one element of the latent variable vector and the element value of the first latent variable vector is greater than or equal to the element value of the second latent variable vector for all remaining elements of the latent variable vector, then, when an output vector obtained by converting the first latent variable vector is designated as a first output vector and an output vector obtained by converting the second latent variable vector is designated as a second output vector, the element value of the first output vector for at least one element of the output vector is smaller than the element value of the second output vector and the element value of the first output vector is less than or equal to the element value of the second output vector for all remaining elements of the output vector.
[0025] For convenience, the third relationship may be expressed as the output vector being in a monotonically increasing relationship with the latent variable, and the fourth relationship may be expressed as the output vector being in a monotonically decreasing relationship with the latent variable. Furthermore, for convenience, the output vector may be expressed as being monotonic with the latent variable to indicate that the output vector has either the third or fourth relationship.
[0026] By learning the latent variables to have the above characteristic 1, a latent variable is provided that satisfies the condition that the greater the magnitude of a certain property contained in the input vector, the larger a certain latent variable contained in the latent variable vector is, or the smaller a certain latent variable contained in the latent variable vector is.
[0027] In the embodiment of the present invention, the latent variables may be learned as having the following feature (hereinafter referred to as feature 2) in addition to the above feature 1.
[0028] [Feature 2] The system learns so that the values that a latent variable can take are within a specified range.
[0029] By training the latent variable to have the above feature 2 in addition to feature 1, a latent variable that satisfies the condition that the greater the magnitude of a certain property contained in the input vector, the greater a certain latent variable contained in the latent variable vector is, or the smaller a certain latent variable contained in the latent variable vector is, is provided as a parameter that is easy for general users to understand.
[0030] We will explain the constraints for training a neural network including an encoder that outputs a latent variable having the above feature 1. Specifically, we will explain the following two constraints.
[0031] [Constraint 1] Learn to minimize a loss function that includes a loss term for monotonicity violations.
[0032] [Constraint 2] Learning is performed by constraining all weight parameters of the decoder to be non-negative values, or by constraining all weight parameters of the decoder to be non-positive values.
[0033] First, the neural network to be learned will be described. For example, the following VAE can be used. The encoder and decoder are each a two-layer neural network, and the first and second layers of the encoder and the first and second layers of the decoder are all fully connected. The input vector that is the input to the first layer of the encoder is, for example, a 60-dimensional vector. The output vector that is the output of the second layer of the decoder is a vector that restores the input vector. In addition, a sigmoid function is used as the activation function of the second layer of the encoder. This makes the values of the elements of the latent variable vector that is the output of the encoder (i.e., each latent variable) 0 to 1. Note that the latent variable vector is a vector with a lower dimensionality than the input vector, for example, a 5-dimensional vector. For example, Adam (see Reference Non-Patent Document 2) can be used as a learning method.
[0034] (Reference Non-Patent Document 2: Kingma, D. P. and Jimmy B., “Adam: A Method for Stochastic Optimaization,” arXiv:1412.6980, 2014) Instead of setting the range of possible values of the latent variable to [0, 1], it can also be set to [m, M] (where m < M), and in this case, for example, the following function s(x) can be used instead of the sigmoid function as the activation function.
Number
Number
Number
[0035] term L mono is the three-term L real , L syn-encoder (p) , L syn-decoder (p) It is the sum of the terms L real is a term for establishing monotonicity between the latent variables and the output vector, that is, a term related to feature 1. In other words, the term L real is a term for establishing a monotonically increasing relationship between the latent variable and the output vector, or a term for establishing a monotonically decreasing relationship between the latent variable and the output vector. On the other hand, the term L syn-encoder (p) and the term L syn-decoder (p) is a term relating to feature 2.
[0036] Below, we use the term L to establish a monotonically increasing relationship between the latent variables and the output vector. realAn example of the method will be described together with the learning method. First, actual data (in the example of FIG. 1, a list of correct answers for each student) is input as an input vector, and a latent variable vector (hereinafter referred to as the original latent variable vector) is obtained as the output of the encoder. Next, a vector is obtained in which the value of at least one element of the original latent variable vector is replaced with a value smaller than the value of the element. The vector obtained here is hereinafter referred to as an artificial latent variable vector. In addition, when limiting the lower limit of the range in which the element values can be taken, a vector in which the value of at least one element of the original latent variable vector is replaced with a value equal to or larger than the lower limit of the range and smaller than the value of the element can be obtained as an artificial latent variable vector. In this specification, the term "artificial" such as "artificial latent variable vector" is used, but this is a term to explain that the artificial latent variable vector is not the original latent variable, and is not intended to be obtained by manual work.
[0037] Here, an example of a process for obtaining an artificial latent variable vector is shown. For example, an artificial latent variable vector is generated by decreasing the value of one element of an original latent variable vector within the range that the value of the element can take. The artificial latent variable vector thus obtained has a smaller value of one element than the original latent variable vector, and the values of the other elements are the same. Note that a plurality of artificial latent variable vectors may be generated by decreasing the values of different elements of a latent variable vector within the range that the value of the element can take. In other words, if the latent variable vector is a five-dimensional vector, five artificial latent variable vectors are generated from one original latent variable vector. In addition, an artificial latent variable vector may be generated by decreasing the values of multiple elements of a latent variable vector within the range that the value of each element can take. In other words, an artificial latent variable vector may be generated in which the values of multiple elements are smaller than those of the original latent variable vector, and the values of the remaining elements are the same. In addition, for multiple combinations of multiple elements of a latent variable vector, the value of each element included in each combination may be decreased within the range that the value of each element can take, thereby generating multiple artificial latent variable vectors.
[0038] As a method for obtaining an element value of an artificial latent variable vector that is smaller than the element value of the original latent variable vector, if the lower limit of the range in which the element value can be is 0, for example, a method for multiplying the element value of the original latent variable vector by a random number in the interval (0, 1) to reduce the value and obtain the element value of the artificial latent variable vector, or a method for multiplying the element value of the original latent variable vector by 1 / 2 to halve the value and obtain the element value of the artificial latent variable vector may be used.
[0039] When using an artificial latent variable vector in which the element values of the original latent variable vector are replaced with values smaller than the element values, it is desirable that the value of each element of the output vector when the original latent variable vector is input is larger than the value of the corresponding element of the output vector when the artificial latent variable vector is input. Therefore, the term L real For example, it can be a term called Margin Ranking Error, which becomes large when the value of each element of the output vector when the original latent variable vector is input is smaller than the value of the corresponding element of the output vector when the artificial latent variable vector is input. Here, the margin ranking error L MRE is defined by the following equation, where Y is the output vector when the original latent variable vector is input, and Y' is the output vector when the artificial latent variable vector is input.
number
[0040] Instead of using a vector in which at least one element of the original latent variable vector is replaced with a value smaller than the value of the element, a vector in which at least one element of the original latent variable vector is replaced with a value larger than the value of the element may be used as the artificial latent variable vector. In this case, it is preferable that the value of each element of the output vector when the original latent variable is input is smaller than the value of the corresponding element of the output vector when the artificial latent variable is input. Therefore, the term L real may be a term that takes a large value when the value of each element of the output vector when the original latent variable vector is input is larger than the value of the corresponding element of the output vector when the artificial latent variable vector is input.
[0041] As a method for obtaining an element value of an artificial latent variable vector that is greater than the value of the element in the original latent variable vector, if the upper limit of the range in which the element value can be taken is limited, a value of the element of the artificial latent variable vector that is less than the upper limit of the range and greater than the value of the element in the original latent variable vector can be obtained. For example, a method for obtaining a value randomly selected from between the value of the element of the original latent variable vector and the upper limit of the range in which the value of the element in the original latent variable vector can be taken, or a method for obtaining an average value of the value of the element of the original latent variable vector and the upper limit of the range in which the value of the element in the original latent variable vector can be taken, can be used.
[0042] term L syn-encoder (p) is a term related to artificial data that is the upper limit of the range of values that all elements of the input vector can take, or the lower limit of the range of values that all elements of the input vector can take. For example, in the example of Figure 1, where each element of the input vector has a value of either 1 or 0, the term L syn-encoder (p) is a term related to artificial data in which the input vector is a vector (1, ..., 1) corresponding to all correct answers, or the input vector is a vector (0, ..., 0) corresponding to all incorrect answers. Specifically, the term Lsyn-encoder (1) is the binary cross entropy between the latent variable vector that is the output of the encoder when the input vector is the vector (1, …, 1) corresponding to all correct answers, and the ideal latent variable vector (1, …, 1) with all elements 1 (i.e., the upper limit of the range of possible values) when the input vector is the vector (1, …, 1) corresponding to all correct answers. In addition, the term L syn-encoder (2) is the binary cross entropy between the latent variable vector that is the output of the encoder when the input vector is the vector (0, …, 0) corresponding to all incorrect answers, and the ideal latent variable vector (0, …, 0) with all elements 0 (i.e., the lower limit of the range of possible values) when the input vector is the vector (0, …, 0) corresponding to all incorrect answers. The term L syn-encoder (1) is based on the requirement that when the input vector is (1, ..., 1), that is, when all elements of the input vector are 1 (i.e., the upper limit of the range of possible values), it is desirable that all elements of the latent variable vector are 1 (i.e., the upper limit of the range of possible values), and the term L syn-encoder (2) is based on the requirement that when the input vector is (0, ..., 0), i.e., when all elements of the input vector are 0 (i.e., the lower limit of the range of possible values), it is desirable that all elements of the latent variable vector be 0 (i.e., the lower limit of the range of possible values).
[0043] On the other hand, the term L syn-decoder (p) is a term related to artificial data that is the upper limit of the range of values that all elements of the output vector can take, or the lower limit of the range of values that all elements of the output vector can take. For example, in the example of Figure 1, where each element of the input vector has a value of either 1 or 0, the term L syn-decoder (p)is a term related to artificial data in which the output vector is a vector (1, ..., 1) corresponding to all correct answers, or the output vector is a vector (0, ..., 0) corresponding to all incorrect answers. Specifically, the term L syn-decoder (1) is the binary cross entropy between the output vector, which is the decoder output when the latent variable vector is a vector (1, …, 1) whose values for all elements are the upper limit of the range of possible values, and the ideal output vector, which is a vector (1, …, 1) whose elements are all 1 (i.e., corresponding to all correct answers) when the values for all elements of the latent variable vector are the upper limit of the range of possible values. In addition, the term L syn-decoder (2) is the binary cross entropy between the output vector, which is the decoder output when the latent variable vector is a vector (0, …, 0) whose values for all elements are the lower limit of the range of possible values, and the ideal output vector, which is a vector (0, …, 0) whose all elements are 0 (i.e., all questions are answered incorrectly) when the values for all elements of the latent variable vector are the lower limit of the range of possible values. syn-decoder (1) is based on the requirement that when the latent variable vector is (1, ..., 1), that is, when all elements of the latent variable vector are 1 (i.e., the upper limit of the range of possible values), it is desirable that all elements of the output vector are 1 (i.e., the upper limit of the range of possible values), and the term L syn-decoder (2) is based on the requirement that when the latent variable vector is (0, ..., 0), i.e., when all elements of the latent variable vector are 0 (i.e., the lower limit of the range of possible values), it is desirable that all elements of the output vector be 0 (i.e., the lower limit of the range of possible values).
[0044] The term L defined above realBy including the loss function, the neural network is trained to have the following characteristics: when two input vectors are taken as a first input vector and a second input vector, and the value of the element of the first input vector is greater than the value of the element of the second input vector for at least one element of the input vector, and the value of the element of the first input vector is equal to or greater than the value of the element of the second input vector for all remaining elements of the input vector, the latent variable vector obtained by converting the first input vector is taken as a first latent variable vector, and the latent variable vector obtained by converting the second input vector is taken as a second latent variable vector, the value of the element of the first latent variable vector is greater than the value of the element of the second latent variable vector for at least one element of the latent variable vector, and the value of the element of the first latent variable vector is equal to or greater than the value of the element of the second latent variable vector for all remaining elements of the latent variable vector. In addition, the term L real In addition to L syn-encoder (p) , L syn-decoder (p) Also, the loss function L includes the term L mono By including this in the loss function L, the neural network is trained so that the values of all elements of the latent variable vector are within the range [0, 1] (i.e., the range of possible values).
[0045] Next, a learning method for Constraint 2 will be described. In the explanation of the learning method for Constraint 2, the number of the input vector used for learning is s (s is an integer between 1 and S, and S is the number of learning data), the number of the element of the latent variable vector is j (j is an integer between 1 and J), the number of the elements of the input vector and the output vector is k (k is an integer between 1 and K, and K is an integer larger than J), and the input vector is X s Let the input vector X s The latent variable vector obtained by transforming s Then, the latent variable vector Z s Let P be the output vector obtained by converting s Let the input vector X s The kth element of x sk Then, the output vector P s The kth element of p skThen, the latent variable vector Z s The jth element of z sj Let us assume that.
[0046] The encoder receives an input vector X s Let Z be the latent variable vector s Any function that converts the input image into a vector image may be used, for example, a general VAE encoder. In addition, the loss function used for learning does not need to be a special one, and can be a conventionally used loss function, for example, the above-mentioned term L RC and the term L prior The sum of these can be used as the loss function.
[0047] The decoder generates a latent variable vector Z s The output vector P s which is trained by constraining all weight parameters of the decoder to be non-negative or constraining all weight parameters of the decoder to be non-positive.
[0048] We will explain the decoder constraints using an example where all weight parameters of a decoder consisting of one layer are constrained to be non-negative. Let us assume that the input vector X1, X2, ..., X is a vector that represents the student's answers to a test question with K questions, with correct answers being 1 and incorrect answers being 0. S If so, the input vector of the sth student is X s =(x s1 , x s2 , ..., x sK ) and the input vector X s The latent variable vector obtained by converting s =(z s1 , z s2 , ..., z sJ ) and the latent variable vector Z s The output vector obtained by converting s =(p s1 , p s2 , ..., p sK) is the probability that a student will correctly answer each test question, and therefore will need various categories of abilities, such as writing ability and diagramming ability, each with their own weighting. In order to make each element of the latent variable vector correspond to each category of ability, and to make the value of the latent variable corresponding to the category larger the greater the ability of each category possessed by the student, then the probability p sk Let j be the latent variable z sj The weight w for the kth test problem is jk It is advisable to express this as equation (5) with a non-negative value.
number
[0049] From the above, in order to make a certain latent variable included in the latent variable vector larger as the magnitude of a certain property included in the input vector increases, learning is performed while restricting all weight parameters of the decoder to be non-negative values. Also, as can be seen from the above explanation, in order to make a certain latent variable included in the latent variable vector smaller as the magnitude of a certain property included in the input vector increases, it is preferable to perform learning while restricting all weight parameters of the decoder to be non-positive values.
[0050] As mentioned above, in the example of Figure 1, each column represents a list of correct answers for each student. Using a trained encoder, the 60-dimensional list of correct answers for students is converted into 5-dimensional secondary data. The trained encoder converts the latent variables to have monotonicity with respect to the input vector, so the secondary data compressed to 5 dimensions reflects the characteristics of the list of correct answers for students. For example, if a list of correct answers for students in a Japanese or arithmetic test is converted to obtain a latent variable vector, the elements of the secondary data, which is the latent variable vector, can be data corresponding to, for example, writing ability or diagramming ability. Therefore, by analyzing secondary data instead of the list of correct answers for students, it is possible to reduce the burden on the analyst.
[0051] First Embodiment The neural network learning device 100 uses learning data to learn parameters of a neural network to be learned. Here, the neural network to be learned includes an encoder that converts an input vector into a latent variable vector and a decoder that converts the latent variable vector into an output vector. The latent variable vector is a vector with a lower dimension than the input vector and the output vector, and is a vector with latent variables as elements. The parameters of the neural network include a weight parameter and a bias parameter of the encoder, and a weight parameter and a bias parameter of the decoder. Learning is performed so that the input vector and the output vector are approximately the same. Learning is also performed so that the latent variables have monotonicity with respect to the input vector.
[0052] Here, it is described that the possible values of the elements of the input vector and the output vector are either 1 or 0, and the range of possible values of the latent variables, which are the elements of the latent variable vector, is [0, 1]. Note that the fact that the possible values of the elements of the input vector and the output vector are either 1 or 0 is just an example. The range of possible values of the elements of the input vector and the output vector may be [0, 1], or further, the range of possible values of the elements of the input vector and the output vector may not be [0, 1]. That is, taking a and b as any numbers satisfying a < b, the range of possible values of the elements of the input vector and the range of possible values of the elements of the output vector can be set as [a, b].
[0053] Hereinafter, the neural network learning device 100 will be described with reference to FIGS. 2 to 3. FIG. 2 is a block diagram showing the configuration of the neural network learning device 100. FIG. 3 is a flowchart showing the operation of the neural network learning device 100. As shown in FIG. 2, the neural network learning device 100 includes an initialization unit 110, a learning unit 120, an end condition determination unit 130, and a recording unit 190. The recording unit 190 is a component configured to appropriately record information necessary for the processing of the neural network learning device 100. The recording unit 190 records, for example, initialization data used for the initialization of the neural network. Here, the initialization data refers to the initial values of the parameters of the neural network, and for example, the initial values of the weight parameters and bias parameters of the encoder, and the initial values of the weight parameters and bias parameters of the decoder.
[0054] The operation of the neural network learning device 100 will be described according to FIG. 3.
[0055] In S110, the initialization unit 110 performs the initialization process of the neural network using the initialization data. Specifically, the initialization unit 110 sets initial values for each parameter of the neural network.
[0056] In S120, the learning unit 120 receives the learning data as input, performs a process of updating each parameter of the neural network using the learning data (hereinafter, referred to as a parameter update process), and outputs the parameters of the neural network together with information (e.g., the number of times the parameter update process has been performed) required for the termination condition determination unit 130 to determine the termination condition. The learning unit 120 uses a loss function to train the neural network, for example, by the backpropagation method. That is, in each parameter update process, the learning unit 120 performs a process of updating each parameter of the encoder and decoder so that the loss function becomes smaller.
[0057] Here, the loss function includes a term for making the latent variable monotonic with respect to the input vector. When the monotonicity is a relationship in which the latent variable increases monotonically with respect to the input vector, the loss function includes a term for making the output vector larger as the latent variable increases, for example, the margin ranking error term described in <Technical Background>. That is, the loss function includes at least one of the following: a term that becomes large when a vector in which the value of at least one element of the latent variable vector is replaced with a value smaller than the value is taken as an artificial latent variable vector, and the value of a corresponding element of the output vector when the latent variable vector is input is smaller than the value of any element of the output vector when the artificial latent variable vector is input; and a term that becomes large when a vector in which the value of at least one element of the latent variable vector is replaced with a value larger than the value is taken as an artificial latent variable vector, and the value of a corresponding element of the output vector when the latent variable vector is input is larger than the value of any element of the output vector when the artificial latent variable vector is input. Furthermore, in the case where the elements of the input vector are either 1 or 0, and the range of possible values of the elements of the latent variable vector is [0, 1], the loss function may include at least one term among the binary cross entropy between the latent variable vector and vector (1, ..., 1) (where the dimension of the vector is equal to the dimension of the latent variable vector) when the input vector is (1, ..., 1), the binary cross entropy between the latent variable vector and vector (0, ..., 0) (where the dimension of the vector is equal to the dimension of the latent variable vector) when the input vector is (0, ..., 0), the binary cross entropy between the output vector and vector (1, ..., 1) (where the dimension of the vector is equal to the dimension of the output vector) when the latent variable vector is (1, ..., 1), and the binary cross entropy between the output vector and vector (0, ..., 0) (where the dimension of the vector is equal to the dimension of the output vector) when the latent variable vector is (0, ..., 0).
[0058] On the other hand, when the monotonicity is a relationship in which the latent variable is monotonically decreasing with respect to the input vector, the loss function includes a term for making the output vector smaller as the latent variable becomes larger. That is, the loss function includes at least one of the following: a term that becomes a large value when a vector in which the value of at least one element of the latent variable vector is replaced with a value smaller than the value is taken as an artificial latent variable vector, and the value of a corresponding element of the output vector when the latent variable vector is input is larger than the value of any element of the output vector when the artificial latent variable vector is input, and a term that becomes a large value when a vector in which the value of at least one element of the latent variable vector is replaced with a value larger than the value is taken as an artificial latent variable vector, and the value of a corresponding element of the output vector when the latent variable vector is input is smaller than the value of any element of the output vector when the artificial latent variable vector is input. Furthermore, in the case where the elements of the input vector are either 1 or 0, and the range of possible values of the elements of the latent variable vector is [0, 1], the loss function may include at least one of the following terms: the binary cross entropy between the latent variable vector and vector (0, ..., 0) (where the dimension of the vector is equal to the dimension of the latent variable vector) when the input vector is (1, ..., 1); the binary cross entropy between the latent variable vector and vector (1, ..., 1) (where the dimension of the vector is equal to the dimension of the latent variable vector) when the input vector is (0, ..., 0); the binary cross entropy between the value of the output vector and vector (0, ..., 0) (where the dimension of the vector is equal to the dimension of the output vector) when the latent variable vector is (1, ..., 1); and the binary cross entropy between the value of the output vector and vector (1, ..., 1) (where the dimension of the vector is equal to the dimension of the output vector) when the latent variable vector is (0, ..., 0).
[0059] In S130, the end condition determination unit 130 takes as input the parameters of the neural network output in S120 and the information necessary to determine the end condition, and determines whether the end condition, which is a condition regarding the end of learning, is satisfied (for example, the number of times the parameter update process has been performed has reached a predetermined number of repetitions). If the end condition is satisfied, it outputs the parameters of the encoder obtained in the last S120 as learned parameters and ends the process. On the other hand, if the end condition is not satisfied, it returns to the process of S120.
[0060] (Modification example) Instead of setting the range of possible values of the latent variables, which are elements of the latent variable vector, to [0, 1], it may be set to [m, M] (where m < M), or as described above, the range of possible values of the elements of the input vector and the output vector may be set to [a, b]. Furthermore, the range of possible values for each element of the latent variable vector may be set individually, or the range of possible values for each element of the input vector and the output vector may be set individually. In this case, let the number of elements of the latent variable vector be j (j is an integer from 1 to J, and J is an integer of 2 or more), the range of possible values of the j-th element be [m j , M j (where m j < M j ), and let the number of elements of the input vector and the output vector be k (k is an integer from 1 to K, and K is an integer greater than J), the range of possible values of the k-th element be [a k , b k (where a k < b k ). Then, the terms included in the loss function may be as follows. When the monotonicity is a relationship in which the latent variable increases monotonically with respect to the input vector, the loss function is the cross-entropy between the latent variable vector and the vector (M1, …, M K ) when the input vector is (b1, …, b J ), and the cross-entropy between the latent variable vector and the vector (m1, …, m K ) when the input vector is (a1, …, a J) and the cross entropy of the latent variable vector (M1, …, M J ) and the output vector when vector (b1, …, b K ), the cross entropy between the latent variable vector (m1, …, m J ) and the output vector when vector (a1, …, a K ), and the cross-entropy with
[0061] On the other hand, if monotonicity is the relationship in which the latent variables are monotonically decreasing with respect to the input vector, the loss function is the input vector (b1, …, b K ) and the latent variable vector (m1, …, m J ) and the cross entropy with the input vector (a1, …, a K ) and the latent variable vector (M1, …, M J ) and the cross entropy of the latent variable vector (M1, …, M J ) and the output vector when vector (a1, …, a K ), the cross entropy between the latent variable vector (m1, …, m J ) and the output vector when vector (b1, …, b K ) and the cross entropy with the vector. The above-mentioned cross entropy is an example of a value corresponding to the magnitude of the difference between vectors, and any value that increases as the difference between vectors increases, such as the mean squared error (MSE), can be used instead of the above-mentioned cross entropy.
[0062] In the above explanation, an example was described in which the number of dimensions of the latent variable vector is two or more, but the number of dimensions of the latent variable vector may be one. That is, the above-mentioned J may be one. When the number of dimensions of the latent variable vector is one, the above-mentioned "latent variable vector" may be read as "latent variable", and "the value of at least one element of the latent variable vector" may be read as "the value of the latent variable", and there is no condition regarding "all remaining elements of the latent variable vector".
[0063] Finally, we will explain the analysis process. Using an encoder with learned parameters (a learned encoder), the data to be analyzed is converted into lower-dimensional secondary data. Here, the secondary data refers to the latent variable vector obtained by inputting the data to be analyzed into the learned encoder. Since this secondary data is lower-dimensional than the data to be analyzed, analyzing the secondary data is easier than directly analyzing the data to be analyzed.
[0064] According to the first embodiment, it is possible to train a neural network including an encoder and a decoder so as to obtain encoder parameters such that a certain latent variable included in a latent variable vector becomes larger or a certain latent variable included in a latent variable vector becomes smaller as the magnitude of a certain property included in an input vector becomes larger. Then, by using the trained encoder to convert high-dimensional analysis target data to obtain low-dimensional secondary data, the burden on the analyst can be reduced.
[0065] <Second embodiment> In the first embodiment, a method for learning an encoder that outputs a latent variable vector in which a certain latent variable included in the latent variable vector is larger or a latent variable vector in which a certain latent variable included in the latent variable vector is smaller as the magnitude of a certain property included in the input vector increases, by learning using a loss function including a term for making the latent variable monotonic with respect to the input vector, has been described. Here, a method for learning an encoder that outputs a latent variable vector in which a certain latent variable included in the latent variable vector is larger or a latent variable vector in which a certain latent variable included in the latent variable vector is smaller as the magnitude of a certain property included in the input vector increases, by learning so that the weight parameters of the decoder satisfy a predetermined condition, is described.
[0066] The neural network training device 100 of this embodiment differs from the neural network training device 100 of the first embodiment only in the operation of the training unit 120. Therefore, only the operation of the training unit 120 will be described below.
[0067] In S120, the learning unit 120 receives the learning data as input, performs a process of updating each parameter of the neural network using the learning data (hereinafter, referred to as a parameter update process), and outputs the parameters of the neural network together with information (e.g., the number of times the parameter update process has been performed) required for the termination condition determination unit 130 to determine the termination condition. The learning unit 120 uses a loss function to train the neural network, for example, by the backpropagation method. That is, in each parameter update process, the learning unit 120 performs a process of updating each parameter of the encoder and decoder so that the loss function becomes smaller.
[0068] The neural network training device 100 of this embodiment trains in such a way that the weight parameters of the decoder satisfy a predetermined condition. When the neural network training device 100 trains so that the latent variables have a monotonically increasing relationship with the input vector, the neural network training device 100 trains in such a way that the weight parameters of the decoder are all non-negative. That is, in this case, in each parameter update process performed by the training unit 120, the parameters of the encoder and the decoder are updated while restricting all the weight parameters of the decoder to be non-negative values. More specifically, the decoder included in the neural network training device 100 includes a layer that obtains a plurality of output values from a plurality of input values, and each output value of the layer includes a term obtained by adding up the plurality of input values by giving a weight parameter to each of the plurality of input values, and each parameter update process performed by the training unit 120 is performed while satisfying the condition that all the weight parameters of the decoder are non-negative values. In addition, a term obtained by adding multiple input values each having a weighting parameter assigned thereto can also be referred to as a term obtained by adding together all of the products of each input value and the weighting parameter corresponding to each input value, or a term obtained by weighting and adding together multiple input values using the weighting parameters corresponding to each input value as weights, etc.
[0069] On the other hand, when the neural network training device 100 trains so that the latent variables have a monotonically decreasing relationship with the input vector, the learning unit 120 trains in a manner that satisfies the condition that all of the decoder weight parameters are non-positive. That is, in this case, in each parameter update process performed by the training unit 120, the encoder and decoder parameters are updated while restricting all of the decoder weight parameters to be non-positive values. More specifically, the decoder included in the neural network training device 100 includes a layer that obtains multiple output values from multiple input values, and each output value of the layer includes a term obtained by adding up the multiple input values by giving a weight parameter to each of the multiple input values, and each parameter update process performed by the training unit 120 is performed while satisfying the condition that all of the decoder weight parameters are non-positive values.
[0070] When neural network training device 100 performs training in a manner that satisfies the condition that all decoder weight parameters are non-negative, it is preferable that the initial values of the decoder weight parameters in the initialization data recorded by recording unit 190 be non-negative values. Similarly, when neural network training device 100 performs training in a manner that satisfies the condition that all decoder weight parameters are non-positive, it is preferable that the initial values of the decoder weight parameters in the initialization data recorded by recording unit 190 be non-positive values.
[0071] In the second embodiment, as in the first embodiment, the number of dimensions of the latent variable vector may be 1. When the number of dimensions of the latent variable vector is 1, the above-mentioned "latent variable vector" may be read as "latent variable".
[0072] (Modification) Although the learning in which the weight parameters of the decoder all satisfy the condition that they are non-negative has been described as learning in which the latent variables have a monotonically increasing relationship with the input vector, by using an encoder having parameters in which the positive and negative signs of all the parameters of the encoder obtained by learning (i.e., all the learned parameters) are inverted, it is possible to obtain an encoder in which the latent variables have a monotonically decreasing relationship with the input vector. Similarly, the learning in which the weight parameters of the decoder all satisfy the condition that they are non-positive has been described as learning in which the latent variables have a monotonically decreasing relationship with the input vector, by using an encoder having parameters in which the positive and negative signs of all the parameters of the encoder obtained by learning (i.e., all the learned parameters) are inverted, it is possible to obtain an encoder in which the latent variables have a monotonically increasing relationship with the input vector.
[0073] That is, the neural network training device 100 may further include a sign inversion unit 140 as shown by a dashed line in Fig. 2, and may also perform S140 shown by a dashed line in Fig. 3. In S140, the sign inversion unit 140 may invert the positive and negative signs of each learned parameter output in S130, that is, obtain and output the learned sign-inverted parameters by leaving the absolute values of each learned parameter as they are, making positive values negative and negative values positive, as the learned sign-inverted parameters. More specifically, when the encoder included in the neural network training device 100 is configured with one or more layers that obtain multiple output values from multiple input values, and each output value of each layer includes a term obtained by adding up each of the multiple input values by giving a weight parameter, the neural network training device 100 may further include a sign inversion unit 140 that outputs a sign-inverted weight parameter obtained by inverting the positive and negative signs of each weight parameter of the encoder obtained by learning (i.e., each learned parameter output by the termination condition determination unit 130).
[0074] In the analysis process, an encoder with learned, sign-inverted parameters is used to convert the data to be analyzed into lower-dimensional, secondary data.
[0075] According to the second embodiment, it is possible to train a neural network including an encoder and a decoder so as to obtain encoder parameters such that a certain latent variable included in a latent variable vector becomes larger or a certain latent variable included in a latent variable vector becomes smaller as the magnitude of a certain property included in an input vector becomes larger. Then, by using the trained encoder to convert high-dimensional analysis target data to obtain low-dimensional secondary data, the burden on the analyst can be reduced.
[0076] <Third embodiment> In the above example of analyzing the test results of students for test questions, if the test results (information on whether the answer is correct or incorrect) of students for all test questions are obtained, the value of the latent variable obtained by converting the list of correct answers of students for the test can be a value corresponding to the level of ability of each student for each ability category, by using the trained encoder of the first or second embodiment. However, if the test results of students for some test questions are not available, such as when the students have taken the Japanese and arithmetic tests but not the science and social studies tests, it is possible to obtain the corresponding latent variable by using a value corresponding to the level of ability of each student for each ability category by using further ingenuity. A neural network learning device 100 including this ingenuity will be described as a third embodiment.
[0077] First, the technical background of the neural network learning device 100 of this embodiment will be described using an example of analyzing a student's test results for test questions. The neural network of this embodiment and its learning have the following characteristics a to c.
[0078] [Feature a] The test result for each question is represented by a correct answer bit and an incorrect answer bit.
[0079] In the neural network of this embodiment, answers to test questions that a student has not taken are treated as no answers, and answers to each question are represented as input vectors using a correct answer bit, where a correct answer is 1 and no answers and incorrect answers are 0, and an incorrect answer bit, where an incorrect answer is 1 and no answers and correct answers are 0. For example, the correct answer bit for the kth test question of the sth student is expressed as x (1) sk The incorrect answer bit is x (0) sk Then, the input vector of the sth student for the test problem with the number of questions being K is the set of correct answers {x (1) s1 , x (1) s2 , ..., x (1) sK} and the incorrect answer bit group {x (0)s1 , x (0) s2 , ..., x (0) sK} is a vector consisting of
[0080] [Feature b] At the beginning of the encoder, a layer is provided that obtains intermediate information from the correct answer bit group and the incorrect answer bit group so that the absence of an answer does not affect the encoder output.
[0081] In the neural network of this embodiment, the first layer of the encoder (a layer that receives an input vector) is set to the intermediate information set {q s1 ,q s2 , ..., q sH} intermediate information q sh The following shall be obtained.
number
[0082] [Feature c] A loss function is used that does not count non-response as a loss.
[0083] In the learning of this embodiment, the decoder is a latent variable vector Z s =(z s1 , z s2 , ..., z sJ ) vector P s =(p s1 , p s2 , ..., p sK ) as the output vector, and the loss L sk , x (1) sk If p is 1 (i.e., the answer is correct), then -log(p sk ), x (0) sk If is 1 (i.e., if the answer is incorrect), then -log(1-p sk ), x(1) sk x (1) sk If the answer is also 0 (i.e., no answer), set it to 0. Then, the loss L for all test questions k=1, ..., K in the training data s=1, ..., S is sk The sum of (equation (7) below) is the above-mentioned term L RC We use a loss function including
number
[0084] Next, the neural network training device 100 of this embodiment will be described with respect to points that differ from the neural network training devices 100 of the first and second embodiments.
[0085] As described above as feature a, the input vector of the encoder is represented by treating answers to test questions that each student has not taken as no answer, and using a correct answer bit in which a correct answer is 1 and no answer and incorrect answers are 0, and an incorrect answer bit in which an incorrect answer is 1 and no answer and correct answers are 0. That is, the learning data is represented by treating answers to K test questions for the sth student as no answer, and using a correct answer bit in which a correct answer is 1 and no answer and incorrect answers are 0, and an incorrect answer bit in which an incorrect answer is 1 and no answer and correct answers are 0. In other words, the learning data is represented by using a correct answer bit and an incorrect answer bit for each question for each student i for learning, and by treating answers to K test questions for each student i as no answer, and using a correct answer bit and an incorrect answer bit in which a correct answer bit is 1 and an incorrect answer bit is 0, and by treating answers to incorrect answers as 0 and no answer and correct answers as 0.
[0086] The first layer of the encoder (the layer that takes the input vector as input) obtains multiple intermediate information pieces from the input vector for the sth student, as described above as feature b. Each intermediate information piece is the sum of the correct bit values to which a weight parameter has been assigned and the incorrect bit values to which a weight parameter has been assigned.
[0087] In the parameter update process performed by the learning unit 120 of the neural network learning device 100 of this embodiment, as described above, when the sth student correctly answers the kth question, the probability p sk The smaller the value, the larger the value. If the sth student answers the kth question incorrectly, the probability p skThe process updates each parameter of the encoder and decoder so that the loss function, which includes the sum of losses over all training data and all test questions, is small, where the smaller the value is, the larger the loss function is, and the larger the loss function is if the sth student does not answer the kth question, which is 0.
[0088] In addition, this embodiment is not limited to the above-mentioned example of analyzing the test results of students for test questions, but can also be applied to the analysis of information acquired by multiple sensors. For example, if a sensor detects the presence or absence of a specific situation, two types of information can be acquired: information that the specific situation has been detected and information that the specific situation has not been detected. However, when collecting and analyzing information acquired by multiple sensors via a communication network, there is a possibility that, due to loss of communication packets, it is not possible to obtain information that the specific situation has been detected or not detected for any of the sensors, and none of the information exists. In other words, there are cases where the information that can be used for analysis is one of three types of information for each sensor: information that the specific situation has been detected, information that the specific situation has not been detected, and information that none of the information exists. This embodiment can also be used in such cases.
[0089] That is, to explain without specializing in the form of use, the neural network training device 100 of this embodiment is a neural network training device that trains a neural network including an encoder that converts an input vector into a latent variable vector having latent variables as elements and a decoder that converts the latent variable vector into an output vector so that the input vector and the output vector become approximately identical, and includes a learning unit 120 that performs training by repeating a parameter update process that updates parameters included in the neural network, and the encoder, when each piece of input information included in a predetermined input information group corresponds to any one of the three cases of positive information, negative information, and no information, converts each piece of input information into a positive information bit that is 1 when the input information corresponds to positive information and 0 when no information exists or the input information corresponds to negative information, and a positive information bit that is 1 when the input information corresponds to negative information and 0 when no information exists or the input information corresponds to positive information. The encoder is configured with a plurality of layers, and the layer receiving the input vector obtains a plurality of output values from the input vector, and each output value is the sum of a value of each positive information bit included in the input vector to which a weight parameter has been applied and a value of each negative information bit included in the input vector to which a weight parameter has been applied. The parameter update process is performed so as to reduce the value of a loss function including the sum of all input information in the input information group for learning losses, which is a value that increases as the probability that the input information obtained by the decoder (i.e., the input information restored by the decoder) corresponds to positive information decreases when the input information corresponds to positive information, and a value that increases as the probability that the input information obtained by the decoder corresponds to negative information decreases when the input information corresponds to negative information, and is approximately 0 when there is no input information.
[0090] In the example of analyzing a student's test results for a test question, a correct answer corresponds to the input information "corresponding to positive information", an incorrect answer corresponds to the input information "corresponding to negative information", and no answer corresponds to "non-existence of information". In the example of analyzing information acquired by a sensor, information that a specific situation has been detected corresponds to the input information "corresponding to positive information", information that a specific situation has not been detected corresponds to the input information "corresponding to negative information", and the absence of any information corresponds to "non-existence of information".
[0091] In the analysis process, for example, in the case of analyzing a student's test results for test questions, as described above for feature a, for the student being analyzed, answers to test questions that the student did not take are treated as no answers, and the answers to each question are represented using correct answer bits where correct answers are 1 and no answers are 0, and incorrect answer bits where incorrect answers are 1 and no answers are 0, and these are used as input vectors for the encoder, and are converted into low-dimensional secondary data using an encoder with learned parameters set.
[0092] <Additional Notes> 4 is a diagram showing an example of the functional configuration of a computer that realizes each of the above-mentioned devices (i.e., each node). The processing in each of the above-mentioned devices can be implemented by having a recording unit 2020 load a program for causing a computer to function as each of the above-mentioned devices, and having the control unit 2010, input unit 2030, output unit 2040, etc. operate.
[0093] The device of the present invention has, as a single hardware entity, an input section to which a keyboard or the like can be connected, an output section to which a liquid crystal display or the like can be connected, a communication section to which a communication device (e.g., a communication cable) capable of communicating with the outside of the hardware entity can be connected, a CPU (which may also have a central processing unit, cache memory, registers, etc.), memories such as RAM and ROM, an external storage device such as a hard disk, and buses connecting the input section, output section, communication section, CPU, RAM, ROM, and external storage device so that data can be exchanged between them. If necessary, the hardware entity may also be provided with a device (drive) capable of reading and writing a recording medium such as a CD-ROM. A physical entity equipped with such hardware resources is, for example, a general-purpose computer.
[0094] The external storage device of the hardware entity stores the programs required to realize the above-mentioned functions and the data required in the processing of these programs (not limited to the external storage device, but for example the programs may be stored in a ROM, which is a read-only storage device). Data obtained by the processing of these programs is stored appropriately in the RAM, the external storage device, etc.
[0095] In the hardware entity, each program stored in an external storage device (or ROM, etc.) and the data required for processing each program are loaded into memory as necessary, and interpreted, executed, and processed by the CPU as appropriate, so that the CPU realizes a predetermined function (each component represented as the above, "unit," "means," etc.).
[0096] The present invention is not limited to the above-described embodiment, and can be modified as appropriate without departing from the spirit of the present invention. In addition, the processes described in the above-described embodiment are not limited to being executed in chronological order according to the order described, but may be executed in parallel or individually depending on the processing capacity of the device that executes the processes or as necessary.
[0097] As described above, when the processing functions of the hardware entities (the devices of the present invention) described in the above embodiments are realized by a computer, the processing contents of the functions that the hardware entities should have are described by a program. Then, by executing this program on a computer, the processing functions of the hardware entities are realized on the computer.
[0098] The program describing the processing contents can be recorded in a non-transitory computer-readable recording medium. The computer-readable recording medium may be, for example, a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, or any other suitable medium. Specifically, for example, a hard disk drive, a flexible disk, a magnetic tape, or the like can be used as the magnetic recording device, a DVD (Digital Versatile Disc), a DVD-RAM (Random Access Memory), a CD-ROM (Compact Disc Read Only Memory), a CD-R (Recordable) / RW (ReWritable), or the like can be used as the optical disk, an MO (Magneto-Optical disc), or the like can be used as the magneto-optical recording medium, and an EEP-ROM (Electronically Erasable and Programmable-Read Only Memory) or the like can be used as the semiconductor memory.
[0099] The program may be distributed, for example, by selling, transferring, lending, etc. portable recording media such as DVDs and CD-ROMs on which the program is recorded. Furthermore, the program may be distributed by storing the program in a storage device of a server computer and transferring the program from the server computer to other computers via a network.
[0100] A computer that executes such a program, for example, first stores the program recorded on a portable recording medium or the program transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored in its own storage device and executes a process according to the read program. As another execution form of the program, the computer may directly read the program from the portable recording medium and execute a process according to the program, or may execute a process according to the received program each time a program is transferred from the server computer to the computer. The server computer may not transfer the program to the computer, but may execute the above-mentioned process by a so-called ASP (Application Service Provider) type service that realizes the processing function only by issuing an execution instruction and obtaining the result. Note that the program in this embodiment includes information used for processing by an electronic computer that is equivalent to a program (data that is not a direct instruction to the computer but has a nature that specifies the processing of the computer, etc.).
[0101] In addition, in this embodiment, a hardware entity is configured by executing a specific program on a computer, but at least a part of the processing contents may be realized by hardware.
[0102] The foregoing description of the embodiments of the invention has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form disclosed. Modifications and variations are possible in light of the above teachings. The embodiments have been chosen and depicted to provide a best illustration of the principles of the invention and to enable those skilled in the art to utilize the invention in various embodiments and with various modifications as may be suitable for the practical use contemplated. All such modifications and variations are within the scope of the invention as defined by the appended claims interpreted in accordance with the breadth to which they are fairly, legally, and equitably entitled.
Claims
1. A neural network learning device that learns a neural network including an encoder that converts an input vector into a latent variable vector having latent variables as elements and a decoder that converts the latent variable vector into an output vector so that the input vector and the output vector become substantially identical, comprising: a learning unit that performs learning by repeating a parameter update process that updates parameters included in the neural network; The encoder comprises: When each piece of input information included in a predetermined input information group is positive information, negative information, or no information exists, Each input information, a positive information bit that is 1 when the input information corresponds to positive information and is 0 when there is no information or when the input information corresponds to negative information; a negative information bit that is 1 when the input information corresponds to negative information and is 0 when there is no information or when the input information corresponds to positive information; The input vector is represented as The encoder is constructed from a plurality of layers, The layer, which receives the input vector, obtains a plurality of output values from the input vector; each of the output values is obtained by adding together a value obtained by applying a weight parameter to each of the positive information bits included in the input vector and a value obtained by applying a weight parameter to each of the negative information bits included in the input vector, The parameter update process includes: When the input information corresponds to positive information, the smaller the probability that the input information obtained by the decoder corresponds to positive information, the larger the value of the loss function including the sum of all input information of the input information group for learning the loss, which is approximately 0 when the input information does not exist, is reduced. Neural network learning device.
2. A neural network training method in which a neural network training device trains a neural network including an encoder that converts an input vector into a latent variable vector having latent variables as elements and a decoder that converts the latent variable vector into an output vector so that the input vector and the output vector become substantially identical, comprising: The neural network learning device includes a learning step of performing learning by repeating a parameter update process for updating parameters included in the neural network, The encoder comprises: When each piece of input information included in a predetermined input information group is positive information, negative information, or no information exists, Each input information, a positive information bit that is 1 when the input information corresponds to positive information and is 0 when there is no information or when the input information corresponds to negative information; a negative information bit that is 1 when the input information corresponds to negative information and is 0 when there is no information or when the input information corresponds to positive information; The input vector is represented as The encoder is constructed from a plurality of layers, The layer, which receives the input vector, obtains a plurality of output values from the input vector; each of the output values is obtained by adding together a value obtained by applying a weight parameter to each of the positive information bits included in the input vector and a value obtained by applying a weight parameter to each of the negative information bits included in the input vector, The parameter update process includes: When the input information corresponds to positive information, the smaller the probability that the input information obtained by the decoder corresponds to positive information, the larger the value of the loss function including the sum of all input information of the input information group for learning the loss, which is approximately 0 when the input information does not exist, is reduced. Neural network training methods.
3. A program for causing a computer to function as the neural network learning device according to claim 1.
Citation Information
Patent Citations
Method and a device for modeling student
CN108446768A
Mechanical equipment deep learning state recognition and diagnosis method
CN111562109A
Expert automatic matching system in education platform
KR1020200089914A
System and method for machine learning architecture with differential privacy
US20210133590A1