Model training device and model training program
Patent Information
- Application Number
- PCT/JP2025/012428
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025012428_01102026_PF_FP_ABST
Abstract
Description
Model learning device and model learning program
[0001] The embodiments relate to a model learning device and a model learning program.
[0002] Shuffled regression (SR) is a problem of training a regression model from data where the correspondence between input features and target output is unknown (shuffled data). Data collected independently by different organizations or devices, or data collected without linking to personal attributes (such as annual income or test scores) to protect privacy, are often represented as shuffled data. Shuffled data is a different data format from data where the correspondence between input and output is known (hereinafter referred to as paired data). Therefore, it is impossible to use standard regression methods (which use paired data) to train a regression model that represents the relationship between input and output.
[0003] Existing research on shuffle regression has focused on approaches using linear models (see, for example, Non-Patent Document 1). This approach has a strong limitation: the relationship between input and output is linear. Therefore, recent research has proposed methods using deep learning-based neural network models to remove this limitation and enable the capture of more complex relationships (see, for example, Non-Patent Document 2).
[0004] Abubakar Abid and James Zou. A stochastic expectation-maximization approach to shuffled linear regression. In Annual Allerton Conference on Communication, Control, and Computing, pp. 470-477, 2018. Masahiro Kohjima. Shuffled deep regression. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, pp. 13238-13245, 2024. Eliane Maria Loiola, Nair Maria Maia DeAbreu, Paulo Oswaldo Boaventura-Netto, Peter Hahn, and Tania Querido. A survey for the quadratic assignment problem. European journal of operational research, Vol. 176, No.2, pp.657-690, 2007. Michalis Titsias. Variational learning of inducing variables in sparse gaussian processes. In Artificial intelligence and statistics, pp. 567-574. PMLR, 2009. Christopher KI Williams and Carl Edward Rasmussen. Gaussian processes for machine learning, Vol.2. MIT press Cambridge, MA, 2006.
[0005] Generally, training neural networks requires a large amount of data and computational resources. Therefore, if a large amount of data is not available, this method may not perform adequately.
[0006] The problem addressed by this invention is to provide a technique that enables the modeling of complex nonlinear relationships even when only relatively small amounts of data are available, which would not allow the neural network model to perform to its full potential.
[0007] The model learning device of the embodiment includes: an acquisition unit that acquires input data and test points, which include: shuffle data, which is a set of multiple input features whose set of function outputs that define the input-output relationship follows a multivariate normal distribution that is a Gaussian process, and output observed values generated by randomly swapping the elements of the function outputs using a permutation matrix which is a double probability matrix; a kernel function which is a function of the multivariate normal distribution and is expressed using hyperparameters; and setting parameters; an estimation unit that estimates the hyperparameters and the permutation matrix based on the input data using an optimization method to maximize the log marginal likelihood, which is the logarithm of the marginal distribution of the output observed values given by a normal distribution; a prediction distribution processing unit that takes the shuffle data, the kernel function, the estimated hyperparameters and permutation matrix, and the test points as input and calculates the predicted distribution derived when the shuffle data is given, thereby calculating the predicted distribution of the output observed values at the test points; and an output control unit that outputs the predicted distribution of the output observed values at the test points.
[0008] According to one aspect of this invention, it becomes possible to model complex nonlinear relationships even when only relatively small amounts of data are available that do not allow the neural network model to perform to its full potential.
[0009] Figure 1 is a diagram showing an example of paired data according to one embodiment. Figure 2 is a diagram showing an example of shuffled data according to one embodiment. Figure 3 is a diagram showing an example of an alternating optimization algorithm according to one embodiment. Figure 4 is a block diagram showing an example of the hardware configuration of a model learning device according to one embodiment. Figure 5 is a block diagram showing the software configuration of a model learning device according to one embodiment in relation to the hardware configuration shown in Figure 4. Figure 6 is a flowchart showing an example of the operation performed by the model learning device 1 according to one embodiment.
[0010] Each embodiment is described below with reference to the drawings. Each embodiment illustrates an apparatus or method for realizing the technical idea of the invention. The drawings are schematic or conceptual. Hereafter, components having substantially the same function and configuration are denoted by the same reference numerals. The numbers following the letters that constitute the reference numerals are used to distinguish elements that are referenced by reference numerals containing the same letters and that have similar configurations.
[0011] [Embodiments] (1) Problems to be solved by the embodiment In one embodiment, we focus on a probabilistic model called a Gaussian process (GP) (see, for example, Non-Patent Document 5), which is widely used for regression and classification problems using paired data. By employing a non-parametric approach, the Gaussian process can model complex nonlinear relationships even when the available data is relatively small. It can also provide a measure of uncertainty in predictions, which can help with predictions in risk-sensitive settings and decision-making in uncertain situations. In one embodiment, we construct a method that inherits these properties by using a Gaussian process in shuffle regression.
[0012] Furthermore, in one embodiment, we propose a new method for shuffled regression based on a Gaussian process, called Shuffled Gaussian Process Regression (SGPR). We describe an algorithm that derives the predictive distribution (conditional probability) of the target output value for an unknown test input given shuffled data based on a Gaussian process, and estimates the latent variable, the permutation matrix, and hyperparameters by maximizing the log marginal likelihood.
[0013] Here, the estimation algorithm is one that alternately estimates the permutation matrix and hyperparameters. For hyperparameter estimation, we demonstrate that by using an objective function defined by pseudo-paired data created from shuffled data, any optimization method such as gradient descent, SCG (Scaled Conjugate Gradient), or L-BFGS can be used, similar to standard Gaussian process regression. Furthermore, for permutation matrix estimation, we demonstrate that by reducing the problem to an optimization problem called the Quadratic Assignment Problem (QAP), any optimization method for QAP, such as simulated annealing, can be used.
[0014] By using SGPR, it becomes possible to estimate regression models with greater accuracy than existing shuffled regression methods that use linear models, such as Shuffled Linear Regression (SLR) (see Non-Patent Document 1). Furthermore, even when only relatively small amounts of data are available, where methods using neural networks, such as Shuffled Deep Regression (SDR) (see, for example, Non-Patent Document 2), cannot perform adequately, it is possible to model complex nonlinear relationships. In addition, because SGPR inherits the properties of Gaussian processes and can provide a measure of uncertainty in predictions, it can be used as a shuffled regression method for predictions in risk-sensitive settings and for decision-making in uncertain situations.
[0015] (2) Preparation First, a standard Gaussian process regression using paired data will be described. First, let X⊂R d be the space of input features, and Y⊂R be the space of target output values. Data D consisting of N pairs of input values and target values PD ={(x i , y i )} N i=1 is referred to as paired data. Here, x i ∈X and y i ∈Y represent the i-th input feature and target value respectively, and N is an arbitrary positive integer.
[0016] FIG. 1 is a diagram showing an example of paired data according to an embodiment. As shown in FIG. 1, paired data D PD includes N input values x i and target values y i wherein each input value x i has a one-to-one correspondence with a target value y i .
[0017] In Gaussian process regression, it is assumed that the function f that defines the input-output relationship follows a Gaussian process. This means that for any set of N input values X={x i} N i=1 , the set of function outputs f={f(x i )} N i=1 follows the N-dimensional multivariate normal distribution represented by the following equation (1).
[0018]
[0019] where ~ K is a variance-covariance matrix whose (i,j)-th component is given by k(x i , x j ) using a kernel function k. Here, " ~The symbol " represents a tilde placed over a character. Note that equation (1) should strictly be written as a conditional probability p(f|X), but following convention, the description of the condition regarding the input value is omitted. Any kernel can be used for the kernel function k, including polynomial kernels, Gaussian kernels (RBF kernels), exponential kernels, Matter kernels, string kernels, Fisher kernels, marginalization kernels, etc. Also, the target value y = {y i} N i=1 This can be modeled as a set of function outputs f to which Gaussian noise has been added, as shown in equation (2) below.
[0020]
[0021] However, β is the precision parameter, I N This represents the identity matrix of size N×N. Therefore, by marginalizing and eliminating the set of function outputs f, the marginal distribution of the target value y is given by the following equation (3).
[0022]
[0023] however, ~ C has an (i, j) component that is k(x i , x j ) + β -1 δ ij This is the variance-covariance matrix given by δ. ij This represents the Kronecker delta. Given a target value y, the predictive distribution of the target value y* at a test point x* is expressed as the conditional probability in equation (3) (when considering the target value as an N+1-dimensional vector), and is therefore given as a normal distribution represented by the following equation (4).
[0024]
[0025] however, ~ k is the i-th element of k(x i It is an N-dimensional vector given by x*. The above prediction distribution calculation uses a vector of size N × N. ~ We need to obtain the inverse matrix of C. This can be done in O(N) time using the usual method. 3The computational complexity of this process can be enormous, which can be problematic when the data size N is large. In such cases, approaches that utilize auxiliary variables (see, for example, Non-Patent Document 4) can be applied.
[0026] In a Gaussian process, the kernel function k (or covariance matrix) ~ K ~ In some cases, more flexible modeling can be achieved by expressing C) using the hyperparameter θ. In such cases, the hyperparameter θ can be estimated by maximizing the log marginal likelihood, which is expressed by equation (5) below, obtained by taking the logarithm of equation (3).
[0027]
[0028] However, to explicitly show that it depends on the hyperparameter θ, we use the variance-covariance matrix. ~ C ~ C θ This is expressed as follows. The partial derivative of the above equation with respect to the hyperparameter θ is given by equation (6) below.
[0029]
[0030] This allows the hyperparameter θ to be optimized using optimization methods such as gradient descent, SCG (Scaled Conjugate Gradient) method, and L-BFGS method. (3) Shuffled Data Next, shuffled data used in one embodiment will be described. Shuffled data is data in which there is no correspondence between input features and output observations. Shuffled Data D SD D is a set of N input features and a set of output observations. SD = {(X i , y i )} N i=1 Defined as: X i and y i Each of these is a set of L input features X. i = {x i1 , x i2 , ..., x iL} and the set of L output observations y i = {y i1 , y i2, ..., y iL}. Each element of the set is the input feature x il ∈X and output observed value y il ∈ Y, where L is any positive integer. This definition can also be extended to assume that the size L of the set is different for each sample i.
[0031] Figure 2 shows an example of shuffled data according to one embodiment. In the example in Figure 2, L = 2 is shown. As shown in Figure 2, unlike the paired data shown in Figure 1, there is no one-to-one correspondence between input and output in the shuffled data. Therefore, the input feature x il The corresponding output observation value is {y i1 , ..., y iL It is one of the above, but it is unknown which one. Since shuffled data is equivalent to paired data when L=1, shuffled data can also be considered a more general representation of paired data.
[0032] In early studies on shuffle regression (see, for example, Non-Patent Document 1), the data size was assumed to be N=1. However, in realistic data collection scenarios, such as collecting survey data, it is common to collect data multiple times from multiple groups of respondents at different times, methods, and locations. In such cases, even if the data is processed in a way that makes it impossible to know which respondent each response came from, as long as it is possible to determine the time, method, and location from which the response and the respondent were obtained, this can be represented as shuffled data with a data size of N>1.
[0033] (3.1) Generation Process of Shuffled Data The shuffled data is modeled as being obtained according to the following generation process. First, similar to Gaussian process regression of normal paired data, all input values {X} of any N shuffled data are obtained. i} N i=1 = {x il} N,L i,l=1 The set of function outputs f = {f i} Ni=1 = {f(x il )} N,L i,l=1 However, assume that it follows an NL-dimensional multivariate normal distribution represented by the following equation (7).
[0034]
[0035] However, the element at the (i-1)L+lth row and (i'-1)L+l'th column of matrix K is k(x) using the kernel function k. il , x i´l´ The given value is an NL×NL variance-covariance matrix. For the kernel function k, any kernel used in normal Gaussian process regression may be used, such as a polynomial kernel, Gaussian kernel (RBF kernel), exponential kernel, Matter kernel, string kernel, Fisher kernel, marginalization kernel, etc.
[0036] Next, independently for each sample i, a permutation matrix of size L×L (a double probability matrix Π whose elements are 0 or 1 and whose row and column sums are all 1) is generated as shown in equation (8) below. i Using f i After randomly rearranging the elements, noise is added, resulting in the output observed value y i = {y il} L l=1 This is generated.
[0037]
[0038] Equation (8) above gives the observed values of all samples y = {y i} N i=1 It can be expressed by the following equation (9), which summarizes the process.
[0039]
[0040] Therefore, from equations (7) and (9) and the properties of the marginal distribution of the Gaussian distribution, the output observed value y i The marginal distribution of is given by a normal distribution represented by the following equation 10.
[0041]
[0042] Provided that matrix C is an NL×NL variance-covariance matrix whose element at the ((i−1)L+l)-th row and the ((i′−1)L+l′)-th column is k(x il , x i´l´ ) + β -1 δ ii´ δ ll´ given by the above formula.
[0043] (3.2) Construction of Predictive Distribution The predictive distribution at any arbitrary test point x* can be derived as follows. y′ = (y T , y*) T represented by an (NL+1)-dimensional vector is generated following an (NL+1)-dimensional multivariate normal distribution obtained by increasing the dimension of formula (7) by one, and for each element of y, formula (8) is used, and the newly added element y* is generated according to the following formula (11) that does not use the permutation matrix Π. T , f*) T
[0044]
[0045] At this time, the marginal distribution of y′ is given by the following formula (12).
[0046]
[0047] Provided that matrix S′ is expressed in a block diagram matrix form represented by the following formula.
[0048]
[0049] Provided that k is an NL-dimensional vector whose ((i−1)L+l)-th element is given by k(x il , x*) using the kernel function k. The predictive distribution of the output observation y* at a certain test point x* under the condition that the observation set y is given is a conditional probability in formula (12), so it is expressed as a normal distribution represented by the following formula (13).
[0050]
[0051] Provided that in the transformation of the above formula (13), S -1 = ΠC -1 Π T and Π T = Π -1 This was used. Furthermore, the output observations rearranged using the permutation matrix Π are expressed as shown in equation (14) below.
[0052]
[0053] moreover - y i The l-th element - y il When written as above, the predictive distribution represented by equation (13) is the input-output pair (with a one-to-one correspondence) {(x il , - y il )} N,L i,l=1 It can be seen that this is an equivalent representation of the predictive distribution expressed by equation (4) in a normal Gaussian process regression where given the values. From this, we can derive the predictive distribution expressed by equation (13) when shuffled data is given based on a Gaussian process. Here, - " refers to a bar placed over a letter.
[0054] (3.3) Optimization of the Permutation Matrix Π and Hyperparameters θ In order to compute the predictive distribution represented by equation (13) derived in the previous section, it is necessary to estimate the permutation matrix Π and the hyperparameters θ (of which the kernel function k etc. are located). For the optimization of these, we can consider maximizing the marginal likelihood represented by equation (10), similar to ordinary Gaussian process regression. Taking the logarithm of equation (10), we obtain the log-marginal likelihood represented by the following equation (15).
[0055]
[0056] In the above transformation, |Π||Π T |=1 and S -1 = ΠC -1 Π T We use the following. Also, to explicitly show that it depends on the hyperparameter θ, the variance-covariance matrix C is used in the last row. θThis is what I wrote. This log marginal likelihood can be efficiently optimized using an alternating optimization algorithm that repeats the steps of fixing the permutation matrix Π and optimizing the hyperparameter θ (Step 1), and fixing the hyperparameter θ and optimizing the permutation matrix Π (Step 2). The details of each step are explained below.
[0057] Step 1: Optimization of hyperparameter θ Focusing on the representation of the last row of the optimized log marginal likelihood expressed by equation (15) for the hyperparameter θ, similar to the case of the predictive distribution, the pairs of output and input are rearranged by the permutation matrix Π defined by equation (14) {(x il , - y il )} N,L i,l=1 It can be seen that this is an equivalent expression to the log marginal likelihood expressed by equation (5) in a normal Gaussian process regression given the given values. Therefore, the same methods as in Gaussian process regression can be used to optimize hyperparameters θ other than the kernel function k and the permutation matrix Π such as the variance. In other words, by calculating the partial derivative of the above log marginal likelihood with respect to the hyperparameter θ expressed by the following equation (16), the hyperparameter θ can be optimized using (continuous) optimization methods such as the gradient method, SCG method, and L-BFGS method.
[0058]
[0059] Step 2: Optimization of the Permutation Matrix Π Since the permutation matrix Π is made up of discrete variables, continuous optimization methods such as gradient descent cannot be directly applied. However, as shown below, optimization can be performed by reducing it to a quadratic assignment problem (QAP). Focusing on the expression of the last row of the log marginal likelihood shown in equation (15), and extracting the term relating to the permutation matrix Π, the equation can be transformed as follows.
[0060]
[0061] However, the expression x using tracing in the algebraic manipulation T Ax = tr(Axx) TThis was used. Therefore, in order to maximize the log marginal likelihood, we need to find the permutation matrix Π that minimizes the objective function L, which is expressed by the following equation (17).
[0062]
[0063] Since this optimization problem is a quadratic assignment problem (QAP), any method for solving QAPs described in Non-Patent Document 3, etc., such as simulated annealing, can be applied. The definition of QAP will be described later.
[0064] Simulated annealing is easy to implement and closely related to the Markov Chain Monte Carlo (MCMC) method used for permutation matrix estimation in existing (linear model) shuffle regression methods (see, for example, Non-Patent Document 1). Therefore, the following describes how to update the permutation matrix Π when using simulated annealing.
[0065] The permutation matrix Π at a given point in time (old) Given a matrix i ∈ {1, ..., N} and randomly selected l, l' ∈ {1, ..., L} (l ≠ l'), this permutation matrix Π (old) The result obtained by swapping the (i-1)L+lth row and (i-1)L+l' is Π (new) This transformation corresponds to an operation that partially swaps the input-output correspondences defined by the permutation matrix Π in the i-th sample. Simulated annealing accepts this transformation (called accepting it) with a probability q expressed by the following equation (19), and Π (old) to Π (new) This is an optimization method that repeatedly updates the data.
[0066]
[0067] However, T is a temperature parameter, and it is set to decrease from T to 0 during the optimization process. The sequence of T is {T s} (where s represents the number of optimization steps by simulated annealing) is, for example, T s = T 0 (T 1 / T0 ) S One possible way to set it is (for example, normally T 1 / T 0 (Often set so that the value is in the range of 0.9 to 0.95). The probability q is the case when the solution is improved (L(Π) (new) ) < L (Π (old) In this case, q=1, so the transformation is always accepted, and if not, the transformation is accepted with probability exp(-Δ / T). Therefore, in the initial stages of optimization, even if the solution deteriorates, the transformation is accepted with a certain probability, resulting in exploratory behavior that escapes the local optimum. As optimization progresses and T approaches 0, the probability of accepting a transformation when the solution deteriorates decreases, leading to behavior that converges to the optimal solution, making it possible to reach the global optimum.
[0068] Figure 3 shows an example of an alternating optimization algorithm according to one embodiment. Note that in the second row of Figure 3, the initial value of the permutation matrix Π may be set randomly. Alternatively, existing shuffle regression methods using simple models such as linear models, like SLR (see Non-Patent Literature 1), can be applied, and the estimated value of the permutation matrix Π that can be calculated from the learning results of the estimated linear model can be used as the initial value. Furthermore, it goes without saying that various extensions can be implemented based on this algorithm (similar to the case of Gaussian processes using ordinary paired data), such as combining it with approaches that utilize auxiliary variables, such as those described in Non-Patent Literature 4.
[0069] Supplement 1: Definition of the Quadratic Assignment Problem (QAP) Next, we will supplement the definition of the quadratic assignment problem. The quadratic assignment problem is defined as a matrix A = {a} of size n × n. kl} n,n k,l=1 , B = {b ij} n,n i,j=1 This can be formalized as shown in equation (19) below.
[0070]
[0071] However, P represents the set of all n×n permutation matrices Π. Element a of matrix A kl The distance between points k and l, and element b of matrix B. ijIf we consider that represents the amount of movement of objects between matrix facilities i and j, then equation (19) can be interpreted as a problem of optimizing the arrangement of facilities so that the total distance of movement of objects is minimized. In QAP, matrices A and B are A←C -1 θ , B←(yy T ) T If this setting is used, the objective function expressed in equation (19) will be the same as the objective function expressed in equation (17) in Step 2. It should be noted that the formulation of QAP is not limited to the above, and various equivalent formulations exist for integer programming problems, etc., and these may, of course, be used (see, for example, Non-Patent Document 3).
[0072] Supplement 2: Regarding the fact that the permutation matrix Π is a block diagonal matrix, as mentioned in Supplement 1 above, in a normal quadratic assignment problem, the entire permutation matrix Π is considered as a candidate solution. However, in the optimization of the objective function expressed by equation (17) in the alternating optimization algorithm, the candidate solution only needs to be the permutation matrix Π which is a block diagonal matrix, as shown in equation (9). This constraint of restricting the matrix to a block diagonal matrix can be addressed by changing the constraint that the row sum and column sum of the permutation matrix Π must be 1 to a constraint that the row sum and column sum of the block diagonal components must be 1. Alternatively, as in the simulated annealing example mentioned above, it is also possible to address this constraint by designing an optimization method that takes it into account, such as designing an operation to (randomly) swap rows so that the permutation matrix Π remains a block diagonal matrix.
[0073] (Configuration) Next, the hardware and software configurations of the model learning device for carrying out the method described above will be explained. Figure 4 is a block diagram showing an example of the hardware configuration of the model learning device 1 according to one embodiment. The model learning device 1 is a computer that analyzes input data and generates and outputs output data. For example, the model learning device 1 is installed in any location set by the user who manages the model learning device 1.
[0074] As shown in Figure 4, the model learning device 1 comprises a control unit 10, a program storage unit 20, a data storage unit 30, a communication interface 40, and an input / output interface 50. The control unit 10, program storage unit 20, data storage unit 30, communication interface 40, and input / output interface 50 are connected to each other via a bus so as to be able to communicate with each other. Furthermore, the communication interface 40 may be connected to an external device so as to be able to communicate with it via a network. In addition, the input / output interface 50 is connected to an input device 2 and an output device 3 so as to be able to communicate with them.
[0075] The control unit 10 controls the model learning device 1. The control unit 10 includes a hardware processor such as a central processing unit (CPU). For example, the control unit 10 may be an integrated circuit capable of executing various programs.
[0076] The program storage unit 20 can use a combination of non-volatile memory that can be written to and read at any time, such as EPROM (Erasable Programmable Read Only Memory), HDD (Hard Disk Drive), and SSD (Solid State Drive), as a storage medium, and non-volatile memory such as ROM (Read Only Memory). The program storage unit 20 stores programs necessary to execute various processes. In other words, the control unit 10 can realize various controls and operations by reading and executing programs stored in the program storage unit 20.
[0077] The data storage unit 30 is a storage device that uses a combination of non-volatile memory, such as an HDD or memory card, which allows for writing and reading at any time, and volatile memory, such as RAM (Random Access Memory), as storage media. The data storage unit 30 is used to store data acquired and generated during the process in which the control unit 10 executes a program and performs various processing.
[0078] The communication interface 40 includes one or more wired or wireless communication modules. For example, the communication interface 40 includes a communication module that connects to an external device via a network, either wired or wirelessly. The communication interface 40 may also include a wireless communication module that connects to an external device wirelessly, such as a Wi-Fi access point and a base station. Furthermore, the communication interface 40 may include a wireless communication module that connects to an external device wirelessly using short-range wireless technology. In other words, the communication interface 40 can be any general communication interface that can communicate with an external device and send and receive various types of information under the control of the control unit 10.
[0079] The input / output interface 50 is connected to the input device 2, output device 3, etc. The input / output interface 50 is an interface that enables the transmission and reception of information between the input device 2, output device 3, etc. The input / output interface 50 may be integrated with the communication interface 40. For example, the model learning device 1 and at least one of the input device 2 or output device 3 may be wirelessly connected using short-range wireless technology, and information may be transmitted and received using said short-range wireless technology.
[0080] The input device 2 may include, for example, a keyboard or pointing device for the user to input various information to the model learning device 1. The input device 2 may also include a reader for reading data to be stored in the program storage unit 20 or data storage unit 30 from a memory medium such as a USB memory, or a disk device for reading such data from a disk medium.
[0081] The output device 3 includes a display, etc., that shows the results calculated by the control unit 10, such as a virtual space. The output device 3 also includes a printer, etc., that prints the information displayed on the display.
[0082] Figure 5 is a block diagram showing the software configuration of a model learning device 1 according to one embodiment, in relation to the hardware configuration shown in Figure 4. The control unit 10 includes an input data acquisition unit 101, a parameter estimation unit 102, a prediction distribution processing unit 103, and an output control unit 104.
[0083] The input data acquisition unit 101 is an acquisition unit that acquires input data and test points. For example, the input data is shuffled data D. SD The kernel function k has hyperparameters θ, and the configuration parameters include the hyperparameters θ, the optimization method used to estimate the permutation matrix Π and the settings of the optimization method such as the learning rate and temperature parameters, fixed parameters of the kernel function k, the initial value of the permutation matrix Π, the maximum number of iterations, etc. The input data acquisition unit 101 acquires input data and test points when the user inputs input data and test points into the input device 2, etc.
[0084] The parameter estimation unit 102 is an estimation unit that estimates the hyperparameter θ and the permutation matrix Π. For example, the parameter estimation unit 102 estimates the input data (shuffled data D) SD The alternating optimization algorithm described above takes the kernel function k and setting parameters as input, fixes the permutation matrix Π, and repeatedly optimizes the hyperparameter θ, then fixes the hyperparameter θ and optimizes the permutation matrix Π, thereby estimating the hyperparameter θ and permutation matrix Π that maximize the log marginal likelihood of equation (15).
[0085] The predictive distribution processing unit 103 is a processing unit that calculates the predictive distribution of the input test points x*. For example, the predictive distribution processing unit 103 calculates the acquired shuffled data D SD The kernel function k, hyperparameter θ, permutation matrix Π, and test points x* are used as inputs to calculate the predicted distribution of the output observed values of the test points x* in equation (13) described above.
[0086] The output control unit 104 is an output unit that outputs the predicted distribution of the output observed values of the test point x* stored in the predicted distribution storage unit 303 (described later) to the output device 3 or an external device.
[0087] The data storage unit 30 comprises an input data storage unit 301, a parameter storage unit 302, and a predictive distribution storage unit 303. The input data storage unit 301 is a storage unit used to store input data and test points, etc. The parameter storage unit 302 is a storage unit used to store estimated hyperparameters θ and permutation matrices Π. The predictive distribution storage unit 303 is a storage unit used to store the predictive distribution of output observed values for test points x*.
[0088] [Operation] Figure 6 is a flowchart showing an example of the operation performed by the model learning device 1 according to one embodiment. The operation shown in this flowchart is realized when the control unit 10 of the model learning device 1 reads and executes the program stored in the program storage unit 20.
[0089] For example, this flowchart starts operating when instructed by a user managing the model learning device 1.
[0090] In step ST101, the input data acquisition unit 101 acquires the input data and test points. For example, when a user inputs input data into the input device 2, the input data acquisition unit 101 acquires the input data. The input data is shuffled data D. SD This includes a kernel function k with hyperparameters θ, and configuration parameters. Shuffled data D SD As described above, the kernel function is a set of multiple input features whose set of function outputs that define the input-output relationship follows a multivariate normal distribution that is a Gaussian process, and an output observation value that is generated by randomly swapping the elements of the function output using a permutation matrix, which is a double probability matrix. Furthermore, the kernel function is a function of the multivariate normal distribution and is a function expressed using hyperparameters. The setting parameters include the hyperparameter θ, the optimization method used to estimate the permutation matrix Π and the setting values of the optimization method such as the learning rate and temperature parameter, fixed parameters of the kernel function k, the initial value of the permutation matrix Π, the maximum number of iterations, etc. The input data acquisition unit 101 stores the input data in the input data storage unit 301.
[0091] Furthermore, the input data acquisition unit 101 acquires test points x* input by the user or an external device and stores these test points x* in the input data storage unit 301.
[0092] In step ST102, the parameter estimation unit 102 estimates the hyperparameter θ and the permutation matrix Π. For example, the parameter estimation unit 102 acquires the input data stored in the input data storage unit 301. Based on the input data, the parameter estimation unit 102 uses an optimization method to estimate the hyperparameter and the permutation matrix so as to maximize the log marginal likelihood, which is the logarithm of the marginal distribution of the output observation value given by a normal distribution. For example, the parameter estimation unit 102 uses the input data (shuffled data D) SD The alternating optimization algorithm described above takes the kernel function k and setting parameters as input, fixes the permutation matrix Π, and repeatedly optimizes the hyperparameter θ, then fixes the hyperparameter θ and optimizes the permutation matrix Π, thereby estimating the hyperparameter θ and permutation matrix Π that maximize the log marginal likelihood of equation (15). The parameter estimation unit 102 then stores the estimated hyperparameter θ and permutation matrix Π in the parameter storage unit 302.
[0093] Here, in the step of optimizing the hyperparameter θ, the parameter estimation unit 102 may use any optimization method such as the (stochastic) gradient method, SCG method, or L-BFGS method. Also, in the step of optimizing the permutation matrix Π, the parameter estimation unit 102 may use any method for solving QAP, such as simulated annealing or the method described in Non-Patent Literature 3, as the optimization method. Furthermore, any kernel function k used in normal Gaussian process regression may be used, such as a polynomial kernel, Gaussian kernel (RBF) kernel, exponential kernel, Matter kernel, string kernel, Fisher kernel, marginalization kernel, etc. Furthermore, it may be extended by combining it with approaches that utilize auxiliary variables as described in Non-Patent Literature 4, etc.
[0094] In step ST103, the predictive distribution processing unit 103 calculates the predicted distribution of test points. The predictive distribution processing unit 103 uses the shuffled data D contained in the input data stored in the input data storage unit 301. SD The kernel function k, the hyperparameters θ and permutation matrix Π stored in the parameter storage unit 302, and the test points x* stored in the input data storage unit 301 are obtained. The prediction distribution processing unit 103 then obtains the shuffled data D SD The predictive distribution processing unit 103 calculates the predictive distribution derived when shuffled data is given, i.e., the predictive distribution of the output observed values of test points x* in equation (13) described above, using the kernel function k, hyperparameter θ, permutation matrix Π, and test points x* as input. The predictive distribution processing unit 103 stores the calculated predictive distribution of the output observed values of test points x* in the predictive distribution storage unit 303.
[0095] In step ST104, the output control unit 104 outputs the predicted distribution of the output observed values of the test point x* stored in the predicted distribution storage unit 303 to the output device 3 or an external device.
[0096] (Effects of the Embodiment) As described above, in one embodiment, the model learning device 1 can model complex nonlinear relationships and learn regression functions with high accuracy even when only small-scale shuffled data is available compared to the data required for training a neural network.
[0097] Furthermore, because it can provide a measure of uncertainty in predictions, the provided data can be used as a shuffle regression method for predictions in risk-sensitive settings or for decision-making in uncertain situations.
[0098] [Other Embodiments] In the above embodiment, the model learning device 1 can also be configured by constructing the operation of each component as a program, and installing and executing it on a computer used as a user device or on a computer used as an external device.
[0099] Furthermore, the method described in the above embodiment can be distributed by storing the program (software means) that can be executed by a computer in a storage medium such as a magnetic disk (floppy disk, hard disk, etc.), an optical disk (CD-ROM, DVD, MO, etc.), or a semiconductor memory (ROM, RAM, flash memory, etc.), and by transmitting it via a communication medium. The program stored on the medium also includes a configuration program that configures the software means (including not only the execution program but also tables and data structures) to be executed by the computer. The computer realizing this device reads the program stored in the storage medium and, if necessary, constructs the software means using the configuration program, and executes the above-described process by controlling its operation with this software means. The storage medium referred to in this specification is not limited to distribution mediums, but also includes storage mediums such as magnetic disks and semiconductor memories provided inside the computer or in devices connected via a network.
[0100] In short, this invention is not limited to the embodiments described above, and can be modified in various ways during implementation without departing from its essence. Furthermore, each embodiment may be combined as appropriately as possible, in which case the combined effects can be obtained. Moreover, the embodiments described above include inventions at various stages, and various inventions can be extracted by appropriate combinations of the multiple constituent elements disclosed.
[0101] 1...Model learning device 10...Control unit 101...Input data acquisition unit 102...Parameter estimation unit 103...Prediction distribution processing unit 104...Output control unit 20...Program storage unit 30...Data storage unit 301...Input data storage unit 302...Parameter storage unit 303...Prediction distribution storage unit 40...Communication interface 50...Input / output interface 2...Input device 3...Output device
Claims
1. A model learning device comprising: an acquisition unit that acquires input data and test points, including: shuffle data which is a set of multiple input features whose set of function outputs that define the input-output relationship follows a multivariate normal distribution which is a Gaussian process, and output observation values which are generated by randomly swapping the elements of the function outputs using a permutation matrix which is a double probability matrix; a kernel function which is a function of the multivariate normal distribution and is expressed using hyperparameters; and setting parameters; an estimation unit which estimates the hyperparameters and the permutation matrix based on the input data using an optimization method to maximize the log marginal likelihood obtained by taking the logarithm of the marginal distribution of the output observation values which are given by a normal distribution; a prediction distribution processing unit which calculates the predicted distribution of the output observation values at the test points by taking the shuffle data, the kernel function, the estimated hyperparameters and permutation matrix and the test points as input and calculating the predicted distribution which is derived when the shuffle data is given; and an output control unit which outputs the predicted distribution of the output observation values at the test points.
2. The model learning device according to claim 1, wherein the estimation unit estimates the hyperparameters and the permutation matrix that maximize the log marginal likelihood using an alternating optimization algorithm that repeats the steps of fixing the permutation matrix and optimizing the hyperparameters, and fixing the hyperparameters and optimizing the permutation matrix.
3. The model learning device according to claim 2, wherein the estimation unit optimizes the hyperparameters by calculating the partial derivative of the log marginal likelihood with respect to the hyperparameters using the optimization method, and estimates the permutation matrix using an optimization method for solving a quadratic assignment problem to find the permutation matrix that minimizes the objective function, which is an expression obtained by transforming the term of the log marginal likelihood with respect to the permutation matrix.
4. A model learning program comprising instructions for execution by the processor of a model learning device, wherein the instructions include: acquiring input data and test points, which are a set of shuffle data consisting of a set of multiple input features whose set of function outputs defining input-output relationships follows a multivariate normal distribution where the set of function outputs is a Gaussian process, and output observation values generated by randomly swapping the elements of the function outputs using a permutation matrix which is a double probability matrix; a kernel function which is a function of the multivariate normal distribution and is expressed using hyperparameters; and setting parameters; estimating the hyperparameters and the permutation matrix based on the input data using an optimization method to maximize the log marginal likelihood obtained by taking the logarithm of the marginal distribution of the output observation values given by a normal distribution; calculating the predicted distribution of the output observation values at the test points by taking the shuffle data, the kernel function, the estimated hyperparameters and permutation matrix, and the test points as input and calculating the predicted distribution derived when the shuffle data is given; and outputting the predicted distribution of the output observation values at the test points.