Model learning device, model learning method, and model learning program

The model learning device improves zero-shot learning by integrating meta-learning techniques, enhancing adaptability to unknown tasks and improving classification performance in data-scarce scenarios.

JP7845498B2Active Publication Date: 2026-04-14NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-11-10
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Conventional zero-shot learning methods fail to adapt effectively to changing tasks when data for certain classes is unavailable, leading to performance degradation.

Method used

A model learning device that incorporates meta-learning to improve zero-shot learning performance by acquiring knowledge useful for unknown tasks through a composite model comprising an encoder, decoder, and projection model, optimizing parameters using a tailored objective function.

Benefits of technology

Enhances zero-shot learning performance by adapting to unknown tasks, enabling effective classification even when data for some classes is not available.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007845498000036
    Figure 0007845498000036
  • Figure 0007845498000037
    Figure 0007845498000037
  • Figure 0007845498000038
    Figure 0007845498000038
Patent Text Reader

Abstract

A model training device according to one embodiment, which achieves zero-shot learning about an unknown task, comprises: an acquisition unit which acquires input data including an observable task set used for meta learning, first auxiliary information that indicates a known class prototype set in the zero-shot learning, and second auxiliary information that indicates an unknown class prototype set, and model data including a synthesis model composed of a plurality of models, and which acquires learning rate parameters of the model data; a model parameter learning unit which learns the model parameters of the synthesis model by using an optimization method so as to minimize, by using the input data, the model data, and the learning rate parameters, an objective function in which a loss function used in the meta learning is evaluated as a loss function used for the zero-shot learning; and an output control unit which controls the model parameters to be displayed on an output device.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a model learning device, a model learning method, and a model learning program. [Background technology]

[0002] Models trained using conventional supervised learning methods suffer from performance degradation when data is scarce or when they need to adapt to changing tasks. This problem becomes even more pronounced, for example, when data for some classes is completely unavailable. This is because, despite performance degradation even with limited data, the model is designed to handle situations where no data is available at all. However, there are situations where a learning method is needed that can adapt to changing tasks even in environments where data for some classes is unavailable.

[0003] For example, consider a training method for a model that estimates the words a person is thinking of based on their brainwaves. Generally, there are individual differences in human biological signals, including brainwaves, so when we consider an individual as a task, the task will also be different for each person. Therefore, it is necessary to acquire data for each individual (task) we want to estimate and train a model that estimates words from brainwaves individually for each one. Furthermore, this setup presents the challenge that "it is not always possible to obtain brainwave data for all the target words." This would require, in the future, every time an individual we want to estimate appears, we would have them recall all the target words and measure the brainwaves for each recall. However, it is clear that such an operation is not practical. In such cases, a training method is needed that is "adaptable even if there are differences in data from person to person" and "can make inferences for other data by training with only a portion of the data."

[0004] One conventional technique specifically designed for situations where only data from certain classes is available is zero-shot learning (see, for example, Non-Patent Documents 1 and 2). Zero-shot learning is a method for training a model that can infer data from classes not present in the training, using labeled data and existing knowledge called auxiliary information. Auxiliary information contains information about the characteristics of each class and is often represented as a set of representative vectors that represent each class. Because this auxiliary information includes information about classes for which data is unavailable, the classification of instances of classes for which data is unavailable is indirectly inferred from the learning of classes for which data is available. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Hugo Larochelle, Dumitru Erhan, and Yoshua Bengio. Zero-data learning of new tasks. In Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 2, AAAI'08, p. 646-651. AAAI Press, 2008. [Non-Patent Document 2] Mark Palatucci, Dean Pomerleau, Geoffrey Hinton, and Tom M. Mitchell. Zero-shot learning with semantic output codes. In Proceedings of the 22nd International Conference on Neural Information Processing Systems, NIPS'09, p. 1410-1418, Red Hook, NY, USA, 2009. Curran Associates Inc. [Non-Patent Document 3] Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic metalearning for fast adaptation of deep networks. In Doina Precup and YeeWhye Teh, editors, Proceedings of the 34th International Conference on Machine 16 Learning, Vol.70 of Proceedings of Machine Learning Research, pp. 1126-1135. PMLR, 06-11 Aug 2017.

Non-Patent Document 4

Non-Patent Document 5

[0006] Conventional zero-shot learning methods are one solution for situations where no data is available at all, but they have the problem of not addressing the challenge of having to adapt to changing tasks.

[0007] This invention was made in view of the above circumstances, and its purpose is to provide a technology that improves the learning performance when adapting zero-shot learning to unknown tasks by learning how to learn from the process of zero-shot learning models for multiple tasks.

[0008] Specifically, the aim is to provide a new learning method that introduces meta-learning, a technique that can adapt to changing tasks, in addition to zero-shot learning. [Means for solving the problem]

[0009] To solve the above problems, one aspect of the present invention is a model learning device that realizes zero-shot learning for an unknown task, comprising: an acquisition unit that acquires input data including a set of observable tasks used in meta-training, first auxiliary information indicating a set of known class prototypes in the zero-shot learning, and second auxiliary information indicating a set of unknown class prototypes, and model data including a composite model consisting of multiple models, and acquires the learning rate parameters of the model data; a model parameter learning unit that learns the model parameters of the composite model using an optimization method so as to minimize an objective function in which the loss function used in the meta-training is evaluated by the loss function used in the zero-shot learning, using the input data, the model data, and the learning rate parameters; and an output control unit that controls the display of the model parameters on an output device. [Effects of the Invention]

[0010] According to one aspect of this invention, it is possible to provide a method for improving the performance of zero-shot learning for future unknown tasks by training a model that acquires knowledge useful for zero-shot learning through meta-learning. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 illustrates the problem setting of the prior art and the present invention. [Figure 2] Figure 2 shows an overview of the Meta zero-shot learning method. [Figure 3] Figure 3 is a block diagram showing an example of the hardware configuration of a model learning device according to this embodiment. [Figure 4] Figure 4 is a block diagram showing the software configuration of the model learning device in the embodiment, in relation to the hardware configuration shown in Figure 1. [Figure 5] Figure 5 is a flowchart illustrating an example of the general operation by which a model learning device obtains model parameters θ. [Figure 6] Figure 6 is a flowchart illustrating an example of the overview operation, which explains the operation of step ST103 in more detail. [Modes for carrying out the invention]

[0012] Embodiments of this invention will be described below with reference to the drawings. Hereafter, elements identical or similar to those already described will be denoted by the same or similar reference numerals, and redundant descriptions will generally be omitted. For example, when multiple identical or similar elements exist, a common reference numeral may be used to describe them without distinction, or a sub-number may be used in addition to the common reference numeral to describe them separately.

[0013] [Embodiment] (Meta-learning) First, meta - learning will be explained. Meta - learning is a method of finding meta - knowledge that improves the learning performance for tasks generated from a task distribution by expanding the learning scope not only to a specific task but also to the distribution of related tasks. Briefly described, meta - learning is a method of learning how to learn. Here, meta - learning is a method of learning knowledge that generalizes to related unknown tasks by training a model not on a single task but on a distribution of tasks similar to that task. With the introduction of this meta - learning, it becomes possible to learn the distribution of the model, and when applying zero - shot learning to future unknown tasks, the model can be learned by referring to that model distribution.

[0014] Meta - learning consists of meta - training, which is a training process for acquiring meta - knowledge, and meta - testing, which actually utilizes the acquired meta - knowledge to evaluate its performance. Meta - training is the phase of learning how to learn. As described above, since meta - learning is performed on the distribution of tasks, for the sake of explanation, a task set for meta - training is

[0015]

Number

[0016] defined as τ i src is a task sampled from a certain task distribution p(τ), and N src is the number of elements of τ src . Generally, since a task is composed of a dataset and a loss function, T i src is expressed as T i src =(D i spt , D i qry , L task ). Here, L task is the loss function, D spt is a dataset called the support set, D qryThis is a dataset called a queryset. Each dataset consists of a set of pairs of input instances x belonging to the D-dimensional real space X and labels y belonging to the class set Y.

[0017] In this case, the objective of metatraining is to obtain metaknowledge ω that is useful for training tasks sampled from the task distribution p(τ). Metaknowledge ω is located in the metaknowledge space W⊂R. K It is represented as a vector belonging to [a specific set]. In metatraining, first, support set D i spt Metaknowledge ω is estimated from the process of training an arbitrary model f(·;θ) for each task using [a specific method / tool]. Then, the obtained metaknowledge ω is used in the queryset D. i qry This is used to calculate the loss, and the loss function L meta By using the following calculations, we search for better meta-knowledge that minimizes the loss. The loss function L used here meta This is a loss function that is uniquely defined for each technology based on the target meta-knowledge.

[0018] Metatesting is the phase in which meta-knowledge gained during meta-training is evaluated. Since meta-knowledge is equivalent to learning techniques, in order to evaluate the performance of meta-knowledge, it is necessary to train the model using that meta-knowledge on a different task. Therefore, although metatesting is called "testing," it performs both "training → testing." In other words, metatesting alone performs operations equivalent to those of normal supervised learning.

[0019] For explanation purposes, the set of tasks to be meta-tested is

[0020]

number

[0021] This is defined as follows: Here, T i tgt These are tasks sampled from a task distribution p(τ), and N tgt is, τ tgtThis is the number of elements. Also, T i src Similarly, T i tgt is T i tgt =(D i tr ,D i te ,L task ) can be expressed as D tr and D te These are the training dataset and test dataset used in metatesting, respectively. In this case, in metatesting, the training dataset D i tr The model is trained using meta-knowledge as input. Then, test dataset D i te We evaluate the trained model against the given data. The symbols defined for explaining meta-learning are summarized in the table below.

[0022] [Table 1]

[0023] Next, we will explain a common approach to meta-learning using two patterns. • Pattern 1: Bi-Level Optimization: This approach involves two stages of metatraining: model training and metaknowledge training. In the first stage, D is trained with given metaknowledge ω, as shown in equation (1) below. i spt L task By minimizing the optimal parameter θ of the model f, i * Obtain (ω).

[0024]

number

[0025] In the second stage, the optimal parameter θ obtained for each task is calculated using equation (2) shown below. i * (ω) and ω are unknown data Di qry By utilizing this and minimizing its loss, we can determine what parameters θ, given ω, yield good performance. i * We search to see if (ω) can be obtained.

[0026]

number

[0027] Non-Patent Document 3 is cited as a specific prior art example of Bi-Level Optimization. For example, Non-Patent Document 3 defines meta-knowledge ω as the initial value of θ in equation (1). In other words, an substitution operation "θ←ω" is performed in equation (1) before training, and the optimal parameter θ is derived from that initial value. i * We begin the search for (ω). Then, we use the obtained optimal parameters to create a model f(·;θ). i * By using (ω)) and calculating the loss from equation (2), we can find an initial value ω that yields better performance for the entire related task. * This explores a better initial value ω when training a model for an unknown task in metatesting. * Because the search can begin from there, it becomes possible to train the model efficiently.

[0028] Pattern 2: Feed-Forward Model: The characteristic of this approach is that the model distribution is obtained by training a feed-forward model that encodes the task. In this approach, the model distribution of the task set is estimated by encoding the support set into a vector. Then, when applying it to an unknown task, the model is sampled from the model distribution by providing the decoder with a vector that encodes the support set of that task. This allows obtaining a model adapted to the task, even if it is an unknown task.

[0029] For illustrative purposes, let h be a feed-forward model that encodes the task from the dataset, and g be a decoding model that outputs the inference result from the input instance and the output of h. A composite model combining h and g is...

[0030]

number

[0031] This is defined as follows. At this point, by solving equation (3) shown below, the optimal parameter θ of f can be found. * To obtain.

[0032]

number

[0033] This is D i spt The task is embedded in a vector space by providing it to the encoder, and that vector and D i qry By inputting instance x into g, f is learned from the loss with y. This allows us to obtain a vector representation of a task simply by providing training data for that task to h when given an unknown task, and then perform task-specific classification by inputting that vector and the instance to be classified into g. In this approach, metaknowledge is not explicitly defined, and the optimal parameters θ of the obtained encoder / decoder are not determined. * This can be considered meta-knowledge.

[0034] The Feed-Forward Model is disclosed, for example, in Non-Patent Document 4. For example, the encoder supports support set D i spt This is a multilayer perceptron that takes a queryset D as input and outputs a vector r. The decoder takes the vector r as input and outputs a queryset D i qry Instance x j qryThis is a multilayer perceptron that takes as input and outputs a categorical distribution. In this case, the loss function that optimizes the synthetic model f is given by L in equation (3). meta It is defined by taking L as the negative log-likelihood. meta The composite model f can be expressed as shown in equations (4) and (5) below, respectively.

[0035]

number

[0036] Here, the label y is a one-hot encoded vector,<f(x;θ),y> θ is the dot product of f(x;θ) and y. The optimal parameter θ of the composite model f obtained in this way. * Using this, metatesting will reveal an unknown task T i tgt In contrast,

[0037]

number

[0038] The composite model f(·,D i tr ;θ * ) using D i te We evaluate the performance for [the specified condition]. However, this method assumes that the support set and query set are sampled from the same distribution and does not account for situations where classes that did not appear at all during training need to be inferred during testing, as in zero-shot learning.

[0039] Zero-shot learning Next, we will explain zero-shot learning in a general setting. Zero-shot learning is a method for training a model that can infer data for classes for which no data is available, by training only on classes for which data is available. However, inferring data for classes for which no data is available without auxiliary information is difficult. Therefore, we use "auxiliary information," which is data that shows the relationships between classes, to adapt the model to classes for which no data is available.

[0040] First, we define the class from which data can be obtained as a known class (Seen class), and the set of such classes is

[0041]

number

[0042] This is how it is expressed. Then, classes for which data is not obtained during training but which are the subject of testing are defined as Unseen Classes, and the set of these classes is

[0043]

number

[0044] This is how it is expressed. Here, N s and N u These are the number of elements in the known class set and the unknown class set, respectively. Next, the input space is X⊂R D N used for training tr Individual instances in the input space

[0045]

number

[0046] Let R be the real space. Here, the label corresponding to X is

[0047]

number

[0048] Let's assume that N tr This is the number of instances of the training dataset. The training dataset is created by combining these.

[0049]

number

[0050] This is how it is expressed.

[0051] Auxiliary information is represented as a set of representative vectors set for each class. These vectors are generally called prototypes, and each represents a characteristic of the class. This prototype is called the auxiliary information space A⊂R. M Defined as the vector above, the set of prototypes of the known class set S is

[0052]

number

[0053] The set of prototypes for the unknown class set U

[0054]

number

[0055] This is expressed as follows. In this case, if there is a model Φ(·;σ):X→A that appropriately projects an instance x in the input space into the auxiliary information space, then it can be estimated that the class with a prototype close to Φ(x) in the auxiliary space is the class corresponding to x. Therefore, zero-shot learning aims to solve equation (6) below.

[0056]

number

[0057] Here, \(g\) is a decoding model \(g:A\rightarrow S\) that estimates the class corresponding to \(x\) with \(\varPhi(x)\) and \(A\) as inputs, and often the kNN method, one-vs-rest method, etc. are used. Once such an optimal parameter \(\sigma\) s is obtained, for an instance \(x\) corresponding to an unknown class set, by calculating \(g(\varPhi(x * ;\sigma u ),A u ), we can estimate which unknown class it corresponds to. The symbols defined for the explanation of zero-shot learning can be summarized as shown in Table 2. * ),A u ) can be used to estimate which unknown class it corresponds to. The symbols defined for the explanation of zero-shot learning can be summarized as shown in Table 2.

[0058]

Table 2

[0059] Next, an example as disclosed in Non-Patent Document 5 of zero-shot learning will be described. In this example, the input instance, label, and auxiliary information are represented as matrices, and a linear model is used for zero-shot learning. The input instance is regarded as a matrix

[0060]

Number

[0061] to be

[0062]

Number

[0063] and becomes

[0064]

Number

[0065] is defined, and this is the class \(c\) i sThis is a sequence of so-called one-hot vectors, where the i-th value is 1 when it corresponds to the class of the input instance, and -1 for all other values. The prototype of the auxiliary information is represented by an M-dimensional binary vector and then converted to a matrix representation.

[0066]

number

[0067] This is the result. At this point, the parameters are estimated by solving the following optimization problem using hinge loss.

[0068]

number

[0069] Here, Φ is a transformation matrix containing the parameters to be optimized, and <(X T ΦA s ) i ,Y i )> is the matrix product X T ΦA s This is the inner product of the i-th matrix of matrix X and the i-th matrix of matrix Y. T The operation of multiplying by the transformation matrix Φ corresponds to Φ(·,σ) in equation (6), and matrix A s The operation of multiplying is g(·,A s ) corresponds to. During inference, N u A prototype of an unknown class.

[0070]

number

[0071] When there is an unknown instance x, u against

[0072]

number

[0073] Solving this will give the inference result c j u This is what is obtained. In such general zero-shot learning, each task is learned independently, and the extraction of common meta-knowledge from related tasks is not considered.

[0074] (Problem setting) In this embodiment, a feed-forward model-type meta-learning method specialized for zero-shot learning is used to achieve high-performance zero-shot learning even for unknown tasks. Therefore, we will clarify the differences between the problem setting of the conventional technology and the present invention.

[0075] Figure 1 illustrates the problem setting of the prior art and the present invention. As shown in Figure 1(a), in typical supervised learning, the data of classes available for training is the subject of inference. In contrast, as shown in Figure 1(b), in the problem setting of the zero-shot learning of the present invention, it is assumed that the class to be inferred is an unknown class that does not appear at all during training.

[0076] Furthermore, as shown in Figure 1(c), in a typical meta-learning problem setting, the learning target is broadened to a task distribution in order to cope with changing tasks, and learning methods effective for multiple tasks are considered. As shown in Figure 1(d), compared to the setting in Figure 1(c), the problem setting targeted by this embodiment is a combination of conventional zero-shot learning and meta-learning problem settings. In other words, this embodiment simultaneously solves two problems: the need to learn about the task distribution as a countermeasure against changing tasks, and the fact that the inference target for those tasks is an unknown class that does not appear during training.

[0077] The differences in these problem settings are summarized in Table 3 below.

[0078] [Table 3]

[0079] Based on the above problem setting, we redefine the symbols explained above. The task set τ for meta-training. src of

[0080]

number

[0081] Therefore, support set D i spt In the input space X⊂R D X×S consists only of instances of and the known class set S, and queryset D i qry This becomes X × U, consisting only of instances of X and labels of the unknown class set U. Similarly, the task set τ for metatesting. tgt teeth,

[0082]

number

[0083] This is defined as, and the training dataset D i tr It consists of X × S, and the test dataset D i te It is composed of X × U. Here, N src and N tgt These are τ src and τ tgt The number of elements is T i src and T i tgt This is sampled from a task distribution p(τ).

[0084] In this type of problem setting, as with the feed-forward model type meta-learning described above, support set D i spt Even if you train an encoder using the data from queryset D, i qryThis approach is not always effective for classifying data. This is because the data in the support set is sampled from a known class set S, while the data in the query set is sampled from an unknown class set U, resulting in different data distributions. Therefore, as a machine learning technique specifically suited to this problem setting, we use a feed-forward model-type meta zero-shot learning method.

[0085] Figure 2 shows an overview of the Meta zero-shot learning method. In the Meta zero-shot learning method shown in Figure 2, a composite of three models is provided as the input model. The three models are the encoder h(·): A×A→W used in Conditional Neural Process, the decoder model g(·): A×A×W→S∪U, and the projection model Φ(·): X→A used in zero-shot learning to the auxiliary information space. Here, the encoder and decoder differ from those in the meta-learning method described above. The input to the encoder model h is an instance x in the input space. j Its label y j Instead, the output of the projection model Φ(x j ) and the prototype of the label π(y j )∈A s The following vector r is taken as input. j Outputs. r j =h(Φ(x j ),π(y j )) Here, π is a function π(·): S∪U→A that represents the correspondence between classes and prototypes. The input to the decoder model g is an unknown instance x u and vector r j Well then ,Φ(x u ) and the first auxiliary information A u and r j The mean r of the function is taken as input, and the following estimated results y^, such as categorical distributions, are output.

[0086]

number

[0087] Here, the first auxiliary information A u This is the auxiliary information defined in the explanation of zero-shot learning.

[0088]

number

[0089] Here, J is the number of elements in the dataset input to the encoder model h. Thus, the proposed method's model differs from the feed-forward model-type meta-learning described above in that it uses a zero-shot learning projection model Φ and auxiliary information.

[0090] The objective of the proposed method is to estimate the parameters of the three models h, g, and Φ mentioned above. However, for the sake of simplicity, the entire model formed by combining g, h, and Φ is given the parameter f(·;θ)=g(Φ(·),A u We will express this as ,h(·)) and estimate the parameters θ of the composite model.

[0091] (Input data, input model, and output) Next, we will describe the input data, input model, and output. Input data: The input data consists of (i) the set of tasks used for metatraining.

[0092]

number

[0093] (ii) Used as supplementary information in zero-shot learning

[0094]

number

[0095] That is the case.

[0096] Input model: The input model is (i) X⊂R D The auxiliary information space A⊂R takes an instance of as input. M (ii) A model Φ(·) that projects onto the dataset, (ii) a model h(·) that encodes the task from the dataset, and (iii) the output of h, the output of Φ, and the first auxiliary information A. u Let f(·;θ) be a composite model combining three models: a decoding model g(·) that estimates the categorical distribution from, and a composite model f(·;θ) that combines these three models. Here, θ is a parameter of the composite model f. Φ is R D →R M Any model that can be projected onto Φ is available, for example, a linear model such as the one disclosed in Non-Patent Document 5, or a nonlinear model such as a Neural Network. Furthermore, Φ may be a pre-trained model for each task. h may be any model, such as a multilayer perceptron as described in Non-Patent Document 4, or a CNN. Also, although g is said to output a categorical distribution, it may be a nonparametric model such as a zero-shot learned kNN or one-vs-rest method (see, for example, Non-Patent Documents 6 and 7).

[0097] Output: The output of this method is the optimal parameter estimation result θ of the composite model f. * That is the case.

[0098] (Objective function) Parameter estimation in the proposed method is performed by optimizing the objective function. The objective of this method is to estimate model parameters that improve the performance of zero-shot learning on multiple unknown tasks. Therefore, the objective function is constructed from the loss of zero-shot learning on unknown tasks. Based on this, the objective function is defined as shown in equations (9) to (11) below.

[0099]

number

[0100] Here, L zsl This is an arbitrary loss function (e.g., equation (7)) whose value is small when g(·) and y are close, such as those used in zero-shot learning. The difference from the objective function described above is that, as shown in equation (10), the meta-learning loss is calculated using the zero-shot learning loss for multiple tasks. In other words, comparing the loss function of a general feed-forward model type meta-learning (e.g., equation (3)) with equation (9) in this embodiment, in this embodiment, auxiliary information is used in the input and L meta is L zsl The difference lies in the fact that it is evaluated using the loss function. Also, the model Φ is pre-trained for each task. i When using this, replace Φ in equation (10) with the pre-trained model Φ i That's all you need to do.

[0101] (Optimization method) Any optimization method can be applied to optimize the objective function, such as gradient descent, stochastic gradient descent, or Adam. When using gradient descent, the parameters should be updated according to the following equation at the kth optimization step.

[0102]

number

[0103] Here is a blank γ k This can be calculated numerically, either by using a function derived from calculating the gradient ∇L(·) of the objective function, which is the learning rate parameter.

[0104] (composition) Figure 3 is a block diagram showing an example of the hardware configuration of the model learning device 1 according to this embodiment. Model learning device 1 is a computer that analyzes input data, generates output data, and outputs it. Model learning device 1 can be installed in any location.

[0105] As shown in Figure 3, the model learning device 1 comprises a control unit 10, a program storage unit 20, a data storage unit 30, a communication interface 40, and an input / output interface 50. The control unit 10, program storage unit 20, data storage unit 30, communication interface 40, and input / output interface 50 are connected to each other via a bus so as to be able to communicate with each other. Furthermore, the communication interface 40 may be connected to an external device so as to be able to communicate with it via a network. In addition, the input / output interface 50 is connected to an input device 2 and an output device 3 so as to be able to communicate with them.

[0106] The control unit 10 controls the model learning device 1. The control unit 10 includes a hardware processor such as a central processing unit (CPU). For example, the control unit 10 may be an integrated circuit capable of executing various programs.

[0107] The program storage unit 20 can use a combination of non-volatile memory that allows writing and reading at any time, such as EPROM (Erasable Programmable Read Only Memory), HDD (Hard Disk Drive), and SSD (Solid State Drive), as a storage medium, and non-volatile memory such as ROM (Read Only Memory). The program storage unit 20 stores programs necessary to execute various processes. In other words, the control unit 10 can realize various controls and operations by reading and executing programs stored in the program storage unit 20.

[0108] The data storage unit 30 is a storage device that uses a combination of non-volatile memory, such as an HDD or memory card, which allows for writing and reading at any time, and volatile memory, such as RAM (Random Access Memory), as storage media. The data storage unit 30 is used to store data acquired and generated during the process in which the control unit 10 executes a program and performs various processing.

[0109] The communication interface 40 includes one or more wired or wireless communication modules. For example, the communication interface 40 includes a communication module that connects to an external device via a network, either wired or wirelessly. The communication interface 40 may also include a wireless communication module that connects to an external device wirelessly, such as a Wi-Fi access point and a base station. Furthermore, the communication interface 40 may include a wireless communication module that connects to an external device wirelessly using short-range wireless technology. In other words, the communication interface 40 can be any general communication interface that can communicate with an external device and send and receive various types of information under the control of the control unit 10.

[0110] The input / output interface 50 is connected to the input device 2 and the output device 3, etc. The input / output interface 50 is an interface that enables the transmission and reception of information between the input device 2 and the output device 3. The input / output interface 50 may be integrated with the communication interface 40. For example, the model learning device 1 and the input device 2 or the output device 3 may be wirelessly connected using short-range wireless technology, and information may be transmitted and received using said short-range wireless technology.

[0111] Input device 2 may include, for example, a keyboard or pointing device for the user to input various information to model learning device 1. Input device 2 may also include a reader for reading data to be stored in program storage unit 20 or data storage unit 30 from a memory medium such as a USB memory, or a disk device for reading such data from a disk medium.

[0112] Output device 3 includes a display that shows model parameters and the like estimated by model learning device 1.

[0113] Figure 4 is a block diagram showing the software configuration of the model learning device 1 in the embodiment, in relation to the hardware configuration shown in Figure 1. The control unit 10 comprises an acquisition unit 101, a model parameter learning unit 102, and an output control unit 103. The data storage unit 30 comprises an acquired data storage unit 301 and a model parameter storage unit 302.

[0114] The acquisition unit 101 includes a data acquisition unit 1011 and a parameter acquisition unit 1012.

[0115] The data acquisition unit 1011 acquires input data and model data. The input data includes a set of observable tasks used in metatraining and first auxiliary information (A) that indicates a set of prototypes of known classes in zero-shot learning. u ), second auxiliary information (A) that shows a set of prototypes of the unknown class s ) includes. Model data is X⊂R D The auxiliary information space A⊂R takes an instance of as input. M The model parameters of a composite model are included, which combines a first model (Φ(·)) that projects onto the dataset, a second model (h(·)) that encodes the task from the dataset, and a third model (g(·)) that is a decoding model that estimates the categorical distribution from the output of the first model, the output of the second model, and the second auxiliary information.

[0116] Furthermore, the parameter acquisition unit 1012 acquires the set parameters. Here, the set parameters are the learning rate parameters used when estimating the optimal model parameters.

[0117] The model parameter learning unit 102 learns the model parameters. The model parameter learning unit 102 includes an initialization unit 1021, a count setting unit 1022, an update unit 1023, and a determination unit 1024.

[0118] The initialization unit 1021 initializes the number of calculation iterations. Before learning the model parameters, the initialization unit 1021 initializes the number of calculation iterations stored in the model parameter storage unit 302, which will be described later.

[0119] The count setting unit 1022 sets the maximum number of repetitions. The count setting unit 1022 sets a value for the maximum number of repetitions, which indicates the maximum number of times the model parameter update process will be repeated. The maximum number of repetitions may be set, for example, according to the administrator's input.

[0120] The update unit 1023 uses the input data, input model, and setting parameters (learning rate parameters) to update the model parameters using the optimization method described above, so that the loss function used in meta-training is evaluated by the loss function used in zero-shot learning. Details of how to update the model parameters will be described later.

[0121] Furthermore, the update unit 1023 updates the number of calculation iterations. For example, the update unit 1023 updates the number of calculation iterations by increasing the value of the number of calculation iterations by 1.

[0122] The determination unit 1024 determines whether the number of calculation iterations > the maximum number of iterations.

[0123] The output control unit 103 outputs model parameters. The output control unit 103 may display the model parameters on the display of the output device 3 via the input / output interface 50. In addition, the output control unit 103 may, at any time in accordance with the administrator's instructions, display the learned model parameters stored in the model parameter storage unit 302 on the display of the output device 3.

[0124] (operation) Figure 5 is a flowchart illustrating an example of the general operation by the model learning device 1 to obtain the model parameters θ. The operation of this flowchart is realized when the control unit 10 of the model learning device 1 reads and executes the program stored in the program storage unit 20.

[0125] This operation may be initiated by the administrator of the model learning device 1 inputting data, setting parameters, etc., into the input device 2. Alternatively, it may be initiated in response to instructions from the administrator.

[0126] In step ST101, the data acquisition unit 1011 acquires input data and model data. Here, the input data is a set of observable tasks τ src Supplementary Information A s ,A u The model data includes a composite model f, which is a combination of the multiple models described above. The data acquisition unit 1011 stores the acquired input data and model data in the data storage unit 3011. The input data and model data may be information entered by the administrator into the input device 2, or information stored in the data storage unit 30.

[0127] In step ST102, the data acquisition unit 1011 acquires the setting parameters. Here, the setting parameters are the learning parameters γ used during optimization. k This includes the following. Furthermore, the setting parameters may be information entered by the administrator into the input device 2, or information stored in the data storage unit 30. The data acquisition unit 1011 stores the acquired setting parameters in the parameter storage unit 3012.

[0128] In step ST103, the model parameter learning unit 102 learns the model parameters θ.

[0129] Figure 6 is a flowchart illustrating an example of the overview operation, which explains the operation of step ST103 in more detail. In step ST201, the initialization unit 1021 initializes the number of calculation iterations. Before learning the model parameters θ, the initialization unit 1021 initializes the number of calculation iterations stored in the model parameter storage unit 302.

[0130] In step ST202, the count setting unit 1022 sets the maximum number of repetitions. The count setting unit 1022 sets a value for the maximum number of repetitions, which indicates the maximum number of times the model parameter update process will be repeated. The maximum number of repetitions may be set, for example, according to the administrator's input.

[0131] In step ST203, the update unit 1023 updates the model parameters θ according to equations (9) and (12) shown above. That is, the update unit 1023 updates the model parameters θ according to the optimization steps shown in equation (12) using the objective function shown in equation (9).

[0132] In step ST204, the update unit 1023 updates the number of calculation iterations. For example, the update unit 1023 updates the number of calculation iterations by increasing the value of the number of calculation iterations by 1.

[0133] In step ST205, the determination unit 1024 determines whether the number of calculation iterations > the maximum number of iterations. If it determines that the number of calculation iterations is less than or equal to the maximum number of iterations, the process returns to step ST203. On the other hand, if the number of calculation iterations is greater than the maximum number of iterations, the model parameter learning unit 102 outputs the learned model parameters θ to the output control unit 103. The model parameter learning unit 102 also stores the learned model parameters θ in the model parameter storage unit 302. Then, the process proceeds to step ST104.

[0134] Returning to Figure 5, in step ST104, the output control unit 103 outputs the model parameter θ. The output control unit 103 may display the model parameter θ on the display of the output device 3 via the input / output interface 50. Alternatively, the output control unit 103 may, at any time according to the administrator's instructions, display the learned model parameter θ stored in the model parameter storage unit 302 on the display of the output device 3.

[0135] (Examples of application) Finally, we will discuss an application example using optimized model parameters. For example, in this embodiment, the process consists of two stages: metatraining and metatesting, similar to conventional meta-learning. In metatraining, the task set τ src and supplementary information A s , A u Then, the model is input into the method described above, and the objective function is minimized.

[0136]

number

[0137] By using an optimization method, the estimated parameter θ * This is obtained. Then, during metatesting, the unknown task T i tgt Training dataset D i tr and instance x of the unknown input space u ∈D i te Since we obtain y^ u =f(x u ,D i tr ,A u ;θ * By calculating ), the estimated value y^ u This can be obtained.

[0138] (Effects and Benefits) According to this embodiment, in zero-shot learning, which involves learning a model capable of classifying instances of a class for which no data is available, the performance of zero-shot learning for future unknown tasks can be improved by incorporating a meta-learning approach that learns the learning process itself. This allows for efficient learning of individual models for unknown individuals, even in cases where data is unavailable, such as estimating what a person is recalling from their brainwaves, and where individual differences exist, requiring the learning of individual models from scratch.

[0139] [Other embodiments] The above embodiment can also be applied to more general cases, such as when it is necessary to perform zero-shot learning individually, not only for individual differences but also for regional differences, differences in data acquisition environments, etc.

[0140] Furthermore, the method described in the above embodiment can be distributed by storing the program (software means) that can be executed by a computer in a storage medium such as a magnetic disk (floppy disk, hard disk, etc.), optical disk (CD-ROM, DVD, MO, etc.), or semiconductor memory (ROM, RAM, flash memory, etc.), and by transmitting it via a communication medium. The program stored on the medium also includes a configuration program that configures the software means (including not only the execution program but also tables and data structures) to be executed by the computer. The computer realizing this device reads the program stored in the storage medium and, if necessary, constructs the software means using the configuration program, and executes the above-mentioned processing by controlling the operation of this software means. The storage medium referred to in this specification is not limited to distribution mediums, but also includes storage mediums such as magnetic disks and semiconductor memory provided inside the computer or in devices connected via a network.

[0141] In short, this invention is not limited to the embodiments described above, and can be modified in various ways during implementation without departing from its essence. Furthermore, each embodiment may be combined as appropriately as possible, in which case the combined effects can be obtained. Moreover, the embodiments described above include inventions at various stages, and various inventions can be extracted by appropriate combinations of the multiple constituent elements disclosed. [Explanation of Symbols]

[0142] 1…Model learning device 2…Input device 3…Output device 10…Control Unit 101…Acquisition Department 1011...Data acquisition unit 1012...Parameter acquisition section 102...Model parameter learning unit 1021...Initialization section 1022...Number of times setting section 1023…Update section 1024…Judgment section 103…Output Control Unit 20...Program memory unit 30...Data storage unit 301...Data Acquisition Storage Unit 3011...Data storage unit 3012...Parameter storage unit 302...Model parameter storage unit 40…Communication Interface 50… Input / Output Interface

Claims

1. A model learning device that enables zero-shot learning for unknown tasks, An acquisition unit acquires input data including a set of observable tasks used in metatraining, first auxiliary information indicating a set of known class prototypes in the zero-shot learning, and second auxiliary information indicating a set of unknown class prototypes, and model data including a composite model consisting of multiple models, and acquires the learning rate parameters of the model data. A model parameter learning unit learns the model parameters of the composite model using an optimization method that minimizes the objective function, which is evaluated by the loss function used in the metatraining, using the input data, the model data, and the learning rate parameters. An output control unit that controls the display of the aforementioned model parameters on an output device, A model learning device comprising: a first model that projects instances of X⊂R D as input onto an auxiliary information space A⊂R M; a second model that encodes a task from a dataset; and a third model that estimates a categorical distribution from the output of the first model, the output of the second model, and the second auxiliary information, wherein X is an input space represented by D-dimensional real numbers, and R is a real number space.

2. The model learning device according to claim 1, further comprising an update unit that updates the model parameters a predetermined number of times using the learning rate parameter and the gradient of the objective function.

3. The model learning apparatus according to claim 1, wherein the second model outputs a vector with the output of the first model and the first auxiliary information as input, and the prototype is a function representing the correspondence between a class and the prototype.

4. The model learning apparatus according to claim 3, wherein the third model estimates the categorical distribution using the output of the first model and the mean of the vector as inputs.

5. The model learning device according to claim 1, wherein learning the model parameters is to estimate the parameters of the first model, the second model, and the third model.

6. A model learning method executed by the processor of a model learning device that achieves zero-shot learning for an unknown task, The process involves obtaining input data including a set of observable tasks used in metatraining, first auxiliary information indicating a set of known class prototypes in the zero-shot learning, and second auxiliary information indicating a set of unknown class prototypes, as well as model data including a composite model consisting of multiple models. To obtain the learning rate parameters of the aforementioned model data, Using the input data, the model data, and the learning rate parameters, the model parameters of the composite model are learned using an optimization method such that the loss function used in metatraining minimizes the objective function evaluated by the loss function used in zero-shot learning. Controlling the model parameters to be displayed on the output device, A model learning method comprising: a first model that projects instances of X⊂R D as input onto an auxiliary information space A⊂R M; a second model that encodes a task from a dataset; and a third model that estimates a categorical distribution from the output of the first model, the output of the second model, and the second auxiliary information, wherein X is an input space represented by D-dimensional real numbers, and R is a real number space.

7. A model learning program executed by the processor of a model learning device that achieves zero-shot learning for an unknown task, The process involves obtaining input data including a set of observable tasks used in metatraining, first auxiliary information indicating a set of known class prototypes in the zero-shot learning, and second auxiliary information indicating a set of unknown class prototypes, as well as model data including a composite model consisting of multiple models. To obtain the learning rate parameters of the aforementioned model data, Using the input data, the model data, and the learning rate parameters, the model parameters of the composite model are learned using an optimization method such that the loss function used in metatraining minimizes the objective function evaluated by the loss function used in zero-shot learning. Controlling the model parameters to be displayed on the output device, A model learning program comprising instructions for executing the above-mentioned processor, wherein the composite model is a composite model that combines a first model that projects instances of X⊂R D as input onto an auxiliary information space A⊂R M, a second model that encodes a task from a dataset, and a third model that estimates a categorical distribution from the output of the first model, the output of the second model, and the second auxiliary information, wherein X is an input space represented by D-dimensional real numbers, and R is a real number space.

Citation Information

Patent Citations

  • Learning device and learning method

    WO2020121378A1

  • Process to learn new image classes without labels

    WO2021118697A1