Learning device, learning system, method and program
Patent Information
- Application Number
- JP2025503298
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-08-01
- Publication Date
- 2025-10-14
AI Technical Summary
Existing PU learning methods face challenges in achieving high accuracy when dealing with data from domains with different feature spaces, particularly in binary classification scenarios where only positive or negative examples are labeled, and unlabeled data is used for inference.
A learning device and system that utilize a combination of data sets from different domains with distinct feature spaces, employing a conversion model to align features into a common space, and an inference model to classify data as positive or negative examples, while optimizing the conversion and inference models using loss calculations to improve domain adaptation and adversarial learning techniques.
This approach enables highly accurate machine learning in hybrid domains with different feature spaces, effectively aligning distributions and improving classification performance even when only positive examples are available in the source domain.
Abstract
Description
Learning device, learning system, method, program, and storage medium
[0001] The present disclosure relates to a learning device, a learning system, a method, a program, and a storage medium.
[0002] Patent Document 1 describes that PU learning is a learning method for performing machine learning when only a portion of positive examples are given as training data, and that the discriminant model learned by PU learning is an estimation model that estimates the probability of a positive example for data whose positive or negative status is unknown.
[0003] Japanese Patent Application Laid-Open No. 2020-173673
[0004] Wenpeng Hu1, Ran Le2, Bing Liu3, Feng Ji, Jinwen Ma, Dongyan Zhao, and Rui Yan2, Predictive Adversarial Learning from Positive and Unlabeled Data, AAAI 2021.
[0005] There is a problem in that it is necessary to improve the accuracy of PU learning in two domains with different feature spaces.
[0006] In view of the above circumstances, an object of this disclosure is to provide a learning device, a learning system, a method, a program, and a storage medium that solve the above-mentioned problems.
[0007] (1) One aspect of the present disclosure is a learning device including: an acquisition unit that acquires a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in a binary classification; and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification; and a model generation unit that generates, based on the first dataset and the second dataset, an inference model for inferring whether second data included in the second dataset is a positive example or a negative example in the binary classification.
[0008] (2) One aspect of the present disclosure is a learning system including a first learning device, a second learning device, and a server device, wherein the first learning device includes a first acquisition unit, a first model generation unit, a first conversion unit, a first inference unit, a first loss calculation unit, a second loss calculation unit, a first model optimization unit, and a first communication unit, wherein the first acquisition unit acquires a first predetermined inference model, a first predetermined conversion model, and a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in binary classification, and the first conversion unit generates a first dataset based on the first predetermined conversion model. the first communication unit transmits the converted first feature to the server device; the first communication unit receives the converted second feature from the server device; the first inference unit generates a first inference value indicating whether the first data corresponding to the converted first feature corresponds to a positive example or a negative example, based on the converted first feature and the first predetermined inference model; and the first loss calculation unit calculates a first inference value indicating whether the first data corresponding to the converted first feature corresponds to a positive example or a negative example, based on the converted first feature and the first predetermined inference model. a first loss calculated by the first predetermined transformation model based on the transformed second feature and the first inference value; a second loss calculation unit calculates a second loss calculated by the first predetermined inference model based on the first inference value; a first model optimization unit updates the first predetermined transformation model and the first predetermined inference model based on the first loss and the second loss; a first communication unit transmits the updated first predetermined inference model and the updated first predetermined transformation model to the server device; the second acquisition unit acquires a second predetermined inference model and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification; the second conversion unit converts second features indicated by second data included in the second dataset into features in the common feature space based on the second predetermined conversion model; and the second communication unit transmits the converted second features to the server device;The second communication unit receives the converted first feature from the server device, the second inference unit generates a second inference value indicating whether second data corresponding to the converted second feature corresponds to a positive example or a negative example, based on the converted second feature and the second predetermined inference model, the third loss calculation unit calculates a third loss according to the second predetermined conversion model, based on the converted first feature, the converted second feature, and the second inference value, and the fourth loss calculation unit calculates a third loss according to the second predetermined conversion model, based on the second inference value. a fourth loss by an inference model is calculated; the second model optimization unit updates the second predetermined conversion model and the second predetermined inference model based on the third loss and the fourth loss; the second communication unit transmits the updated second predetermined inference model and the updated second predetermined conversion model to the server device; the server device includes a third communication unit and a processing unit; the third communication unit receives the converted first feature from the first learning device and the converted second feature from the second learning device; transmits the converted second feature quantity to the first learning device and transmits the converted first feature quantity to the second learning device; the third communication unit acquires the updated first predetermined inference model, the updated first predetermined conversion model, the updated second predetermined inference model, and the updated second predetermined conversion model; the processing unit generates a third updated inference model based on the updated first predetermined inference model and the updated second predetermined inference model and generates a third updated conversion model based on the updated first predetermined conversion model and the updated second predetermined conversion model; the third communication unit transmits the third updated inference model and the third updated conversion model to the first learning device and the second learning device; the first communication unit and the second communication unit acquire the third updated inference model and the third updated conversion model; the first model optimization unit updates the first predetermined inference model based on the first predetermined inference model and the third updated inference model, andThe learning system updates the first predetermined transformation model, the second model optimization unit updates the second predetermined inference model based on the second predetermined inference model and the third updated inference model, and updates the second predetermined transformation model based on the second predetermined transformation model and the third transformation model.
[0009] (3) One aspect of the present disclosure is a learning system including a first learning device, a second learning device, and a server device, wherein the first learning device includes a first acquisition unit, a first conversion unit, a first discrimination unit, a fifth loss calculation unit, a first model optimization unit, and a first communication unit, wherein the first acquisition unit acquires a first predetermined transformation model, a first predetermined discrimination model, and a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in binary classification, and the first conversion unit performs a transformation on the first dataset based on the first predetermined transformation model. a first feature indicated by the included first data being converted into a feature in a common feature space; the first discrimination unit generating a first discrimination value based on the converted first feature and the first predetermined discrimination model; the fifth loss calculation unit calculating a fifth loss by the first predetermined transformation model based on the discrimination value and adversarial learning; the first model optimization unit updating the first predetermined transformation model and the first predetermined discrimination model based on the fifth loss; and the first communication unit transmitting the updated first predetermined discrimination model to the server device. and transmitting the second learning device, the second learning device comprising a second acquisition unit, a second conversion unit, a second inference unit, a second identification unit, a sixth loss calculation unit, a second model optimization unit, and a second communication unit, the second acquisition unit acquiring a second predetermined conversion model, a second predetermined inference model, a second predetermined identification model, and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification, and the second conversion unit converts a second feature indicated by second data included in the second dataset into a feature in the common feature space based on the second predetermined conversion model. the second inference unit converts the second feature into a second feature and generates an inference value based on the converted second feature and the second predetermined inference model; the sixth loss calculation unit calculates a sixth loss using the second predetermined conversion model based on the converted second feature and the second inference value; the second model optimization unit updates the second predetermined conversion model, the second predetermined inference model, and the second predetermined discrimination model based on the sixth loss; the second communication unit transmits the updated second predetermined discrimination model to the server device; and the server devicea learning system including a third communication unit and a processing unit, wherein the third communication unit acquires the updated first predetermined discriminative model and the updated second predetermined discriminative model; the processing unit generates a third updated discriminative model based on the updated first predetermined discriminative model and the updated second predetermined discriminative model; the third communication unit transmits the third updated discriminative model to the first learning device and the second learning device; the first communication unit and the second communication unit acquire the third updated discriminative model; the first model optimization unit updates the first predetermined discriminative model based on the first predetermined discriminative model and the third updated discriminative model; and the second model optimization unit updates the second predetermined discriminative model based on the second predetermined discriminative model and the third updated discriminative model.
[0010] (4) One aspect of the present disclosure is a method executed by a computer, the method including the steps of acquiring a first dataset belonging to a first domain including a first feature space and labeled as only one of positive examples and negative examples in a binary classification, and a second dataset belonging to a second domain including a second feature space and not labeled by a binary classification, and generating an inference model based on the first dataset and the second dataset to infer whether second data included in the second dataset is a positive example or a negative example in a binary classification.
[0011] (5) One aspect of this disclosure is a program that causes a computer to execute the steps of acquiring a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in a binary classification, and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification, and generating an inference model based on the first dataset and the second dataset to infer whether second data included in the second dataset is a positive example or a negative example in the binary classification.
[0012] (6) One aspect of this disclosure is a storage medium having stored thereon a program that causes a computer to execute the steps of acquiring a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in a binary classification, and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification, and generating an inference model based on the first dataset and the second dataset to infer whether second data included in the second dataset is a positive example or a negative example in a binary classification.
[0013] It is possible to learn with high accuracy even in domains with different feature spaces.
[0014] 1 is a diagram showing a learning flow by a learning device in the first embodiment. FIG. 2 is a diagram showing hybrid domain adaptation in the first embodiment. FIG. 3 is a block diagram showing an example functional configuration of the learning device 10 in the first embodiment. FIG. 4 is a flowchart showing the processing flow of the learning device 10 in the first embodiment. FIG. 5 is another example configuration of the feature converter in the first embodiment. FIG. 6 is another example configuration of the feature converter in the first embodiment. FIG. 7 is a block diagram showing an example functional configuration of the learning device 10 in the second embodiment. FIG. 8 is a flowchart showing the processing flow of the learning device 10 in the second embodiment. FIG. 9 is a block diagram showing an example functional configuration of the learning device 10 in the third embodiment. FIG. 10 is a flowchart showing the processing flow of the learning device 10 in the third embodiment. FIG. 11 is a flowchart showing the processing flow of the learning device 10 in the third embodiment. FIG. 12 is a block diagram showing an example functional configuration of a learning system S in the fourth embodiment. FIG. 13 is a block diagram showing an example functional configuration of a first learning device 11 in the fourth embodiment. FIG. 14 is a block diagram showing an example functional configuration of a second learning device 12 in the fourth embodiment. FIG. 15 is a block diagram showing an example functional configuration of a server device 30 in the fourth embodiment. FIG. 16 is a flowchart showing the processing flow of the first learning device 11 in the fourth embodiment. FIG. 17 is a flowchart showing the processing flow of the second learning device 12 in the fourth embodiment. FIG. 18 is a flowchart showing the processing flow of the learning device 30 in the fourth embodiment. 10 is a block diagram showing an example functional configuration of a first learning device 11 in a fifth embodiment. FIG. 11 is a block diagram showing an example functional configuration of a second learning device 12 in a fifth embodiment. FIG. 12 is a flowchart showing the processing flow of a first learning device 11 in a fifth embodiment. FIG. 13 is a flowchart showing the processing flow of a second learning device 12 in a fifth embodiment. FIG. 14 is a flowchart showing the processing flow of a learning device 30 in a fifth embodiment. FIG. 15 is a diagram showing the minimum configuration of a learning device 10 in this disclosure. FIG. 16 is a flowchart showing the processing flow in an embodiment with the minimum configuration of a learning device 10 in this disclosure. FIG. 17 is a diagram explaining the hardware configuration of each device according to the present embodiment.
[0015] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the disclosure according to the claims, and not all combinations of features described in the embodiments are necessarily required. Two or more features among the multiple features described in the embodiments may be arbitrarily combined. Furthermore, the same reference numerals are used for the same or similar components, and redundant description will be omitted.
[0016] (First embodiment) Figure 1 is a diagram showing the flow of learning by a learning device in the first embodiment. In this embodiment, the domain in which data is handled is logically and / or physically classified into a source domain and a target domain. In this embodiment, the source domain may be referred to as a first domain, and the target domain may be referred to as a second domain.
[0017] The source domain includes a first dataset, and the first dataset includes only one of binary classifications, i.e., positive and negative examples. For example, the first dataset includes positive examples but not negative examples. Alternatively, the first dataset includes negative examples but not positive examples. The second dataset does not include data labeled with a binary classification.
[0018] 1, the first data set of the source domain includes data with features specific to the source domain and data with features common to the source and target domains, and the second data set of the target domain similarly includes data with features specific to the target domain and data with features common to the source and target domains.
[0019] Feature transformer F in the source domain s transforms the first dataset into features in a feature space common to the target domain based on a predetermined transformation model. tThe second dataset is transformed into features in a feature space common to the source domain based on a predetermined transformation model. The predetermined transformation model may be a linear transformation that maps to the common feature space, or a nonlinear transformation such as a neural network. The transformation of the features may involve estimating features not included in the dataset based on the data and features included in the dataset.
[0020] The inference unit C (inference unit) is shared by the source domain and the target domain. The inference unit C uses a predetermined inference model to estimate whether unlabeled data included in the second dataset is a positive example or a negative example in a binary classification. The predetermined inference model is based on machine learning using PU learning. PU learning is a learning method for learning an inference model that infers whether a data item is P or N in an environment where labels for positive examples (P: positive) and negative examples (N: negative) in a binary classification exist, in a situation where only data U (Unlabeled) that is not labeled as either P or N and P are available.
[0021] An inference model is a machine learning model that infers whether input data is a positive or negative example. Examples of inference models include logistic regression and neural networks. The output of an inference operation based on an inference model is a two-dimensional vector, where the first component represents the probability that the input data is a positive example and the second component represents the probability that the input data is a negative example.
[0022] Although details will be described later, in this embodiment, a predetermined conversion model and a predetermined inference model are updated by machine learning so that the loss in inference by PU learning and the loss in feature conversion are reduced.
[0023] Figure 2 is a diagram illustrating hybrid domain adaptation in the first embodiment. Figure 2 shows an example of a first dataset in the source domain and an example of a second dataset in the target domain. The first dataset includes data such as sample ID, marital status, gender, and age of the individuals listed in the table. The items shown in each column are examples of the first data.
[0024] The second dataset includes data such as the sample ID, gender, age, deposit balance, and loan balance of the individuals listed in this table. The items shown in each of these columns are examples of the second data. If data such as "whether or not the annual income is 10 million yen" (not shown) is to be labeled, the first dataset includes only data labeled with P, and the second dataset includes data labeled neither with P nor with N.
[0025] In Figure 2, the attributes gender and age exist in both the source domain and the target domain. Such features are called "common features." On the other hand, "marital status" exists only in the target domain, not in the source domain. Furthermore, "bank balance" and "loan balance" exist only in the source domain, not in the target domain. Two domains with different attributes (features) like this are also called "hybrid domains."
[0026] 3 is a block diagram showing an example of the functional configuration of the learning device 10 according to the first embodiment. The learning device 10 includes an acquisition unit 110, a data processing unit 130, a conversion unit 131, an inference unit 132, a loss calculation unit 140, a first loss calculation unit 141, a second loss calculation unit 142, a model optimization unit 150, and a storage unit 160.
[0027] The acquisition unit 110 acquires a predetermined conversion model, a predetermined inference model, a first dataset, and a second dataset, and stores them in the storage unit 160. The conversion unit 131 reads the predetermined conversion model from the storage unit 160 and converts the first dataset and the second dataset into datasets in a common feature space. The conversion unit 131 may convert the first dataset and the second dataset into data in a common feature space based on the predetermined conversion model so that the first dataset and the second dataset contain features that were not included in the first dataset and the second dataset at the stage when the acquisition unit 110 acquired them.
[0028] Converting the data into data in a common feature space may mean, for example, converting the first and second data sets into a feature space that includes all items shown in the table of Fig. 2 based on a predetermined estimation model. In this case, the first and second data sets after conversion will both include common attributes, and there will be no features that are not included in either the first or second data set. In other words, features (or attributes) that were not included in the data sets before conversion may be estimated based on the predetermined estimation model and other features that were included in the data sets before conversion.
[0029] The inference unit 132 reads a predetermined inference model from the storage unit 160, and infers whether unlabeled data included in the converted first data set and second data set is a positive example or a negative example based on PU learning for the converted first data set and second data set. The predetermined conversion model and the predetermined inference model may be preset in the conversion unit 131 and the inference unit 132, respectively.
[0030] The first loss calculation unit 141 calculates a loss, a domain loss L, that brings the distributions of the transformed features of the source data (first data set) and the target data (second data set) closer together. d The first loss calculation unit 141 calculates the domain loss L d For the calculation, MMD (maximum mean discrepancy) as shown in the following equation 1 may be used.
[0031] where x s i and x t j are the transformed features of the source data and target data for samples i and j, respectively, and n s and n t are the number of samples in the source data and target data, respectively. Φ is the mapping.
[0032] Also, L d As shown in Equation 2, the covariance matrix M for the transformed features of the source data and target data iss and M t The squared norm of the difference may also be used.
[0033] While unlabeled target data contains both positive and negative examples, the source data only contains positive examples, so the above method does not guarantee that the distributions are sufficiently similar. Therefore, by calculating the domain loss for the entire source data and for data that are likely to be positive examples in the target data, it is possible to optimize the conversion model and inference model to achieve better performance.
[0034] As an example of this method, the second data set determined as a positive case by the inference unit 132 based on a predetermined inference model is used as the domain loss (the first loss calculated by the first loss calculation unit). Let the inference unit 132 be C, and let the set determined as a positive case (i.e., 1) by the inference unit 132 be Sp = {j|C (x t j ) = 1}, then, for example, Equation 1 is modified to the following Equation 3.
[0035] The second loss calculation unit 142 calculates the loss for executing PU learning, the PU loss L pu (second loss). For example, Equation 4 calculates the PU loss L pu This is a typical example.
[0036] L p + and L p - are the losses when the correct labels of the source data consisting only of positive examples are considered as positive and negative examples, respectively. u - is the loss when the correct label of the unlabeled target data is considered as a negative example. p is the prior probability of a positive example in the target data.
[0037] The model optimization unit 150 optimizes the loss L calculated by the first loss calculation unit 141 and the second loss calculation unit 142. d and L puBased on the above, the parameters of the conversion unit 131 and the inference unit 132, i.e., the predetermined conversion model and the predetermined inference model, are updated. For example, by introducing a parameter λ, the model optimization unit 150 calculates the loss L shown in Equation 5.
[0038] The model optimization unit 150 may update the parameters of the predetermined transformation model and the predetermined inference model using the gradient descent method based on the loss L.
[0039] As described above, in this embodiment, in the data processing unit 130, source data consisting only of positive examples and unlabeled target data are input to the conversion unit 131, whereby the respective features mapped onto a common feature space are calculated, and the converted features are input to the inference unit 132 to calculate a prediction score.
[0040] Based on these, the first loss calculation unit 141 and the second loss calculation unit 142 calculate the domain loss and the PU loss, respectively, and the model optimization unit 150 learns a predetermined transformation model and a predetermined inference model by minimizing the losses. In this way, by simultaneously learning the predetermined transformation model and the predetermined inference model, it becomes possible to learn a binary classification inference machine (inference model) that shows excellent performance in the target domain.
[0041] FIG. 4 is a flowchart showing the flow of processing by the learning device 10 in the first embodiment.
[0042] In step S101, the acquisition unit 110 acquires a first dataset and a second dataset. The first dataset belongs to a first domain including a first feature space and is labeled as either a positive example or a negative example in a binary classification. The second dataset belongs to a second domain including a second feature space and includes data that is not labeled by a binary classification. The first domain may be referred to as a source domain, and the second domain may be referred to as a target domain.
[0043] The first feature is included in the first feature space but not in the second feature space, and the second feature is included in the second feature space but not in the first feature space. The first feature space and the second feature space may include common features (attributes).
[0044] In step S102 , the acquisition unit 110 acquires a predetermined transformation model and a predetermined inference model, and stores them in the storage unit 160 .
[0045] In step S103, the conversion unit 131 reads a predetermined conversion model from the storage unit 160, and converts each of the first feature amount and the second feature amount into a feature amount in a common feature amount space based on the predetermined conversion model. The common feature amount space has a correspondence relationship between the first feature amount space and the second feature amount space.
[0046] In step S104, the inference unit 132 reads a predetermined inference model from the storage unit 160 and generates an inference value based on the predetermined inference model, the converted first feature quantity, and the converted second feature quantity. The inference value indicates whether the second data is a positive example or a negative example.
[0047] In step S105, the first loss calculation unit 141 calculates a first loss between the converted first feature quantity and the converted second feature quantity, based on the converted first feature quantity and the converted second feature quantity.
[0048] In step S106, the second loss calculation unit 142 calculates a second loss according to a predetermined inference model based on the inferred value.
[0049] In step S107, the model optimization unit 150 updates the parameters of the predetermined transformation model and the predetermined inference model based on the first loss and the second loss. The storage unit 160 stores the updated transformation model and the inference model.
[0050] As described above, the learning device according to this embodiment includes an acquisition unit that acquires a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in binary classification, and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification; a model generation unit that generates, based on the first dataset and the second dataset, an inference model for inferring whether second data included in the second dataset is a positive example or a negative example in binary classification, and a conversion unit that converts, based on a predetermined conversion model, first features indicated by first data included in the first dataset into features in a common feature space and converts second features indicated by the second data into features in the common feature space; an inference unit that generates, based on the converted first features, the converted second features, and the inference model, an inference value indicating whether the second data corresponds to a positive example or a negative example; a first loss calculation unit that calculates a first loss between the transformed first feature and the transformed second feature based on the inference value; a second loss calculation unit that calculates a second loss by the inference model based on the inference value; and a model optimization unit that updates the inference model and the predetermined conversion model based on the first loss and the second loss, wherein the common feature space has a correspondence relationship between the first feature space and the second feature space, and the first feature is included in the first feature space but not in the second feature space, and the second feature is included in the second feature space but not in the first feature space.
[0051] This makes it possible to provide a learning device that trains a model for binary classification into positive and negative examples using source data consisting only of positive examples with different feature spaces and unlabeled target data when adapting to a hybrid domain in a PU learning setting. Even when there are only P (positive examples) on the source domain side, the distributions of the source and target domains can be aligned with high accuracy. Furthermore, even when the feature spaces of the source and target domains are different in PU learning, high-accuracy machine learning is possible.
[0052] Further, the inference value indicates whether the second data corresponds to a positive example or a negative example, and the converted second feature of the second data corresponding to either the positive example or the negative example is used as an input when the first loss calculation unit calculates the first loss.
[0053] The predetermined transformation model may also be in other forms.
[0054] 5 shows another example of the configuration of the feature converter (conversion model) according to the first embodiment. As shown in Fig. 5, the conversion model on which the conversion operation of the conversion unit 131 is based may be an identity conversion in which either the source or the target does nothing to the input feature.
[0055] Fig. 6 shows another example of the configuration of the feature converter (conversion model) according to the first embodiment. As shown in Fig. 6, the conversion model on which the conversion operation of the conversion unit 131 is based may have a common feature connected next to the converted feature. The conversion model may also be a combination of Figs. 5 and 6.
[0056] Second Embodiment Next, a second embodiment of the present disclosure will be described. PU learning used in the first embodiment requires knowledge of the prior probability of positive examples in target data. Predictive Adversarial Learning (PAN), a PU learning method proposed in Non-Patent Document 1, uses adversarial learning to eliminate the need for such prior knowledge and demonstrates excellent performance.
[0057] By using PAN as a PU learning method and adversarial learning as a method for approximating the distribution of source domain data and target domain data, more compatible and effective learning becomes possible. In this embodiment, a learning device that enables such learning will be described.
[0058] 7 is a block diagram showing an example of the functional configuration of a learning device 10 according to the second embodiment. In addition to the components of the learning device 10 according to the first embodiment, the learning device 10 further includes an identification unit 133 in the data processing unit 130, and a third loss calculation unit 143 in the loss calculation unit 140. The learning device 10 does not necessarily have to include the first loss calculation unit 141 and the second loss calculation unit 142.
[0059] In adversarial learning, the inference unit 132 learns, based on a predetermined inference model, to determine as a positive example from unlabeled data data that is similar to positive example data so that the classification unit 133 cannot correctly classify it. The classification unit 133 learns, based on a predetermined classification model, to mutually adversarially distinguish between data that the inference unit 132 has determined to be a positive example and true positive example data.
[0060] The third loss calculation unit 143 calculates the loss L when performing adversarial learning between the output of the conversion unit 131 based on a predetermined conversion model, the output of the inference unit 132 based on a predetermined inference model, and the output of the identification unit 170 based on a predetermined identification model.
[0061] In this embodiment, the conversion model is updated by learning to deceive the classifier 133, that is, to make it unable to distinguish between positive examples in the target data and source data.
[0062] In other words, the classifier used in PU learning and the classifier used to approximate the distribution of features after conversion of source data and target data are shared, which enables effective learning of inference models and discrimination models.
[0063] The acquisition unit 110 acquires data to be identified and a predetermined identification model, and stores them in the storage unit 160. The identification unit 133 reads the predetermined identification model from the storage unit 160. The predetermined identification model is a machine learning model that infers whether input data is a source or a target. For example, the identification model may be a logistic regression or a neural network, similar to the inference model.
[0064] The conversion unit 131 calculates post-conversion features based on the source data (first data set), the target data (second data set), and a predetermined conversion model. At this time, the post-conversion features of the source data and the target data are vectors in the same space.
[0065] The inference unit 132 generates an inference value based on the feature quantities of the converted source data and target data and a predetermined inference model. The identification unit 133 generates an identification value based on the feature quantities of the converted source data and target data and a predetermined identification model. The identification value indicates whether the data to be identified corresponds to the first data or the second data.
[0066] The third loss calculation unit 143 calculates the loss L when performing adversarial learning for the conversion unit 131, the inference unit 132, and the classification unit 133. The conversion unit for the source and the target is F s and F t If the inference unit (inference section) is C and the classifier (classification unit) is D, L is expressed as in Equation 6.
[0067] KL is the KL divergence and represents the distance between the two probability distributions received as arguments. P0 and P1 are vectors P0 = (1, 0) and P1 = (0, 1), respectively. That is, the first and second terms in Equation 7 represent the negative distance between the output of classifier D and the correct label indicating whether the input data is target data or source data.
[0068] Classifier D learns to increase the value of Equation 7, and therefore learns to decrease this distance. The third term in Equation 7 represents the distance between inferer C and classifier D for the target data. Inferer C learns to decrease this distance.
[0069] If classifier D mistakenly outputs a score that indicates a high probability of a positive example for target data, then inference device C learns to imitate that output, since the probability of that data being a positive example is high. In contrast, classifier D learns to deviate from the output of inference device C, since it wants to determine that data that inference device C inferred as a positive example is a negative example.
[0070] is obtained by swapping the first and second components of the vector output by the inference unit C, so the fourth term in Equation 7 is a term added to make the third term symmetrical, which contributes to further improving accuracy. As will be described later, when pseudo labels are assigned, the classification loss L for the pseudo labels is calculated in the same way as in the first embodiment. c (fourth loss) is multiplied by the parameter η and added to the loss expressed in Equation 7.
[0071] The model optimization unit 150 performs optimization of the following equation (9) based on the loss L, thereby updating each parameter of the inference device (inference model), the classifier (classification model), and the converter (conversion model) by, for example, gradient descent.
[0072] Adversarial learning is performed by the inference unit (inference unit 132) and the classifier (classification unit 133), and by the converter (conversion unit 131) and the classifier (classification unit 133), respectively.
[0073] FIG. 8 is a flowchart showing the flow of processing by the learning device 10 in the second embodiment.
[0074] In step S1101, the acquisition unit 110 acquires a first data set and a second data set.
[0075] In step S1102, the acquisition unit 110 acquires a predetermined transformation model, a predetermined inference model, and a predetermined identification model. The storage unit 160 may store the first data set, the second data set, the predetermined transformation model, the predetermined inference model, and the predetermined identification model.
[0076] The processes in steps S1103 and S1104 are similar to those in steps S103 and S104.
[0077] In step S1105, the identification unit 133 generates an identification value based on the converted first feature amount, the converted second feature amount, the inference value, and a predetermined identification model. The identification value indicates whether the data to be identified corresponds to the first data or the second data. In other words, the identification value indicates whether the data to be identified corresponds to the inference result by the inference unit 132, or whether it corresponds to one of the positive examples and negative examples included in the first dataset rather than the inference result.
[0078] In step S1106, the third loss calculation unit 143 calculates a loss value for adversarial learning (third loss) based on the inference value, the discrimination value, and the adversarial machine learning.
[0079] In step S1107, the model optimization unit 150 updates the parameters of the predetermined transformation model, the inference model, and the predetermined discriminative model based on the loss value for adversarial learning.
[0080] As described above, in the second embodiment of the present disclosure, the acquisition unit further acquires a predetermined discriminative model, and includes a conversion unit that converts, based on a predetermined transformation model, first features indicated by first data included in the first dataset into features in a common feature space and converts second features indicated by the second data into features in the common feature space; an inference unit that generates an inference value indicating whether the second data corresponds to one of a positive example or a negative example, based on the transformed first features, the transformed second features, and the inference model; a discrimination unit that generates a discrimination value indicating whether the second data corresponds to the first data or the second data, based on the transformed second features, the transformed first features, and the predetermined discriminative model; a third loss calculation unit that calculates a loss value for adversarial learning based on the inference value, the discrimination value, and adversarial machine learning; and an optimization unit that updates the predetermined transformation model, the inference model, and the predetermined discriminative model based on the loss value for adversarial learning.
[0081] Furthermore, the optimization unit updates the predetermined transformation model and the inference model so as to minimize the loss value for adversarial learning, and updates the predetermined discriminative model so as to maximize the loss value for adversarial learning.
[0082] In this way, training of the feature transformer and PU learning can be performed simultaneously. In addition, by using PAN as a PU learning method and adversarial learning as a method to align the distributions of source domain data and target domain data, more compatible and effective learning becomes possible.
[0083] Third Embodiment FIG. 9 is a block diagram showing an example of the functional configuration of a learning device 10 according to a third embodiment.
[0084] In this embodiment, the learning device 10 includes a pseudo label generation unit 120, a pseudo label inference unit 135, an inference unit learning unit 136, and a fourth loss calculation unit 144 in addition to the components of the first embodiment. A pseudo label is a label assigned to second data to pseudo-indicate whether the data is a positive example or a negative example, or to indicate the probability of the data being a positive example or a negative example. Details of the operation of assigning pseudo labels, the function of the pseudo labels, and the operation of the fourth loss calculation unit 144 will be described later.
[0085] 10 and 11 are flowcharts showing the flow of processing by the learning device 10 in the third embodiment.
[0086] In step S2101, the acquisition unit 110 acquires a first data set and a second data set. The setting of whether to enable the pseudo-label function may be configured in advance, or may be configured to be set semi-statically or dynamically in response to some kind of notification.
[0087] In step S2102, the acquisition unit 110 acquires a predetermined transformation model and a predetermined inference model. The storage unit 160 may store the first data set, the second data set, the pseudo-label configuration information, the predetermined transformation model, and the predetermined inference model.
[0088] In step S2103, the inference unit training unit 136 trains the pseudo-label inference unit 135, which generates an inference value indicating whether the second data corresponds to a positive example or a negative example, based on features of the first data set that correspond to both the first feature space and the second feature space and features of the second data set that correspond to both the first feature space and the second feature space. In this case, the inference unit may be trained using a standard PU learning method, with the first data as a positive example and the second data as unlabeled data. The inference value represents a "certainty level," which indicates the likelihood that the data is a positive example.
[0089] In step S2104, the pseudo-label inference unit 135 generates an inference value for the second data set based on the inference unit trained in step S2103.
[0090] In step S2105, if the inference value is greater than the first predetermined threshold, the process proceeds to step S2106, otherwise, the process proceeds to step S2107.
[0091] In step S2106, the pseudo label generation unit 120 assigns a pseudo label of a positive example to the second data. Note that in step S2105, if the difference is smaller than the second predetermined threshold, the pseudo label generation unit 120 may assign a pseudo label of a negative example to the second data. The pseudo label generation unit 120 does not assign a pseudo label to data that does not fit either of these categories. The predetermined thresholds may be preset arbitrarily or may be set appropriately. For example, the first predetermined threshold may be 0.9, and the second predetermined threshold may be 0.1. Alternatively, as a method of assigning pseudo labels, the inferred values may be directly assigned as labels to all or part of the second dataset.
[0092] In step S2107, the conversion unit 131 converts, based on a predetermined conversion model, each of the first feature indicated by the first data included in the first dataset and the second feature indicated by the second data included in the second dataset into a feature in a common feature space.
[0093] In step S2108, the inference unit 132 generates an inferred value based on a predetermined inference model, the converted first feature amount, and the converted second feature amount.
[0094] Moving on to FIG. 11 , in step S2109, the first loss calculation unit 141 calculates the first loss of the converted first feature quantity and the converted second feature quantity based on the converted first feature quantity, the converted second feature quantity, and the inferred value.
[0095] In step S2110, the second loss calculation unit 142 calculates the second loss due to the inference model based on the inferred value.
[0096] In step S2111, the fourth loss calculation unit calculates a fourth loss due to the pseudo label based on the inference value.
[0097] In step S2112, the model optimization unit updates the parameters of the specified inference model and the specified conversion model based on the first loss, the second loss, and the fourth loss so as to minimize these loss values, and stores them in the memory unit 160.
[0098] In step S2113, the pseudo label generation unit 120 determines whether a series of estimation and model optimization processes based on pseudo labeling is the first time. If the determination result is true, the process returns to step S2104; otherwise, the process proceeds to step S2114.
[0099] In step S2114, the pseudo label generation unit 120 determines whether the accuracy of the estimation process has improved since the previous series of processes. If the determination result is true, the process returns to step S2104; if not, the process ends.
[0100] As described above, in the third embodiment of the present disclosure, the learning device 10 further includes an inference unit, a pseudo-label inference unit, an inference unit learning unit, a pseudo-label generation unit, a conversion unit, a first loss calculation unit, a second loss calculation unit, a fourth loss calculation unit, and a model optimization unit, wherein the pseudo-label inference unit generates an inference value indicating whether the second data corresponds to a positive example or a negative example based on the pseudo-label generation inference model and common features common to both the first feature space and the second feature space of the first dataset and the second dataset, the pseudo-label generation unit assigns to the second data a pseudo label that pseudo-indicates that the second data corresponds to a positive example when the inference value is higher than a predetermined first threshold and a negative example when the inference value is lower than a predetermined second threshold, or assigns the inference value directly as a pseudo label, and the conversion unit converts first features indicated by first data included in the first dataset into features in the common feature space based on a predetermined conversion model. the second loss calculation unit calculates a second loss by the inference model based on the inference value; the fourth loss calculation unit calculates a fourth loss by the pseudo label based on the inference value; the model optimization unit updates the inference model and the predetermined conversion model based on the first loss, the second loss, and the fourth loss; the common feature space has a correspondence relationship between the first feature space and the second feature space, and the first feature is included in the first feature space but not in the second feature space.
[0101] This makes it possible to improve the learning accuracy of the conversion model and the estimation model.
[0102] 12 is a block diagram showing an example of the functional configuration of a learning system S in a fourth embodiment. The learning system S includes a first learning device 11, a second learning device 12, and a server device 30. The first learning device 11, the second learning device 12, and the server device 30 are connected to each other via a communication network NW.
[0103] The first learning device 11 belongs to the source domain. The second learning device 12 belongs to the target domain. For example, it is assumed that the first learning device 11 manages customer dataset A of company A. Company A's customer dataset A includes company A's trade secrets. When considering labels based on binary classification of whether or not the customer is interested in company A's products, it is assumed that the customer dataset of company A has already been assigned a positive example label because the customer is a customer.
[0104] An example would be a case where company A partners with company B to find potential customers for company A's products from company B's customer database based on its trade secrets, but company B's customer database does not contain labels based on binary classification.
[0105] The communication network NW may be a wired network based on an optical fiber network or metal wires, or a wireless communication network. The wired network may conform to standards such as Ethernet (registered trademark). The wireless communication network may be realized by a wireless LAN conforming to the IEEE 802.11 standard, or may be realized by a third-generation mobile communication network, fourth-generation mobile communication network, fifth-generation mobile communication network, etc. conforming to the 3GPP (registered trademark) standard.
[0106] FIG. 13 is a block diagram showing an example of the functional configuration of a first learning device 11 in the fourth embodiment.
[0107] FIG. 14 is a block diagram showing an example of the functional configuration of a second learning device 11 in the fourth embodiment.
[0108] The configurations of the first learning device 11 and the second learning device 12 are generally similar to the configuration of the learning device 10 according to the first embodiment, but the first learning device 11 further includes a first communication unit 180-1, and the second learning device 12 further includes a second communication unit 180-2.
[0109] 15 is a block diagram showing an example of the functional configuration of the server device 30 according to the fourth embodiment. The server device 30 includes an acquisition unit 310, a processing unit 320, a storage unit 330, and a communication unit 340.
[0110] FIG. 16 is a flowchart showing the flow of processing by the first learning device 11 in the fourth embodiment of the present disclosure.
[0111] In step S3101, the first acquisition unit 110-1 acquires a first dataset. The first dataset belongs to a first domain including a first feature space, and is labeled as either a positive example or a negative example in binary classification.
[0112] In step S3102, the first acquisition unit 110-1 acquires a first predetermined transformation model and a first predetermined inference model. The first storage unit 160-1 may store the first data set, the first predetermined transformation model, and the first predetermined inference model.
[0113] In step S3103, the first conversion unit 131-1 converts the first feature into a feature in a common feature space based on a first predetermined conversion model. It is assumed that the first learning device 11 and the second learning device 12 share in advance information about the specific features in the common feature space via the first communication unit 180-1 and the second communication unit 180-2, respectively. For example, they share information about gender and age, but do not share raw data such as who is how old they are and whether they are male or female.
[0114] In step S3104, when MMD is used to calculate the domain loss, first communication unit 180-1 transmits the converted first feature to server device 30. At this time, it is preferable that the converted feature to be transmitted to server device 30 is encrypted based on a predetermined algorithm. Note that when the squared norm of the difference of covariance matrices is used to calculate the domain loss, first communication unit 180-1 transmits the covariance matrix calculated based on the converted first feature to server device 30. Encryption is not required at this time.
[0115] In step S3105, when MMD is used to calculate the domain loss, the first communication unit 180-1 receives the converted second feature from the server device 30. The converted second feature is generated by the second learning device 12. Note that when the squared norm of the difference of covariance matrices is used to calculate the domain loss, the first communication unit 180-1 receives from the server device 30 the covariance matrix calculated based on the converted second feature.
[0116] In step S3106, the first inference unit 132-1 generates a first inference value indicating whether the first data corresponding to the converted first feature corresponds to a positive example or a negative example based on the converted first feature and a first predetermined inference model.
[0117] In step S3107, the first loss calculation unit 141-1 calculates a first loss using a first predetermined transformation model based on the transformed first feature and the transformed second feature, or based on a covariance matrix calculated based on the transformed first feature and a covariance matrix calculated based on the transformed second feature.
[0118] In step S3108, the second loss calculation unit 142-1 calculates a second loss according to a first predetermined inference model based on the first inference value.
[0119] In step S3109, the first model optimization unit 150-1 updates the parameters of the first predetermined transformation model and the first predetermined inference model based on the first loss and the second loss.
[0120] In step S3110, the first communication unit 180-1 transmits the updated first predetermined inference model to the server device 30.
[0121] In step S3111, the first communication unit 180-1 receives the third updated inference model from the server device 30.
[0122] In step S3112, the first model optimization unit 150-1 updates the first predetermined inference model based on the third updated inference model and the updated first predetermined inference model read from the first storage unit 160-1, and stores them in the first storage unit 160-1. The process then ends.
[0123] FIG. 17 is a flowchart showing the flow of processing by the second learning device 12 in the fourth embodiment of the present disclosure.
[0124] In step S3201, the second acquisition unit 110-2 acquires a second data set. The second data set belongs to a second domain including a second feature space and includes data that has not been labeled by binary classification.
[0125] In step S3202, the second acquisition unit 110-2 acquires a second predetermined transformation model and a second predetermined inference model. The second storage unit 160-2 may store the second data set, the second predetermined transformation model, and the second predetermined inference model.
[0126] In step S3203, the second conversion unit 131-2 converts the second feature amount into a feature amount in a common feature amount space based on a second predetermined conversion model.
[0127] In step S3204, when MMD is used to calculate the domain loss, the second communication unit 180-2 transmits the converted second feature to the server device 30. Note that when the squared norm of the difference of covariance matrices is used to calculate the domain loss, the second communication unit 180-2 transmits the covariance matrix calculated based on the converted second feature to the server device 30.
[0128] In step S3205, when MMD is used to calculate the domain loss, the second communication unit 180-2 receives the converted first feature from the server device 30. The converted first feature is generated by the first learning device 11. When the squared norm of the difference of covariance matrices is used to calculate the domain loss, the second communication unit 180-2 transmits the covariance matrix calculated based on the converted first feature to the server device 30.
[0129] In step S3206, the second inference unit 132-2 generates a second inference value indicating whether the second data corresponding to the converted second feature corresponds to a positive example or a negative example based on the converted second feature and a second predetermined inference model.
[0130] In step S3207, the third loss calculation unit 141-3 calculates a third loss using a second predetermined transformation model based on the transformed first feature quantity and the transformed second feature quantity, or based on a covariance matrix calculated based on the transformed first feature quantity and a covariance matrix calculated based on the transformed second feature quantity.
[0131] In step S3208, the fourth loss calculation unit 141-4 calculates a fourth loss according to a second predetermined inference model based on the second inference value.
[0132] In step S3209, the second model optimization unit 150-2 updates the parameters of the second predetermined transformation model and the second predetermined inference model based on the third loss and the fourth loss.
[0133] In step S3210, the second communication unit 180-2 transmits the updated second predetermined inference model to the server device 30.
[0134] In step S3211, the second communication unit 180-2 receives the third updated inference model from the server device 30.
[0135] In step S3212, the second model optimization unit 150-2 updates the second predetermined inference model based on the third updated inference model and the updated second predetermined inference model read from the second storage unit 160-2, and stores them in the second storage unit 160-2. The process then ends.
[0136] FIG. 18 is a flowchart showing the flow of processing by the learning device 30 according to the fourth embodiment of the present disclosure.
[0137] In step S3301, when MMD is used to calculate the domain loss, the communication unit 340 receives the transformed first feature from the first learning device 11 and receives the transformed second feature from the second learning device 12. Note that when the squared norm of the difference of covariance matrices is used to calculate the domain loss, the communication unit 340 receives the covariance matrix calculated based on the transformed first feature from the learning device 11 and receives the covariance matrix calculated based on the transformed second feature from the learning device 12.
[0138] In step S3302, when MMD is used to calculate the domain loss, the communication unit 340 transmits the transformed second feature to the first learning device 11 and transmits the transformed first feature to the second learning device 12. Note that when the squared norm of the difference of covariance matrices is used to calculate the domain loss, the communication unit 340 transmits the covariance matrix calculated based on the transformed second feature to the learning device 11 and transmits the covariance matrix calculated based on the transformed first feature to the learning device 12.
[0139] In step S3303, the communication unit 340 receives the updated first predetermined inference model from the first learning device 11. The communication unit 340 receives the updated second predetermined inference model from the second learning device 12.
[0140] The storage unit 330 may store the updated first predetermined inference model and the updated second predetermined inference model. Note that the acquisition unit 310 may take over the function of the communication unit 340.
[0141] In step S3304, the processing unit 320 generates a third updated inference model based on the updated first predetermined inference model and the updated second predetermined inference model.
[0142] The third updated inference model may be the average value of the matrix parameters included in the updated first predetermined inference model and the updated second predetermined inference model, respectively.
[0143] In step S3305, the communication unit 340 transmits the third updated inference model to the first learning device 11 and the second learning device 12.
[0144] In this way, the first learning device 11 and the second learning device 12 perform the learning process independently of each other, so that the learning process can be performed efficiently even when the data exists independently of each other and the details of the data are not shared.
[0145] 19 is a block diagram showing an example of the functional configuration of a first learning device 11 according to a fifth embodiment of this disclosure. In the fifth embodiment, the properties of the first data set and the second data set are the same as those in the fourth embodiment.
[0146] The first learning device 11 includes a first acquisition unit 110-1, a first data processing unit 130-1, a first loss calculation unit 140-1, a first model optimization unit 150-1, a first storage unit 160-1, and a first communication unit 180-1. The first data processing unit 130-1 further includes a first conversion unit 131-1 and a first identification unit 133-1. The first loss calculation unit 140-1 further includes a fifth loss calculation unit 143-1.
[0147] FIG. 20 is a block diagram showing an example of the functional configuration of the second learning device 12 in the fifth embodiment of this disclosure.
[0148] The second learning device 12 includes a second acquisition unit 110-2, a second data processing unit 130-2, a second loss calculation unit 140-2, a second model optimization unit 150-2, a second storage unit 160-2, and a second communication unit 180-2. The second data processing unit 130-2 includes a first conversion unit 131-1, a first inference unit 132-1, and a first identification unit 133-1. The second loss calculation unit 140-2 further includes a sixth loss calculation unit 143-2. The second loss calculation unit 140-2 further includes a sixth loss calculation unit 143-2.
[0149] An example of the functional configuration of the server device 30 is the same as that of the fourth embodiment.
[0150] FIG. 21 is a flowchart showing the flow of processing by the first learning device 11 in the fifth embodiment.
[0151] In step S4101, the first acquisition unit 110-1 acquires a first data set.
[0152] In step S4102, the first acquisition unit 110-1 acquires a first predetermined transformation model and a first predetermined discrimination model. The first storage unit 160-1 may store the first data set, the predetermined transformation model, and the predetermined discrimination model.
[0153] In step S4103, the first conversion unit 131-1 converts the first feature into a feature in a common feature space based on a first predetermined conversion model. The properties of the common feature space are the same as those in the fourth embodiment.
[0154] In step S4104, the first discrimination unit 133-1 generates a discrimination value based on the converted first feature amount and a first predetermined discrimination model. The discrimination value is similar to that in the second embodiment.
[0155] In step S4105, the fifth loss calculation unit 143-1 calculates a loss value for adversarial learning (a fifth loss) based on the classification value and the adversarial machine learning. Note that the calculation of the fifth loss by the fifth loss calculation unit 143-1 may be based on calculating the first term of Equation 6.
[0156] In step S4106, the first model optimization unit 150-1 updates the parameters of the first predetermined transformation model and the first predetermined discriminative model based on the fifth loss based on adversarial learning.
[0157] In step S4107, the first communication unit 180-1 transmits the updated first identification model to the server device 30.
[0158] In step S 4108 , first communication unit 180 - 1 receives the third updated identification model from server device 30 .
[0159] In step S4109, the first model optimization unit 150-1 updates the first predetermined discriminant model based on the third updated discriminant model. The first storage unit 160-1 may store the updated first predetermined transformation model and each parameter of the first predetermined discriminant model. The process then ends.
[0160] FIG. 22 is a flowchart showing the flow of processing by the second learning device 12 in the fifth embodiment.
[0161] In step S4201, the second acquisition unit 110-2 acquires a second data set.
[0162] In step S4202, the second acquisition unit 110-2 acquires a second predetermined transformation model, a second predetermined inference model, and a second predetermined identification model. The second storage unit 160-2 may store the second data set, the predetermined transformation model, the predetermined inference model, and the predetermined identification model.
[0163] In step S4203, the second conversion unit 131-2 converts the second feature amount into a feature amount in a common feature amount space based on a second predetermined conversion model.
[0164] In step S4204, the second inference unit 132-2 generates an inferred value based on a second predetermined inference model and the converted second feature amount.
[0165] In step S4205, the second identification unit 133-2 generates an identification value based on the converted second feature, the inference value, and a second predetermined identification model. The content indicated by the identification value in the calculation process for the second data set is the same as in the second embodiment.
[0166] In step S4206, the sixth loss calculation unit 143-2 calculates a classifier loss value (sixth loss) based on the inference value, the classification value, and machine learning. Note that the calculation of the sixth loss by the sixth loss calculation unit 143-2 may be based on calculating the second-fourth term of Equation 6.
[0167] In step S4207, the second model optimization unit 150-2 updates the parameters of the second predetermined transformation model, the second predetermined inference model, and the second predetermined discrimination model based on the sixth loss.
[0168] In step S4208, the second communication unit 180-2 transmits the updated second identification model to the server device 30.
[0169] In step S4209, second communication unit 180-2 receives the third updated identification model from server device 30.
[0170] In step S4210, the second model optimization unit 150-2 updates the second predetermined discriminant model based on the third updated discriminant model. The second storage unit 160-2 may store the updated parameters of the second predetermined transformation model, the second predetermined inference model, and the second predetermined discriminant model. The process then ends.
[0171] FIG. 23 is a flowchart showing the flow of processing by the learning device 30 in the fifth embodiment.
[0172] In step S4301, the communication unit 340 receives the updated first predetermined discrimination model from the first learning device 11 and receives the updated second predetermined discrimination model from the second learning device 12. The storage unit 330 may store the updated first predetermined discrimination model and the updated second predetermined discrimination model. Note that the acquisition unit 310 may perform the function of the communication unit 340.
[0173] In step S4302, the processing unit 320 generates a third updated discriminative model based on the updated first and second predetermined discriminative models. The third updated discriminative model may be an average value of matrix parameters included in the updated first and second predetermined discriminative models.
[0174] In step S4303, the communication unit 340 transmits the third updated discriminative model to the first learning device 11 and the second learning device 12. The process then ends.
[0175] In this way, the first learning device 11 and the second learning device 12 perform the learning process independently of each other, so that the learning process can be performed efficiently even when the data exists independently of each other and the details of the data are not shared.
[0176] FIG. 24 is a diagram showing the minimum configuration of the learning device 10 in this disclosure.
[0177] The learning device 10 includes an acquisition unit 110 and a model generation unit 134. The function of the acquisition unit 110 is as described above. The model generation unit 134 generates an inference model for inferring whether second data included in the second data set is a positive example or a negative example in binary classification, based on the first data set and the second data set. This inference model is applied to the inference unit 132, the first inference unit 132-1, and the second inference unit in any of the first to fifth embodiments.
[0178] FIG. 25 is a flowchart showing the flow of processing in an embodiment of the minimum configuration of the learning device 10 in this disclosure.
[0179] In step S4101, the acquisition unit 110 acquires a first dataset and a second dataset. The first dataset belongs to a first domain including a first feature space and is labeled as either a positive example or a negative example in a binary classification. The second dataset belongs to a second domain including a second feature space and is not labeled by a binary classification.
[0180] In step S4102, the model generation unit 134 generates an inference model for inferring whether the second data set included in the second data set is a positive example or a negative example in binary classification, based on the first data set and the second data set. Then, the process ends.
[0181] This makes it possible to accurately infer an inference model for inferring whether data that is not labeled by binary classification is a positive example or a negative example, based on PU learning and data that is labeled only as a positive example or a negative example in binary classification.
[0182] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.
[0183] (Supplementary Note 1) A learning device comprising: an acquisition unit that acquires a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in a binary classification, and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification; and a model generation unit that generates, based on the first dataset and the second dataset, an inference model for inferring whether second data included in the second dataset is a positive example or a negative example in the binary classification.
[0184] (Supplementary Note 2) The system further comprises: a conversion unit that converts, based on a predetermined conversion model, a first feature indicated by first data included in the first dataset into a feature in a common feature space, and converts a second feature indicated by the second data into a feature in the common feature space; an inference unit that generates an inference value indicating whether the second data corresponds to a positive example or a negative example, based on the converted first feature, the converted second feature, and the inference model; a first loss calculation unit that calculates a first loss by the conversion model based on the converted first feature and the converted second feature; a second loss calculation unit that calculates a second loss by the inference model based on the inference value; and a model optimization unit that updates the inference model and the predetermined conversion model based on the first loss and the second loss, wherein the common feature space has a correspondence relationship between the first feature space and the second feature space, and the first feature is included in the first feature space but not in the second feature space, The learning device according to claim 1, wherein the second feature is included in the second feature space and is not included in the first feature space.
[0185] (Supplementary Note 3) The learning device according to Supplementary Note 1 or 2, wherein the inference value indicates whether the second data corresponds to a positive example or a negative example, and the converted second feature of the second data corresponding to one of the positive example and the negative example is used as an input when the first loss calculation unit calculates the first loss.
[0186] (Supplementary Note 4) The learning device according to any one of Supplementary Notes 1 to 3, wherein the acquisition unit further acquires a predetermined discriminative model; and a conversion unit that converts, based on a predetermined transformation model, first features indicated by first data included in the first dataset into features in a common feature space, and converts second features indicated by the second data into features in the common feature space; an inference unit that generates an inference value indicating whether the second data corresponds to a positive example or a negative example, based on the transformed first features, the transformed second features, and the inference model; a discrimination unit that generates a discrimination value indicating whether input data corresponds to the first data or the second data, based on the transformed second features, the transformed first features, and the predetermined discriminative model; a third loss calculation unit that calculates a loss value for adversarial learning based on the inference value, the discrimination value, and adversarial machine learning; and an optimization unit that updates the predetermined transformation model, the inference model, and the predetermined discriminative model based on the loss value for adversarial learning.
[0187] (Supplementary Note 5) The learning device according to any one of Supplementary Notes 1 to 4, wherein the third loss calculation unit calculates the loss value for adversarial learning based on: a first KL divergence function having the transformed first feature as an input and a discrimination value as an argument; a second KL divergence function having the transformed second feature as an input and a discrimination value as an argument; a third KL divergence function having the transformed second feature as an input and an estimated value and a discrimination value as arguments; and a fourth KL divergence function having the transformed second feature as an input and an estimated value and a discrimination value as arguments.
[0188] (Supplementary Note 6) The learning device according to any one of Supplementary Notes 1 to 5, wherein the optimization unit updates the predetermined transformation model and the inference model so as to minimize the loss value for adversarial learning, and updates the predetermined discriminative model so as to maximize the loss value for adversarial learning.
[0189] (Supplementary Note 7) The data processing system further includes an inference unit, a pseudo-label inference unit, an inference unit learning unit, a pseudo-label generation unit, a conversion unit, a first loss calculation unit, a second loss calculation unit, a fourth loss calculation unit, and a model optimization unit, wherein the pseudo-label generation inference unit learning unit generates an inference unit that generates an inference value indicating whether data of a second dataset corresponds to a positive example or a negative example, based on the inference model, common features of the first dataset that are common to both the first feature space and the second feature space, and common features of the second dataset that are common to both the first feature space and the second feature space, the pseudo-label generation inference unit uses the inference unit to generate an inference value indicating whether data of the second dataset corresponds to a positive example or a negative example, the pseudo-label generation unit assigns pseudo labels to the data of the second dataset based on the inference value, and the conversion unit performs the following steps based on a predetermined conversion model: the first feature indicated by first data included in the first dataset is transformed into a feature in a common feature space; the second feature indicated by the common data is transformed into a feature in the common feature space; the inference unit generates the inference value based on the transformed first feature, the transformed second feature, and the inference model; the first loss calculation unit calculates a first loss between the transformed first feature and the transformed second feature based on the transformed first feature and the transformed second feature; the second loss calculation unit calculates a second loss by the inference model based on the inference value; the fourth loss calculation unit calculates a fourth loss by the pseudo label based on the inference value; the model optimization unit updates the inference model and the predetermined conversion model based on the first loss, the second loss, and the fourth loss; the common feature space has a correspondence relationship between the first feature space and the second feature space; The learning device according to any one of appendixes 1 to 6, wherein the first feature is included in the first feature space and is not included in the second feature space.
[0190] (Supplementary Note 8) A learning system including a first learning device, a second learning device, and a server device, wherein the first learning device includes a first acquisition unit, a first model generation unit, a first conversion unit, a first inference unit, a first loss calculation unit, a second loss calculation unit, a first model optimization unit, and a first communication unit, wherein the first acquisition unit acquires a first predetermined inference model, a first predetermined conversion model, and a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in binary classification, wherein the first conversion unit converts first features indicated by first data included in the first dataset into features in a common feature space based on the first predetermined conversion model, wherein the first communication unit transmits the converted first features to the server device, and wherein the first communication unit receives the converted second features from the server device, the first inference unit generates a first inference value indicating whether first data corresponding to the converted first feature corresponds to a positive example or a negative example, based on the converted first feature and the first predetermined inference model; the first loss calculation unit calculates a first loss using the first predetermined conversion model, based on the converted first feature, the converted second feature, and the first inference value; the second loss calculation unit calculates a second loss using the first predetermined inference model, based on the first inference value; the first model optimization unit updates the first predetermined conversion model and the first predetermined inference model, based on the first loss and the second loss; the first communication unit transmits the updated first predetermined inference model to the server device; and the second learning device comprises a second acquisition unit, a second model generation unit, a second conversion unit, a second inference unit, a second loss calculation unit, a second model optimization unit, and a second communication unit, the second acquisition unit acquires a second predetermined inference model and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification; the second conversion unit converts second features indicated by second data included in the second dataset into features in the common feature space based on the second predetermined conversion model; and the second communication unit transmits the converted second features to the server device.the second communication unit receives the converted first feature from the server device; the second inference unit generates a second inference value indicating whether second data corresponding to the converted second feature corresponds to a positive example or a negative example, based on the converted second feature and the second predetermined inference model; the third loss calculation unit calculates a third loss using the second predetermined conversion model based on the converted first feature, the converted second feature, and the second inference value; the fourth loss calculation unit calculates a fourth loss using the second predetermined inference model based on the second inference value; the second model optimization unit updates the second predetermined conversion model and the second predetermined inference model based on the third loss and the fourth loss; and the second communication unit transmits the updated second predetermined inference model to the server device; the server device comprises a third communication unit and a processing unit, the third communication unit receives the converted first feature amount from the first learning device and receives the converted second feature amount from the second learning device; the third communication unit transmits the converted second feature amount to the first learning device and transmits the converted first feature amount to the second learning device; the third communication unit acquires the updated first predetermined inference model and the updated second predetermined inference model; the processing unit generates a third updated inference model based on the updated first predetermined inference model and the updated second predetermined inference model; the third communication unit transmits the third updated inference model to the first learning device and the second learning device; the first communication unit and the second communication unit acquire the third updated inference model; the first model optimization unit updates the first predetermined inference model based on the first predetermined inference model and the third updated inference model; and the second model optimization unit updates the second predetermined inference model based on the second predetermined inference model and the third updated inference model. Learning system.
[0191] (Supplementary Note 9) A learning system including a first learning device, a second learning device, and a server device, wherein the first learning device includes a first acquisition unit, a first conversion unit, a first identification unit, a fifth loss calculation unit, a first model optimization unit, and a first communication unit; the first acquisition unit acquires a first predetermined transformation model, a first predetermined identification model, and a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in binary classification; the first conversion unit converts first features indicated by first data included in the first dataset into features in a common feature space based on the first predetermined transformation model; the first identification unit generates a first identification value based on the converted first features and the first predetermined identification model; and the fifth loss calculation unit calculates a fifth loss using the first predetermined transformation model based on the identification value and adversarial learning. the first model optimization unit updates the first predetermined transformation model and the first predetermined discrimination model based on the fifth loss; the first communication unit transmits the updated first predetermined discrimination model to the server device; the second learning device includes a second acquisition unit, a second conversion unit, a second inference unit, a second discrimination unit, a sixth loss calculation unit, a second model optimization unit, and a second communication unit; the second acquisition unit acquires a second predetermined transformation model, a second predetermined inference model, a second predetermined discrimination model, and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification; the second conversion unit converts second features indicated by second data included in the second dataset into features in the common feature space based on the second predetermined transformation model; and the second inference unit generates an inferred value based on the converted second features and the second predetermined inference model. the sixth loss calculation unit calculates a sixth loss by the second predetermined transformation model based on the transformed second feature and the second inference value; the second model optimization unit updates the second predetermined transformation model, the second predetermined inference model, and the second predetermined discrimination model based on the sixth loss;the second communication unit transmits the updated second predetermined discriminative model to the server device; the server device includes a third communication unit and a processing unit; the third communication unit acquires the updated first predetermined discriminative model and the updated second predetermined discriminative model; the processing unit generates a third updated discriminative model based on the updated first predetermined discriminative model and the updated second predetermined discriminative model; the third communication unit transmits the third updated discriminative model to the first learning device and the second learning device; the first communication unit and the second communication unit acquire the third updated discriminative model; the first model optimization unit updates the first predetermined discriminative model based on the first predetermined discriminative model and the third updated discriminative model; and the second model optimization unit updates the second predetermined discriminative model based on the second predetermined discriminative model and the third updated discriminative model.
[0192] (Supplementary Note 10) A method executed by a computer, comprising the steps of: acquiring a first dataset belonging to a first domain including a first feature space and labeled as only one of positive examples and negative examples in a binary classification; and a second dataset belonging to a second domain including a second feature space and not labeled by a binary classification; and generating an inference model based on the first dataset and the second dataset to infer whether second data included in the second dataset is one of a positive example and a negative example in a binary classification.
[0193] (Supplementary Note 11) A program that causes a computer to execute the steps of: acquiring a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in a binary classification; and acquiring a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification; and generating an inference model based on the first dataset and the second dataset to infer whether second data included in the second dataset is one of the positive examples and negative examples in the binary classification.
[0194] (Supplementary Note 12) A storage medium storing a program that causes a computer to execute the steps of: acquiring a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in binary classification; and acquiring a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification; and generating an inference model based on the first dataset and the second dataset, for inferring whether second data included in the second dataset is one of the positive examples and negative examples in binary classification.
[0195] 21 is an explanatory diagram illustrating the hardware configuration of a learning device 10 according to this embodiment. The learning device 10 includes an input / output module I, a memory module M, and a control module P. The input / output module I is implemented by including some or all of the following: a communication module H11, a connection module H12, a pointing device H21, a keyboard H22, a display H23, a button H3, a microphone H41, a speaker H42, a camera H51, or a sensor H52.
[0196] The storage module M is realized by including a drive H7. The storage module M may further be configured to include a part or all of a memory H8. The control module P is realized by including a memory H8 and a processor H9. These hardware components are connected to each other so as to be able to communicate with each other via a bus, and are supplied with power from a power supply H6.
[0197] The connection module H12 is a digital input / output port such as a USB (Universal Serial Bus). In the case of a portable device, the pointing device H21, keyboard H22, and display H23 are touch panels. The sensor H52 is an acceleration sensor, a gyro sensor, a GPS receiving module, a proximity sensor, etc. The power source H6 is a power supply unit that supplies the electricity necessary to operate each device. In the case of a portable device, the power source H6 is a battery.
[0198] The drive H7 is an auxiliary storage medium such as a hard disk drive or a solid state drive. The drive H7 may be a nonvolatile memory such as an EEPROM or a flash memory, or a magneto-optical disk drive or a flexible disk drive. Furthermore, the drive H7 is not limited to being built into each device, but may also be an external storage device connected to a connector of the IF module H12.
[0199] The memory H8 is a main storage medium such as a random access memory. The memory H8 may be a cache memory. The memory H8 stores instructions when the instructions are executed by one or more processors H9. The processor H9 is a CPU (Central Processing Unit). The processor H9 may be an MPU (Microprocessing Unit) or a GPU (Graphics Processing Unit). The processor H9 reads programs and various data from the drive H7 via the memory H8 and performs calculations to execute the instructions stored in the one or more memories H8.
[0200] The input / output module I realizes the acquisition unit 110 and the acquisition unit 310. The storage module M realizes the storage unit 160, the first storage unit 160-1, the second storage unit 160-2, and the storage unit 330. The communication module H11 realizes the communication unit 180, the first communication unit 180-1, the second communication unit 180-2, and the communication unit 340.
[0201] The control module P includes a data processing unit 130, a first data processing unit 130-1, a second data processing unit 130-2, a conversion unit 131, a first conversion unit 131-1, a second conversion unit 131-2, an inference unit 132, a first inference unit 132-1, a second inference unit 132-2, a classification unit 133, a model generation unit 134, a pseudo-label inference unit 135, an inference unit learning unit 136, a loss calculation unit 140, a first loss calculation unit 141, a first loss calculation unit 142, a first loss calculation unit 143, a second loss calculation unit 144, a first loss calculation unit 145, a second loss calculation unit 146, a first loss calculation unit 147, a second loss calculation unit 148, a first loss calculation unit 149, a second loss calculation unit 150, a second loss calculation unit 151, a third loss calculation unit 152, a fourth loss calculation unit 153, a fifth loss calculation unit 154, a fifth loss calculation unit 155, a sixth loss calculation unit 156, a fifth loss calculation unit 157, a sixth loss calculation unit 158, a sixth loss calculation unit 159, a sixth loss calculation unit 160, a sixth loss calculation unit 161, a sixth loss calculation unit 162, a sixth loss calculation unit 163, a sixth loss calculation unit 164, a sixth loss calculation unit 165, a seventh loss calculation unit 166, a seventh loss calculation unit 167, a eighth loss calculation unit 168, a eighth loss calculation unit 169 ... The system includes a first loss calculation unit 141-1, a second loss calculation unit 142, a second loss calculation unit 142-1, a second loss calculation unit 142-2, a third loss calculation unit 143, a third loss calculation unit 141-2, a fifth loss calculation unit 143-1, a sixth loss calculation unit 143-2, a fourth loss calculation unit 144, a fourth loss calculation unit 142-2, a model optimization unit 150, a first model optimization unit 150-1, a second model optimization unit 150-2, and an identification unit 170. Note that in this specification and the like, the descriptions of the learning device 10, the first learning device 11, the second learning device 12, and the server device 30 may be replaced with descriptions of a control unit P10, a control unit P11, a control unit P12, and a control unit P30, respectively, and descriptions of these devices may be replaced with descriptions of a control module P.
[0202] While the embodiments and modifications have been described above in detail with reference to the drawings as one aspect of the present invention, the specific configuration is not limited to the embodiments and modifications, and design changes within the scope of the present invention are also included. Furthermore, various modifications of one aspect of the present invention are possible within the scope of the claims, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. Furthermore, configurations in which elements described in the above embodiments and modifications are substituted with elements that achieve the same effect are also included.
[0203] For example, one aspect of the present invention may be realized by combining some or all of the above-described embodiments.
[0204] As used herein, including in the claims, "or" in lists of terms (e.g., lists of terms ending with phrases such as "at least one of" or "one or more of") indicates an inclusive list, such as, for example, a list of at least one of A, B, or C means A or B or C or AB or AC or BC or ABC (i.e., A and B and C). Also, the phrase "based on" as used herein is not to be construed as a reference to a closed set of conditions. For example, an example step described as "based on condition A" may be based on both condition A and condition B without departing from the scope of the present disclosure. In other words, the phrase "based on" as used herein is to be construed the same as the phrase "based at least in part on."
[0205] 10 Learning device 11 First learning device 12 Second learning device 30 Server device 110 Acquisition unit 110-1 First acquisition unit 110-2 Second acquisition unit 120 Pseudo label generation unit 130 Data processing unit 130-1 First data processing unit 133-1 First identification unit 131 Conversion unit 131-1 First conversion unit 132 Inference unit 132-2 Second inference unit 133 Identification unit 134 Model generation unit 140 Loss calculation unit 140-1 First loss calculation unit 141 First loss calculation unit 142 Second loss calculation unit 143 Third loss calculation unit 144 Fourth loss calculation unit 143-1 Fifth loss calculation unit 143-2 Sixth loss calculation unit 150 Model optimization unit 150-1 First model optimization unit 150-2 Second model optimization unit 160 Storage unit 160-1 First storage unit 160-2 Second storage unit 170 Identification unit 180 Communication unit 180-1 First communication unit 180-2 Second communication unit 310 Acquisition unit 320 Processing unit 330 Storage unit 340 Communication unit S Learning system
Claims
1. an acquisition unit that acquires a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in binary classification, and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification; a model generation unit that generates an inference model for inferring whether second data included in the second dataset is a positive example or a negative example in binary classification, based on the first dataset and the second dataset; A learning device comprising:
2. Based on a given transformation model, converting a first feature indicated by first data included in the first dataset into a feature in a common feature space; converting the second feature indicated by the second data into a feature in the common feature space; A conversion unit; an inference unit that generates an inference value indicating whether the second data corresponds to a positive example or a negative example based on the transformed first feature, the transformed second feature, and the inference model; a first loss calculation unit that calculates a first loss due to the conversion model based on the converted first feature amount and the converted second feature amount; a second loss calculation unit that calculates a second loss by the inference model based on the inference value; a model optimization unit that updates the inference model and the predetermined transformation model based on the first loss and the second loss, the common feature space has a correspondence relationship between the first feature space and the second feature space, the first feature is included in the first feature space and not included in the second feature space; The second feature is included in the second feature space and is not included in the first feature space. The learning device according to claim 1.
3. the inference value indicates whether the second data corresponds to a positive example or a negative example; The converted second feature amount of the second data corresponding to one of the positive example and the negative example is used as an input when the first loss calculation unit calculates the first loss. The learning device according to claim 2.
4. the acquisition unit further acquires a predetermined discrimination model; Based on a given transformation model, converting a first feature indicated by first data included in the first dataset into a feature in a common feature space; converting the second feature indicated by the second data into a feature in the common feature space; A conversion unit; an inference unit that generates an inference value indicating whether the second data corresponds to a positive example or a negative example based on the transformed first feature, the transformed second feature, and the inference model; a discrimination unit that generates a discrimination value indicating whether input data corresponds to the first data or the second data, based on the converted second feature amount, the converted first feature amount, and the predetermined discrimination model; a third loss calculation unit that calculates a loss value for adversarial learning based on the inference value, the discrimination value, and adversarial machine learning; an optimization unit that updates the predetermined transformation model, the inference model, and the predetermined discriminative model based on the loss value for adversarial learning. The learning device according to claim 1.
5. The third loss calculation unit a first KL divergence function that takes as an argument a classification value that receives the transformed first feature as an input; a second KL divergence function that takes as an argument a classification value that receives the converted second feature as an input; a third KL divergence function having an estimated value and a classification value as arguments, the converted second feature amount being input; a fourth KL divergence function having an estimated value and a classification value as arguments, the converted second feature amount being input; Calculate the loss value for the adversarial training based on The learning device according to claim 4.
6. The system further includes an inference unit, a pseudo-label inference unit, an inference unit learning unit, a pseudo-label generation unit, a conversion unit, a first loss calculation unit, a second loss calculation unit, a fourth loss calculation unit, and a model optimization unit, the inference unit learning unit generates an inference model that generates an inference value indicating whether data in the second dataset corresponds to a positive example or a negative example, based on the inference model, common features in the first dataset that are common to both the first feature space and the second feature space, and common features in the second dataset that are common to both the first feature space and the second feature space; The pseudo-label inference unit generates an inference value indicating whether the data in the second data set corresponds to a positive example or a negative example using the early inference unit; The pseudo-label generator assigns pseudo-labels to the data of the second dataset based on the inferred values; The conversion unit, based on a predetermined conversion model, converting a first feature indicated by first data included in the first dataset into a feature in a common feature space; converting a second feature indicated by second data included in the second data set into a feature in the common feature space; the inference unit generates the inferred value based on the transformed first feature amount, the transformed second feature amount, and the inference model; the first loss calculation unit calculates a first loss between the converted first feature amount and the converted second feature amount based on the converted first feature amount and the converted second feature amount; the second loss calculation unit calculates a second loss by the inference model based on the inference value; the fourth loss calculation unit calculates a fourth loss by the pseudo label based on the inference value; the model optimization unit updates the inference model and the predetermined transformation model based on the first loss, the second loss, and the fourth loss; the common feature space has a correspondence relationship between the first feature space and the second feature space, The first feature is included in the first feature space and is not included in the second feature space. The learning device according to claim 1.
7. A learning system including a first learning device, a second learning device, and a server device, the first learning device includes a first acquisition unit, a first conversion unit, a first inference unit, a first loss calculation unit, a second loss calculation unit, a first model optimization unit, and a first communication unit; the first acquisition unit acquires a first predetermined inference model, a first predetermined transformation model, and a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in binary classification; the first conversion unit converts, based on the first predetermined conversion model, a first feature indicated by first data included in the first data set into a feature in a common feature space; the first communication unit transmits the converted first feature to the server device; the first communication unit receives the converted second feature from the server device; the first inference unit generates a first inference value indicating whether first data corresponding to the converted first feature corresponds to a positive example or a negative example, based on the converted first feature and the first predetermined inference model; the first loss calculation unit calculates a first loss according to the first predetermined transformation model based on the transformed first feature amount, the transformed second feature amount, and the first inferred value; the second loss calculation unit calculates a second loss according to the first predetermined inference model based on the first inference value; the first model optimization unit updates the first predetermined transformation model and the first predetermined inference model based on the first loss and the second loss; the first communication unit transmits the updated first predetermined inference model to the server device; the second learning device includes a second acquisition unit, a second conversion unit, a second inference unit, a third loss calculation unit, a fourth loss calculation unit, a second model optimization unit, and a second communication unit; the second acquisition unit acquires a second predetermined inference model and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification; the second conversion unit converts, based on the second predetermined conversion model, second features indicated by second data included in the second data set into features in the common feature space; the second communication unit transmits the converted second feature to the server device; the second communication unit receives the converted first feature amount from the server device; the second inference unit generates a second inference value indicating whether second data corresponding to the converted second feature corresponds to a positive example or a negative example, based on the converted second feature and the second predetermined inference model; the third loss calculation unit calculates a third loss according to the second predetermined transformation model based on the transformed first feature amount, the transformed second feature amount, and the second inferred value; the fourth loss calculation unit calculates a fourth loss according to the second predetermined inference model based on the second inference value; the second model optimization unit updates the second predetermined transformation model and the second predetermined inference model based on the third loss and the fourth loss; the second communication unit transmits the updated second predetermined inference model to the server device; the server device includes a third communication unit and a processing unit; the third communication unit receives the converted first feature amount from the first learning device and receives the converted second feature amount from the second learning device; the third communication unit transmits the converted second feature to the first learning device and transmits the converted first feature to the second learning device; the third communication unit acquires the updated first predetermined inference model and the updated second predetermined inference model; the processing unit generates a third updated inference model based on the updated first predetermined inference model and the updated second predetermined inference model; the third communication unit transmits the third updated inference model to the first learning device and the second learning device; the first communication unit and the second communication unit acquire the third updated inference model; the first model optimization unit updates the first predetermined inference model based on the first predetermined inference model and the third updated inference model; The second model optimization unit updates the second predetermined inference model based on the second predetermined inference model and the third updated inference model. Learning system.
8. A learning system including a first learning device, a second learning device, and a server device, the first learning device includes a first acquisition unit, a first conversion unit, a first identification unit, a fifth loss calculation unit, a first model optimization unit, and a first communication unit; the first acquisition unit acquires a first predetermined transformation model, a first predetermined discrimination model, and a first dataset that belongs to a first domain including a first feature space and is labeled as only one of positive examples and negative examples in binary classification; the first conversion unit converts, based on the first predetermined conversion model, a first feature indicated by first data included in the first data set into a feature in a common feature space; the first discrimination unit generates a first discrimination value based on the converted first feature amount and the first predetermined discrimination model; the fifth loss calculation unit calculates a fifth loss using the first predetermined transformation model based on the first classification value and adversarial learning; the first model optimization unit updates the first predetermined transformation model and the first predetermined discrimination model based on the fifth loss; the first communication unit transmits the updated first predetermined identification model to the server device; the second learning device includes a second acquisition unit, a second conversion unit, a second inference unit, a second identification unit, a sixth loss calculation unit, a second model optimization unit, and a second communication unit; the second acquisition unit acquires a second predetermined transformation model, a second predetermined inference model, a second predetermined discrimination model, and a second dataset that belongs to a second domain including a second feature space and is not labeled by binary classification; the second conversion unit converts, based on the second predetermined conversion model, second features indicated by second data included in the second data set into features in the common feature space; the second inference unit generates an inference value based on the converted second feature amount and the second predetermined inference model; the sixth loss calculation unit calculates a sixth loss according to the second predetermined transformation model based on the transformed second feature amount and the inference value; the second model optimization unit updates the second predetermined transformation model, the second predetermined inference model, and the second predetermined discrimination model based on the sixth loss; the second communication unit transmits the updated second predetermined identification model to the server device; the server device includes a third communication unit and a processing unit; the third communication unit acquires the updated first predetermined identification model and the updated second predetermined identification model; the processing unit generates a third updated discriminant model based on the updated first predetermined discriminant model and the updated second predetermined discriminant model; the third communication unit transmits the third updated discriminant model to the first learning device and the second learning device; the first communication unit and the second communication unit acquire the third updated identification model; the first model optimization unit updates the first predetermined discrimination model based on the first predetermined discrimination model and the third updated discrimination model; The second model optimization unit updates the second predetermined discrimination model based on the second predetermined discrimination model and the third updated discrimination model. Learning system.
9. 1. A computer-implemented method comprising: A step of acquiring a first dataset belonging to a first domain including a first feature space and labeled with only one of positive examples and negative examples in a binary classification, and a second dataset belonging to a second domain including a second feature space and not labeled with a binary classification; generating an inference model based on the first dataset and the second dataset to infer whether second data included in the second dataset is one of a positive example and a negative example in binary classification; method.
10. On the computer, A step of acquiring a first dataset belonging to a first domain including a first feature space and labeled with only one of positive examples and negative examples in a binary classification, and a second dataset belonging to a second domain including a second feature space and not labeled with a binary classification; generating an inference model for inferring whether second data included in the second dataset is one of a positive example and a negative example in binary classification, based on the first dataset and the second dataset. program.