Information processing method, information processing device, and program
Patent Information
- Application Number
- JP2024544001
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2023-07-07
- Filing Date
- 2023-07-07
- Publication Date
- 2025-05-12
AI Technical Summary
Existing data-based classification or prediction models often suffer from performance deterioration when trained in one hospital facility and introduced to another due to domain shifts, where data distribution differences occur, making it challenging to achieve domain generalizability.
An information processing method that extracts a domain-generalizable subset from a dataset, evaluating its domain generalizability by assessing the change in joint probability distribution of explanatory and objective variables, and using this subset to construct a classification or prediction model that maintains performance across different domains.
The method effectively reduces classification or prediction performance degradation due to domain shifts, enabling the construction of models with high domain generalizability, ensuring consistent accuracy across varying data conditions.
Smart Images

Figure 2024048078000001
Abstract
Description
Information processing method, information processing device, and program
[0001] The present disclosure relates to an information processing method, an information processing device, and a program, and in particular to a technique for constructing a model with domain generalizability.
[0002] Data-based classification or prediction models are widely used, for example, to classify diseases based on medical images.
[0003] However, a model trained on data from one hospital facility may not achieve the expected accuracy when introduced to another facility. This is often due to a domain shift, where the data distribution differs between the training facility and the facility where the model is introduced.
[0004] Research into improving robustness against domain shifts is called domain generalization, and research into this has been active in recent years. Known methods include learning feature representations that are invariant regardless of domain, and meta-learning methods that learn from evaluation performance after pseudo-domain shifts.
[0005] For example, Non-Patent Document 1 proposes a method for selecting domain-universal features. Specifically, for training data of multiple domains, the correlation coefficient between each feature (explanatory variable) and the objective variable is calculated for each domain, and features whose absolute values of the correlation coefficients are equal to or greater than a threshold value in all domains in the training data are selected.
[0006] Vikas Garg, Adam Tauman Kalai, Katrina Ligett, Steven Wu "Learn to Expect the Unexpected: Probably Approximately Correct Domain Generalization" (2021 AISTATS)Ivan Cantador, Ignacio Fenandez-Tobias, Shlomo Bwrkovsky, Paolo Cremonesi, Chapter 27:"Cross-domain Recommender System" (2015 Springer)G Chandrashekar, F Sahin, Computers & Electrical Engineering,"A survey on feature selection methods"(2014 Elsevier)Y Saeys, I Inza, P Larranaga "A review of feature selection techniques in bioinformatics"(bioinformatics, 2007)
[0007] The method in Non-Patent Document 1 is based on the premise that it is possible to construct a classification model or prediction model with domain generalizability for the entire data. However, in reality, this is often not possible, resulting in a problem of reduced performance.
[0008] The present invention has been made in view of the above circumstances, and has as its object to provide an information processing method, an information processing device, and a program for extracting a subset with domain generalizability.
[0009] To achieve the above object, a first aspect of the present disclosure provides an information processing method executed by one or more processors, the information processing method including: extracting a subset from a dataset to be analyzed under specified conditions; and evaluating the domain generalizability of the extracted subset. The domain generalizability of a subset refers to a relatively small decrease in classification or prediction performance due to a change in the joint probability distribution of explanatory variables and a target variable (domain shift) caused by changes in various conditions in the data generation process in a classification or prediction model (classification / prediction model), or a relatively high classification performance or prediction accuracy in the event of a domain shift. According to this aspect, a subset with domain generalizability can be extracted, thereby enabling the construction of a classification / prediction model with domain generalizability.
[0010] An information processing method according to a second aspect of the present disclosure is preferably the information processing method according to the first aspect, wherein the dataset includes datasets of multiple domains, and the evaluating includes, by one or more processors, extracting a training subset from the dataset of one or more learning domains among the multiple domains under specified conditions and training a learning model using the training subset; extracting an evaluation subset from the dataset of one or more evaluation domains among the multiple domains that are different from the learning domain under specified conditions and evaluating the learning model using the evaluation subset; and evaluating domain generalizability using the evaluation results of the learning model.
[0011] An information processing method according to a third aspect of the present disclosure is preferably the information processing method according to the first or second aspect, wherein the dataset includes datasets of multiple domains, and the evaluating step preferably includes: extracting, by one or more processors, a training subset from the dataset of one or more learning domains among the multiple domains under specified conditions, and training a learning model using the training subset; extracting a first evaluation subset different from the training subset from the dataset of the learning domain under specified conditions, and evaluating the learning model using the first evaluation subset; extracting a second evaluation subset from the dataset of one or more evaluation domains among the multiple domains under specified conditions, and evaluating the learning model using the second evaluation subset; and evaluating domain generalizability using a difference between the evaluation result using the first evaluation subset and the evaluation result using the second evaluation subset.
[0012] An information processing method according to a fourth aspect of the present disclosure is the information processing method according to any one of the first to third aspects, wherein the dataset includes datasets of a plurality of domains, and the evaluating preferably includes: one or more processors evaluating the degree of association between the feature of the subset and the objective variable for each domain; and determining that the higher the degree of association across more domains, the higher the domain generalizability of the feature; and determining that the subset with more feature with relatively high domain generalizability has relatively high domain generalizability.
[0013] An information processing method according to a fifth aspect of the present disclosure is the information processing method according to any of the first to fourth aspects, wherein the dataset includes datasets of a plurality of domains, and the evaluating preferably includes: evaluating, by one or more processors, the degree of association between the feature of the subset and the objective variable for each domain; determining, as feature having a degree of association equal to or greater than a certain threshold in a certain number of domains, feature having relatively high domain generality; and evaluating the domain generalizability of the subset using the number of feature having relatively high domain generality.
[0014] An information processing method according to a sixth aspect of the present disclosure is preferably the information processing method according to any of the first to fifth aspects, wherein the dataset includes datasets of multiple domains, and the evaluating includes: extracting, by one or more processors, a training subset from the dataset of one or more learning domains among the multiple domains under specified conditions and extracting a feature set from the training subset; extracting, by one or more processors, an evaluation subset from the dataset of one or more evaluation domains among the multiple domains that are different from the learning domain under specified conditions and evaluating the extracted feature set using the evaluation subset; and evaluating the domain generalizability of the training subset using a proportion of features of the extracted feature set that are also valid in the evaluation subset.
[0015] An information processing method according to a seventh aspect of the present disclosure is the information processing method according to the sixth aspect, further comprising: one or more processors evaluating the degree of association between the features of the training subset and the objective variable for each domain; and determining that the domain generality of the features is relatively high when the degree of association is relatively high across more domains; and it is preferable that the extracted feature set includes features with relatively high domain generality.
[0016] An information processing method according to an eighth aspect of the present disclosure is the information processing method according to the sixth aspect, further comprising: one or more processors evaluating the degree of association between the features of the learning subset and the objective variable for each domain; and determining, as features having a degree of association equal to or greater than a certain threshold in a certain number of domains, the features with relatively high domain universality; and it is preferable that the extracted feature set includes features with relatively high domain universality.
[0017] In an information processing method according to a ninth aspect of the present disclosure, in the information processing method according to any one of the first to eighth aspects, the evaluating preferably includes one or more processors evaluating the usefulness of the subset from known usefulness information for each sample in the dataset and the samples included in the subset, and combining the usefulness and domain generalization to form an evaluation of the subset.
[0018] An information processing method according to a tenth aspect of the present disclosure is preferably the information processing method according to any one of the first to ninth aspects, wherein the dataset includes datasets of a plurality of domains, and the evaluating includes, by one or more processors, evaluating domain uniformity, which indicates the closeness of distribution between the number of data items in the dataset for each domain and the number of data items in the subset, and combining the domain uniformity and domain generalizability to evaluate the subset.
[0019] An information processing method according to an eleventh aspect of the present disclosure is preferably an information processing method according to any one of the first to tenth aspects, wherein one or more processors learn a subset classification model that classifies data in a dataset as being a subset or not.
[0020] In an information processing method according to a twelfth aspect of the present disclosure, in the information processing method according to the eleventh aspect, it is preferable that the evaluating includes, by one or more processors, evaluating the subset classification performance of the subset classification model, and combining the subset classification performance and domain generalization to obtain an evaluation of the subset.
[0021] An information processing method according to a thirteenth aspect of the present disclosure is the information processing method according to any one of the first to twelfth aspects, wherein the extracting and evaluating preferably include searching for a subset with higher domain generalization by one or more processors repeatedly adding or deleting samples from a starting subset to the subset. The extracting may include repeatedly adding or deleting samples from a starting subset to the subset, and the evaluating may include searching for a subset with higher domain generalization.
[0022] An information processing method according to a fourteenth aspect of the present disclosure is preferably the information processing method according to the thirteenth aspect, wherein the dataset includes datasets of a plurality of domains, and the searching includes, by one or more processors, evaluating and searching by also evaluating any one of the following: usefulness of the subset evaluated from known usefulness information for each sample in the dataset and the samples included in the subset; domain evenness indicating the closeness of distribution between the number of data items in the dataset and the number of data items in the subset for each domain; and subset classification performance of a subset classification model that classifies data items in the dataset as being a subset or not.
[0023] An information processing method according to a fifteenth aspect of the present disclosure is preferably an information processing method according to any of the first to fourteenth aspects, wherein one or more processors present a plurality of different subset conditions, evaluate the subsets extracted under each of the plurality of different subset conditions, and extract a subset under the subset condition that provides the best evaluation result among the plurality of different subset conditions.
[0024] In order to achieve the above object, an information processing device according to a sixteenth aspect of the present disclosure is an information processing device including one or more processors and one or more memories that store instructions to be executed by the one or more processors, wherein the one or more processors extract subsets of a dataset to be analyzed under specified conditions and evaluate the domain generalizability of the extracted subsets. According to this aspect, it is possible to extract subsets with domain generalizability, thereby constructing a classification / prediction model with domain generalizability.
[0025] To achieve the above object, a seventeenth aspect of the present disclosure provides a program that causes a computer to perform the following functions: extract a subset of a dataset to be analyzed under specified conditions; and evaluate the domain generalizability of the extracted subset. According to this aspect, it is possible to extract a subset with domain generalizability, thereby constructing a classification / prediction model with domain generalizability.
[0026] According to the present disclosure, a subset with domain generalizability can be extracted.
[0027] FIG. 1 is an explanatory diagram showing the implementation flow of a classification / prediction system. FIG. 2 is an explanatory diagram for performing model learning through domain adaptation. FIG. 3 is an explanatory diagram showing examples of training data and evaluation data used in machine learning. FIG. 4 is a graph schematically showing differences in model performance due to differences in datasets. FIG. 5 is a block diagram generally showing an example of the hardware configuration of an information processing device according to an embodiment. FIG. 6 is a functional block diagram showing the functional configuration of the information processing device. FIG. 7 is an explanatory diagram showing classification or prediction according to embodiment 1. FIG. 8 is an explanatory diagram showing processing according to embodiment 2. FIG. 9 is an explanatory diagram showing processing according to embodiment 4. FIG. 10 is an explanatory diagram showing processing according to embodiment 5. FIG. 11 is an explanatory diagram showing processing according to embodiment 6. FIG. 12 is a graph for explaining the usefulness of subsets. FIG. 13 is an explanatory diagram showing processing according to embodiment 7. FIG. 14 is an explanatory diagram showing processing according to embodiment 8. FIG. 15 is an explanatory diagram showing an algorithm for extracting subsets according to embodiment 9. FIG. 16 is a table showing an example of a dataset. Fig. 17 is a table showing an example of a data set obtained by abstracting the data set shown in Fig. 16. Fig. 18 is a table showing another example of an abstracted data set.
[0028] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings.
[0029] [Construction and Operation of Classification / Prediction Model] FIG. 1 is an explanatory diagram showing the implementation flow of a classification system or prediction system (classification / prediction system). This diagram shows a typical flow for implementing a classification / prediction system in a facility. The implementation of a classification / prediction system involves first constructing a model 14 that performs the desired classification or prediction task (classification / prediction task) (Step 1), and then deploying and operating the constructed model 14 (Step 2). In the case of a machine learning model, "constructing" the model 14 includes training the model 14 using learning (training) data to create a classification / prediction model that meets practical classification or prediction performance. "Operating" the model 14 means, for example, inputting a dataset related to employees and obtaining an output of employee turnover risk from the trained model 14.
[0030] Building the model 14 requires data for training. Generally, the model 14 of a classification / prediction system is trained based on data collected at the facility where it is installed. By training using data collected from the facility where it is installed, the model 14 learns the behavior of employees at the facility where it is installed, and is able to accurately predict the risk of employees at the facility where it is installed.
[0031] However, for various reasons, data from the facility where the system will be implemented may not be available. For example, in the case of a document information classification system for a company's in-house system or a hospital's in-house system, the company developing the classification model often does not have access to data from the facility where the system will be implemented. When data from the facility where the system will be implemented is not available, the system must instead be trained using data collected at a different facility.
[0032] Figure 1 shows the classification flow or prediction flow when data from the facility where the system is being introduced is unavailable. If a model 14 trained using data collected at a facility other than the facility where the system is being introduced is used at the facility where the system is being introduced, the prediction accuracy of the model 14 may be reduced due to differences in user behavior between the facilities.
[0033] The problem of a machine learning model not performing well at unknown facilities different from the facilities it was trained on can be broadly understood as a technical challenge of improving robustness against the problem of domain shift, where the source domain in which the model 14 was trained differs from the target domain to which the model 14 is applied. A problem setting related to domain generalization is domain adaptation, which is a learning method that uses data from both the source domain and the target domain. The purpose of using data from a different domain even when data from the target domain exists is to compensate for the small amount of data in the target domain that is insufficient for learning.
[0034] 2 is an explanatory diagram of domain adaptation for learning the model 14. Although the amount of data collected at the target domain, which is the facility where the model is introduced, is relatively small compared to the amount of data collected at a different facility, by learning using both sets of data, the model 14 can predict with a certain degree of accuracy the behavior of users at the facility where the model is introduced.
[0035] Domain generalization is a more difficult problem than domain adaptation because data from the target domain is not accessible during training.
[0036] [Description of Domains] In Non-Patent Document 2, which is a document on research into domain adaptation in information recommendation, differences in domains are classified into the following four types.
[0037] [1] Item attribute level: For example, comedy movies and horror movies are separate domains.
[0038] [2] Item type level: For example, movies and TV series are separate domains.
[0039] [3] Item level: For example, movies and books are separate domains.
[0040] [4] System level: For example, movies in cinemas and movies broadcast on television are separate domains.
[0041] The difference between the "facilities" shown in Figures 1 and 2 corresponds to the system level domain [4] among the four classifications mentioned above.
[0042] The domain is defined by the joint probability distribution P(X,Y) of the objective variable Y and the explanatory variable X. If Pd1(X,Y)≠Pd2(X,Y), then d1 and d2 are different domains. In other words, a domain shift has occurred.
[0043] The joint probability distribution P(X, Y) can be expressed as the product of the distribution of the explanatory variables P(X) and the conditional probability distribution P(Y|X), or the product of the distribution of the objective variable P(Y) and the conditional probability distribution P(Y|X).
[0044] P(X,Y) = P(Y|X) P(X) = P(X|Y) P(Y) Therefore, a change in one or more of P(X), P(Y), P(Y|X) and P(X|Y) results in a different domain.
[0045] [Reasons for the Impact of Domain Shift] Classification / prediction models that perform prediction or classification tasks make inferences based on the relationship between explanatory variable X and target variable Y. Therefore, changes in P(Y|X) naturally result in a decline in classification or prediction performance. Furthermore, when machine learning a classification / prediction model, the classification error or prediction error within the training data is minimized. However, for example, if the frequency with which the explanatory variable X = X_1 is greater than the frequency with which X = X_2 is achieved (i.e., P(X = X_1) > P(X = X_2)), there is more data for X = X_1 than for X = X_2, and therefore error reduction for X = X_1 is prioritized over error reduction for X = X_2 in training. Therefore, classification or prediction performance also declines when P(X) changes between facilities.
[0046] [Example of domain shift] Domain shift can be a problem for models of various tasks. For example, when a model for predicting the risk of employee resignation is trained using data from one company, domain shift can become a problem when it is used in another company.
[0047] Furthermore, for a model that predicts antibody production by cells, domain shift can become a problem when a model trained using data on one antibody is used with another antibody.Furthermore, for a model that classifies Voice of Customer (VOC), for example, a model that classifies VOC into "product features," "support responses," and "other," domain shift can become a problem when a classification model trained using data on one product is used with another product.
[0048] [Typical patterns of domain shift] [Covariate shift] When the distribution of explanatory variables P(X) differs, it is called a covariate shift. For example, when the distribution of user attributes differs between datasets, or more specifically, when the ratio of men to women differs, this corresponds to a covariate shift.
[0049] Prior probability shift: When the distribution of the objective variable P(Y) differs, this is called a prior probability shift. For example, a prior probability shift occurs when the average browsing rate and average purchase rate differ between datasets.
[0050] [Concept Shift] When the conditional probability distributions P(Y|X) and P(X|Y) differ, this is called a concept shift. For example, the probability that a company's R&D department will read a data analysis document is P(Y|X), but this differs between data sets, which is an example of a concept shift.
[0051] Research on domain adaptation or domain generalization can be divided into two types: one that assumes one of the above patterns as the main factor, and the other that considers how to deal with changes in P(X, Y) without considering which pattern is the main factor. In the former case, many studies assume covariate shifts in particular.
[0052] [Subset Extraction] When it is impossible to build a classification / prediction model with domain generalizability for the entire data, by appropriately extracting a subset of the data, it becomes possible to build a classification / prediction model with domain generalizability within the subset. This disclosure proposes an appropriate subset extraction method.
[0053] A subset is a subset of the dataset to be analyzed. For example, when training a model to predict the risk of employee turnover, a subset can be used that extracts only the data of employees in the "sales department" from the data of a certain company. If a predictive model trained using the "sales department" subset is used in another company, the subset is domain-universal if it is the sales department.
[0054] Furthermore, for a model that predicts antibody production by cells, a subset can be used that extracts only the data for antibodies with a "certain cell size or larger" from the data for a certain antibody. When a model trained using a subset with a "certain cell size or larger" is used with a different antibody, the subset is domain-universal as long as the "certain cell size or larger" is used.
[0055] Furthermore, for a model that classifies customer feedback into "product features," "support responses," and "other," it is possible to use a subset of data about a certain product that extracts only the data of "cat owners" customers. When a classification model trained using the "cat owners" subset is applied to a different product, the subset is domain-universal if it is "cat owners."
[0056] [Regarding Generalizability] FIG. 3 is an explanatory diagram showing examples of training data and evaluation data used in machine learning. A data set obtained from a joint probability distribution Pd1(X,Y) of a certain domain d1 is divided into training data and evaluation data. Evaluation data from the same domain as the training data is called "first evaluation data" and is denoted as "evaluation data 1" in FIG. 3. In addition, a data set obtained from a joint probability distribution Pd2(X,Y) of a domain d2 different from domain d1 is prepared and used as evaluation data. Evaluation data from a domain different from the training data is called "second evaluation data" and is denoted as "evaluation data 2" in FIG. 3.
[0057] The model 14 is trained using the training data of domain d1, and the performance of the trained model 14 is evaluated using the first evaluation data of domain d1 and the second evaluation data of domain d2.
[0058] 4 is a graph schematically illustrating differences in model performance due to differences in data sets. If the performance of model 14 in the training data is performance A, the performance of model 14 in the first evaluation data is performance B, and the performance of model 14 in the second evaluation data is performance C, then the relationship between performance A and performance B is typically greater than performance C, as shown in FIG.
[0059] The high generalization performance of the model 14 generally refers to a high performance B or a small difference between performance A and B. In other words, the aim is to achieve high predictive performance even for untrained data without overfitting to the training data.
[0060] In the context of domain generalization in this specification, this refers to high performance C or a small difference between performance B and performance C. In other words, the goal is to achieve consistently high performance even in domains different from the domain used for learning.
[0061] Furthermore, in this specification, the domain generalizability of training data refers to the fact that, in a classification / prediction model trained using training data, the degradation of classification performance or prediction performance due to changes in the joint probability distribution of explanatory variables and target variables caused by changes in the conditions of the data generation process (domain shift) is relatively small, or that the classification performance or prediction accuracy is relatively high in the event of a domain shift. In other words, if the domain generalizability of model 14 is relatively high, the training data used to train that model 14 has relatively high domain generalizability.
[0062] Strictly speaking, the domain generalizability of training data means that there exists a classification / prediction model with relatively little performance degradation. In other words, the performance of a classification / prediction model with the smallest performance degradation in response to a domain shift among candidate classification / prediction models is domain generalizability. For example, if one gene is selected from gene A, gene B, and gene C and cancer is determined if its expression level is equal to or greater than the average value of all training data, the classification / prediction model is uniquely determined by selecting the gene (= feature), so there are three candidate models: model A using gene A, model B using gene B, and model C using gene C.
[0063] Here, assume that without domain shift, Model A, Model B, and Model C all have classification accuracies of 90%, but with domain shift, the classification accuracies are 70%, 80%, and 50%, respectively. In this case, the performance degradation of Model A, Model B, and Model C due to domain shift is -20%, -10%, and -40%, respectively, and the domain generalization is -10% of Model B, the best model.
[0064] However, if the classification accuracies of Model A, Model B, and Model C without the domain shift are 75%, 90%, and 80%, respectively, the performance degradation due to the domain shift will be -5%, -10%, and -30%, respectively. In this case, the domain generalization performance may be defined as -5% for Model A, which has the smallest performance degradation, or as 80% for Model B, which has the highest accuracy upon domain shift.
[0065] Furthermore, since universal characteristics are considered to be applicable to a wide range of subjects, if a subset is domain-universal, it will lead to domain generalization. For example, if the characteristic that the expression of a certain gene indicates cancer is universal regardless of race (domain), then a method for detecting cancer based on the expression of this gene can be used universally regardless of race.
[0066] 5 is a block diagram illustrating an example of a hardware configuration of an information processing device 100 according to an embodiment. The information processing device 100 has a function of extracting a subset of a dataset to be analyzed under specified conditions and a function of evaluating domain generalizability.
[0067] The information processing device 100 can be realized using computer hardware and software. The physical form of the information processing device 100 is not particularly limited, and it may be a server computer, a workstation, a personal computer, a tablet terminal, or the like. Here, an example is described in which the processing functions of the information processing device 100 are realized using a single computer, but the processing functions of the information processing device 100 may also be realized by a computer system configured using multiple computers.
[0068] The information processing device 100 includes a processor 102 , a non-transitory tangible computer-readable medium 104 , a communication interface 106 , an input / output interface 108 , and a bus 110 .
[0069] The processor 102 includes a CPU (Central Processing Unit). The processor 102 may also include a GPU (Graphics Processing Unit). The processor 102 is connected to a computer-readable medium 104, a communication interface 106, and an input / output interface 108 via a bus 110. The processor 102 reads various programs, data, etc. stored in the computer-readable medium 104 and executes various processes. The term "program" includes the concept of a program module and includes instructions equivalent to a program.
[0070] The computer-readable medium 104 is, for example, a storage device including a memory 112 serving as a main storage device and a storage 114 serving as an auxiliary storage device. The storage 114 is configured using, for example, a hard disk drive (HDD), a solid state drive (SSD), an optical disk, a magneto-optical disk, or a semiconductor memory, or an appropriate combination of these. The storage 114 stores various programs, data, and the like.
[0071] The memory 112 is used as a working area for the processor 102, and as a storage unit that temporarily stores programs and various data read from the storage 114. When a program stored in the storage 114 is loaded into the memory 112 and the processor 102 executes the instructions of the program, the processor 102 functions as a means for performing various processes defined by the program.
[0072] The memory 112 stores an extraction program 130, a learning program 132, an evaluation program 134, various programs, various data, and the like, which are executed by the processor 102.
[0073] The extraction program 130 is a program that acquires a dataset and executes a process of extracting a subset, which is a subset of the dataset to be analyzed, from the acquired dataset.
[0074] The learning program 132 is a program that executes a process of learning a classification / prediction model that performs classification or prediction for each subset extracted by the extraction program 130 .
[0075] The evaluation program 134 is a program that executes a process of evaluating the domain generalizability of the subset extracted by the extraction program 130. The evaluation program 134 may also execute a process of evaluating the domain generalizability of the classification / prediction model trained by the training program 132.
[0076] The memory 112 includes a dataset storage unit 140 and a learning model storage unit 142. The dataset storage unit 140 is a storage area for storing datasets collected at the learning facility. The learning model storage unit 142 is a storage area for storing classification / prediction models learned by the learning program 132.
[0077] The communication interface 106 performs communication processing with an external device via a wired or wireless connection, and exchanges information with the external device. The information processing device 100 is connected to a communication line (not shown) via the communication interface 106. The communication line may be a local area network, a wide area network, or a combination of these. The communication interface 106 can serve as a data acquisition unit that accepts input of various data such as an original data set.
[0078] The information processing device 100 may include an input device 152 and a display device 154. The input device 152 and the display device 154 are connected to the bus 110 via the input / output interface 108. The input device 152 may be, for example, a keyboard, a mouse, a multi-touch panel, or other pointing device, or an audio input device, or an appropriate combination thereof. The display device 154 may be, for example, a liquid crystal display, an organic electro-luminescence (OEL) display, a projector, or an appropriate combination thereof. Note that the input device 152 and the display device 154 may be integrally configured, such as a touch panel, or the information processing device 100, the input device 152, and the display device 154 may be integrally configured, such as a touch panel tablet terminal.
[0079] 6 is a functional block diagram showing the functional configuration of the information processing device 100. The information processing device 100 includes a dataset acquisition unit 160, a subset extraction unit 162, a learning unit 164, and a domain generalization evaluation unit 166.
[0080] The dataset acquisition unit 160 acquires a dataset collected at a learning facility from the dataset storage unit 140. A dataset is a collection of data for each domain, and includes an explanatory variable and a target variable.
[0081] The subset extraction unit 162 extracts a subset from the dataset acquired by the dataset acquisition unit 160 under specified conditions. The subset includes all feature quantities (explanatory variables) and objective variables of the dataset. In other words, the subset extraction unit 162 does not narrow down the feature quantities. The extraction conditions may be specified by the user.
[0082] The learning unit 164 generates a classification / prediction model that performs classification or prediction for each subset.
[0083] The domain generalization evaluation unit 166 is a program that executes a process of evaluating the domain generalization of the subset extracted by the subset extraction unit 162. The domain generalization evaluation unit 166 may evaluate the domain generalization of the classification / prediction model trained by the training unit 164.
[0084] First Embodiment An information processing device 100 executes an information processing method in which a subset is extracted from a dataset to be analyzed under specified conditions, and the domain generalizability of the extracted subset is evaluated.
[0085] 7 is an explanatory diagram showing classification or prediction according to embodiment 1. The training or evaluation dataset (training / evaluation dataset), which is the dataset to be analyzed in embodiment 1, includes a dataset DS1 of domain 1, a dataset DS2 of domain 2, and a dataset DS3 of domain 3. Domain 1, domain 2, and domain 3 are each different domains.
[0086] The subset extraction unit 162 extracts subsets from the datasets under specified conditions. That is, the subset extraction unit 162 extracts subset SS1 from dataset DS1, subset SS2 from dataset DS2, and subset SS3 from dataset DS3.
[0087] The domain generalizability evaluation unit 166 evaluates the domain generalizability of the subsets SS1, SS2, and SS3.
[0088] If the subsets SS1, SS2, and SS3 have domain generalization, the learning unit 164 generates a model 14 that performs classification or prediction using the subsets SS1, SS2, and SS3. Because the subsets SS1, SS2, and SS3 have domain generalization, a model 14 with domain generalization can be constructed.
[0089] In this embodiment, there are also data sets DS4 and DS5 of the domain to which the learning model is applied. This domain is unknown when the model 14 is trained.
[0090] The subset extraction unit 162 extracts subsets from the datasets under specified conditions. That is, the subset extraction unit 162 extracts subset SS4 from dataset DS4 and subset SS5 from dataset DS5. The specified conditions here are the same conditions as those used to extract subsets SS1 to SS3 from datasets DS1 to DS3.
[0091] Applying model 14 to subsets SS4 and SS5 enables accurate classification or prediction.
[0092] [Embodiment 2] The information processing device 100 may divide a dataset of multiple domains by domain, extract a learning subset from the dataset of one or more learning domains among the multiple domains under specified conditions, train a learning model using the learning subset, extract an evaluation subset from the dataset of one or more evaluation domains among the multiple domains that are different from the learning domain under specified conditions, evaluate the learning model using the evaluation subset, and evaluate domain generalizability using the evaluation result of the learning model.
[0093] 8 is an explanatory diagram showing the processing according to the second embodiment. Here, domain 1 and domain 2 are learning domains, and domain 3 is an evaluation domain.
[0094] The subset extraction unit 162 extracts a subset SS1 from a data set DS1 of a domain 1, which is a learning domain, and a subset SS2 from a data set DS2, under specified conditions.
[0095] Furthermore, the subset extraction unit 162 divides the subsets SS1 and SS2 into training data TD and first evaluation data ED, respectively. That is, the subset extraction unit 162 divides the subset SS1 into training data TD1 and evaluation data ED1, and divides the subset SS2 into training data TD2 and evaluation data ED2. The training data TD includes the training data TD1 and the training data TD2, and the first evaluation data ED includes the evaluation data ED1 and the evaluation data ED2.
[0096] Furthermore, the subset extraction unit 162 extracts a subset SS3 from the data set DS3 of the evaluation domain, domain 3, under the specified conditions. The subset SS3 corresponds to the second evaluation data.
[0097] The learning unit 164 learns the model 14 using the training data TD (an example of a "training subset").
[0098] The domain generalization evaluation unit 166 evaluates the domain generalization using the performance C shown in Fig. 4. That is, the domain generalization evaluation unit 166 evaluates the domain generalization of the trained model 14 using the subset SS3 (an example of an "evaluation subset"), which is the second evaluation data.
[0099] The information processing device 100 may divide datasets of multiple domains by domain, extract a learning subset from the dataset of one or more learning domains among the multiple domains under specified conditions, train a learning model using the learning subset, extract a first evaluation subset different from the learning subset from the dataset of the learning domain under specified conditions, evaluate the learning model using the first evaluation subset, extract a second evaluation subset from the dataset of one or more evaluation domains among the multiple domains under specified conditions, evaluate the learning model using the second evaluation subset, and evaluate domain generalizability using the difference between the evaluation result using the first evaluation subset and the evaluation result using the second evaluation subset.
[0100] That is, the domain generalization evaluation unit 166 evaluates domain generalization based on the difference between performance B and performance C shown in Figure 4. In this case, the domain generalization evaluation unit 166 may evaluate domain generalization based on the difference between the performance of the model 14 trained using the first evaluation data ED (an example of a "first evaluation subset") and the performance of the model 14 trained using the subset SS3 (an example of a "second evaluation subset"), which is the second evaluation data. Here, the smaller the difference, the higher the domain generalization.
[0101] The domain generalizability evaluation unit 166 may evaluate the domain generalizability by comprehensively determining both the performance C shown in FIG. 4 and the difference between the performance B and the performance C shown in FIG.
[0102] [Embodiment 3] The information processing device 100 may divide a dataset of multiple domains into domains, evaluate the degree of association between feature amounts (explanatory variables) of subsets of the dataset divided into domains and the objective variable, define feature amounts that have a degree of association equal to or greater than a certain threshold in a certain number of domains as feature amounts with relatively high domain universality, and evaluate the domain generalizability of the subsets using the number of feature amounts with relatively high domain universality.
[0103] In this embodiment, S is a set of samples of data, denoted as S={1, 2,..., i,..., |S|}.
[0104] F is a set of features, expressed as F = {1, 2, ..., k, ..., |F|}.
[0105] D is the set of domains, denoted as D = {1, 2,..., d,..., |D|}.
[0106] Each sample i is a feature vector x contained in an |F|-dimensional real vector. i (x i ∈R |F| ), and the target variable yi. The value of the kth feature of sample i is
[0107] Each sample is included in one of the domains. The domain of sample i is denoted as di, where di∈D. A subset G is extracted from the dataset S, and G⊂S.
[0108] indicates the correlation between the kth feature of domain d and the objective variable. For example, G∩[i| d i =d], the target variable yi and the value of the kth feature of sample i are
[0109] There is a Pearson correlation coefficient between
[0110] A subset of domain-universal features F DG ⊂G is defined as follows:
[0111] where
[0112] is an indicator function that is 1 when the correlation between the feature and the objective variable is greater than or equal to θ or less than -θ, and is 0 otherwise. θ is a constant threshold, for example, θ = 0.8. m is a constant number, for example, the total number of domains is 5, and m = 4. The domain generalization of subset G is expressed as F DG The number of features in (G), i.e., |F DG (G) is evaluated by |
[0113] The information processing device 100 may divide a dataset of multiple domains into domains, evaluate the degree of association between the features of subsets of the dataset divided into domains and the target variable, and determine that the higher the degree of association in more domains, the higher the domain universality of the features, and evaluate that the subset with more features with relatively high domain universality has relatively high domain generalizability.
[0114] That is, more generally, for larger θ, m, |F DG The larger (G)| is, the better. Such an evaluation method can be expressed as the following equation (2).
[0115]
[0116] Here, α and β are scaling factors that give more weight to larger values of m and θ, respectively.
[0117] There are various feature selection methods, which can be classified into filter methods, wrapper methods, and embedded methods, as described in Non-Patent Documents 3 and 4.
[0118] The filter method evaluates the correlation between features and target variables independently of the classification model. The wrapper method evaluates features based on the performance of a specific classification model. The embedding method incorporates feature selection inherently into the classification model algorithm. Examples include decision trees and Lasso.
[0119] Filtering methods are classified into univariate methods, which evaluate features one by one, and multivariate methods, which evaluate a set of features. The method using Pearson correlation is one of the univariate methods of filtering. Of course, other univariate methods can also be used. It is also easy to extend to multivariate methods and wrapper methods. In such cases, the relevance v is defined for the feature set, and a feature set F' that satisfies the condition in Equation 3 below is a domain-universal feature set.
[0120]
[0121] [Embodiment 4] The information processing device 100 may divide datasets of multiple domains by domain, extract a training subset from the dataset of one or more learning domains among the multiple domains under specified conditions, extract a feature set from the training subset, extract an evaluation subset from the dataset of one or more evaluation domains different from the learning domain among the multiple domains under specified conditions, evaluate the extracted feature set using the evaluation subset, and evaluate the domain generalizability of the training subset using the proportion of features of the extracted feature set that are also valid in the evaluation subset.
[0122] The information processing device 100 evaluates the degree of association between the features of the learning subset and the objective variable for each domain, and determines that the higher the degree of association in more domains, the higher the domain universality of the feature, and the extracted feature set may include features with relatively high domain universality.
[0123] The information processing device 100 evaluates the degree of association between the features of the learning subset and the objective variable for each domain, and defines features that have an association degree equal to or greater than a certain threshold in a certain number of domains as features with relatively high domain universality, and the extracted feature set may include features with relatively high domain universality.
[0124] 9 is an explanatory diagram showing the processing according to the fourth embodiment. Here, domain 1 and domain 2 are learning domains, and domain 3 is an evaluation domain.
[0125] The subset extraction unit 162 extracts a subset SS1 from a dataset DS1 of domain 1, which is the learning domain, and a subset SS2 from a dataset DS2, under specified conditions. The subset extraction unit 162 also extracts a subset SS3 from a dataset DS3 of domain 3, which is the evaluation domain, under specified conditions.
[0126] The domain generalization evaluation unit 166 extracts a feature set VS from all features included in the subsets SS1 and SS2 (examples of "learning subsets"). The feature set VS is a subset of the explanatory variables of the subsets SS1 and SS2, or features created therefrom.
[0127] The domain generalization evaluation unit 166 basically extracts a feature set VS based on a criterion based on the degree of association between the objective variable and the explanatory variables. For example, if height is the objective variable and the expression levels of 10,000 genes are the explanatory variables, the Pearson correlation coefficient between each gene expression level and height is calculated, and those with a correlation coefficient above a certain value are extracted. If the explanatory variable is not a continuous value like height but a binary variable indicating whether or not there is cancer, selection is made based on criteria such as the p-value of a t-test or AUC (Area Under the Curve). These correspond to the filtering method described above.
[0128] In the case of the wrapper method, the evaluation is based on the performance of the classification model. For example, in a logistic regression model, the classification accuracy is compared when each gene is added to the explanatory variables and when it is not added, and the model with the largest difference (the model with the largest improvement in accuracy when added to the explanatory variables) is selected.
[0129] However, when extracting a feature set VS taking into consideration domain universality within the learning domain, for example, a set with a Pearson correlation coefficient equal to or greater than a certain value is extracted in both domain 1 and domain 2.
[0130] In addition, the domain generalization evaluation unit 166 verifies the validity of the feature set VS using the subset SS3 (an example of an "evaluation subset").
[0131] One of the following methods can be considered for validating the effectiveness: - Extract the same features in the evaluation domain as in the learning domain and create an overlap. - Ensure that the positive and negative correlations between the learning domain and the features are the same. In other words, if there is a positive correlation in the learning domain, there should also be at least a positive correlation in the evaluation domain.
[0132] 9 shows the percentage of feature sets VS that are valid in the evaluation domain, shaded in. For example, when 15 extracted features and 9 valid features are extracted, the percentage is higher than when 20 extracted features and 10 valid features are extracted, and therefore the latter subset has relatively higher domain generalizability.
[0133] Fifth Embodiment The information processing device 100 may present a plurality of different subset conditions, evaluate the subsets extracted under each of the plurality of different subset conditions, and extract a subset under the subset condition that provides the best evaluation result among the plurality of different subset conditions.
[0134] 10 is an explanatory diagram showing processing according to the fifth embodiment. The information processing device 100 includes a subset condition presentation unit 168. The subset condition presentation unit 168 presents a plurality of different subset conditions (an example of "designation"). Here, the learning / evaluation datasets include a dataset DS1 for domain 1, a dataset DS2 for domain 2, and a dataset DS3 for domain 3.
[0135] 10 , the subset condition presenting unit 168 presents two different subset conditions: a subset condition A and a subset condition B. For a subset used in a model for predicting antibody production by cells, for example, subset condition A is "cell size is 10 μm or more," and subset condition B is "cell size is 15 μm or more." In other words, the different subset conditions A and B are defined by different values (attributes), "10 μm or more" and "15 μm or more," for the same feature item, "cell size."
[0136] The subset extraction unit 162 extracts subsets SS1A, SS2A, and SS3A from the data sets DS1, DS2, and DS3, respectively, under subset condition A. The subset extraction unit 162 also extracts subsets SS1B, SS2B, and SS3B from the data sets DS1, DS2, and DS3, respectively, under subset condition B.
[0137] The domain generalizability evaluation unit 166 adopts the subset condition that provides the better evaluation result from among the subset condition A and the subset condition B. The subset extraction unit 162 extracts a subset under the adopted subset condition.
[0138] Here, domain 1 and domain 2 are set as learning domains, and domain 3 is set as an evaluation domain. The learning unit 164 generates model 14 using subsets SS1A and SS2A from among the subsets of subset condition A as learning data, and generates model 14 using subsets SS1B and SS2B from among the subsets of subset condition B as learning data.
[0139] Furthermore, the domain generalization evaluation unit 166 evaluates the model 14 under subset condition A using subset SS3A from among the subsets under subset condition A as evaluation data, and evaluates the model 14 under subset condition B using subset SS3B from among the subsets under subset condition B as evaluation data. In this way, subset conditions A and B can be evaluated.
[0140] As in the fourth embodiment, subset conditions A and B may be evaluated using feature sets. In this case, the domain generalization evaluation unit 166 extracts feature sets from subsets SS1A and SS2A and verifies the effectiveness of the feature sets using subset SS3A. Also, the domain generalization evaluation unit 166 extracts feature sets from subsets SS1B and SS2B and verifies the effectiveness of the feature sets using subset SS3B.
[0141] [Embodiment 6] The information processing device 100 may evaluate the usefulness of a subset based on known usefulness information for each sample in a dataset and the samples included in the subset, and may evaluate the subset by combining the usefulness and domain generalization.
[0142] 11 is an explanatory diagram showing processing according to the sixth embodiment. The domain generalization evaluation unit 166 includes a subset usefulness evaluation unit 166A. The subset usefulness evaluation unit 166A evaluates the usefulness of a subset based on known usefulness information for each sample in the dataset and the samples included in the subset.
[0143] 11, the learning / evaluation datasets provided here are a dataset DS1 for domain 1, a dataset DS2 for domain 2, and a dataset DS3 for domain 3. The subset extraction unit 162 extracts subsets SS1, SS2, and SS3 from the datasets DS1, DS2, and DS3, respectively.
[0144] The domain generalizability evaluation unit 166 evaluates the domain generalizability of the subsets SS1, SS2, and SS3.
[0145] The subset usefulness evaluation unit 166A also evaluates the usefulness of the subsets SS1, SS2, and SS3. FIG. 12 is a graph illustrating the usefulness of the subsets. Here, in the camera VOC classification, it is assumed that the more photo prints a user makes per year, the better the user, i.e., the sample has relatively high usefulness. For the subsets SS1, SS2, and SS3, the subset usefulness evaluation unit 166A calculates the average number of photo prints made per year by each of the users in the "cat owners," "dog owners," "other pets," and "no pets" categories, and evaluates the usefulness of the subset relatively higher the higher the average. In the example shown in FIG. 12, the usefulness increases in the order of "cat owners," "dog owners," "other pets," and "no pets."
[0146] In the case of VOC classification, the subset usefulness evaluation unit 166A may evaluate the usefulness of a subset to be relatively higher the more loyal customers the subset has.In the case of antibody production prediction, the subset usefulness evaluation unit 166A may evaluate the usefulness of a subset to be relatively higher the faster the cell proliferation rate.
[0147] The subset usefulness evaluation unit 166A calculates the average of the usefulness of each sample for the samples belonging to the subsets SS1, SS2, and SS3, and sets this average as the usefulness of the subsets SS1, SS2, and SS3.
[0148] The domain generalizability evaluation unit 166 combines the usefulness and domain generalizability evaluated by the subset usefulness evaluation unit 166A to form an evaluation of the subsets SS1, SS2, and SS3.
[0149] [Embodiment 7] The dataset includes datasets of multiple domains, and the information processing device 100 evaluates domain uniformity, which indicates the closeness of the distribution between the number of data in the dataset for each domain and the number of data in the subset, and may evaluate the subset by combining the domain uniformity and domain generalization.
[0150] 13 is an explanatory diagram showing processing according to the seventh embodiment. The domain generalization evaluation unit 166 includes a domain uniformity evaluation unit 166B. The domain uniformity evaluation unit 166B evaluates domain uniformity, which indicates the closeness of the distribution between the number of data items in a dataset and the number of data items in a subset for each domain.
[0151] 13, the learning / evaluation datasets provided here are a dataset DS1 for domain 1, a dataset DS2 for domain 2, and a dataset DS3 for domain 3. The numbers of data in the datasets DS1, DS2, and DS3 are n1, n2, and n3, respectively.
[0152] The subset extraction unit 162 extracts subsets SS1, SS2, and SS3 from the data sets DS1, DS2, and DS3, respectively, under specified conditions. The numbers of data in the subsets SS1, SS2, and SS3 are m1, m2, and m3, respectively.
[0153] The domain generalizability evaluation unit 166 evaluates the domain generalizability of the subsets SS1, SS2, and SS3.
[0154] The domain uniformity evaluation unit 166B also evaluates domain uniformity. The domain uniformity evaluation unit 166B evaluates the data set as uniform if the abundance ratio of subsets in domains 1, 2, and 3 is the same, 40%. The domain uniformity evaluation unit 166B evaluates the data set as unequal if the abundance ratio of subsets in domain 1 is 90%, 50% in domain 2, and 10% in domain 3.
[0155] Here, the domain uniformity evaluation unit 166B evaluates the distribution closeness between the numbers of data n1, n2, and n3 in the data sets DS1, DS2, and DS3 and the numbers of data m1, m2, and m3 in the subsets SS1, SS2, and SS3. For example, the domain uniformity evaluation unit 166B evaluates that the smaller the KL divergence (Kullback-Leibler divergence) is, the better the domain uniformity is. The KL divergence can be expressed as follows:
[0156] Kullback-Leibler divergence=Σ k P1(d=k) (log(P1(d=k)-log(P2(d=k))) where k is the domain ID, here 1, 2, and 3. P1 is the distribution over n1, n2, and n3, and P1(d=k) = n1 / (n1+n2+n3). Also, P2 is the distribution over m1, m2, and m3, and P2(d=k) = m1 / (m1+m2+m3).
[0157] Eighth Embodiment The information processing device 100 may train a subset classification model. The information processing device 100 may evaluate the subset classification performance of the subset classification model, and may evaluate the subset by combining the subset classification performance and domain generalization.
[0158] 14 is an explanatory diagram showing processing according to the eighth embodiment. Here, the learning / evaluation datasets include a dataset DS1 for domain 1, a dataset DS2 for domain 2, and a dataset DS3 for domain 3. The subset extraction unit 162 extracts subsets SS1, SS2, and SS3 from the datasets DS1, DS2, and DS3, respectively, based on the subset classification model 162B.
[0159] Here, the subset extraction unit 162 according to the eighth embodiment includes a subset classification learning unit 162A. The subset classification learning unit 162A learns a subset classification model 162B that classifies data in a dataset as being a subset or not.
[0160] For example, let's say that antibody production prediction is performed using gene expression as a feature, and subset extraction is performed based on DNA mutations. During operation, it is more convenient to be able to perform subset extraction using gene expression, so a subset classification model from gene expression is also trained separately. Naturally, the higher the performance of the classification model, the better, so the ease of classification is also added to the evaluation of the domain generalizability of the subset.
[0161] F14A in FIG. 14 is a diagram showing subsets extracted from a certain dataset. In F14A, the vertical and horizontal axes represent features or indices calculated from the features, and the samples of the dataset are plotted according to the values of the vertical and horizontal axes. This example shows an example in which the samples of the dataset are classified into samples belonging to the subset indicated by black circles and samples not belonging to the subset indicated by white circles. For example, in VOC classification, samples belonging to the subset are users who "keep cats," and samples not belonging to the subset are users other than "keep cats."
[0162] Whether or not a subject is a subset is defined by the subset extraction unit 162. F14A shows an example in which the subset extraction condition is "keeping a cat." In the stability prediction example described below, the subset extraction condition is that the amount of antibody production is equal to or greater than the average value, and if the amount of antibody production is equal to or greater than the average value, it is defined as a subset.
[0163] The subset classification learning unit 162A uses the data shown in F14A to learn the subset classification model 162B.
[0164] The domain generalization evaluation unit 166 evaluates the domain generalization of the subsets SS1, SS2, and SS3. The domain generalization evaluation unit 166 also evaluates the subset classification performance of the subset classification model 162B, and evaluates the subsets by combining the subset classification performance and the domain generalization.
[0165] Subset classification performance is evaluated using the same method as for evaluating general machine learning models: data is divided into "training" and "evaluation" data, and a model trained using the training data is evaluated using the evaluation data. For example, 80% of the training data is divided into "cat keeping" and 20% of the evaluation data, and a model trained using the former is evaluated based on whether it can correctly classify the latter data into "cat keeping" and "non-cat keeping." Training and evaluation of subset classification can be performed using a combination of data from Domain 1, Domain 2, and Domain 3.
[0166] [Embodiment 9] The information processing device 100 may search for a subset with higher domain generalizability by repeatedly adding or deleting samples from a starting subset (an example of a "specified condition"). The information processing device 100 may also search for a subset with higher domain generalizability by evaluating any one of the following: usefulness of the subset evaluated based on known usefulness information for each sample in the dataset and the samples included in the subset; domain evenness indicating the closeness of the distribution between the number of data items in the dataset and the number of data items in the subset for each domain; and subset classification performance of a subset classification model that classifies data items in the dataset as being a subset or not.
[0167] 15 is an explanatory diagram showing an algorithm for extracting a subset according to the ninth embodiment. In this algorithm, the inputs are θ, which is a threshold for the degree of association between a feature and a target variable, n, which is the maximum size of the subset, and e, which is a convergence condition for terminating the process when the improvement in domain generalizability falls below a certain value, and the output is a subset G with domain generalizability. Here, the domain generalizability of the subset is evaluated by the number of features with domain generalizability.
[0168] 15, the subset G is initialized to the starting subset. The starting subset can be defined arbitrarily, and may be the entire data set or an empty set.
[0169] 15 defines |G|, which represents the number of samples in the subset G for processing, and the range of the increase amount Δ of domain generalization. That is, if |G|<n and Δ>e are satisfied, processing continues, and if not, processing ends.
[0170] In the third line of Fig. 15, a subset G is created by adding one sample from the subset G. F In the fourth line of FIG. 15, one sample is removed from the subset G to create a subset G B Here, the samples are increased or decreased one by one, but multiple samples can be increased or decreased at once. The target of argmax is |F DG |In addition to |, it can also be a weighted sum of usefulness evaluation, domain uniformity, and subset classification performance.
[0171] In the fifth line of FIG. 15, the subset G F The subset G B The subset G is judged to have relatively higher domain generality than the subset G (whether the number of domain-generalizable features is large or not). F If the domain universality is relatively higher, then in the sixth line of Figure 15, Δ is set to |F DG (G F ) |-|F DG (G) | Update G to G F Update to.
[0172] Subset GF If the domain universality of is not relatively high, then in the 9th line of Figure 15, Δ is changed to |F DG (G B ) |-|F DG (G) | Update G to G B Update to.
[0173] The above process is repeated until |G| or Δ falls outside the range defined in the second line of Fig. 15. This makes it possible to search for a subset with high domain generalizability.
[0174] [Embodiment 10] Fig. 16 is a table showing an example of a dataset. The dataset shown in Fig. 16 is a dataset for a certain company, "Company 1." As shown in Fig. 16, the dataset has the following items as explanatory variables for each sample (employee): "job type," "age," "years of service," "years since last promotion," "annual income," "highest level of education," "commuting time," "average overtime hours," and "age difference with superior," and has the item "whether or not the employee has left the company" as a dependent variable. Note that "whether or not the employee has left the company" is represented by "1" for "yes" and "0" for "no."
[0175] For example, the employee in the top row of Figure 16 has the following "job type," "age," "years of service," "years since last promotion," "annual income," "highest level of education," "commuting time," "average overtime hours," "age difference between superior and employee," and "whether or not the employee has left the company" fields: "sales," "32," "10," "5," "600," "university," "40," "20," "3," and "1," respectively.
[0176] 17 is a table showing an example of a dataset obtained by abstracting the dataset shown in FIG. 16. Here, "occupation" is a "variable for subset extraction," and "age," "years of service," "years since last promotion," "annual income," "highest level of education," "commuting time," "average overtime hours," and "age difference with superior" are "feature A," "feature B," "feature C," "feature D," "feature E," "feature F," "feature G," and "feature H," respectively. Furthermore, "whether or not the employee has left the company" is a "target variable."
[0177] We will now explain the resignation risk prediction model. In this example, each company is a domain. There are five data sets: "Company 1," "Company 2," "Company 3," "Company 4," and "Company 5." We want to build a robust prediction model for companies (other domains) other than "Company 1" to "Company 5." We will build a prediction model that predicts whether or not a person will resign (1 / 0) based on the eight feature quantities A-H shown in Figure 17. The feature quantities are converted to numerical values and are all treated as continuous variables. Here, a subset is extracted based on "occupation type." No feature refinement is performed when extracting the subset.
[0178] The relevance of each feature and objective variable is evaluated for each of the five companies (domains), and features with high relevance in four or more of the five domains are considered to have high domain generality. Relevance is evaluated based on AUC, with an AUC of 0.8 or higher being considered to be high relevance. More precisely, since AUC is symmetric around 0.5, relevance is evaluated as high if |AUC-0.5| is 0.3 or higher. The domain generalization of the extracted subset is evaluated based on the number of features with high domain generality.
[0179] First, we evaluated domain universality for all data without extracting subsets, and found that there was only one feature with high domain universality. Next, we extracted subsets for each job type and evaluated domain universality. When subsets were extracted for "sales," "development," "research," and "production," the number of features with high domain universality was 5, 2, 1, and 2, respectively. This shows that extracting "sales" as a subset results in high universality.
[0180] Therefore, we prepare a subset of data from each company that extracts only sales positions, and use this to train a turnover risk prediction model.The trained model is expected to have high domain generalizability for the "sales" subset, and can be used with confidence for other unknown companies.
[0181] [Another Example of Data Set] Fig. 18 is a table showing another example of an abstracted dataset. The dataset shown in Fig. 18 is a dataset from a certain company, and it is assumed that a VOC classification model constructed from this dataset will be applied to other companies.
[0182] 18 has the following items: "word A," "word B," "word C," "word D," "word E," "word F," "word G," "word H," and "target variable," in addition to "subset extraction variable." "word A," "word B," "word C," "word D," "word E," "word F," "word G," and "word H" each correspond to a feature.
[0183] [Embodiment 11] Pharmaceutically useful antibodies are widely produced using CHO (Chinese hamster ovary) cells, etc. In such production methods, it is desirable that the production amount of cells per unit time is large and that the production is stable over a long period of time (without significant changes in the production amount).
[0184] Therefore, we will build a model that predicts stability labels (1 or 0) based on gene expression values. We want to build a model that can predict stability regardless of the antibody species produced by the cells. In other words, predictive performance for unknown antibody species is important, and we need a prediction model with domain generalizability that uses antibody species as a domain.
[0185] Here, there are five domain datasets: antibody types a, b, c, d, and e. Gene expression values are count values that take positive integers and are logarithmically transformed before being used as features. The relevance between each feature and the objective variable (stability label) is evaluated for each of the five antibody types (domains), and features with high relevance in four or more of the five domains (m = 4) are considered to have high domain generality. The relevance is evaluated based on the absolute value of the difference between the feature mean value and the objective variable, and a difference of 0.2 (θ = 0.2) or greater is considered to have high relevance. The domain generalization of the extracted subset is evaluated based on the number of features with high domain generality |F(DG)|.
[0186] First, when domain universality was evaluated for all data without extracting subsets, only one feature was found to have high domain universality. Next, data with antibody production levels above the average were extracted as a subset, and similar features with high domain universality were extracted. 224 features were found to have high domain universality. Furthermore, when data with viable cell density above the average were extracted as a subset, 33 features were found to have high domain universality. Therefore, we decided to extract subsets based on antibody production levels.
[0187] Using data from the extracted subset (antibody production levels above the average), a stability prediction model is machine-learned. Here, a logistic regression model is used. Feature values are sequentially selected from 224 features with high domain generality, and 50 are trained to maximize the predictive accuracy of logistic regression. Because training is limited to subsets with high domain generality, the trained model is capable of robustly predicting stability for other antibody types using this subset.
[0188] When using a trained model to predict stability (for untrained antibody types), first extract a subset of samples under the same conditions (antibody production levels above the average) from the samples you want to predict, and then predict the stability label for that subset from gene expression levels. During operation, samples predicted to be stable are sent to the next step.
[0189] In the above example, the domain generalizability of the subset is evaluated based on the number of features with high domain generalizability, but it may also be evaluated based on the ratio of extracted features that are also effective in other domains.
[0190] Specifically, for example, using data from four of the five domains, the top 100 features with the highest AUC are extracted, and whether these features are equally effective for prediction in the remaining domain is evaluated (here, the condition is AUC of 0.6 or higher). If the number of subsets extracted is 52 (a proportion of 0.52) when the subset is based on viable cell density, and 83 (a proportion of 0.83) when the subset with the highest antibody production level is extracted, the subset with the highest antibody production level is evaluated as having higher domain generalizability.
[0191] As another evaluation method, as described in embodiment 2, the domain generalizability of a subset may be evaluated by evaluating the predictive performance of a model trained in four domains in the remaining one domain.
[0192] [Regarding the program that operates the computer] A program that causes a computer to realize some or all of the processing functions of the information processing device 100 can be recorded on a computer-readable medium, such as an optical disk, a magnetic disk, a semiconductor memory, or other tangible, non-transitory information storage medium, and the program can be provided through this information storage medium.
[0193] In addition, instead of providing the program by storing it on such a tangible, non-transitory computer-readable medium, it is also possible to provide the program signal as a download service using a telecommunications line such as the Internet.
[0194] Furthermore, some or all of the processing functions of the information processing device 100 may be realized by cloud computing, and may also be provided as SaaS (Software as a Service).
[0195] [Hardware Configuration of Each Processing Unit] The hardware configuration of the processing units that execute various processes in the information processing device 100, such as the dataset acquisition unit 160, subset extraction unit 162, learning unit 164, and domain generalization evaluation unit 166, is, for example, various processors as shown below.
[0196] The various types of processors include CPUs, which are general-purpose processors that execute programs and function as various processing units, GPUs, programmable logic devices (PLDs) such as FPGAs (Field Programmable Gate Arrays) that are processors whose circuit configuration can be changed after manufacture, and dedicated electrical circuits such as ASICs (Application Specific Integrated Circuits) that are processors with a circuit configuration designed specifically for executing specific processing.
[0197] A single processing unit may be composed of one of these various processors, or two or more processors of the same or different types. For example, a single processing unit may be composed of multiple FPGAs, a combination of a CPU and an FPGA, or a combination of a CPU and a GPU. Multiple processing units may also be composed of a single processor. Examples of multiple processing units composed of a single processor include, first, a configuration in which a single processor is composed of a combination of one or more CPUs and software, as typified by client and server computers, and this processor functions as multiple processing units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a system-on-chip (SoC). In this way, the various processing units are composed of one or more of the above-mentioned various processors as a hardware structure.
[0198] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.
[0199] Advantages of the Embodiment According to the present embodiment, a subset with domain generalizability can be extracted, and therefore a classification / prediction model with domain generalizability can be constructed.
[0200] [Other Application Examples] In the above-described embodiments, examples have been described using a model for predicting the risk of employee resignation, a model for predicting the amount of antibody production in cells, and a model for classifying VOCs, but the technology of the present disclosure can be applied to the construction of various classification / prediction models.
[0201] The domain generalizability evaluation unit 166 may evaluate the domain generalizability of the subset by combining the subset evaluation methods according to the respective embodiments.
[0202] [Others] The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the technical idea of the present disclosure.
[0203] 14 Model 100 Information processing device 102 Processor 104 Computer readable medium 106 Communication interface 108 Input / output interface 110 Bus 112 Memory 114 Storage 130 Extraction program 132 Learning program 134 Evaluation program 140 Dataset storage unit 142 Learning model storage unit 152 Input device 154 Display device 160 Dataset acquisition unit 162 Subset extraction unit 162A Subset classification learning unit 162B Subset classification model 164 Learning unit 166 Domain generalizability evaluation unit 166A Subset usefulness evaluation unit 166B Domain uniformity evaluation unit 168 Subset condition presentation unit DS1 Dataset DS2 Dataset DS3 Dataset DS4 Dataset DS5 Dataset ED First evaluation data ED1 Evaluation data ED2 Evaluation data SS1 subset SS1A subset SS1B subset SS2 subset SS2A subset SS2B subset SS3 subset SS3A subset SS3B subset SS4 subset SS5 subset TD Training data TD1 Training data TD2 Training data
Claims
1. An information processing method executed by one or more processors, comprising: extracting a subset from a dataset to be analyzed using specified conditions; and evaluating the domain generalizability of the extracted subset.
2. The information processing method of claim 1, wherein the dataset includes datasets of multiple domains, and the evaluating step includes the one or more processors extracting a learning subset from a dataset of one or more learning domains among the multiple domains under the specified conditions and training a learning model using the learning subset; extracting an evaluation subset from a dataset of one or more evaluation domains among the multiple domains that are different from the learning domain under the specified conditions and evaluating the learning model using the evaluation subset; and evaluating the domain generalizability using the evaluation results of the learning model.
3. The information processing method of claim 1, wherein the datasets include datasets of multiple domains, and the evaluating step includes, by the one or more processors: extracting a learning subset from the dataset of one or more learning domains among the multiple domains under the specified conditions and training a learning model using the learning subset; extracting a first evaluation subset different from the learning subset from the dataset of the learning domain under the specified conditions and evaluating the learning model using the first evaluation subset; extracting a second evaluation subset from the dataset of one or more evaluation domains among the multiple domains different from the learning domain under the specified conditions and evaluating the learning model using the second evaluation subset; and evaluating the domain generalizability using a difference between the evaluation result using the first evaluation subset and the evaluation result using the second evaluation subset.
4. The information processing method of claim 1, wherein the dataset includes datasets of multiple domains, and the evaluating step includes: one or more of the processors evaluating the degree of association between the feature of the subset and the target variable for each domain; and determining that the domain generalizability of the feature is relatively high when the degree of association is relatively high in more domains, and that the subset with more feature with relatively high domain generalizability has relatively high domain generalizability.
5. The information processing method according to claim 1, wherein the dataset includes datasets of multiple domains, and the evaluating step includes: evaluating, by the one or more processors, the degree of association between the features of the subset and the objective variable for each domain; and defining features having a degree of association equal to or greater than a certain threshold in a certain number of domains as features with relatively high domain generality, and evaluating the domain generalizability of the subset using the number of features with relatively high domain generality.
6. The information processing method of claim 1, wherein the dataset includes datasets of multiple domains, and the evaluating step includes: extracting, by the one or more processors, a learning subset from a dataset of one or more learning domains among the multiple domains under the specified conditions and extracting a feature set from the learning subset; extracting, by the one or more processors, an evaluation subset from a dataset of one or more evaluation domains among the multiple domains that are different from the learning domain under the specified conditions and evaluating the extracted feature set using the evaluation subset; and evaluating the domain generalizability of the learning subset using a proportion of features of the extracted feature set that are also valid features in the evaluation subset.
7. The information processing method of claim 6, wherein the one or more processors evaluate the degree of association between the features of the learning subset and the target variable for each domain, and determine that the domain generality of the features is relatively high the higher the degree of association in more domains, and the extracted feature set includes features with relatively high domain generality.
8. The information processing method of claim 6, wherein the one or more processors evaluate the degree of association between the features of the learning subset and the target variable for each domain, and determine that features having a degree of association equal to or greater than a certain threshold in a certain number of domains are features with relatively high domain universality, and the extracted feature set includes the features with relatively high domain universality.
9. The information processing method of claim 1, wherein the evaluating step includes: one or more of the processors evaluating the usefulness of the subset from known usefulness information for each sample in the dataset and the samples included in the subset; and combining the usefulness and the domain generalizability to form an evaluation of the subset.
10. The information processing method of claim 1, wherein the dataset includes datasets of multiple domains, and the evaluating step includes the one or more processors evaluating domain uniformity, which indicates the closeness of distribution between the number of data in the dataset for each domain and the number of data in the subset, and combining the domain uniformity and the domain generalizability to evaluate the subset.
11. The information processing method according to claim 1, further comprising: training, by the one or more processors, a subset classification model that classifies data in the dataset as being a subset or not.
12. The information processing method of claim 11, wherein the evaluating step includes: evaluating the subset classification performance of the subset classification model by the one or more processors; and combining the subset classification performance and the domain generalization to form an evaluation of the subset.
13. The information processing method of claim 1, wherein the extracting and evaluating includes the one or more processors repeatedly adding or deleting samples from a starting subset to the subset to search for a subset with higher domain generalization.
14. The information processing method according to claim 13, wherein the dataset includes datasets of multiple domains, and the searching includes the one or more processors searching by evaluating any one of the following: usefulness of the subset evaluated from known usefulness information for each sample in the dataset and the samples included in the subset; domain evenness indicating the closeness of distribution between the number of data items in the dataset and the number of data items in the subset for each domain; and subset classification performance of a subset classification model that classifies data items in the dataset as being subsets or not.
15. An information processing method according to any one of claims 1 to 14, wherein the one or more processors include: presenting a plurality of different subset conditions; evaluating the subsets extracted under each of the plurality of different subset conditions; and extracting a subset under the subset condition that provides the best evaluation result among the plurality of different subset conditions.
16. An information processing device comprising: one or more processors; and one or more memories storing instructions to be executed by the one or more processors, wherein the one or more processors extract a subset of a dataset to be analyzed under specified conditions; and evaluate the domain generalizability of the extracted subset.
17. A program that enables a computer to perform the following functions: extract a subset of a dataset to be analyzed based on specified conditions; and evaluate the domain generalizability of the extracted subset.
18. A non-transitory computer-readable recording medium on which the program according to claim 17 is recorded.