Information processing system, information processing method, information processing program, and molecular compound production method
The information processing system enhances the accuracy of molecular property predictions by selecting input feature amounts with consistent tendencies across training and test data, addressing the accuracy challenges faced by existing models.
Patent Information
- Application Number
- PCT/JP2023/045988
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-26
AI Technical Summary
Existing machine learning models for predicting molecular properties face challenges in accuracy due to differences in feature amount tendencies between training and test data.
An information processing system that selects input feature amounts by calculating an index based on probability distributions in training and test data, excluding feature amounts with significant differences in tendency, and using the remaining feature amounts as input parameters for a prediction model.
Improves the accuracy of machine learning models for predicting molecular characteristics by selecting feature amounts with similar tendencies in both training and test data.
Smart Images

Figure JP2023045988_26062025_PF_FP_ABST
Abstract
Description
Information processing system, information processing method, information processing program, and method for producing molecular compounds
[0001] One aspect of the present disclosure relates to an information processing system, an information processing method, an information processing program, and a method for producing a molecular compound.
[0002] Patent Literature 1 describes a method for identifying amino acid sequences of antibodies that have affinity for an antigen, which includes the steps of querying a machine learning engine for proposed amino acid sequences of antibodies that have high affinity for an antigen, and obtaining the proposed amino acid sequences from the machine learning engine.
[0003] Non-Patent Document 1 describes a method for selecting features for training data and test data. In this method, if the score of an adversarial classifier classifying training data and test data is higher than a predetermined threshold, features that are important in the classification are removed and the adversarial classifier is retrained. If the score falls below the threshold, a predictive model is trained using the remaining features.
[0004] International Publication No. 2018 / 132752
[0005] Pan, Jing, et al. "Adversarial validation approach to concept drift problem in user targeting automation systems at uber." arXiv:2004.03045 (2020).
[0006] It is desirable to improve the accuracy of machine learning models for predicting molecular properties.
[0007] An information processing system according to one aspect of the present disclosure includes at least one processor, which acquires, for each of a plurality of first molecules, training data indicating a plurality of feature quantities related to the first molecules, acquires, for each of a plurality of second molecules, test data indicating a plurality of feature quantities related to the second molecules, calculates an index for each of the plurality of feature quantities based on a first probability distribution that is a probability distribution in the training data and a second probability distribution that is a probability distribution in the test data, and selects, based on the respective indexes of the plurality of feature quantities, one or more of the plurality of feature quantities as one or more input feature quantities to be used as input parameters of a prediction model that predicts a characteristic value of the molecule based on machine learning.
[0008] An information processing method according to one aspect of the present disclosure is executed by an information processing system including at least one processor. The information processing method includes the steps of: acquiring, for each of a plurality of first molecules, training data indicating a plurality of feature quantities related to the first molecules; acquiring, for each of a plurality of second molecules, test data indicating a plurality of feature quantities related to the second molecules; calculating, for each of the plurality of feature quantities, an index based on a first probability distribution that is a probability distribution in the training data and a second probability distribution that is a probability distribution in the test data; and selecting, based on the respective indexes of the plurality of feature quantities, one or more of the plurality of feature quantities as one or more input feature quantities to be used as input parameters of a prediction model that predicts a characteristic value of the molecule based on machine learning.
[0009] An information processing program according to one aspect of the present disclosure causes a computer to execute the following steps: acquiring, for each of a plurality of first molecules, training data indicating a plurality of feature quantities related to the first molecule; acquiring, for each of a plurality of second molecules, test data indicating a plurality of feature quantities related to the second molecule; calculating, for each of the plurality of feature quantities, an index based on a first probability distribution that is a probability distribution in the training data and a second probability distribution that is a probability distribution in the test data; and selecting, based on the respective indexes of the plurality of feature quantities, one or more of the plurality of feature quantities as one or more input feature quantities to be used as input parameters of a prediction model that predicts a characteristic value of the molecule based on machine learning.
[0010] In this aspect, features to be used as input parameters of a machine learning-based prediction model are selected as input features based on the trends of individual features in both the training data and the test data. By selecting features in this manner, the accuracy of the machine learning model (prediction model) for predicting molecular properties can be improved.
[0011] According to one aspect of the present disclosure, the accuracy of machine learning models for predicting molecular properties can be improved.
[0012] FIG. 1 is a diagram for explaining selection of input feature quantities. FIG. 2 is a diagram for explaining an example of the functional configuration of an information processing system. FIG. 3 is a diagram for explaining selection of input feature quantities. FIG. 4 is a diagram for explaining selection of input feature quantities. FIG. 5 is a diagram for explaining selection of input feature quantities. FIG. 6 is a diagram for explaining selection of input feature quantities. FIG. 7 is a diagram for explaining selection of input feature quantities. FIG. 8 is a diagram for explaining selection of input feature quantities. FIG. 9 is a diagram for explaining selection of input feature quantities. FIG. 10 is a diagram for explaining selection of input feature quantities. FIG. 11 is a diagram for explaining selection of input feature quantities. FIG. 12 is a diagram for explaining selection of input feature quantities. FIG. 13 is a diagram for explaining selection of input feature quantities. FIG. 14 is a diagram for explaining selection of input feature quantities. FIG. 15 is a diagram for explaining selection of input feature quantities. FIG. 16 is a diagram for explaining selection of input feature quantities. FIG. 17 is a diagram for explaining selection of input feature quantities. FIG. 18 is a diagram for explaining selection of input feature quantities. FIG. 19 is a diagram for explaining selection of input feature quantities.
[0013] Various examples of the present disclosure will be described in detail below with reference to the accompanying drawings. In the description of the drawings, the same or equivalent elements are designated by the same reference numerals, and redundant description will be omitted.
[0014] [System Overview] The information processing system according to the present disclosure is a computer system that selects one or more feature quantities of a molecule to be used as input parameters of a prediction model that predicts a characteristic value of the molecule based on machine learning. In the present disclosure, the feature quantities used as input parameters are referred to as "input feature quantities." The information processing system excludes some feature quantities from a plurality of feature quantities that are candidates for input feature quantities, and selects one or more remaining feature quantities as one or more input feature quantities. The number of feature quantities is also referred to as the number of dimensions. The process of selecting one or more input feature quantities from a plurality of feature quantities can also be considered a process of reducing the number of dimensions of the feature quantities.
[0015] In one example, the information processing system performs machine learning using one or more selected input features to generate a prediction model. This process corresponds to the learning phase of machine learning. In one example, the information processing system inputs one or more input features of a molecule whose characteristic values are unknown into the generated prediction model to predict the characteristic value of the molecule. In the present disclosure, a molecule whose characteristic value is predicted by the information processing system is also referred to as a "target molecule." The information processing system may predict characteristic values of each of multiple target molecules and perform screening to select at least one target molecule from the multiple target molecules based on each predicted characteristic value. Predicting such characteristic values corresponds to the prediction phase or operation phase of machine learning. In the learning phase and prediction phase, the information processing system generates and uses a prediction model using one or more selected input features.
[0016] Machine learning is a technique for autonomously discovering laws or rules by repeatedly learning based on given information. A predictive model generated by machine learning is a machine learning model built using algorithms and data structures, and is also called a trained model. In one example, the predictive model is built using a neural network such as a convolutional neural network (CNN).
[0017] In the present disclosure, a feature refers to a numerical value that quantitatively represents a feature of a molecule. In one example, multiple feature values are set for each molecule. Each feature value may represent a feature regarding the relationship between multiple structural units that make up the molecule, a feature regarding the arrangement of the multiple structural units, or a feature regarding a specific structural unit.
[0018] In the present disclosure, a characteristic value refers to a numerical value that quantitatively represents a characteristic of a molecule. For example, the characteristic value may represent various characteristics related to the binding ability, affinity, pharmacological activity, physical properties, kinetics, or safety of a molecule.
[0019] The information processing system selects input features. A predictive model is generated by machine learning using training data related to molecules. Based on the predictive model, predicted values of the properties of the molecules represented by the test data are calculated. In this disclosure, the generated predictive model is also referred to as a "property prediction model."
[0020] In the present disclosure, an individual molecule represented by the training data is referred to as a "first molecule." An individual molecule represented by the test data is referred to as a "second molecule." The modalities of the first and second molecules may be the same or different. As an example, the first molecule is an antibody, and the second molecule is also an antibody. As another example, the first molecule is a cyclic peptide, and the second molecule is also a cyclic peptide. As yet another example, when affinity is predicted using the interaction energy between a ligand and a protein as a feature, the first molecule may be a small molecule, and the second molecule may be a cyclic peptide.
[0021] The first molecule group consists of a plurality of first molecules. The training data includes, for each of the plurality of first molecules, a plurality of feature quantities related to the first molecule. The second molecule group consists of a plurality of second molecules. The test data includes, for each of the plurality of second molecules, a plurality of feature quantities related to the second molecule. The plurality of first molecules included in the first molecule group may have the same or different modalities. The plurality of second molecules included in the second molecule group may have the same or different modalities.
[0022] For example, if the molecule is an antibody, training data representing the antibody sequence is used. In the case of molecules having sequence information, the first group of molecules may be referred to as a training sequence group, and the second group of molecules may be referred to as a prediction sequence group.
[0023] The training data includes, for each of a plurality of first molecules, a plurality of feature amounts of the first molecule and at least one characteristic value of the first molecule. The test data includes, for each of a plurality of second molecules, a plurality of feature amounts of the second molecule. If there are a relatively large number of feature amounts whose trends differ between the training data and the test data, the accuracy of the prediction model decreases. In the present disclosure, the information processing system selects one or more input feature amounts so that feature amounts whose trends differ between the training data and the test data do not affect machine learning and the prediction model. In one example, the information processing system selects feature amounts whose trends differ relatively little between the training data and the test data. In one example, the information processing system excludes feature amounts whose trends differ greatly between the training data and the test data.
[0024] In one example, the information processing system calculates, as a selection index, an inter-distribution distance, which is the distance between a first probability distribution, which is a probability distribution in the training data, and a second probability distribution, which is a probability distribution in the test data, for each of a plurality of feature quantities. The probability distribution of a certain feature quantity refers to a distribution that represents the probability of each value that the feature quantity takes. Since the probability distribution indicates the tendency of the feature quantity, a feature quantity in which the difference between the first probability distribution and the second probability distribution is relatively large can be said to be a feature quantity whose tendency differs between the training data and the test data. The selection index refers to an index for selecting input feature quantities.
[0025] FIG. 1 is a diagram for explaining the selection of input features. In this example, the information processing system calculates, for each of Q features, the inter-distribution distance, which is the distance between a first probability distribution 210 in the training data and a second probability distribution 220 in the test data, as a selection index. The information processing system selects a feature (e.g., feature F) that has a relatively large difference between the first probability distribution 210 and the second probability distribution 220. 1 , F 3 , F 4) and the remaining features (e.g., feature F 2 , F 5 , F Q ) are selected as input features. The selected individual input features are features in which the difference between the first probability distribution 210 and the second probability distribution 220 is relatively small and the trends are relatively similar. Therefore, a prediction model generated by machine learning based on the input features can predict the characteristic value of the target molecule with high accuracy.
[0026] 2 is a diagram showing the functional configuration of an information processing system 10 according to an example. In this example, the information processing system 10 accesses a database 20 that stores training data used for machine learning. The database 20 may be provided in a computer system separate from the information processing system 10, or may be a component of the information processing system 10. In one example, the information processing system 10 accesses the database 20 via a communication network such as the Internet or an intranet.
[0027] The information processing system 10 includes functional modules: a feature selection unit 11, a learning unit 12, and a prediction unit 13. The feature selection unit 11 is a functional module that selects one or more input features from a plurality of features. The learning unit 12 is a functional module that generates a prediction model 30 by machine learning based on the selected one or more input features. The prediction unit 13 is a functional module that executes predictions regarding target molecules using the generated prediction model 30.
[0028] FIG. 3 is a diagram showing an example of the hardware configuration of a computer 100 functioning as the information processing system 10. For example, the computer 100 includes a processor 101, a main memory unit 102, an auxiliary memory unit 103, a communication control unit 104, an input device 105, and an output device 106. The processor 101 executes an operating system and application programs. The main memory unit 102 is composed of, for example, ROM and RAM. The auxiliary memory unit 103 is composed of, for example, a hard disk or flash memory, and generally stores larger amounts of data than the main memory unit 102. The communication control unit 104 is composed of, for example, a network card or a wireless communication module. The input device 105 is composed of, for example, a keyboard, a mouse, a touch panel, etc. The output device 106 is composed of, for example, a monitor and speakers.
[0029] Each functional module of the information processing system 10 is realized by an information processing program 110 pre-stored in the auxiliary storage unit 103. Specifically, each functional module is realized by loading the information processing program 110 onto the processor 101 or the main storage unit 102 and causing the processor 101 to execute the information processing program 110. The processor 101 operates the communication control unit 104, the input device 105, or the output device 106 in accordance with the information processing program 110, and reads and writes data from and to the main storage unit 102 or the auxiliary storage unit 103. Data, prediction models, or databases required for processing may be stored in the main storage unit 102 or the auxiliary storage unit 103.
[0030] The information processing program 110 may be provided in the form of being recorded on a non-transitory computer-readable storage medium such as a CD-ROM, a DVD-ROM, a semiconductor memory, etc. Alternatively, the information processing program 110 may be provided via a communication network as a data signal superimposed on a carrier wave.
[0031] The information processing system 10 may be configured with one computer 100 or multiple computers 100. When multiple computers 100 are used, these computers 100 are connected via a communication network such as the Internet or an intranet, thereby logically constructing a single information processing system 10.
[0032] FIG. 4 is a diagram illustrating an example of training data stored in the database 20. The training data is composed of multiple data records corresponding to multiple molecules. In one example, each data record includes, as data items, a molecule ID, which is an identifier that uniquely identifies each molecule, multiple feature quantities of the molecule, and a characteristic value of the molecule. When the molecules are proteins, the multiple feature quantities used as training data may be obtained, for example, by the Tasks Assessment Protein Embeddings method (TAPE). In the example of FIG. 4, the database 20 stores R data records, each representing Q feature quantities (i.e., Q-dimensional feature quantities). The number of dimensions Q of the feature quantities may be less than 10, or may be on the order of tens, hundreds, thousands, or tens of thousands. The information processing system 10 excludes some of the Q feature quantities and selects one or more remaining feature quantities as one or more input feature quantities. In the example of FIG. 4, physical property values are shown as the characteristic values of molecules, but the characteristic values are not limited to physical property values.
[0033] [System Operation] FIG. 5 is a flowchart illustrating an example of processing executed by the information processing system 10 as a processing flow S1. The processing flow S1 is an example of an information processing method according to the present disclosure. In step S10, the feature selection unit 11 selects one or more input features to be used for machine learning from a plurality of features indicated by training data. In one example, the feature selection unit 11 calculates evaluation indices for a plurality of cases in which some features are selected or excluded, while changing the selected or excluded features. The feature selection unit 11 then selects, as input features, a set of features for which the evaluation indices satisfy a predetermined condition. In step S20, the learning unit 12 performs machine learning based on the selected one or more input features to generate a prediction model 30. In step S30, the prediction unit 13 performs prediction using the prediction model 30. Step S10 corresponds to preprocessing in the learning phase, step S20 corresponds to the learning phase, and step S30 corresponds to the prediction phase. The one or more input features selected in step S10 are used in steps S20 and S30.
[0034] (Selection of Input Feature Amount) FIG. 6 is a flowchart showing in detail an example of the process of selecting input feature amounts, that is, an example of step S10.
[0035] In step S11, the feature selection unit 11 acquires training data and test data from the database 20. The training data includes, for each of a plurality of first molecules, a plurality of feature amounts related to the first molecule. The test data includes, for each of a plurality of second molecules, a plurality of feature amounts related to the second molecule.
[0036] In step S12, the feature selection unit 11 calculates a selection index for selecting input features for each of the multiple features based on a first probability distribution in the training data and a second probability distribution in the test data. In one example, the feature selection unit 11 calculates, as a selection index, a distribution distance between the first probability distribution and the second probability distribution for each feature. The distribution distance indicates the degree of difference between the first probability distribution and the second probability distribution. The feature selection unit 11 may use integral probability metrics (IPMs) as the distribution distance. Examples of IPMs that the feature selection unit 11 may calculate include the Wasserstein distance, the maximum mean discrepancy (MMD), or the Dudley metric.
[0037] The feature quantity selection unit 11 may calculate a first probability distribution and a second probability distribution in order to calculate the selection index. i The number of data records in the training data is j, and the number of data records in the test data is k. The feature selection unit 11 selects the feature F i For j, a probability distribution of j values may be calculated as a first probability distribution, and a probability distribution of k values may be calculated as a second probability distribution.
[0038] In step S13, the feature selection unit 11 initializes the feature exclusion number n. The exclusion number n is the number of features to be excluded from the multiple features based on the selection indexes for each of the multiple features. The exclusion number n can also be considered the number of feature dimension reductions. The feature selection unit 11 changes the exclusion number n in later processing. When the exclusion number n is incremented, the initial value of the reduction number n may be 0 or a value equal to or greater than 1. When the exclusion number n is decremented, the initial value of the reduction number n may be {(total number of features)-1} or a smaller number.
[0039] In step S14, the feature selector 11 excludes n feature quantities from the plurality of feature quantities based on the selection index for each feature quantity. When a distribution distance such as the Wasserstein distance is used, the feature selector 11 excludes n feature quantities in descending order of the distribution distance. In other words, the feature selector 11 excludes n feature quantities for which the difference between the first probability distribution and the second probability distribution is relatively large.
[0040] In step S15, the feature selector 11 generates and evaluates a provisional prediction model based on one or more features remaining after the exclusion process. In the present disclosure, the generation and evaluation of the provisional prediction model is also referred to as an "evaluation process." The provisional prediction model is not the prediction model 30 generated by the learning unit 12 and used by the prediction unit 13, but a machine learning model (trained model) temporarily used to select one or more input features. The feature selector 11 may use cross-validation or a holdout method as a method for the evaluation process.
[0041] Cross-validation will be described as an example of step S15. In cross-validation, the feature selection unit 11 divides the training data into multiple groups, selects one of the multiple groups as validation data, and selects the remaining group as narrow-sense training data. Each divided group is also called a "fold." A combination of validation data and narrow-sense training data is also called a "split." The feature selection unit 11 generates a provisional prediction model using the narrow-sense training data and evaluates the provisional prediction model using the validation data. The feature selection unit 11 generates and evaluates a provisional prediction model while changing the group (fold) used as validation data, and calculates an evaluation index of the provisional prediction model for each split. The feature selection unit 11 obtains statistics of multiple evaluation indexes obtained from multiple splits as evaluation indexes for one cross-validation.
[0042] FIG. 7 is a diagram for explaining cross-validation. In the example shown in this figure, the feature selection unit 11 divides the training data into five groups. For the first split, the feature selection unit 11 selects group "Fold 1" as validation data and selects the remaining four groups as narrowly defined training data. The feature selection unit 11 generates a provisional prediction model by machine learning using the narrowly defined training data, and calculates an evaluation index E of the provisional prediction model using the validation data. 1 The feature selection unit 11 generates and evaluates a provisional prediction model for the second to fifth splits while changing the validation data, and calculates four evaluation indices E 2 , E 3 , E 4 , E 5 The feature selection unit 11 calculates five evaluation indices E 1 ~E 5 The statistic is calculated as the final evaluation index of one cross-validation. Examples of the statistic include the mean and the median, but the statistic is not limited to these.
[0043] 8 is a flowchart illustrating an example of cross-validation. In step S151, the feature selection unit 11 sets a split for cross-validation. The feature selection unit 11 divides the training data acquired from the database 20 into multiple groups, selects one of the multiple groups as validation data, and selects the remaining groups as narrowly defined training data.
[0044] In step S152, the feature selection unit 11 performs machine learning based on one or more features remaining after the exclusion process to generate a provisional prediction model. The feature selection unit 11 performs machine learning (supervised learning) using narrow-sense training data to generate a provisional prediction model. The feature selection unit 11 inputs the remaining one or more features as input parameters (e.g., input vectors) to the machine learning model without using the excluded n features. In one example, the feature selection unit 11 performs backpropagation (error backpropagation) based on the error between the predicted value calculated by the machine learning model and the correct answer (label) to update the parameter set in the machine learning model. The learning unit 12 repeats this process until a given termination condition is met to obtain a provisional prediction model. The termination condition may be that all data records of the narrow-sense training data are processed.
[0045] In step S153, the feature selection unit 11 calculates an evaluation index for the provisional prediction model based on the one or more remaining feature values. This evaluation index is a value indicating how accurately the provisional prediction model can calculate the molecular characteristic values. For each data record of the validation data, the feature selection unit 11 inputs the remaining one or more feature values to the provisional prediction model as input parameters (e.g., input vectors) without using the excluded n feature values. For each data record, the provisional prediction model calculates a characteristic value based on the input parameters. Hereinafter, the characteristic value calculated by the provisional prediction model is also referred to as a predicted characteristic value.
[0046] The feature selection unit 11 calculates an evaluation index for the provisional prediction model based on the predicted characteristic values and correct answers (labels) for each data record in the validation data. The correct answers (labels) in this process are characteristic values included in the training data used as validation data. For example, the feature selection unit 11 may calculate, as the evaluation index, the mean square error between the predicted characteristic values and the correct answers, or a correlation coefficient indicating the degree of correlation between the predicted characteristic values and the correct answers.
[0047] As shown in step S154, if there is a split that has not been executed (NO in step S154), the process returns to step S151. In the repeated step S151, the feature selection unit 11 changes the group (fold) used as validation data and sets the next split. In the repeated step S152, the feature selection unit 11 performs machine learning (supervised learning) using narrow-sense training data based on the one or more remaining features to generate a provisional prediction model. In the repeated step S153, the feature selection unit 11 calculates an evaluation index for the provisional prediction model based on the one or more remaining features.
[0048] If all splits have been processed (YES in step S154), the process proceeds to step S155. In step S155, the feature selection unit 11 calculates statistics of the evaluation indices for each split as the final evaluation indices of the cross-validation. For example, the feature selection unit 11 calculates the average value of the mean squared errors or the average value of the correlation coefficients as the final evaluation indices.
[0049] 6, in step S16, the feature quantity selection unit 11 determines whether or not to change the exclusion number n. The exclusion number n may be changed by incrementing or decrementing.
[0050] In one example, the feature selection unit 11 determines whether to increment the exclusion number n based on the number of features excluded by the exclusion process or the number of features not excluded by the exclusion process. For example, the feature selection unit 11 determines to increment n if the number of effective dimensions, which is the number of features remaining after the exclusion process, is equal to or greater than a predetermined threshold, and determines not to increment n if the number of effective dimensions is less than the threshold. Alternatively, the feature selection unit 11 determines to increment n if the number of effective dimensions is greater than a predetermined threshold, and determines not to increment n if the number of effective dimensions is equal to or less than the threshold. In another example, the feature selection unit 11 determines to increment n if the number of features excluded by the exclusion process is equal to or less than a predetermined threshold, and determines not to increment n if the number of features is greater than the threshold. Alternatively, the feature selection unit 11 determines to increment n if the number of features excluded by the exclusion process is smaller than a predetermined threshold, and determines not to increment n if the number of features is equal to or greater than the threshold.
[0051] In one example, the feature selection unit 11 determines whether to decrement the exclusion number n based on the number of features excluded by the exclusion process or the number of features not excluded by the exclusion process. For example, the feature selection unit 11 determines to decrement n when the effective number of dimensions, which is the number of features remaining after the exclusion process, is equal to or less than a predetermined threshold, and determines not to decrement n when the effective number of dimensions is greater than the threshold. Alternatively, the feature selection unit 11 determines to decrement n when the effective number of dimensions is less than a predetermined threshold, and determines not to decrement n when the effective number of dimensions is equal to or greater than the threshold. In another example, the feature selection unit 11 determines to decrement n when the number of features excluded by the exclusion process is equal to or greater than a predetermined threshold, and determines not to decrement n when the number of features is less than the threshold. Alternatively, the feature selection unit 11 determines to decrement n if the number of features excluded by the exclusion process is greater than a predetermined threshold, and determines not to decrement n if the number of features is equal to or less than the threshold.
[0052] If the exclusion number n is to be changed (YES in step S16), the process proceeds to step S17. In step S17, the feature selection unit 11 changes the exclusion number n by a predetermined number. This change is an increment or decrement. The predetermined number is any natural number. For example, the feature selection unit 11 increases the exclusion number n by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. As another example, the feature selection unit 11 increases the number of excluded features by the same number as the initial value of n. Alternatively, the feature selection unit 11 decreases the exclusion number n by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. After step S17, the process returns to step S14. In the repeated step S14, the feature selection unit 11 excludes n feature quantities from the plurality of feature quantities based on the selection index for each feature quantity. In step S15, the feature selection unit 11 generates and evaluates a provisional prediction model based on the one or more remaining features. The feature selection unit 11 calculates a final evaluation index for the provisional prediction model for each of the exclusion numbers n while gradually increasing or decreasing the exclusion number n.
[0053] If the number of exclusions n is not to be changed (NO in step S16), the process proceeds to step S18. In step S18, the feature selector 11 selects one or more input feature quantities from the plurality of feature quantities. In one example, the feature selector 11 selects the final number of exclusions n based on the final evaluation indexes obtained for each of the number of exclusions. FINAL The feature selection unit 11 determines the number of exclusions that yields the best final evaluation index as the number of exclusions n FINAL For example, the feature quantity selection unit 11 may determine the number of exclusions that results in the highest or lowest final evaluation index as the number of exclusions n FINAL For example, when the statistic of the mean square error is the evaluation index, the feature quantity selection unit 11 may determine the number of exclusions that gives the lowest final evaluation index as the number of exclusions n FINAL In addition, when the statistical quantity of the correlation coefficient is the evaluation index, the feature quantity selection unit 11 may determine the number of excluded features that has the highest final evaluation index as the number of excluded features n FINALThat is, the feature quantity selection unit 11 may determine the number of exclusions n to be the number of exclusions with the smallest statistical value of mean square error or the number of exclusions with the highest statistical value of correlation coefficient as the number of exclusions n FINAL Alternatively, the feature quantity selection unit 11 may select one of the one or more exclusion numbers that has a higher final evaluation index than when all of the plurality of feature quantities are used as the exclusion number n FINAL Alternatively, the feature quantity selection unit 11 may select one of the one or more exclusion numbers for which the extent of the final decline in the evaluation index compared to when all of the plurality of feature quantities are used is less than a predetermined threshold, as the exclusion number n FINAL The feature quantity selection unit 11 may determine n features from among a plurality of feature quantities based on the selection index for each feature quantity. FINAL When a distance between distributions such as the Wasserstein distance is used, the feature selection unit 11 selects n feature quantities in descending order of the distance between distributions. FINAL feature quantities are excluded, and one or more remaining feature quantities are selected as one or more input feature quantities. In this case, the feature quantity selection unit 11 selects one or more feature quantities whose inter-distribution distance is smaller than a predetermined criterion determined based on the final evaluation index as one or more input feature quantities. The predetermined criterion may be determined by a user or administrator of the information processing system 10, or may be determined automatically by any method such as machine learning.
[0054] As described above, the feature selection unit 11 excludes some of the multiple feature quantities based on the selection index for each of the multiple feature quantities. The feature selection unit 11 selects one or more feature quantities remaining after the exclusion process as one or more input feature quantities. In one example, the feature selection unit 11 calculates, for each of the multiple feature quantities, a distribution distance that is the distance between the first probability distribution and the second probability distribution, and selects one or more input feature quantities from the multiple feature quantities based on the distribution distance for each of the multiple feature quantities. In one example, the feature selection unit 11 calculates, for each of the multiple feature quantities, the distribution distance that is the distance between the first probability distribution and the second probability distribution as a selection index, and selects, as input feature quantities, feature quantities for which the distribution distance for each of the multiple feature quantities is relatively small.
[0055] As described above, in one example, the feature selector 11 repeats the evaluation process of generating a provisional prediction model and calculating an evaluation index for the provisional prediction model based on one or more remaining features while changing the number of feature quantities to be excluded from the plurality of feature quantities based on the selection indexes for each of the plurality of feature quantities (repeating steps S14 to S17). The feature selector 11 may perform the evaluation process by cross-validation. The feature selector 11 determines the number of exclusions based on the plurality of evaluation indexes obtained by the repetition (step S18). The feature selector 11 determines the number of exclusions n from the plurality of feature quantities based on the selection indexes for each of the plurality of feature quantities. FINAL Then, one or more input feature quantities are selected (step S18).
[0056] Another example of the process of selecting input features will be described with reference to FIG. 9 . FIG. 9 is a flowchart showing this example in detail as step S10A. Step S10A can also be considered a modification of step S10. Step S10A differs from step S10 in that the generation and evaluation of a provisional prediction model are repeated while changing the number of selected features m instead of the number of excluded features n. This difference will be particularly described below.
[0057] As in step S10, the feature selection unit 11 executes the processes of steps S11 and S12.
[0058] In step S13A, the feature quantity selection unit 11 initializes the selection number m of selection quantities. The selection number m is the number of feature quantities selected from a plurality of feature quantities based on the selection indexes for each of the plurality of feature quantities. When the selection number m is incremented, the initial value of the selection number m may be 1 or a value of 2 or greater. When the selection number m is decremented, the initial value of the selection number m may be the total number of feature quantities or a smaller number.
[0059] In step S14A, the feature selection unit 11 selects m feature quantities from the plurality of feature quantities based on the selection index for each feature quantity. When a distribution distance such as the Wasserstein distance is used, the feature selection unit 11 selects m feature quantities in ascending order of the distribution distance. In other words, the feature selection unit 11 selects m feature quantities for which the difference between the first probability distribution and the second probability distribution is relatively small.
[0060] In step S15A, the feature selection unit 11 generates and evaluates a tentative prediction model based on the one or more selected features. That is, the feature selection unit 11 performs an evaluation process. The "one or more selected features" are substantially the same as the "one or more remaining features" in step S10. Therefore, in step S15A, cross-validation such as that shown in FIG. 8 can also be performed.
[0061] In step S16A, the feature selection unit 11 determines whether to change the selection number m. The selection number m may be changed by incrementing or decrementing.
[0062] In one example, the feature selection unit 11 determines whether to increment the selection number m based on the number of selected features or the number of unselected features. In another example, the feature selection unit 11 determines whether to decrement the selection number m based on the number of selected features or the number of unselected features.
[0063] If the selection number m is to be changed (YES in step S16A), the process proceeds to step S17A. In step S17A, the feature selection unit 11 changes the selection number m by a predetermined number. This change is an increment or decrement. The predetermined number is any natural number. For example, the feature selection unit 11 increases the selection number m by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. As another example, the feature selection unit 11 increases the number of selected features by the same number as the initial value of m. Alternatively, the feature selection unit 11 decreases the selection number m by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. After step S17A, the process returns to step S14A. In the repeated step S14A, the feature selection unit 11 selects m feature quantities from the plurality of feature quantities based on the selection index for each feature quantity. In the repeated step S15A, the feature selection unit 11 generates and evaluates a provisional prediction model based on one or more selected feature quantities. The feature selection unit 11 calculates a final evaluation index of the provisional prediction model for each selection number m while gradually increasing or decreasing the selection number m.
[0064] If the selection number m is not changed (NO in step S16A), the process proceeds to step S18. In step S18, the feature selection unit 11 selects one or more input feature quantities from the plurality of feature quantities. In one example, the feature selection unit 11 selects the final selection number m based on the final evaluation index obtained for each selection number. FINAL The feature selection unit 11 determines m from a plurality of feature quantities based on the selection index of each feature quantity. FINAL When a distance between distributions such as the Wasserstein distance is used, the feature selection unit 11 selects m feature quantities in ascending order of the distance between distributions. FINAL The feature selection unit 11 selects, as input features, one or more feature quantities whose inter-distribution distance is smaller than a predetermined criterion determined based on the final evaluation index. As described above, the predetermined criterion can be determined manually or automatically.
[0065] As described above, the feature selection unit 11 repeats the evaluation process of generating a provisional prediction model and calculating an evaluation index for the provisional prediction model based on one or more selected features while changing the number of features selected from the plurality of features based on the selection indexes for each of the plurality of features (repeating steps S14A to S17A). The feature selection unit 11 may perform the evaluation process by cross-validation. The feature selection unit 11 selects one or more input features based on the plurality of evaluation indexes obtained by the repetition.
[0066] (Learning Phase) The process of generating the prediction model 30 based on the one or more selected input features, i.e., step S20, will be described. The learning unit 12 performs machine learning (supervised learning) using one or more input features of the training data acquired by the feature selection unit 11 to generate the prediction model 30. In this machine learning, the feature selection unit 11 selects the n input features of the training data that have been excluded. FINAL Instead of using individual features, one or more input features of the training data are input to the machine learning model as input parameters (e.g., an input vector). In one example, the feature selection unit 11 performs backpropagation based on the error between the predicted value calculated by the machine learning model and the correct answer (label) to update the parameter set in the machine learning model. The learning unit 12 repeats this process until a given termination condition is met to obtain the prediction model 30. The termination condition may be that all data records of the training data (first molecule group) have been processed. It should be noted that the generated prediction model 30 is a computational model estimated to be optimal, and is not necessarily a "computational model that is actually optimal."
[0067] (Prediction Phase) Prediction using the prediction model 30, i.e., step S30, will now be described. The prediction unit 13 acquires one or more input features for each of one or more target molecules. Each target molecule may be a molecule designated by an input operation or a selection operation by a user of the information processing system 10. The prediction unit 13 may read the input features for each target molecule from a predetermined database or may receive the input features from another computer such as a user terminal. The prediction unit 13 inputs the one or more input features of the target molecule into the prediction model 30. If one or more additional features are acquired for the target molecule in addition to the one or more input features, the prediction unit 13 inputs the one or more input features to the prediction model 30 as input parameters (e.g., input vectors) without using the one or more additional features. The prediction model 30 calculates a characteristic value of the target molecule based on the input features, and the prediction unit 13 acquires the characteristic value from the prediction model 30. The prediction unit 13 uses the prediction model 30 in this manner to acquire a characteristic value for each of one or more target molecules.
[0068] The prediction unit 13 may output the characteristic values of each of one or more target molecules as a prediction result. When characteristic values for each of multiple target molecules are acquired, the prediction unit 13 may select at least one target molecule from the multiple target molecules based on these characteristic values. For example, the prediction unit 13 selects a target molecule whose characteristic value meets a predetermined standard. Such selection can also be considered screening. The prediction unit 13 may output information about the selected at least one target molecule as a prediction result. The information may include at least one of the name, structure, sequence information, and characteristic value of the target molecule.
[0069] In step S30, the prediction unit 13 may process at least one second molecule of the multiple second molecules as a target molecule. The prediction unit 13 may acquire one or more input feature amounts for each target molecule, and output a characteristic value of the target molecule obtained by inputting the acquired one or more input feature amounts for each target molecule into the generated prediction model. When characteristic values have been acquired for each of the multiple second molecules, the prediction unit 13 may select at least one target molecule of the multiple second molecules as a candidate molecule based on these characteristic values, and output information on the selected at least one candidate molecule.
[0070] The prediction unit 13 may store the prediction result in a storage device such as the auxiliary storage unit 103, may display the prediction result on the output device 106, or may transmit the prediction result to another computer such as a user terminal.
[0071] [Method for Producing Molecular Compounds] Based on the information on at least one target molecule or candidate molecule output in the prediction phase, a molecular compound having the molecular sequence of the at least one target molecule or candidate molecule may be generated.
[0072] When the molecular compound is an antibody, the antibody can be produced using recombinant methods or constructs, for example, as described in U.S. Patent No. 4,816,567. An example of a production method is a method for producing an antibody, comprising culturing a host cell containing a nucleic acid encoding the antibody under conditions suitable for expression of the candidate molecular compound antibody described therein, or recovering the antibody from the host cell (or host cell culture medium). The isolated nucleic acid encoding the antibody may encode an amino acid sequence comprising the VL and / or an amino acid sequence comprising the VH of the antibody (e.g., the light chain and / or the heavy chain of the antibody). A host cell containing such a nucleic acid contains (1) a vector containing a nucleic acid encoding an amino acid sequence comprising the VL and an amino acid sequence comprising the VH of the antibody, or (2) a first vector containing a nucleic acid encoding an amino acid sequence comprising the VL of the antibody and a second vector containing a nucleic acid encoding an amino acid sequence comprising the VH of the antibody (e.g., the host cell is transformed). In one example, the host cell is eukaryotic (e.g., Chinese hamster ovary (CHO) cells, or lymphoid cells (e.g., Y0, NS0, Sp2 / 0 cells)). Suitable host cells for cloning or expressing antibody-encoding vectors include prokaryotic or eukaryotic cells. For example, antibodies may be produced in bacteria, particularly if glycosylation and Fc effector functions are not required. For the expression of antibody fragments and polypeptides in bacteria, see, e.g., U.S. Patent Nos. 5,648,237, 5,789,199, and 5,840,523. Additionally, for the expression of antibody fragments in E. coli, see Charlton, Methods in Molecular Biology, Vol. 248 (BKC Lo, ed., Humana Press, Totowa, NJ, 2003), pp. 245-254. Following expression, the antibody may be isolated in a soluble fraction from the bacterial cell paste and further purified.
[0073] When the molecular compound is a peptide compound or a cyclic peptide compound, such a compound can be produced by liquid phase synthesis, solid phase synthesis using Fmoc synthesis, Boc synthesis, or the like, or a combination of these. Solid phase synthesis is a method in which a compound is bound to a solid and then the compound is chemically reacted with a reagent on the solid resin to synthesize a target compound. Solid phase peptide synthesis is a method in which a desired amino acid or peptide is bound to a solid resin, and further desired amino acids or peptides are sequentially linked to the amino acids or peptides bound to the solid resin to elongate the peptide chain, thereby synthesizing a peptide. The target peptide is obtained by cleaving the peptide bound to the solid resin from the solid resin.
[0074] [Molecules] Molecules may be low molecular weight (low molecular weight compounds), medium molecular weight (medium molecular weight compounds), or polymers (polymer compounds). In the present disclosure, low molecular weight or low molecular weight compounds refer to compounds with a molecular weight of less than 500 g / mol. In the present disclosure, medium molecular weight or medium molecular weight compounds refer to compounds with a molecular weight of 500 g / mol or more and less than 30,000 g / mol. In the present disclosure, polymers or polymer compounds refer to compounds with a molecular weight of 30,000 g / mol or more.
[0075] The molecule may be a biomolecule or a non-biomolecule. The molecule may be an antigen-binding molecule such as a nucleic acid, peptide, cyclic peptide, protein, or antibody, or may be a molecule that binds to a target molecule (target molecule-binding molecule). The molecule may also be a drug candidate molecule. When the molecule is a peptide, cyclic peptide, protein, or antibody, the building blocks of the molecule are amino acids. When the molecule is a nucleic acid, the building blocks of the molecule are nucleosides or nucleotides.
[0076] In the present disclosure, the desired property is a property required for a new target suitable for a drug candidate, and may be set arbitrarily. Examples of properties include, but are not limited to, binding ability to a predetermined in vivo target, pharmacological activity, physical properties, kinetics, and safety. Examples of physical properties include thermal stability, chemical stability, solubility, viscosity, light stability, long-term storage stability, nonspecific adsorption, lipid solubility, and membrane permeability. For example, if the molecule is messenger RNA (mRNA), the property is the translation ability of the protein. If the molecule is a molecule that binds to a target molecule, such as an antigen-binding molecule, the property may be binding ability to the target molecule.
[0077] Antigen-binding molecules: In the present disclosure, the term "antigen-binding molecule containing an antigen-binding domain" is used in the broadest sense. Specifically, antigen-binding molecules include various molecular types as long as they contain an antigen-binding domain. An antigen-binding molecule may be a molecule consisting of only an antigen-binding domain, or may be a molecule containing an antigen-binding domain and other domains. For example, when an antigen-binding molecule is a molecule in which an antigen-binding domain is linked to an Fc region, examples include complete antibodies and antibody fragments. Antibodies may include single monoclonal antibodies (including agonist and antagonist antibodies), human antibodies, humanized antibodies, chimeric antibodies, etc. The antigen-binding molecules of the present disclosure may also include scaffold molecules in which a pre-existing stable α / β barrel protein structure or other three-dimensional structure is used as a scaffold, and only a partial structure of the scaffold is compiled into a library for constructing an antigen-binding domain.
[0078] - Binding Ability Evaluation The method for evaluating the binding ability of a target molecule-binding molecule to a target molecule is not particularly limited. Binding ability evaluation is possible by quantitatively evaluating the binding of a target molecule-binding molecule to a target molecule. The target molecule is, for example, a target protein. The target molecule-binding molecule is, for example, an antigen-binding molecule, and the target molecule is, for example, an antigen. For example, when the target molecule is an antigen, evaluation can be performed by measuring the binding activity between the antigen-binding molecule and the antigen. Binding activity refers to the total strength of non-covalent interactions between one or more binding sites of a molecule (e.g., an antibody) and the binding partner of the molecule (e.g., an antigen). Here, "binding activity" is not strictly limited to a 1:1 interaction between members of a binding pair (e.g., an antibody and an antigen). For example, when the members of a binding pair reflect a monovalent 1:1 interaction, binding activity refers to the intrinsic binding affinity (sometimes simply referred to as "affinity"). When the members of a binding pair are capable of both monovalent and multivalent binding, the binding activity is the sum of their binding strengths. The binding activity of molecule X to partner Y is generally expressed by the dissociation constant (KD) or "the amount of analyte bound per unit amount of ligand." For example, the octet value is an index of binding ability and is measured as the amount of analyte bound per unit amount of ligand. Binding activity can be measured by conventional methods known in the art, including those described herein. Conditions other than the concentration of the target tissue-specific compound can be appropriately determined by those skilled in the art.
[0079] In one embodiment, the binding activity of an antibody is measured using surface plasmon resonance analysis, for example, a ligand capture method using BIACORE (registered trademark) T200 or BIACORE (registered trademark) 4000 (GE Healthcare, Uppsala, Sweden).
[0080] In one embodiment, the measurement results are analyzed using BIACORE® Evaluation Software. Kinetic parameters are calculated by simultaneously fitting the binding and dissociation sensorgrams using a 1:1 binding model. This process allows the binding rate (k or k), dissociation rate (k or k), and equilibrium dissociation constant (K) to be calculated.
[0081] As the value of antigen-binding activity, KD (dissociation rate constant) can be used when the antigen is a soluble molecule, and apparent kD (apparent dissociation rate constant) can be used when the antigen is a membrane-type molecule. kD (dissociation rate constant) and apparent KD (apparent dissociation rate constant) can be measured by methods known to those skilled in the art. For example, Biacore (GE healthcare), a flow cytometer, etc. can be used.
[0082] Another aspect of characterization is a method for selecting antigen-binding molecules using a display library. One aspect includes panning using phage display. Taking affinity evaluation as an example, a phage library displaying multiple different antigen-binding molecules is prepared, and the target antigen is contacted with the prepared phages. After contacting the target antigen with the prepared phages, unbound phages are washed away, allowing phages displaying antigen-binding molecules that interact with the target antigen to be enriched. Sequences with affinity for the target antigen can be identified by analyzing the nucleic acid sequences encoding the antigen-binding molecules contained in the enriched phages. One aspect includes panning using mammalian cell display. In pharmacological activity evaluation using this display system, a library containing multiple different antigen-binding molecules is expressed in target mammalian cells, and reporter activity, etc., is altered depending on the effect the molecules have on the same cells, allowing cells carrying antigen-binding molecule genes with the desired pharmacological activity to be isolated using a flow cytometer or the like. In property evaluation using this display system, a library containing multiple different antigen-binding molecules is expressed in target mammalian cells, and the expression level is stained with an antibody specific to the antigen-binding molecule, allowing cells carrying antigen-binding molecule genes that can be stably and highly expressed to be isolated using a flow cytometer or the like. Characterization of antigen-binding molecules by panning is not limited to the above-mentioned methods using phages or mammalian cells, and various methods may be used as long as they can display antigen-binding molecules. For example, methods using ribosomes, mRNA, viruses other than phages, or bacteria such as Escherichia coli may also be used.
[0083] Other aspects of characterization include methods for obtaining antibody gene sequences from immune cells derived from an individual, or methods for obtaining antibody protein sequences from serum. In affinity evaluations in which antibody gene sequences are extracted from immune cells, a target antigen protein is administered to an individual to induce immune sensitization, and genes are extracted from immune cells that have antibody genes that bind to the target antigen, allowing sequences with affinity for the target antigen to be identified.
[0084] In addition to the above-mentioned methods using proteins as antigens that induce immune sensitization, methods using genes encoding the proteins or cells that express the proteins may also be used.
[0085] Examples of subject individuals include, but are not limited to, humans, mice, rats, hamsters, rabbits, monkeys, chickens, camels, llamas, and alpacas.
[0086] Methods for analyzing the above-mentioned nucleic acid sequences or occurrence frequencies include, but are not limited to, a method in which a genetically modified organism having the nucleic acid sequence of each antigen-binding molecule is cloned and analyzed by the Sanger method using capillary electrophoresis, and a method in which the nucleic acid sequence is analyzed using a next-generation sequencer.
[0087] When analyzing the above nucleic acid sequences, the strength of the characteristics may be determined based on the frequency of occurrence. For example, by analyzing the nucleic acid sequences after enrichment, antigen-binding molecules encoded by sequences with high frequency of occurrence can be predicted to have high characteristics. On the other hand, antigen-binding molecules encoded by sequences with low frequency of occurrence after enrichment can be predicted to have lower characteristics than antigen-binding molecules encoded by sequences with high frequency of occurrence.
[0088] The above-mentioned display library or the method for obtaining information on antigen-binding molecules derived from an individual can be applied to various characteristic evaluations other than those described above.
[0089] - Pharmacological activity evaluation The method for evaluating the pharmacological activity of a molecule is not particularly limited. For example, the pharmacological activity can be evaluated by measuring the neutralizing activity, agonist activity, or cytotoxic activity exhibited by the molecule. Examples of cytotoxic activity evaluation, which is one type of pharmacological activity evaluation, include antibody-dependent cell-mediated cytotoxicity (ADCC) activity, complement-dependent cytotoxicity (CDC) activity, T-cell-dependent cytotoxicity (TDCC) activity, and antibody-dependent cellular phagocytosis (ADCP) activity. CDC activity refers to cytotoxic activity mediated by the complement system. ADCC activity refers to the activity of immune cells binding to the Fc region of an antigen-binding molecule comprising an antigen-binding domain that binds to a membrane-type molecule expressed on the cell membrane of a target cell via an Fcγ receptor expressed on the immune cell, causing damage to the target cell. TDCC activity refers to the activity of T cells damaging the target cell by bringing the target cell and T cell into close proximity using a bi-specific antibody comprising an antigen-binding domain that binds to a membrane-type molecule expressed on the cell membrane of the target cell and an antigen-binding domain (particularly an antigen-binding domain that binds to the CD3 epsilon chain) that binds to either a constituent subunit of the T cell receptor (TCR) complex on the T cell. Whether an antigen-binding molecule of interest has ADCC activity, CDC activity, TDCC activity, or ADCP activity can be measured by known methods.
[0090] Neutralizing activity refers to the activity of inhibiting the biological activity of a ligand (e.g., a virus or toxin) that has biological activity against a cell. That is, a substance with neutralizing activity refers to a substance that binds to a ligand or to a receptor to which the ligand binds, and inhibits the binding of the ligand to the receptor. A receptor whose binding to a ligand is prevented by neutralizing activity is no longer able to exert biological activity through the receptor. Neutralizing activity is not limited to inhibiting the binding of a ligand to a receptor, but also includes the activity of inhibiting the function of a protein with biological activity. An example of the function of the above protein is enzymatic activity.
[0091] Physical property evaluation: For example, stability evaluation such as thermal stability, chemical stability, light stability, stability against mechanical stimuli, and long-term storage stability can be performed by measuring molecular decomposition, chemical modification, and aggregation before and after the treatments targeted for the stability evaluation, such as heat treatment, exposure to a low pH environment, light exposure, mechanical stirring, and long-term storage. Examples of measurement methods for performing such stability evaluation include chromatographic techniques such as ion exchange chromatography and size exclusion chromatography, mass spectrometry, and electrophoresis. Other measurement methods may also be used.
[0092] Other examples of physical property evaluations include evaluation of protein solubility by polyethylene glycol precipitation, evaluation of viscosity by small-angle X-ray scattering, and evaluation of nonspecific binding based on binding to extracellular matrix (ECM).
[0093] Further examples of physical property evaluation include evaluation of protein expression level, evaluation of binding to a purification resin or a purification ligand, and evaluation of surface charge.
[0094] Kinetic evaluation Kinetic evaluation of a molecule can be performed by administering the molecule to animals such as mice, rats, monkeys, and dogs and measuring the amount of the molecule in the blood over time after administration. Alternatively, kinetic evaluation can be performed by pharmacokinetics (PK) evaluation. As a method other than direct evaluation of PK, kinetic behavior can be predicted from the amino acid sequence of the molecule by calculating the surface charge, isoelectric point, etc. of the molecule using software.
[0095] Safety evaluation Examples of molecule safety evaluation include immunogenicity prediction tools such as ISPRI Web-Based Immunogenicity Screening (EpiVax), HLA binding evaluation of fragment peptides of antigen-binding molecules, MHC-Associated Peptide Proteomics (MAPPs), detection of T cell epitopes using T cell proliferation evaluation, etc., and evaluation of immunogenicity. Safety evaluation can be performed as long as it can be measured by techniques such as binding to rheumatoid factor (RF), evaluation of immune responses using PBMC and whole blood, platelet aggregation evaluation, etc.
[0096] In one example, a drug discovery system including the information processing system 10 provides a method for generating new targets with predetermined properties, such as specific physiological activity (e.g., binding to a specific protein). Examples of drugs include potential active agents such as small molecule drugs, medium molecule drugs, biological drugs, cells, nucleic acid drugs, biopharmaceuticals, or other active agents. The targets include molecular structures with desired or defined biological activity. The biological activity may be, for example, preferential binding to a specific protein over other proteins. Molecules that serve as drug candidates include biomolecules or compounds, including various molecules such as nucleic acids, peptides, cyclic peptides, proteins, antibodies, target molecule-binding molecules, polymeric compounds, medium molecule compounds, and small molecule compounds.
[0097] The drug discovery system may include a device for selecting molecules that interact with a drug target, a device for creating lead molecules, etc. The drug discovery system may be, for example, an information processing system configured including the disclosure of WO2020 / 246617.
[0098] The drug discovery system may include a molecular design device that searches for candidate molecules having desired properties and outputs information on the identified candidate molecules. The drug discovery system uses the information on the candidate molecules output from the molecular design device to select new targets suitable for drug candidates.
[0099] [Verification Example] (First Verification Example) Next, a verification example of the information processing system according to the present disclosure will be described.
[0100] In a first verification example, the effectiveness of the information processing system according to the present disclosure was verified using a dataset showing the actual measured values of the binding strength of each of multiple antibody sequences. The dataset was composed of sequences and characteristic values of binding strength of bispecific antibodies (bispecific antibodies) whose antigens were Marvel D3 and CD3. In the first verification example, Marvel D3 was assumed to be the antigen, and the subject was improving the binding strength of the antibody. Octet values were used as the characteristic values of binding strength.
[0101] Marvel D3 is a tight junction protein with four transmembrane domains. In the first validation example, Marvel D3 was selected as a candidate target for anticancer drugs. The development of bispecific antibodies that cross-link cancer antigens with antigens on T cells is expected to be applied to cancer treatment. In the first validation example, the trained prediction model was used to identify candidate anti-Marvel D3 sequences with superior properties from the lead antibody.
[0102] In the first verification example, the octet values of antibody sequences were measured multiple times, and the antibody sequences and the octet values of the antibody sequences obtained by the measurements were used as inputs to the prediction model. By batch measurement, the octet values of antibody sequences of typically 100 or less are obtained in one measurement.
[0103] The antibodies used for the assay were obtained as follows. First, plasmids encoding predesigned heavy or light chains were prepared, and the recombinant antibodies were transiently expressed in Expi293F cells. Antibodies were captured from the culture supernatant using protein A and eluted into buffer. The eluted buffers were mixed under reducing conditions to prepare Marvel D3 / CD3 bispecific antibodies. In this preparation, selective heavy chain heterodimerization was achieved by applying charge repulsion between the same heavy chains. The antibody concentration in the buffer was determined by absorbance at 280 nm. The buffer containing the bispecific antibody was then subjected to ion exchange chromatography to confirm the preparation of the desired Marvel D3 / CD3 antibody.
[0104] The Octet HTX system was used to measure the Octet value. Extracellular vesicles bearing CD81 and human Marvel D3 proteins on their surfaces were captured on a sensor chip using an anti-CD81 antibody. After a 600-second baseline step in D-PBS(-) containing 0.1% BSA, the association and dissociation responses were measured for 900 and 1500 seconds, respectively, in the same buffer containing 20 nM antibody. The antibody binding activity was measured as a shift in wavelength between the baseline step and the end of the association phase. Measurements were performed during the baseline step, association, and dissociation phases at 30°C, with the sample plate vibrating at a rate of 1,000 vibrations per minute.
[0105] The specific method will be described below. First, 1,144 antibody sequences were extracted as a training sequence group. Next, a property prediction model for predicting binding strength was created using training data representing the training sequence group. Then, a virtual sequence group consisting of 196,608 hypothetical sequences created by combining amino acid modification candidates and having no actual measured values was input into the property prediction model as a prediction sequence group. The prediction sequence group was screened based on the predicted values of the property prediction model for each antibody sequence included in the prediction sequence group, and it was verified whether antibody sequences with high binding strength could be selected.
[0106] When creating the property prediction model, the protein language model TAPE was used to convert each antibody sequence into a 768-dimensional vector for each VH region and VL region, and the combined 1,536-dimensional vector data was defined as the feature of each antibody sequence.
[0107] To verify the effectiveness of the present invention, input features to be used for generating a prediction model were selected from 1,536-dimensional features by the following two methods. A property prediction model was then constructed using the selected features, and the binding avidity of antibody sequences selected by the model was compared. The first method uses LightGBM (Light Gradient-Boosting Machine). The first method creates a classification model that discriminates between a training sequence group and a prediction sequence group, and calculates the importance of each dimension in the 1,536-dimensional features based on the feature importance obtained from the model. The first method then selects features by increasing the number of features by 10, starting from the least important feature, i.e., from the feature estimated to have the least difference between the training sequence group and the prediction sequence group. The second method defines the distribution distance using the Wasserstein distance for each dimension of the features possessed by the training sequence group and the prediction sequence group. The second method selects features by increasing the number of features, starting with features with the smallest Wasserstein distance, i.e., features with relatively similar distributions between the training sequence group and the prediction sequence group. Cross-validation, using mean squared error as an index, was applied to each set of features selected stepwise by the first and second methods, and the feature set with the highest prediction performance for the training sequence group was identified. The feature set was then used to evaluate the performance of LightGBM property prediction targeting binding strength.
[0108] 10 shows the distribution of actual values of the selected antibody sequences when the top 16 antibody sequences with the highest predicted binding strength were selected from a group of predictive sequences in descending order of predicted binding strength for each of the following cases: when all features are used as input features without feature selection (Baseline); when features selected using feature importance obtained from a classification model (Classifier); and when features selected using Wasserstein distance (Wasserstein). Figure 10 shows that when features selected using Wasserstein distance were used, antibody sequences with higher actual values of binding strength were screened compared to when all features were used and when features selected using feature importance were used. In other words, Figure 10 demonstrates the effectiveness of the information processing system (feature selection method) according to the present disclosure.
[0109] (Second Verification Example) Next, a second verification example will be described. In the second verification example, the effectiveness of the information processing system according to the present disclosure was verified using AqSolDB, a dataset of the solubility of 9,982 small molecular compounds downloaded using Therapeutics Data Commons, a platform that provides download functions for public datasets. A property prediction model for predicting solubility was created using 1,000 compounds randomly extracted from all compounds as a first molecule group. Next, the remaining 8,982 compounds were used as a second molecule group, and the second molecule group was screened based on the predicted values of the property prediction model. It was then verified whether molecules with high measured solubility values could be selected. For this verification, 10 random number seeds were set when randomly extracting molecules, and the verification was performed 10 times. In the verification, features were selected according to step S10A described above. In step S11, training data and test data containing 2048-dimensional features extracted using MolCLR (Molecular Contrastive Learning of Representations via Graph Neural Networks) were obtained. In steps S12 and S14A, Wasserstein distance and the feature importance of the model classifying the first molecular group and the second molecular group were used as selection indices for each feature, respectively, and the two types of selection indices were compared. In step S13A, the initial value of the number m of selected features was set to 18. In step S15A, cross-validation was used as a method for evaluating the performance of the provisional prediction model, mean squared error was used as an index to evaluate its performance, and LightGBM (Light Gradient-Boosting Machine) was used as a supervised learning method. In step S16A, it was determined that the process proceeded to step S17A if the number of selected features m was less than 2048, and proceeded to step S18 if the number of selected features m was 2048. In step S17A, the number of selected features m was incremented by 10. In step S18, the feature set that minimized the average value of the mean squared error of the provisional prediction model was selected as the input features.
[0110] 11 shows the distribution of the average values of the actual solubilities of the selected compounds when the top 100 compounds are selected in descending order of predicted solubility for each of the following cases: when all features are used as input features without feature selection (Baseline); when features selected using feature importance as a feature selection index are used as input features (Classifier); and when features selected using Wasserstein distance as a feature selection index are used as input features (Wasserstein). Figure 11 shows that when features selected using Wasserstein distance are used, compounds with higher actual solubility values can be screened compared to when all features are used and when features selected using feature importance are used. In other words, Figure 11 demonstrates the effectiveness of the information processing system (feature selection method) according to the present disclosure.
[0111] [Modifications] The technology according to the present disclosure has been described in detail above based on various examples. However, the present disclosure is not limited to the above examples. The technology according to the present disclosure can be modified in various ways without departing from the spirit of the present disclosure.
[0112] In the above example, the information processing system 10 includes the feature selection unit 11, the learning unit 12, and the prediction unit 13. However, the information processing system does not necessarily need to include at least one of the learning unit and the prediction unit. Prediction models are portable between computer systems. Therefore, a prediction model generated by an information processing system may be used in another computer system. Alternatively, another computer system may generate a prediction model by machine learning using one or more input features selected by the information processing system, and the information processing system may execute the prediction phase using the prediction model.
[0113] The information processing system may be configured as a server in a client-server system, may be implemented in a stand-alone computer, or may be implemented in a user terminal that can access a database storing training data and test data via a communication network.
[0114] The processing steps of the method executed by at least one processor are not limited to the above examples. For example, some of the steps or processes described above may be omitted, or the steps may be executed in a different order. Furthermore, any two or more of the steps described above may be combined, or some of the steps may be modified or deleted. Alternatively, other steps may be executed in addition to the steps described above.
[0115] In the present disclosure, when comparing the magnitude of two numerical values, either of the two criteria "greater than or equal to" and "greater than" may be used, or either of the two criteria "less than or equal to" and "less than" may be used.
[0116] In the present disclosure, the expression "at least one processor executes a first process, executes a second process, ... executes an nth process" or an expression corresponding thereto indicates a concept including a case where the processor that executes n processes from the first process to the nth process changes midway. In other words, this expression indicates a concept including both a case where all n processes are executed by the same processor and a case where the processor changes among the n processes according to an arbitrary policy.
[0117] In the present disclosure, the term "to" indicating a range is an inclusive expression. For example, "A to B" means a range equal to or greater than A and equal to or less than B.
[0118] In this disclosure, the term "about" when used in conjunction with a numerical value means a range of plus or minus 10% of that numerical value.
[0119] The term "and / or" is used herein to refer to each of the objects listed before and after "and / or" or any combination thereof. For example, "A, B and / or C" includes each of the objects "A," "B," and "C," as well as the combinations "A and B," "A and C," "B and C," and "A and B and C."
[0120] [Supplementary Notes] As can be seen from the various examples above, the present disclosure includes the following aspects. (Supplementary Note 1) An information processing system including at least one processor, wherein the at least one processor: acquires, for each of a plurality of first molecules, training data indicating a plurality of feature quantities related to the first molecule; acquires, for each of a plurality of second molecules, test data indicating the plurality of feature quantities related to the second molecule; calculates, for each of the plurality of feature quantities, an index based on a first probability distribution that is a probability distribution in the training data and a second probability distribution that is a probability distribution in the test data; and selects, based on the index for each of the plurality of feature quantities, one or more of the plurality of feature quantities as one or more input feature quantities to be used as input parameters of a prediction model that predicts a characteristic value of a molecule based on machine learning. (Supplementary Note 2) The information processing system according to Supplementary Note 1, wherein the at least one processor: calculates, for each of the plurality of feature quantities, a distribution-to-distribution distance that is the distance between the first probability distribution and the second probability distribution as the index; and selects the one or more input feature quantities from the plurality of feature quantities based on the distribution-to-distribution distance for each of the plurality of feature quantities. (Supplementary Note 3) The information processing system according to Supplementary Note 2, wherein the at least one processor selects the one or more feature quantities for which the inter-distribution distance is smaller than a predetermined criterion as the one or more input feature quantities. (Supplementary Note 4) The information processing system according to Supplementary Note 2 or 3, wherein the at least one processor calculates a Wasserstein distance as the inter-distribution distance. (Supplementary Note 5) The information processing system according to any one of Supplementary Notes 2 to 4, wherein the at least one processor excludes some of the plurality of feature quantities based on the inter-distribution distance for each of the plurality of feature quantities, and selects the remaining one or more feature quantities as the one or more input feature quantities.(Supplementary Note 6) The information processing system described in Supplementary Note 5, wherein the at least one processor repeats an evaluation process of generating a provisional prediction model by machine learning using the training data and calculating an evaluation index for the provisional prediction model based on one or more remaining features, while changing the number of features excluded from the plurality of features based on the inter-distribution distance for each of the plurality of features; determines the number of features to be excluded from the plurality of features as an exclusion number based on the plurality of evaluation indexes; and selects the one or more input features by excluding the number of features to be excluded from the plurality of features based on the inter-distribution distance for each of the plurality of features. (Supplementary Note 7) The information processing system according to Supplementary Note 6, wherein the plurality of feature quantities are N feature quantities, where N is a natural number greater than or equal to 2, and the at least one processor: sets a number of feature quantities to be excluded from the N feature quantities to n, where n is a natural number, excludes n feature quantities from the plurality of feature quantities based on the inter-distribution distance of each of the N feature quantities, performs an evaluation process to generate a first provisional prediction model by machine learning using the training data and calculate an evaluation index for the first provisional prediction model based on the (N-n) feature quantities, excludes a further n feature quantities from the (N-n) feature quantities based on the inter-distribution distance of each of the (N-n) feature quantities, and performs an evaluation process to generate a second provisional prediction model by machine learning using the training data and calculate an evaluation index for the second provisional prediction model based on the (N-2n) feature quantities. (Supplementary Note 8) The information processing system according to Supplementary Note 6 or 7, wherein the at least one processor determines the number for which the highest evaluation index is obtained as the number to be excluded. (Supplementary Note 9) The information processing system according to Supplementary Note 6 or 7, wherein the at least one processor determines the number of items for which the lowest evaluation index is obtained as the number to be excluded.(Supplementary Note 10) The information processing system described in Supplementary Note 5, wherein the at least one processor repeats an evaluation process of generating a provisional prediction model by machine learning using the training data and calculating an evaluation index for the provisional prediction model based on one or more selected features, while changing the number of features selected from the plurality of features based on the inter-distribution distance for each of the plurality of features; determines the number of features selected from the plurality of features as a selection number based on the plurality of evaluation indexes; and selects the number of features selected from the plurality of features as the selection number, based on the inter-distribution distance for each of the plurality of features, as the one or more input features. (Supplementary Note 11) The information processing system according to Supplementary Note 10, wherein the plurality of feature quantities are N feature quantities, where N is a natural number greater than or equal to 2, and the at least one processor: sets a number of feature quantities selected from the N feature quantities to m, where m is a natural number, selects m feature quantities from the plurality of feature quantities based on the inter-distribution distance of each of the N feature quantities, performs an evaluation process to generate a first provisional prediction model by machine learning using the training data and calculate an evaluation index for the first provisional prediction model based on the m feature quantities, adds m more feature quantities to the m feature quantities based on the inter-distribution distance of each of the m feature quantities, and performs an evaluation process to generate a second provisional prediction model by machine learning using the training data and calculate an evaluation index for the second provisional prediction model based on 2m feature quantities. (Supplementary Note 12) The information processing system according to Supplementary Note 11, wherein the at least one processor determines the number for which the highest evaluation index is obtained as the selection number. (Supplementary Note 13) The information processing system according to Supplementary Note 11, wherein the at least one processor determines the number of items for which the lowest evaluation index is obtained as the selection number. (Supplementary Note 14) The information processing system according to any one of Supplements 6 to 13, wherein the at least one processor performs the evaluation process by cross-validation.(Supplementary Note 15) The information processing system according to any one of Supplements 1 to 14, wherein the at least one processor performs the machine learning using the one or more input features of the training data to generate the prediction model. (Supplementary Note 16) The information processing system according to Supplementary Note 15, wherein the at least one processor acquires the one or more input features for a target molecule whose characteristic value is unknown, and outputs the characteristic value of the target molecule obtained by inputting the acquired one or more input features into the generated prediction model. (Supplementary Note 17) The information processing system according to Supplementary Note 15, wherein the at least one processor acquires the one or more input features for each of a plurality of target molecules whose characteristic value is unknown, inputs the acquired one or more input features for each of the plurality of target molecules into the generated prediction model to acquire the characteristic value of the target molecule from the prediction model, selects at least one target molecule from the plurality of target molecules based on the characteristic value of each of the plurality of target molecules, and outputs information on the selected at least one target molecule. (Supplementary Note 18) The information processing system according to any one of Supplementary Notes 1 to 17, wherein each of the first molecule and the second molecule is one selected from the group consisting of a nucleic acid, a peptide, a cyclic peptide, a protein, an antibody, and a low molecular weight compound. (Supplementary Note 19) The information processing system according to Supplementary Note 15, wherein the at least one processor acquires the one or more input features for at least one second molecule among the plurality of second molecules, and outputs the characteristic value of the second molecule obtained by inputting the acquired one or more input features for each of the at least one second molecule into the generated prediction model.(Supplementary Note 20) The information processing system according to Supplementary Note 15, wherein the at least one processor: acquires the one or more input feature amounts for each of the plurality of second molecules, inputs the acquired one or more input feature amounts for each of the plurality of second molecules into the generated prediction model to acquire the characteristic value of the second molecule from the prediction model, selects at least one second molecule from the plurality of second molecules as a candidate molecule based on the characteristic value of each of the plurality of second molecules, and outputs information on the selected at least one candidate molecule. (Supplementary Note 21) The information processing system according to any one of Supplements 1 to 20, wherein the characteristic value is selected from at least one of affinity, pharmacological activity, physical properties, kinetics, and safety. (Supplementary Note 22) The information processing system according to any one of Supplements 1 to 20, wherein the first molecule and the second molecule are antigen-binding molecules, and the characteristic value is a value regarding the antigen-binding ability of the antigen-binding molecule. (Supplementary Note 23) An information processing method executed by an information processing system having at least one processor, comprising: a step of acquiring, for each of a plurality of first molecules, training data indicating a plurality of feature quantities related to the first molecule; a step of acquiring, for each of a plurality of second molecules, test data indicating the plurality of feature quantities related to the second molecule; a step of calculating, for each of the plurality of feature quantities, an index based on a first probability distribution that is a probability distribution in the training data and a second probability distribution that is a probability distribution in the test data; and a step of selecting, based on the index for each of the plurality of feature quantities, one or more of the plurality of feature quantities as one or more input feature quantities to be used as input parameters of a prediction model that predicts a characteristic value of a molecule based on machine learning.(Supplementary Note 24) An information processing program that causes a computer to execute the steps of: acquiring, for each of a plurality of first molecules, training data indicating a plurality of feature quantities related to the first molecule; acquiring, for each of a plurality of second molecules, test data indicating the plurality of feature quantities related to the second molecule; calculating, for each of the plurality of feature quantities, an index based on a first probability distribution that is a probability distribution in the training data and a second probability distribution that is a probability distribution in the test data; and selecting, based on the index for each of the plurality of feature quantities, one or more of the plurality of feature quantities as one or more input feature quantities to be used as input parameters of a prediction model that predicts a characteristic value of a molecule based on machine learning. (Supplementary Note 25) A method for producing a molecular compound, comprising: a generation step of generating a molecular compound having a molecular sequence of at least one target molecule, based on the information of the at least one target molecule output by the information processing system described in Supplementary Note 17. (Supplementary Note 26) A method for producing a molecular compound, comprising: a generation step of generating a molecular compound having a molecular sequence of the at least one candidate molecule, based on the information of the at least one candidate molecule output by the information processing system described in Supplementary Note 20. (Supplementary Note 27) A non-transitory computer-readable recording medium storing an information processing program that causes a computer to execute the following steps: acquiring, for each of a plurality of first molecules, training data indicating a plurality of feature quantities related to the first molecule; acquiring, for each of a plurality of second molecules, test data indicating the plurality of feature quantities related to the second molecule; calculating, for each of the plurality of feature quantities, an index based on a first probability distribution that is a probability distribution in the training data and a second probability distribution that is a probability distribution in the test data; and selecting, based on the index for each of the plurality of feature quantities, one or more of the plurality of feature quantities as one or more input feature quantities to be used as input parameters of a prediction model that predicts a characteristic value of a molecule based on machine learning.
[0121] According to Supplements 1, 23, 24, and 27, features used as input parameters of a machine learning-based prediction model are selected as input features based on the trends of individual features in both the training data and the test data. By selecting features in this manner, the accuracy of a machine learning model (prediction model) for predicting molecular properties can be improved. The improvement in the accuracy of the machine learning model can lead to more accurate predictions of molecular properties.
[0122] According to Supplementary Note 2, features are selected based on the inter-distribution distance between the first probability distribution and the second probability distribution. By introducing the inter-distribution distance, which quantitatively represents the difference in the trends between the two probability distributions, input features can be selected based on objective criteria. As a result, further improvement in the accuracy of machine learning models can be expected.
[0123] According to Supplementary Note 3, features with relatively small intermolecular distances are selected as input features. The input features selected in this way tend to be relatively similar between the training data and the test data. Therefore, machine learning based on these input features can improve the accuracy of the machine learning model (prediction model).
[0124] According to Supplementary Note 4, by introducing the Wasserstein distance as the intermolecular distance, the intermolecular distance can be calculated appropriately. As a result, it becomes possible to select the input feature amount more appropriately.
[0125] According to Supplementary Note 5, input features can be selected by focusing on features that are estimated to be inappropriate as input features.
[0126] According to Supplementary Note 6, the generation and evaluation of a provisional prediction model are repeated while changing the number of features to be excluded, and the number of features to be excluded is dynamically determined based on multiple evaluation indices obtained through this repetition. Then, the determined number of features to be excluded is excluded to select input features. In one example, if too many features with relatively large distribution distances are excluded, features that may actually contribute to predicting characteristic values may also be excluded. By dynamically determining the number of features to be excluded as described above, it is possible to appropriately select one or more input features to improve the accuracy of the machine learning model.
[0127] According to Supplementary Note 7, the rate of increase of the number of excluded items is constant during repetition of the evaluation process, so that the evaluation process can be efficiently repeated using a simple method.
[0128] According to Supplementary Note 8, the number of features to be excluded when the highest evaluation index is obtained is determined as the final number to be excluded. According to Supplementary Note 9, the number of features to be excluded when the lowest evaluation index is obtained is determined as the final number to be excluded. Therefore, it is possible to select one or more input features that are expected to most improve the accuracy of the machine learning model.
[0129] According to Supplementary Note 10, the generation and evaluation of a provisional prediction model are repeated while changing the number of selected features, and the number of selections is dynamically determined based on multiple evaluation indices obtained by this repetition. Then, input features equal to the determined number of selections are selected. In one example, if too many features with relatively small distribution distances are selected, features that do not actually contribute to predicting characteristic values may be selected. By dynamically determining the number of selections as described above, one or more input features can be appropriately selected to improve the accuracy of the machine learning model.
[0130] According to Supplementary Note 11, the rate of increase in the number of selections is constant during repetition of the evaluation process, so that the evaluation process can be efficiently repeated using a simple method.
[0131] According to Supplementary Note 12, the number of selected features when the highest evaluation index is obtained is determined as the final number of selections. According to Supplementary Note 13, the number of selected features when the lowest evaluation index is obtained is determined as the final number of selections. Therefore, it is possible to select one or more input features that are expected to most improve the accuracy of the machine learning model.
[0132] According to Appendix 14, cross-validation is introduced in each evaluation process for obtaining each evaluation index, so even if the amount of training data is limited, it is possible to accurately evaluate each provisional prediction model while effectively utilizing the limited data.
[0133] According to Supplementary Note 15, machine learning is performed to generate a predictive model using one or more input features selected based on the trends of individual features in both the training data and the test data. Therefore, a highly accurate predictive model can be obtained.
[0134] According to Supplementary Note 16, by using a prediction model generated by machine learning using one or more selected input features, the properties of a target molecule can be predicted with high accuracy.
[0135] According to Supplementary Note 17, the properties of individual target molecules can be predicted with high accuracy by using a prediction model generated by machine learning using one or more selected input features. Therefore, at least one target molecule can be appropriately selected from multiple target molecules.
[0136] According to Appendix 18, the accuracy of machine learning models (prediction models) for predicting the properties of nucleic acids, peptides, cyclic peptides, proteins, antibodies, and low molecular weight compounds can be improved.
[0137] According to Supplementary Note 19, by using a prediction model generated by machine learning using one or more selected input features, the properties of the second molecule can be predicted with high accuracy.
[0138] According to Supplementary Note 20, by using a prediction model generated by machine learning using one or more selected input features, it is possible to accurately predict the properties of each second molecule. Therefore, it is possible to appropriately select at least one candidate molecule from a plurality of second molecules.
[0139] According to Supplementary Note 21, the accuracy of a machine learning model (prediction model) for predicting molecular affinity, pharmacological activity, physical properties, kinetics, or safety can be improved.
[0140] According to Supplementary Note 22, the accuracy of a machine learning model (prediction model) for predicting the antigen-binding ability of an antigen-binding molecule can be improved.
[0141] According to Supplementary Note 25, a molecular compound that is expected to have desired properties can be generated based on information on at least one target molecule that is appropriately selected using a predictive model.
[0142] According to Appendix 26, a molecular compound that is expected to have desired properties can be generated based on information on at least one candidate molecule that is appropriately selected using a predictive model.
[0143] 10...information processing system, 11...feature selection unit, 12...learning unit, 13...prediction unit, 20...database, 30...prediction model, 110...information processing program, 210...first probability distribution, 220...second probability distribution
Claims
1. An information processing system comprising at least one processor, wherein the at least one processor: obtains training data indicating a plurality of feature quantities for each of a plurality of first molecules; obtains test data indicating the plurality of feature quantities for each of a plurality of second molecules; calculates an index based on a first probability distribution which is a probability distribution in the training data and a second probability distribution which is a probability distribution in the test data for each of the plurality of feature quantities; and based on the index of each of the plurality of feature quantities, selects one or more input feature quantities from the plurality of feature quantities to be used as input parameters of a prediction model for predicting a characteristic value of a molecule based on machine learning.
2. The information processing system according to claim 1, wherein the at least one processor: calculates a distance between distributions, which is a distance between the first probability distribution and the second probability distribution, as the index for each of the plurality of feature quantities; and selects the one or more input feature quantities from the plurality of feature quantities based on the distance between distributions of each of the plurality of feature quantities.
3. The information processing system according to claim 2, wherein the at least one processor selects the one or more feature quantities having a distance between distributions smaller than a predetermined criterion as the one or more input feature quantities.
4. The information processing system according to claim 2 or 3, wherein the at least one processor calculates a Wasserstein distance as the distance between distributions.
5. The information processing system according to any one of claims 2 to 4, wherein the at least one processor: excludes a part of the plurality of feature quantities based on the distance between distributions of each of the plurality of feature quantities, and selects the remaining one or more feature quantities as the one or more input feature quantities.
6. The at least one processor repeats an evaluation process of generating a provisional prediction model by the machine learning using the training data and calculating an evaluation index of the provisional prediction model, while changing the number of feature quantities to be excluded from the plurality of feature quantities based on the distance between the distributions of the respective plurality of feature quantities, and executing the evaluation process based on the remaining one or more feature quantities; determines the number of feature quantities to be excluded from the plurality of feature quantities as an exclusion number based on a plurality of the evaluation indexes; and excludes the number of feature quantities from the plurality of feature quantities based on the distance between the distributions of the respective plurality of feature quantities, and selects the one or more input feature quantities. The information processing system according to claim 5.
7. The plurality of feature quantities are N feature quantities, where N is a natural number of 2 or more; the at least one processor sets the number of feature quantities to be excluded from the N feature quantities to n, where n is a natural number; excludes n feature quantities from the plurality of feature quantities based on the distance between the distributions of the respective N feature quantities; performs an evaluation process of generating a first provisional prediction model by the machine learning using the training data and calculating an evaluation index of the first provisional prediction model, based on (N−n) feature quantities; excludes another n feature quantities from the (N−n) feature quantities based on the distance between the distributions of the respective (N−n) feature quantities; and performs an evaluation process of generating a second provisional prediction model by the machine learning using the training data and calculating an evaluation index of the second provisional prediction model, based on (N−2n) feature quantities. The information processing system according to claim 6.
8. The at least one processor determines the number obtained when the highest evaluation index is obtained as the exclusion number, or determines the number obtained when the lowest evaluation index is obtained as the exclusion number. The information processing system according to claim 6 or 7.
9. The at least one processor repeats an evaluation process of generating a provisional prediction model by the machine learning using the training data and calculating an evaluation index of the provisional prediction model while changing the number of feature amounts selected from the plurality of feature amounts based on the distance between the distributions of the respective plurality of feature amounts, determines the number of feature amounts selected from the plurality of feature amounts as a selection number based on the plurality of evaluation indexes, and selects the number of feature amounts corresponding to the selection number from the plurality of feature amounts as the one or more input feature amounts based on the distance between the distributions of the respective plurality of feature amounts. The information processing system according to any one of claims 2 to 4.
10. The at least one processor executes the evaluation process by cross-validation. The information processing system according to any one of claims 6 to 9.
11. The at least one processor executes the machine learning using the one or more input feature amounts of the training data to generate the prediction model. The information processing system according to any one of claims 1 to 10.
12. The at least one processor acquires the one or more input feature amounts for a target molecule whose characteristic value is unknown, and outputs the characteristic value of the target molecule obtained by inputting the acquired one or more input feature amounts into the generated prediction model. The information processing system according to claim 11.
13. The at least one processor acquires the one or more input feature amounts for each of a plurality of target molecules whose characteristic values are unknown, inputs the acquired one or more input feature amounts for each of the plurality of target molecules into the generated prediction model, and acquires the characteristic value of the target molecule from the prediction model, selects at least one target molecule from the plurality of target molecules based on the characteristic values of the respective plurality of target molecules, and outputs information on the selected at least one target molecule. The information processing system according to claim 11.
14. Each of the first molecule and the second molecule is one selected from nucleic acids, peptides, cyclic peptides, proteins, antibodies, and low molecular weight compounds. The information processing system according to any one of claims 1 to 13.
15. An information processing method executed by an information processing system including at least one processor, the method comprising: obtaining training data indicating a plurality of feature quantities for each of a plurality of first molecules; obtaining test data indicating the plurality of feature quantities for each of a plurality of second molecules; calculating an index based on a first probability distribution which is a probability distribution in the training data and a second probability distribution which is a probability distribution in the test data for each of the plurality of feature quantities; and selecting, based on the index for each of the plurality of feature quantities, one or more feature quantities out of the plurality of feature quantities as one or more input feature quantities used as input parameters of a prediction model for predicting a characteristic value of a molecule based on machine learning.
16. An information processing program for causing a computer to execute: obtaining training data indicating a plurality of feature quantities for each of a plurality of first molecules; obtaining test data indicating the plurality of feature quantities for each of a plurality of second molecules; calculating an index based on a first probability distribution which is a probability distribution in the training data and a second probability distribution which is a probability distribution in the test data for each of the plurality of feature quantities; and selecting, based on the index for each of the plurality of feature quantities, one or more feature quantities out of the plurality of feature quantities as one or more input feature quantities used as input parameters of a prediction model for predicting a characteristic value of a molecule based on machine learning.
17. A method for manufacturing a molecular compound, comprising a generating step of generating a molecular compound having a molecular sequence of the at least one target molecule based on the information of the at least one target molecule output by the information processing system according to claim 13.
Citation Information
Patent Citations
Recombinant immunoglobin preparations
US4816567A
Expression of functional antibody fragments
US5648237A
Process for bacterial production of polypeptides
US5789199A
Methods and compositions for secretion of heterologous polypeptides
US5840523A
Machine learning based antibody design
WO2018132752A1