Information processing system, information processing method, program, and method for producing molecular compound
Patent Information
- Application Number
- JP2024566017
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-08-10
- Filing Date
- 2024-07-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-07-31
AI Technical Summary
【0020】 本開示の一側面によれば創薬に要する負担をより軽減することができる。
Smart Images

Figure 00000046_0000 
Figure 00000046_0001 
Figure 00000047_0000
Abstract
Description
[Technical field]
[0001] One aspect of the present disclosure relates to an information processing system, a molecular design device, an information processing method, a program, and a method for producing a molecular compound. [Background technology]
[0002] In recent years, in the pharmaceutical field, information processing technology based on machine learning has been utilized to reduce the burden of drug discovery (Patent Document 1). Attempts are being made to efficiently discover molecules with properties required as drugs using predictions based on machine learning. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2018 / 132752 Summary of the Invention [Problem to be solved by the invention]
[0004] Since predictions made by machine learning involve uncertainty, using predictions made by machine learning without taking this into consideration may lead to unexpected results. In addition, there are many targets to be analyzed in drug discovery, and it is difficult to apply predictions made by machine learning to all targets for which predictions can be implemented.
[0005] In view of the above circumstances, one aspect of the present disclosure aims to provide a technology that further reduces the burden required for drug discovery. [Means for solving the problem]
[0006] In one aspect of the present disclosure, the following embodiments are provided: [A1] An information processing system for identifying molecules suitable as drug candidates, comprising: a property prediction unit that calculates a property prediction value of a molecule, which is an element of a building block combination information set that is a set of building block combination information of a plurality of different molecules, using a prediction model for predicting the property of the molecule from the building block combination information of the molecule, and estimates the uncertainty of the prediction; a candidate molecule identifying unit that searches for candidates of molecules having desired properties based on the predicted property values and an estimated value of the uncertainty of the prediction; An information processing system comprising: [A2] The building block combination information is the molecular sequence information. The information processing system described in A1. [A3] The molecule is at least one of a nucleic acid, a peptide, a cyclic peptide, a protein, an antibody, and a small molecule compound. The information processing system according to A1 or A2. [A4] The molecule is a protein, an antibody, a peptide, or a cyclic peptide, and the building block combination information is amino acid sequence information. An information processing system according to any one of A1 to A3. [A5] The property is at least one of binding ability, pharmacological activity, physical properties, kinetics, and safety. An information processing system according to any one of A1 to A4. [A6] The molecule is a molecule that binds to a target molecule, and the property is the ability to bind to the target molecule. An information processing system according to any one of A1 to A5. [A7] a prediction information processing unit that calculates a characteristic quality value based on the characteristic prediction value and an estimate of the uncertainty of the prediction, the candidate molecule identifying unit identifies at least one candidate molecule from the elements of the building block combination information set based on the characteristic quality value; An information processing system according to any one of A1 to A6. [A8] the characteristic quality value increases with increasing characteristic prediction value and decreases with decreasing prediction uncertainty estimate; An information processing system according to any one of A1 to A7. [A9] The characteristic quality value is an output value of a predetermined function having two input variables, the characteristic prediction value and an estimate of the uncertainty of the prediction. An information processing system described in A8. [A10] The prediction information processing unit calculates an average variance as an objective function that gives the characteristic quality value. The information processing system according to any one of A7 to A9. [A11] the prediction uncertainty estimate is the standard deviation of the property prediction; An information processing system according to any one of A1 to A10. [A12] The property quality value increases in response to an increase in the property prediction value, and the building block combination information set is obtained using a combinatorial optimization algorithm in response to a decrease in the prediction uncertainty estimate, and an extraction parameter of the combinatorial optimization algorithm is updated based on the property prediction value of the molecule or the property quality value. An information processing system according to any one of A1 to A11. [A13] The combinatorial optimization algorithm is a tree-structured Parzen estimator. The information processing system according to A12. [A14] The prediction model is a prediction model generated by learning based on building block combination information of a plurality of molecules for training and the results of the property evaluation of the molecules. An information processing system according to any one of A1 to A13.
[0007] [B1] A molecular design apparatus comprising a control unit for inferring a property of a molecule from building block combination information of the molecule using a predetermined prediction model, The control unit is A building block combination information set is obtained, which is a set of building block combination information of a plurality of different molecules; calculating a predicted property value of a molecule that is an element of the building block combination information set from the building block combination information of the molecule using the prediction model, and estimating the uncertainty of the prediction; searching for candidate molecules having desired properties based on the predicted property values and an estimate of the uncertainty of the predictions; Molecular design equipment. [B2] The molecular design apparatus according to B1, wherein the building block combination information is sequence information of the molecule. [B3] The molecule is at least one of a nucleic acid, a peptide, a cyclic peptide, a protein, an antibody, and a small molecule compound; The molecular design apparatus according to B1 or B2. [B4] The molecular design apparatus according to any one of B1 to B3, wherein the molecule is a protein, an antibody, a peptide, or a cyclic peptide, and the building block combination information is information on an amino acid sequence. [B5] The property is at least one of binding ability, pharmacological activity, physical properties, kinetics, and safety. The molecular design apparatus according to any one of B1 to B4. [B6] The molecule is a molecule that binds to a target molecule, and the property is the ability to bind to the target molecule. The molecular design apparatus according to any one of B1 to B5. [B7] The control unit further calculates a characteristic quality value based on the characteristic prediction value and an estimate of the uncertainty of the prediction; identifying at least one candidate molecule from the elements of the building block combination information set based on the characteristic quality value; The molecular design apparatus according to any one of B1 to B6. [B8] the characteristic quality value increases with increasing characteristic prediction value and decreases with decreasing prediction uncertainty estimate; The molecular design apparatus according to any one of B7. [B9] The characteristic quality value is an output value of a predetermined function having two input variables, the characteristic prediction value and an estimate of the uncertainty of the prediction. A molecular design apparatus as described in B8. [B10] The control unit calculates an average variance as an objective function that gives the characteristic quality value. The molecular design apparatus according to any one of B7 to B9. [B11] the prediction uncertainty estimate is the standard deviation of the property prediction; The molecular design apparatus according to any one of B1 to B10. [B12] the control unit acquires the building block combination information set using a combinatorial optimization algorithm, and updates an extraction parameter of the combinatorial optimization algorithm based on the characteristic predicted value or the characteristic quality value. The molecular design apparatus according to any one of B1 to B11. [B13] The combinatorial optimization algorithm is a tree-structured Parzen estimator. A molecular design apparatus as described in B12. [B14] An output unit that outputs building block combination information of the candidate molecule is further provided. The molecular design apparatus according to any one of B1 to B13. [B15] The prediction model is a prediction model generated by learning based on building block combination information of a plurality of molecules for training and the results of the property evaluation of the molecules. An information processing system according to any one of B1 to B14.
[0008] [C1] An information processing system for identifying molecules suitable as drug candidates, comprising: a building block combination information processing unit that acquires a building block combination information set, which is a set of building block combination information of a plurality of different molecules, in accordance with a combinatorial optimization algorithm; a property prediction unit that calculates a property prediction value of a molecule that is an element of the building block combination information set, using a prediction model for predicting the property of the molecule from the building block combination information of the molecule; a candidate molecule identifying unit that searches for candidates of molecules having desired properties based on the predicted property values; The building block combination information processing unit includes: Based on the predicted property values of the molecules, the extraction parameters of the combinatorial optimization algorithm are updated so as to include molecules having more desirable properties. Information processing system. [C2] the candidate molecule identifying unit searches for candidate molecules having desired properties based on a first property predicted value calculated for molecules that are elements of a first building block combination information set, which is a building block combination information set obtained before the extraction parameters are updated, and a second property predicted value calculated for molecules that are elements of a second building block combination information set, which is a building block combination information set obtained after the extraction parameters are updated. C1's information processing system. [C3] The first building block combination information set and the second building block combination information set are sets including building block combination information of different molecules. 3. The information processing system according to claim 2 . [C4] The building block combination information is sequence information of the molecule. An information processing system according to any one of C1 to C3. [C5] The molecule is at least one of a nucleic acid, a peptide, a cyclic peptide, a protein, an antibody, and a small molecule compound. An information processing system according to any one of C1 to C4. [C6] The molecule is a protein, an antibody, a peptide, or a cyclic peptide, and the building block combination information is amino acid sequence information. An information processing system according to any one of C1 to C5. [C7] The property is at least one of binding ability, pharmacological activity, physical properties, kinetics, and safety. An information processing system according to any one of C1 to C6. [C8] The molecule is a molecule that binds to a target molecule, and the property is the ability to bind to the target molecule. An information processing system according to any one of C1 to C7. [C9] The characteristic prediction unit calculates the characteristic prediction value and calculates an estimate of the uncertainty of the prediction. An information processing system according to any one of C1 to C8. [C10] The building block combination information processing unit includes: updating extraction parameters of a combinatorial optimization algorithm based on the predicted property values of the molecules and the estimate of the uncertainty of the predictions to include molecules having more desirable properties; An information processing system according to any one of C1 to C9. [C11] a prediction information processing unit that processes prediction information output from the characteristic prediction unit, the prediction information including the characteristic prediction value and / or an estimate of the uncertainty of the prediction, the prediction information processing unit calculates a characteristic quality value based on the characteristic prediction value and an estimate of the uncertainty of the prediction; The candidate molecule identifying unit identifies at least one molecule from the building block combination information set based on the characteristic quality value. 2. An information processing system according to claim 1, wherein said information processing system is a system for processing information according to claim 1. [C12] the characteristic quality value increases with increasing characteristic prediction value and decreases with decreasing prediction uncertainty estimate; The information processing system according to C11. [C13] The characteristic quality value is an output value of a predetermined function having two input variables, the characteristic prediction value and an estimate of the uncertainty of the prediction. 13. The information processing system according to claim 12. [C14] The prediction information processing unit calculates an average variance as an objective function that gives the characteristic quality value. An information processing system according to any one of C11 to C13. [C15] the prediction uncertainty estimate is the standard deviation of the property prediction; An information processing system according to any one of C9 to C14. [C16] A tree-structured Parzen estimator is used as the combinatorial optimization algorithm. An information processing system according to any one of C1 to C15.
[0009] [D1] A molecular design apparatus comprising a control unit for inferring a property of a molecule from building block combination information of the molecule using a predetermined prediction model, The control unit is A building block combination information set is obtained according to a combinatorial optimization algorithm, the building block combination information set being a set of building block combination information of a plurality of different molecules; calculating the predicted property value for a molecule that is an element of the building block combination information set using the prediction model; updating extraction parameters of the combinatorial optimization algorithm so as to include molecules having more desirable properties based on the property prediction values for each molecule obtained by the prediction model; searching for candidate molecules having desired properties based on the predicted property values; Molecular design equipment. [D2] The control unit is Obtain a first building block combination information set which is a set of building block combination information of a plurality of different molecules before updating the extraction parameters; calculating a first property prediction value for a molecule that is an element of the first building block combination information set using the prediction model; After updating the extraction parameters, a second building block combination information set is further obtained, which is a set of building block combination information of a plurality of different molecules; calculating a second property prediction value for a molecule that is an element of the second building block combination information set using the prediction model; searching for candidate molecules having desired properties based on the first property predicted value and the second property predicted value; The molecular design apparatus described in D1. [D3] The first building block combination information set and the second building block combination information set are sets including sequence information of different molecules. The molecular design apparatus described in D2. [D4] The building block combination information is sequence information of the molecule. The molecular design apparatus according to any one of D1 to D3. [D5] The molecule is at least one of a nucleic acid, a peptide, a cyclic peptide, a protein, an antibody, and a small molecule compound. A molecular design apparatus according to any one of D1 to D5. [D6] The molecule is a protein, an antibody, a peptide, or a cyclic peptide, and the building block combination information is amino acid sequence information. A molecular design apparatus according to any one of D1 to D5. [D7] The property is at least one of binding ability, pharmacological activity, physical properties, kinetics, and safety. A molecular design apparatus according to any one of D1 to D6. [D8] The molecule is a molecule that binds to a target molecule, and the property is the ability to bind to the target molecule. A molecular design apparatus according to any one of D1 to D7. [D9] The control unit calculates the characteristic prediction value and further calculates an estimate of the uncertainty of the prediction. A molecular design apparatus according to any one of D1 to D8. [D10] the control unit acquires the building block combination information set so as to include molecules having more desirable properties based on a predicted value of the molecular properties and an estimated value of the uncertainty of the prediction. An information processing system according to any one of D1 to D9. [D11] The control unit further calculates a characteristic quality value based on the characteristic prediction value and an estimate of the uncertainty of the prediction; The candidate molecule identifying unit identifies at least one molecule from the building block combination information set based on the characteristic quality value. Molecular design apparatus described in D10. [D12] the characteristic quality value increases with increasing characteristic prediction value and decreases with decreasing prediction uncertainty estimate; A molecular design apparatus as described in D11. [D13] The characteristic quality value is an output value of a predetermined function having two input variables, the characteristic prediction value and an estimate of the uncertainty of the prediction. A molecular design apparatus as described in D12. [D14] The control unit calculates an average variance as an objective function that gives the characteristic quality value. The molecular design apparatus according to any one of D11 to D13. [D15] the prediction uncertainty estimate is the standard deviation of the property prediction; The molecular design apparatus according to any one of D9 to D14. [D16] A tree-structured Parzen estimator is used as the combinatorial optimization algorithm. A molecular design apparatus according to any one of D1 to D15. [D17] The prediction model is a model trained based on building block combination information that differs for each molecule and training data that indicates characteristics of the molecule. The molecular design apparatus according to any one of D1 to D16.
[0010] [E1] An information processing method in an information processing system for identifying molecules suitable as drug candidates, comprising the steps of: calculating a predicted property value of a molecule that is an element of a building block combination information set, which is a set of building block combination information of a plurality of different molecules, using a prediction model for predicting the properties of the molecule from the building block combination information of the molecule, and estimating the uncertainty of the prediction; and searching for candidate molecules having more desirable properties based on the predicted property values and the uncertainty of the predictions. Information processing methods. [E2] The building block combination information is information on the sequence of the molecule. The information processing method described in E1. [E3] The molecule is at least one of a nucleic acid, a peptide, a cyclic peptide, a protein, an antibody, and a small molecule compound. An information processing method according to E1 or E2. [E4] The molecule is a protein, an antibody, a peptide, or a cyclic peptide, and the building block combination information is information of an amino acid sequence. An information processing method according to any one of E1 to E3. [E5] The property is at least one of binding ability, pharmacological activity, physical properties, kinetics, and safety. An information processing method according to any one of E1 to E4. [E6] The molecule is a molecule that binds to a target molecule, and the property is the ability to bind to the target molecule. An information processing method according to any one of E1 to E5. [E7] further comprising the step of calculating a characteristic quality value based on the characteristic prediction and an estimate of the uncertainty of the prediction; The searching step includes identifying at least one molecule from the building block combination information set based on the characteristic quality value. An information processing method according to any one of E1 to E6. [E8] the characteristic quality value increases with increasing characteristic prediction value and decreases with decreasing prediction uncertainty estimate; The information processing method described in E7. [E9] The characteristic quality value is an output value of an arbitrary function having two input variables, the characteristic prediction value and an estimate of the uncertainty of the prediction. The information processing method described in E8. [E10] The step of calculating the characteristic quality value calculates a mean variance as an objective function that gives the characteristic quality value. An information processing method according to any one of E7 to E9. [E11] the estimate of the uncertainty of the prediction is the standard deviation of the prediction; An information processing method according to any one of E1 to E10. [E12] outputting building block combination information of the identified candidate molecule; An information processing method according to any one of E1 to E11. [E13] Further outputting information on the predicted property values of the identified molecules. The information processing method described in E12. [E14] The searching step includes a step of identifying molecules having the characteristic quality value equal to or greater than a predetermined value from the building block combination information set. An information processing method according to any one of E7 to E13. [E15] the searching step includes the steps of: determining a ranking of the characteristic quality values for each molecule; and identifying molecules within a predetermined ranking. An information processing method according to any one of E7 to E13. [E16] The searching step includes a step of selecting at least one molecule from the building block combination information set, the predicted property value and the uncertainty of the prediction each satisfy a predetermined condition. An information processing method according to any one of E1 to E15. [E17] Using a combinatorial optimization algorithm to obtain the building block combination information set, and updating an extraction parameter of the combinatorial optimization algorithm based on the characteristic prediction value or the characteristic quality value; An information processing method according to any one of E1 to E16. [E18] The combinatorial optimization algorithm is a tree-structured Parzen estimator. The information processing method described in E17. [E19] The prediction model is a model trained based on building block combination information that differs for each molecule and training data that indicates characteristics of the molecule. An information processing method according to any one of E1 to E18.
[0011] [F1] On the computer, a step of calculating a predicted property value of a molecule, which is an element of a building block combination information set, which is a set of building block combination information of a plurality of different molecules, using a prediction model for predicting properties of the molecule from the building block combination information of the molecule, and estimating the uncertainty of the prediction; searching for candidate molecules having more desirable properties based on the predicted property values and the uncertainty of the predictions; A program for executing the above. [F2] The building block combination information is sequence information of the molecule. Program described in F1. [F3] The molecule is at least one of a nucleic acid, a peptide, a cyclic peptide, a protein, an antibody, and a small molecule compound. Programs listed in F1 or F2. [F4] The molecule is a protein, an antibody, a peptide, or a cyclic peptide, and the building block combination information is information of an amino acid sequence. A program listed in any of F1 to F3. [F5] The property is at least one of binding ability, pharmacological activity, physical properties, kinetics, and safety. A program listed in any of F1 to F4. [F6] The molecule is a molecule that binds to a target molecule, and the property is the ability to bind to the target molecule. A program described in any of F1 to F5. [F7] The method further comprises the step of calculating a characteristic quality value based on the characteristic prediction value and an estimate of the uncertainty of the prediction, The searching step includes a step of identifying at least one molecule from the building block combination information set based on the characteristic quality value. A program described in any one of F1 to F6. [F8] the characteristic quality value increases with increasing characteristic prediction value and decreases with decreasing prediction uncertainty estimate; Program listed in F7. [F9] The characteristic quality value is an output value of a predetermined function having two input variables, the characteristic prediction value and an estimate of the uncertainty of the prediction. Program listed in F8. [F10] The step of calculating the characteristic quality value includes calculating a mean variance as an objective function that gives the characteristic quality value. A program described in any of F7 to F9. [F11] the uncertainty estimate is the standard deviation of the predicted value; A program described in any one of F1 to F10. [F12] a step of searching for candidate molecules having more desired properties and outputting building block combination information of the identified candidate molecules; A program described in any one of F1 to F11. [F13] Searching for candidate molecules having more desired properties and further outputting information on predicted property values of the identified candidate molecules; The program described in F12. [F14] The searching step includes a step of identifying a molecule having a characteristic quality value equal to or greater than a predetermined value from the building block combination information set. A program described in any of F7 to F13. [F15] The searching step includes the steps of: determining a ranking of the characteristic quality values for each molecule; and identifying molecules within a predetermined ranking. A program described in any of F7 to F13. [F16] The searching step includes a procedure of selecting at least one molecule from the building block combination information set, the predicted property value and the estimated value of the uncertainty of the prediction each satisfy a predetermined condition. A program listed in any of F1 to F15. [F17] The method further includes a step of updating an extraction parameter for obtaining the building block combination information set based on the predicted property value or the property quality value of the molecule using a combinatorial optimization algorithm. A program described in any one of F1 to F16. [F18] The combinatorial optimization algorithm is a tree-structured Parzen estimator. Program described in F17. [F19] The prediction model is a model trained based on building block combination information that differs for each molecule and training data that indicates characteristics of the molecule. A program described in any one of F1 to F18.
[0012] [G1] 1. A method for producing a molecular compound, comprising: accessing a building block combination information set which is a set of building block combination information of a plurality of different molecules; an input step of inputting the building block combination information set into a prediction model; an inference step of searching for molecules having more desired properties from the building block combination information set based on the property prediction value and the prediction uncertainty estimate value for each molecule included in the building block combination information set output from the prediction model, and identifying the molecules as candidate molecules; an output step of outputting building block combination information relating to the candidate molecule; a generating step of generating the molecular compound having a molecular sequence indicated in the building block combination information; The method according to claim 1, [G2] The candidate molecule has a biological sequence, and the building block combination information is sequence information. The method described in G1. [G3] The molecule is at least one of a nucleic acid, a peptide, a cyclic peptide, a protein, an antibody, and a small molecule compound. The method according to G1 or G2. [G4] The molecule is a protein, an antibody, a peptide, or a cyclic peptide, and the building block combination information is amino acid sequence information. The method according to any one of G1 to G3. [G5] The property is at least one of binding ability, pharmacological activity, physical properties, kinetics, and safety. A method according to any one of G1 to G4. [G6] The molecule is a molecule that binds to a target molecule, and the property is the ability to bind to the target molecule. A method according to any one of G1 to G5. [G7] The method further comprises the steps of: calculating a characteristic quality value based on the characteristic prediction value and an estimate of the uncertainty of the prediction; and identifying at least one candidate molecule from the building block combination information set based on the characteristic quality value. A method according to any one of G1 to G6. [G8] the characteristic quality value increases with increasing characteristic prediction value and decreases with decreasing prediction uncertainty estimate; The method described in G7. [G9] The characteristic quality value is an output value of a predetermined function having two input variables, the characteristic prediction value and an estimate of the uncertainty of the prediction. The method described in G8. [G10] The step of calculating the characteristic quality value calculates a mean variance as an objective function that gives the characteristic quality value. A method according to any one of G7 to G9. [G11] the uncertainty estimate is the standard deviation of the predicted value; A method according to any one of G1 to G10. [G12] The information on the candidate molecule includes building block combination information of the candidate molecule. The method according to any one of G1 to G11. [G13] The information on the candidate molecule includes a predicted characteristic value of the candidate molecule. A method according to any one of G1 to G12. [G14] The step of searching for a candidate molecule includes a step of identifying a molecule having a characteristic quality value equal to or greater than a predetermined value as the candidate molecule from the building block combination information set. A method according to any one of G7 to G13. [G15] The step of searching for candidate molecules includes a step of determining a rank of the characteristic quality value for each molecule, and determining a molecule within a predetermined rank as the candidate molecule. A method according to any one of G7 to G13. [G16] The step of searching for a candidate molecule includes a step of selecting, from the building block combination information set, at least one molecule for which the predicted property value and the estimated value of the prediction uncertainty each satisfy a predetermined condition. A method according to any one of G1 to G15. [G17] The building block combination information set is obtained using a combinatorial optimization algorithm; The method further includes updating an extraction parameter of the combinatorial optimization algorithm based on the predicted property value or the property quality value of the molecule. A method according to any one of G1 to G16. [G18] The combinatorial optimization algorithm is a tree-structured Parzen estimator. The method described in G17. [G19] A selection step is provided for selecting a candidate molecule having the best experimental value from among experimental values of the properties of each of the plurality of candidate molecules. The method according to any one of G1 to G18.
[0013] [H1] An information processing method in an information processing system for identifying molecules suitable as drug candidates, comprising: A step of obtaining a building block combination information set, which is a set of building block combination information of a plurality of different molecules, according to a combinatorial optimization algorithm; A step of calculating a predicted property value of a molecule that is an element of the building block combination information set, using a prediction model for predicting the property of the molecule from the building block combination information of the molecule; searching for molecules having more desirable properties based on the predicted property values; updating extraction parameters of the combinatorial optimization algorithm based on the predicted property values of the molecules so that molecules having more desirable properties are included; Information processing methods. [H2] and searching for candidate molecules having desired properties based on a first property predicted value calculated for molecules that are elements of a first building block combination information set, which is a building block combination information set obtained before the extraction parameters are updated, and a second property predicted value calculated for molecules that are elements of a second building block combination information set, which is a building block combination information set obtained after the extraction parameters are updated. The information processing method described in H1. [H3] The first building block combination information set and the second building block combination information set are sets including building block combination information of different molecules. The information processing method described in H2. [H4] The building block combination information is information on the sequence of the molecule. The information processing method according to any one of H1 to H3. [H5] The molecule is at least one of a nucleic acid, a peptide, a cyclic peptide, a protein, an antibody, and a small molecule compound. An information processing system according to any one of H1 to H4. [H6] The molecule is a protein, an antibody, a peptide, or a cyclic peptide, and the building block combination information is amino acid sequence information. The information processing method according to any one of H1 to H5. [H7] The property is at least one of binding ability, pharmacological activity, physical properties, kinetics, and safety. The information processing method according to any one of claims H6 to H7. [H8] The molecule is a molecule that binds to a target molecule, and the property is the ability to bind to the target molecule. The information processing method according to any one of H1 to H7. [H9] calculating the predicted characteristic and an estimate of the uncertainty of the prediction; The information processing method according to any one of H1 to H8. [H10] calculating a characteristic quality value based on the characteristic prediction value and an estimate of the uncertainty of the prediction; the searching step includes identifying at least one molecule from the building block combination information set based on the characteristic quality value; The information processing method described in H9. [H11] the characteristic quality value increases with increasing characteristic prediction value and decreases with decreasing prediction uncertainty estimate; An information processing method described in H10. [H12] The characteristic quality value is an output value of a predetermined function having two input variables, the characteristic predicted value and the prediction uncertainty. An information processing method described in H11. [H13] The step of calculating the characteristic quality value includes a step of calculating a mean variance as an objective function that gives the characteristic quality value. The information processing method according to any one of H10 to H12. [H14] the prediction uncertainty estimate is the standard deviation of the property prediction; An information processing method according to any one of H9 to H13. [H15] A tree-structured Parzen estimator is used as a combinatorial optimization algorithm for updating the extraction parameters. An information processing method according to any one of H1 to H14.
[0014] [I1] On the computer, an acquisition step of acquiring a building block combination information set, which is a set of building block combination information of a plurality of different molecules, according to a combinatorial optimization algorithm; a prediction step of calculating a predicted property value of a molecule that is an element of the building block combination information set, using a prediction model for predicting the property of the molecule from the building block combination information of the molecule; an updating step of updating extraction parameters of the combinatorial optimization algorithm based on the predicted property values for each molecule so that molecules having more desired properties are included; a search procedure for searching for molecules having desired properties based on the predicted property values; A program for executing the above. [I2] The acquisition step includes: A step of acquiring a first building block combination information set which is a set of building block combination information of a plurality of different molecules before updating the extraction parameters; and a step of further acquiring a second building block combination information set, which is a set of building block combination information of a plurality of different molecules, after updating the extraction parameters, The prediction step comprises: calculating a first property prediction value for a molecule that is an element of the first building block combination information set using the prediction model; and calculating a second predicted property value for a molecule that is an element of the second building block combination information set using the prediction model, The search procedure includes: a search procedure for searching for a candidate molecule having a desired property from the first building block combination information set and the second building block combination information set based on the first property predicted value and the second property predicted value. Program described in I1. [I3] The first building block combination information set and the second building block combination information set are sets including building block combination information of different molecules. The program described in I2. [I4] The building block combination information is information on the sequence of the molecule. A program according to any one of I1 to I3. [I5] The molecule is at least one of a nucleic acid, a peptide, a cyclic peptide, a protein, an antibody, and a small molecule compound. A program according to any one of I1 to I4. [I6] The molecule is a protein, an antibody, a peptide, or a cyclic peptide, and the building block combination information is amino acid sequence information. A program according to any one of I1 to I5. [I7] The property is at least one of binding ability, pharmacological activity, physical properties, kinetics, and safety. The program according to any one of I1 to I6. [I8] The molecule is a molecule that binds to a target molecule, and the property is the ability to bind to the target molecule. A program according to any one of I1 to I7. [I9] calculating the predicted characteristic value and calculating an estimate of the uncertainty of the prediction; A program according to any one of I1 to I8. [I10] The method further comprises the step of calculating a characteristic quality value based on the characteristic prediction value and an estimate of the uncertainty of the prediction, The search procedure includes: identifying at least one molecule from the building block combination information set based on the characteristic quality value; The program described in I9. [I11] the characteristic quality value increases with increasing characteristic prediction value and decreases with decreasing prediction uncertainty estimate; Program described in I10. [I12] The characteristic quality value is an output value of a predetermined function having two input variables, the characteristic prediction value and an estimate of the uncertainty of the prediction. Program described in I11. [I13] The step of calculating the characteristic quality value includes calculating a mean variance as an objective function that gives the characteristic quality value. A program described in any one of I10 to I12. [I14] the prediction uncertainty estimate is the standard deviation of the property prediction; A program described in any one of I9 to I13. [I15] A tree-structured Parzen estimator is used as a combinatorial optimization algorithm for updating the building block combination information set. A program according to any one of I1 to I14. [I16] The prediction model is a model trained based on building block combination information that differs for each molecule and training data that indicates characteristics of the molecule. A program according to any one of I1 to I15.
[0015] [J1] 1. A method for producing a molecular compound, comprising: An acquisition step of acquiring a building block combination information set, which is a set of building block combination information of a plurality of different molecules, according to a combinatorial optimization algorithm; an input step of inputting the building block combination information set into a prediction model; an updating step of updating extraction parameters of the combinatorial optimization algorithm based on the property prediction values of each molecule included in the building block combination information set output from the prediction model so that molecules having more desired properties are included; searching for molecules having more desirable properties from the building block combination information set based on the property prediction value, and identifying the molecules as candidate molecules; an output step of outputting building block combination information relating to the candidate molecule; a generating step of generating the molecular compound having a molecular sequence indicated in the building block combination information; The method according to claim 1, [J2] and searching for the candidate molecule based on a first property predicted value calculated for a molecule that is an element of a first building block combination information set, which is a building block combination information set obtained before the extraction parameters are updated, and a second property predicted value calculated for a molecule that is an element of a second building block combination information set, which is a building block combination information set obtained after the extraction parameters are updated. The method described in J1. [J3] The first building block combination information set and the second building block combination information set are sets including building block combination information of different molecules. The method described in J2. [J4] The candidate molecule has a biological sequence. The method according to any one of J1 to J3. [J5] The building block combination information is sequence information of the molecule. The method according to any one of J1 to J4. [J6] The molecule is at least one of a nucleic acid, a peptide, a cyclic peptide, a protein, an antibody, and a small molecule compound. A method according to any one of J1 to J5. [J7] The molecule is a protein, an antibody, a peptide, or a cyclic peptide, and the building block combination information is amino acid sequence information. A method according to any one of J1 to J6. [J8] The property is at least one of binding ability, pharmacological activity, physical properties, kinetics, and safety. A method according to any one of J1 to J7. [J9] The molecule is a molecule that binds to a target molecule, and the property is the ability to bind to the target molecule. A method according to any one of J1 to J8. [J10] calculating the predicted characteristic and an estimate of the uncertainty of the prediction; A method according to any one of J1 to J9. [J11] calculating a characteristic quality value based on the characteristic prediction value and an estimate of the uncertainty of the prediction; The inference step of searching for the candidate molecule includes a step of identifying at least one molecule from the building block combination information set as the candidate molecule based on the characteristic quality value. The method described in J10. [J12] the characteristic quality value increases with increasing characteristic prediction value and decreases with decreasing prediction uncertainty estimate; The method described in J11. [J13] The characteristic quality value is an output value of a predetermined function having two input variables, the characteristic prediction value and an estimate of the uncertainty of the prediction. The method described in J12. [J14] The step of calculating the characteristic quality value includes a step of calculating a mean variance as an objective function that gives the characteristic quality value. A method according to any one of J11 to J13. [J15] the prediction uncertainty estimate is the standard deviation of the property prediction; A method according to any one of J10 to J14. [J16] A tree-structured Parzen estimator is used as a combinatorial optimization algorithm for updating the building block combination information set. A method according to any one of J1 to J15. [J17] A step of selecting a candidate molecule having the best experimental value from among experimental values of the properties of each of the plurality of candidate molecules. A method according to any one of J1 to J16.
[0016] [K] A computer system comprising a processor and a memory, the memory configured to store one or more instructions; The instructions include causing the processor to: calculating a predicted property value of a molecule that is an element of a building block combination information set, which is a set of building block combination information of a plurality of different molecules, using a prediction model for predicting properties of the molecule from the building block combination information of the molecule, and estimating the uncertainty of the prediction; searching for candidate molecules having more desirable properties based on the predicted property values and the uncertainty of the predictions; Computer system.
[0017] [L] A non-transitory computer-readable storage medium storing one or more instructions, comprising: The instructions are sent to the computer: calculating a predicted property value of a molecule that is an element of a building block combination information set, which is a set of building block combination information of a plurality of different molecules, using a prediction model for predicting properties of the molecule from the building block combination information of the molecule, and estimating the uncertainty of the prediction; searching for candidate molecules having more desirable properties based on the predicted property values and the uncertainty of the predictions; A non-transitory computer-readable storage medium.
[0018] [M] A computer system comprising a processor and a memory, the memory configured to store one or more instructions; The instructions include causing the processor to: obtaining a building block combination information set, which is a set of building block combination information of a plurality of different molecules, according to a combinatorial optimization algorithm; calculating a predicted property value of a molecule that is an element of the building block combination information set using a prediction model for predicting the property of the molecule from the building block combination information of the molecule; updating extraction parameters of the combinatorial optimization algorithm based on the predicted property values for each molecule so that molecules having more desirable properties are included; searching for molecules having desired properties based on the predicted property values; Computer system.
[0019] [N] A non-transitory computer-readable storage medium storing one or more instructions, comprising: The instructions are sent to the computer: obtaining a building block combination information set, which is a set of building block combination information of a plurality of different molecules, according to a combinatorial optimization algorithm; calculating a predicted property value of a molecule that is an element of the building block combination information set using a prediction model for predicting the property of the molecule from the building block combination information of the molecule; updating extraction parameters of the combinatorial optimization algorithm based on the predicted property values for each molecule so that molecules having more desirable properties are included; searching for molecules having desired properties based on the predicted property values; A non-transitory computer-readable storage medium. Effect of the Invention
[0020] According to one aspect of the present disclosure, the burden required for drug discovery can be further reduced. [Brief description of the drawings]
[0021] [Figure 1] FIG. 1 is a diagram showing an example of a drug discovery system including a molecular design device according to a first embodiment. [Diagram 2] FIG. 13 is a diagram showing an example of a drug discovery system including a molecular design device according to a second embodiment. [Diagram 3] FIG. 2 is an explanatory diagram for explaining an example of a plurality of pieces of molecular sequence information according to an embodiment of the present application. [Figure 4] FIG. 1 is a diagram showing an example of a hardware configuration of a molecular design apparatus according to an embodiment of the present application. [Diagram 5] FIG. 2 is a diagram showing an example of the configuration of a control unit according to an embodiment of the present application. [Figure 6] 4 is a flowchart showing an example of the flow of a process executed by the molecular design device according to the first embodiment. [Figure 7] 10 is a flowchart showing an example of a flow of a process executed by a molecular design device according to a second embodiment. [Figure 8] 13 is a flowchart showing an example of a combinatorial optimization process in the second embodiment. [Figure 9] 13 is a flowchart showing another example of the combinatorial optimization process in the second embodiment. [Figure 10] FIG. 1 is a schematic diagram of an optimization problem setting according to an application example of an embodiment of the present application. [Figure 11] FIG. 13 is a diagram showing an example of calculation of an objective function value by optimization in TPE according to the first verification example. [Figure 12] FIG. 13 is a graph showing the relationship between the predicted mean value and the predicted standard deviation value of the sequence obtained by sampling using TPE in the first verification example. [Figure 13] FIG. 13 is a graph showing the density distribution of edit distances of sequences obtained by sampling using TPE in the first verification example. [Figure 14] FIG. 13 is a graph showing the distribution of pseudo-correct model scores for the proposed sequence in the first validation example. [Figure 15] FIG. 13 is a diagram showing edit distances of proposed sequences in the first verification example. [Figure 16] FIG. 13 is a graph showing predicted standard deviation values of proposed sequences in the first verification example. [Figure 17] 13 is a table showing examples of amino acid candidates for each mutation candidate site in a search space according to a first verification example. [Figure 18] 13 is a table showing an example of parameter settings of a TPE according to a first verification example. [Figure 19]FIG. 13 is a diagram illustrating the distribution of predicted mean values and predicted standard deviation values of sequences obtained by sampling using TPE in the second verification example. [Figure 20] t-SNE visualization diagram illustrating sequences by sampling of TPE for the second validation example. [Figure 21] FIG. 13 is a diagram illustrating a predicted average value of a proposed sequence according to a second verification example. [Figure 22] FIG. 13 is a diagram illustrating the predicted variance values of the proposed sequences in the second verification example. [Figure 23] FIG. 13 is a diagram illustrating the distribution of expression levels of proposed sequences in the second verification example. [Figure 24] FIG. 13 is a diagram illustrating the distribution of octet values for each sequence according to the second verification example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0022] I. Definition The term "and / or" is used herein to refer to each of the objects listed before and after "and / or" or any combination thereof. For example, "A, B and / or C" includes each of the objects "A", "B", and "C", as well as the combinations "A and B", "A and C", "B and C", and "A and B and C".
[0023] ·amino acid As used herein, natural amino acids and non-natural amino acids may be included. In the case of natural amino acids, the amino acids are represented by one-letter code or three-letter code, or both, for example, Ala / A, Leu / L, Arg / R, Lys / K, Asn / N, Met / M, Asp / D, Phe / F, Cys / C, Pro / P, Gln / Q, Ser / S, Glu / E, Thr / T, Gly / G, Trp / W, His / H, Tyr / Y, Ile / I, Val / V.
[0024] Amino acid modification For modifying amino acids in the amino acid sequence of an antigen-binding molecule, known methods such as site-directed mutagenesis (Kunkel et al. (Proc. Natl. Acad. Sci. USA (1985) 82, 488-492)) and overlap extension PCR can be appropriately used. In addition, as a method for modifying amino acids by substituting amino acids other than natural amino acids, several known methods can also be used (Annu. Rev. Biophys. Biomol. Struct. (2006) 35, 225-249, Proc. Natl. Acad. Sci. USA (2003) 100 (11), 6353-6357). For example, a cell-free translation system (Clover Direct (Protein Express)) containing a tRNA in which a non-natural amino acid is bound to a complementary amber suppressor tRNA of the UAG codon (amber codon), which is one of the stop codons, may also be used.
[0025] ·antigen As used herein, the "antigen" is not limited to a specific structure as long as it contains an epitope to which an antigen-binding domain binds. In some embodiments, the antigen is a peptide, polypeptide, or protein of 4 or more amino acids. Examples of the above antigens include membrane molecules that are expressed on the cell membrane, and soluble molecules that are secreted outside the cells.
[0026] Antigen-binding domain In the present specification, the term "antigen-binding domain" may refer to any domain of any structure as long as it binds to the target antigen. Examples of such domains include the variable regions of the heavy and light chains of an antibody, a module called an A domain of about 35 amino acids contained in Avimer, a cell membrane protein present in the body (International Publication WO2004 / 044011, WO2005 / 040229), Adnectin (International Publication WO2002 / 032925) containing the 10Fn3 domain, which is a domain that binds to a protein in fibronectin, a glycoprotein expressed in the cell membrane, Affibody (International Publication WO1995 / 001937) using an IgG-binding domain consisting of a bundle of three helices consisting of 58 amino acids of Protein A as a scaffold, and DARPins (Designed Ankyrin Repeat (AR)) which are regions exposed on the molecular surface of ankyrin repeats (AR) having a structure in which a turn containing 33 amino acid residues, two antiparallel helices, and a loop subunit are repeatedly stacked. Examples of such structures include lipocalin molecules such as lipocalin proteins (International Publication No. WO 2002 / 020565), anticalin, which is a four loop region supporting one side of a barrel structure in which eight highly conserved antiparallel strands twist toward the center in lipocalin molecules such as neutrophil gelatinase associated lipocalin (NGAL) (International Publication No. WO 2003 / 029462), and a concave region of a parallel sheet structure inside a horseshoe-shaped structure in which leucine-rich-repeat (LRR) modules are repeatedly stacked in the variable lymphocyte receptor (VLR) that does not have an immunoglobulin structure and serves as the adaptive immune system of jawless fish such as lampreys and hagfish (International Publication No. WO 2008 / 016854). Examples of antigen-binding domains of the present disclosure include antigen-binding domains comprising the variable regions of the heavy and light chains of an antibody. Examples of such antigen-binding domains include "single chain Fv (scFv)", "single chain antibody", "Fv", "single chain Fv 2 (scFv2)", "Fab" or "F(ab')2".
[0027] ·Antigen-binding molecules In the present disclosure, an antigen-binding molecule containing an antigen-binding domain is used in the broadest sense, and specifically, various molecular types are included as long as they contain an antigen-binding domain. The antigen-binding molecule may be a molecule consisting of only an antigen-binding domain, or may be a molecule containing an antigen-binding domain and other domains. For example, when the antigen-binding molecule is a molecule in which an antigen-binding domain and an Fc region are bound, examples include a complete antibody and an antibody fragment. The antibody may include a single monoclonal antibody (including agonist and antagonist antibodies), a human antibody, a humanized antibody, a chimeric antibody, and the like. The antigen-binding molecule of the present disclosure may also include a scaffold molecule in which a three-dimensional structure such as an existing stable α / β barrel protein structure is used as a scaffold, and only a part of the structure is compiled into a library for the construction of an antigen-binding domain.
[0028] The term "antibody" is used herein in the broadest sense and encompasses a variety of antibody structures, including, but not limited to, monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments, so long as they exhibit the desired antigen-binding activity. Antibodies may be isolated from natural sources such as plasma or serum in which they naturally occur, or from the culture supernatant of hybridoma cells that produce the antibodies, or may be partially or completely synthesized by using techniques such as genetic recombination. Examples of antibodies include immunoglobulin isotypes and their isotype subclasses. Nine classes (isotypes) of human immunoglobulins are known: IgG1, IgG2, IgG3, IgG4, IgA1, IgA2, IgD, IgE, and IgM. The antibodies of the present disclosure may include IgG1, IgG2, IgG3, and IgG4 among these isotypes. As for the constant regions of human IgG1, human IgG2, human IgG3, and human IgG4, multiple allotype sequences due to genetic polymorphism are described in Sequences of proteins of immunological interest, NIH Publication No. 91-3242, and any of these may be used in the present disclosure. In particular, as the sequence of human IgG1, the amino acid sequence at positions 356-358 as represented by EU numbering may be DEL or EEM. Furthermore, for the human Igκ (Kappa) constant region and the human Igλ (Lambda) constant region, multiple allotype sequences due to genetic polymorphisms are described in Sequences of proteins of immunological interest, NIH Publication No. 91-3242, and any of these may be used in the present disclosure.
[0029] "Antibody fragment" refers to a molecule other than a complete antibody that contains a portion of the complete antibody that binds to the antigen to which the complete antibody binds. Examples of antibody fragments include, but are not limited to, Fv, Fab, Fab', Fab'-SH, F(ab')2; diabodies; linear antibodies; single-chain antibody molecules (e.g., scFv); and multispecific antibodies formed from antibody fragments.
[0030] The terms "full length antibody," "complete antibody," and "whole antibody" are used interchangeably herein and refer to an antibody having a structure substantially similar to a native antibody structure or having a heavy chain that includes an Fc region as defined herein.
[0031] The term "variable region" or "variable domain" refers to the domain of an antibody's heavy or light chain that is involved in binding the antibody to an antigen. The variable domains of the heavy and light chains (VH and VL, respectively) of natural antibodies usually have a similar structure, with each domain containing four conserved framework regions (FR) and three hypervariable regions (HVR). (See, for example, Kindt et al. Kuby Immunology, 6th ed., WH Freeman and Co., page 91 (2007).) One VH or VL domain may be sufficient to confer antigen-binding specificity. Furthermore, antibodies that bind to a particular antigen may be isolated by screening a complementary library of VL or VH domains, respectively, with a VH or VL domain from an antibody that binds to the antigen. See, e.g., Portolano et al., J. Immunol. 150:880-887 (1993); Clarkson et al., Nature 352:624-628(1991).
[0032] In this specification, the term "molecular weight" refers to the sum of the atomic weights of the atoms constituting a compound molecule (unit: "g / mol"), and is obtained by calculating the sum of the atomic weights of the atoms contained in the molecular structure. In this specification, the unit of molecular weight may be omitted. The molecular weight can be measured, for example, by liquid chromatography mass spectrometry (LC / MS).
[0033] As used herein, a medium molecular weight compound, or a medium molecule, is a compound having a molecular weight of 500 g / mol or more and less than 30,000 g / mol. The medium molecular weight compound is, for example, a compound having a molecular weight of 500 g / mol or more and less than 6000 g / mol, may be a compound having a molecular weight of 500 g / mol or more and less than 4000 g / mol, may be a compound having a molecular weight of 600 g / mol or more and less than 4000 g / mol, or may be a compound having a molecular weight of 700 g / mol or more and less than 3000 g / mol. The medium molecular weight compound is, for example, a peptide compound containing a peptide chain, a nucleic acid, or a sugar chain, and may be a peptide compound, a peptide compound containing 5 to 30 amino acid residues, a peptide compound containing 7 to 25 amino acid residues, or a peptide compound containing 9 to 20 amino acid residues. The medium molecular weight compound is, for example, a peptide compound having a molecular weight of 500 g / mol or more and less than 6000 g / mol, or may be a peptide compound having a molecular weight of 500 g / mol or more and less than 4000 g / mol, or may be a peptide compound having a molecular weight of 600 g / mol or more and less than 4000 g / mol, or may be a peptide compound having a molecular weight of 700 g / mol or more and less than 3000 g / mol.
[0034] In this specification, a low molecular weight compound is a compound having a molecular weight of less than 500 g / mol, and a high molecular weight compound is a compound having a molecular weight of 30,000 g / mol or more.
[0035] (Definition of Experimental Method) [Characteristics evaluation] In one embodiment of the present disclosure, a prediction model is generated by performing machine learning based on molecular sequence information and evaluation result information of the molecular property evaluation. Non-limiting examples of molecular property evaluation include, but are not limited to, molecular binding ability evaluation, pharmacological activity evaluation, physical property evaluation, kinetic evaluation, and safety evaluation.
[0036] Binding Ability Assessment The method for evaluating the binding ability of a target molecule-binding molecule to a target molecule is not particularly limited, but it can be performed by quantitatively evaluating the binding of the target molecule-binding molecule to the target molecule. The target molecule is, for example, a target protein. The target molecule-binding molecule is, for example, an antigen-binding molecule, and the target molecule is, for example, an antigen. For example, when the target molecule is an antigen, it can be evaluated by measuring the binding activity of the antigen-binding molecule and the antigen. Binding activity refers to the total strength of non-covalent interactions between one or more binding sites of a molecule (e.g., an antibody) and the binding partner of the molecule (e.g., an antigen). Here, "binding activity" is not strictly limited to 1:1 interactions between members of a binding pair (e.g., an antibody and an antigen). For example, when the members of a binding pair reflect a monovalent 1:1 interaction, the binding activity refers to the inherent binding affinity (sometimes simply referred to as "affinity"). When the members of a binding pair are capable of both monovalent and multivalent binding, the binding activity is the sum of these binding forces. The binding activity of a molecule X to its partner Y can generally be expressed by a dissociation constant (KD) or "amount of analyte bound per unit amount of ligand". Binding activity can be measured by a conventional method known in the art, including those described herein. Conditions other than the concentration of the target tissue-specific compound can be appropriately determined by those skilled in the art. In a specific embodiment, the antigen-binding molecule provided herein is an antibody, and the binding activity of the antibody is ≦1 μM, ≦100 nM, ≦10 nM, ≦1 nM, ≦0.1 nM, ≦0.01 nM, or ≦0.001 nM (e.g., 10 -8 M or less, e.g. 10 -8 M~10 -13 M, for example 10 -9 M~10 -13 is the dissociation constant (KD) of the
[0037] In one embodiment, the binding activity of the antibody is measured by a ligand capture method using, for example, BIACORE® T200 or BIACORE® 4000 (GE Healthcare, Uppsala, Sweden), which uses surface plasmon resonance analysis as a measurement principle. For example, BIACORE® Control Software is used to operate the instrument. In one embodiment, an amine coupling kit (GE Healthcare, Uppsala, Sweden) is used according to the supplier's instructions to immobilize a ligand capture molecule, such as an anti-tag antibody, an anti-IgG antibody, or protein A, on a carboxymethyldextran-coated sensor chip (GE Healthcare, Uppsala, Sweden). The ligand capture molecule is diluted with a 10 mM sodium acetate solution of appropriate pH and injected at an appropriate flow rate and injection time. The binding activity is measured using a buffer containing 0.05% polysorbate 20 (also known as Tween (registered trademark) 20) as the measurement buffer, at a flow rate of 10-30 μL / min, and at a measurement temperature of, for example, 25°C or 37°C. When performing measurement by having the ligand capture molecule capture an antibody as a ligand, the antibody is injected to capture a desired amount, and then a serial dilution (analyte) of the antigen and / or Fc receptor prepared using the measurement buffer is injected. When performing measurement by having the ligand capture molecule capture an antigen and / or Fc receptor as a ligand, the antigen and / or Fc receptor is injected to capture a desired amount, and then a serial dilution (analyte) of the antibody prepared using the measurement buffer is injected.
[0038] In one embodiment, the measurement results are analyzed using BIACORE® Evaluation Software. Kinetics parameter calculation is performed by simultaneously fitting the binding and dissociation sensorgrams using a 1:1 binding model, and the binding rate (kon or ka), dissociation rate (koff or kd), and equilibrium dissociation constant (KD) can be calculated. When the binding activity is weak, particularly when dissociation is rapid and kinetic parameter calculation is difficult, the equilibrium dissociation constant (KD) may be calculated using a steady state model. As another parameter of binding activity, the "analyte binding amount per unit amount of ligand" can also be calculated by dividing the binding amount (RU) of a specific concentration of analyte by the capture amount (RU) of the ligand.
[0039] As the value of antigen-binding activity, KD (dissociation rate constant) can be used when the antigen is a soluble molecule, and apparent kd (apparent dissociation rate constant) can be used when the antigen is a membrane molecule. kd (dissociation rate constant) and apparent KD (apparent dissociation rate constant) can be measured by methods known to those skilled in the art, for example, Biacore (GE Healthcare), a flow cytometer, etc.
[0040] As a different aspect of the characteristic evaluation, a selection method of antigen-binding molecules using a display library can be mentioned. In one aspect, panning using phage display can be mentioned. As an example of affinity evaluation, a phage library displaying multiple different antigen-binding molecules can be prepared, and the target antigen can be contacted with the prepared phage, and then unbound phages can be washed to concentrate the phages displaying antigen-binding molecules that interact with the target antigen. By analyzing the nucleic acid sequence encoding the antigen-binding molecule contained in the concentrated phage, it is possible to identify a sequence having affinity for the target antigen. In another aspect, panning using mammalian cell display can be mentioned. As an example of pharmacological activity evaluation using the display system, a library containing multiple different antigen-binding molecules can be expressed in a target mammalian cell, and reporter activity, etc. can be changed depending on the action that the library shows on the same cell, so that cells having antigen-binding molecule genes with desired pharmacological activity can be isolated using a flow cytometer or the like. As an example of physical property evaluation using the display system, a library containing multiple different antigen-binding molecules is expressed in target mammalian cells, and the expression level is stained with an antibody specific to the antigen-binding molecule, allowing cells having an antigen-binding molecule gene that can be stably and highly expressed to be isolated using a flow cytometer, etc. Evaluation of the characteristics of antigen-binding molecules by panning is not limited to techniques using phages or mammalian cells, and various techniques can be used as long as the antigen-binding molecule can be displayed, including, but not limited to, techniques in which the antigen-binding molecule is displayed on ribosomes, on mRNA, on viruses other than phages, and on bacteria such as Escherichia coli.
[0041] As a different aspect of the characteristic evaluation, there is a method of obtaining an antibody gene sequence from immune cells derived from an individual, or a method of obtaining an antibody protein sequence from serum. Taking affinity evaluation in which an antibody gene sequence is extracted from immune cells as an example, it is possible to identify a sequence having affinity for the target antigen by inducing immune sensitization by administering a target antigen protein to an individual and extracting genes from immune cells having an antibody gene that binds to the target antigen. The antigen that induces immune sensitization is not limited to the above-mentioned protein, but may be a gene encoding the protein or a cell expressing the protein. The target individual may be, but is not limited to, a human, a mouse, a rat, a hamster, a rabbit, a monkey, a chicken, a camel, a llama, or an alpaca. Furthermore, examples of methods for analyzing the nucleic acid sequence or occurrence frequency include, but are not limited to, a method of cloning a genetically modified organism having the nucleic acid sequence of each antigen-binding molecule and analyzing the cloned organism by the Sanger method using capillary electrophoresis, and a method of analyzing the cloned organism using a next-generation sequencer. When analyzing the nucleic acid sequence, it is also possible to judge the strength of a characteristic based on the frequency of occurrence. For example, it is possible to estimate that antigen-binding molecules encoded by sequences that occur frequently after enrichment have high characteristics, and that antigen-binding molecules encoded by sequences that occur less frequently after enrichment have lower characteristics than antigen-binding molecules encoded by sequences that occur more frequently. Furthermore, the techniques for obtaining information on antigen-binding molecules derived from the display library or an individual can be applied to various property evaluations and are not limited to those mentioned above.
[0042] Pharmacological activity evaluation The method for evaluating the pharmacological activity of a molecule is not particularly limited, and can be, for example, evaluated by measuring the neutralizing activity, agonist activity, or cytotoxic activity exhibited by the molecule. When cytotoxic activity evaluation is taken as an example of pharmacological activity evaluation, examples include antibody-dependent cell-mediated cytotoxicity (ADCC) activity, complement-dependent cytotoxicity (CDC) activity, T-cell-dependent cytotoxicity (TDCC) activity, and antibody-dependent cellular phagocytosis (ADCP) activity. CDC activity means cytotoxic activity by the complement system. ADCC activity means an activity in which immune cells or the like bind to the Fc region of an antigen-binding molecule containing an antigen-binding domain that binds to a membrane-type molecule expressed on the cell membrane of a target cell via an Fcγ receptor expressed on the immune cell, and the immune cell inflicts damage on the target cell. Furthermore, TDCC activity refers to the activity of a T cell to damage a target cell by bringing the target cell and the T cell into close proximity using a bi-specific antibody that contains an antigen-binding domain that binds to a membrane molecule expressed on the cell membrane of the target cell and an antigen-binding domain for any of the constituent subunits of the T cell receptor (TCR) complex on the T cell, particularly an antigen-binding domain that binds to the CD3 epsilon chain. Whether or not an antigen-binding molecule of interest has ADCC activity, CDC activity, TDCC activity, or ADCP activity can be measured by known methods. Neutralizing activity refers to an activity that inhibits the biological activity of a ligand that has biological activity against cells, such as a virus or a toxin. That is, a substance that has neutralizing activity refers to a substance that binds to the ligand or a receptor to which the ligand binds, and inhibits the binding of the ligand to the receptor. A receptor that has been prevented from binding to a ligand by neutralizing activity cannot exert biological activity through the receptor. When the antigen-binding molecule is an antibody, an antibody having such neutralizing activity is generally called a neutralizing antibody, and the neutralizing activity can be measured by measuring the inhibitory activity of the binding of the ligand to the receptor. Ligands that have biological activity against cells are not limited to viruses and toxins, and the inhibitory activity of physiological actions caused by the binding of endogenous ligands such as cytokines and chemokines to receptors is also understood as neutralizing activity. In addition, neutralizing activity is not limited to the case of inhibiting the binding of a ligand to a receptor, but also includes the activity of inhibiting the function of a protein having biological activity, and an example of the function of the protein is enzyme activity.
[0043] Physical property evaluation The method for evaluating the physical properties of a molecule is not particularly limited, and examples of the physical properties include thermal stability, chemical stability, solubility, viscosity, light stability, long-term storage stability, non-specific adsorption, lipophilicity, and membrane permeability, and the various physical properties exemplified above can be measured by methods known to those skilled in the art. The evaluation method is not particularly limited, and stability evaluations such as thermal stability, chemical stability, light stability, stability against mechanical stimuli, and long-term storage stability can be evaluated by measuring the decomposition, chemical modification, and association of the molecule before and after the treatments such as heat treatment, exposure to a low pH environment, light exposure, mechanical stirring, and long-term storage that are the object of the stability evaluation. Non-limiting examples of the measurement method for performing such stability evaluation include, but are not limited to, methods using chromatography such as ion exchange chromatography and size exclusion chromatography, mass spectrometry, and electrophoresis, and can be measured by various methods known to those skilled in the art. Other examples of physical property evaluations include, but are not limited to, evaluation of protein solubility using polyethylene glycol precipitation method, evaluation of viscosity using small angle X-ray scattering method, and evaluation of non-specific binding based on binding to extra cellular matrix (ECM). In addition, physical property evaluations such as protein expression level, binding to a purification resin or purification ligand, and surface charge can also be performed as long as they are measurable by methods known to those skilled in the art.
[0044] Dynamic evaluation The method for evaluating the kinetics of a molecule is not particularly limited, but it can be evaluated by administering the molecule to animals such as mice, rats, monkeys, dogs, etc., and measuring the amount of the molecule in the blood over time after administration, and can be evaluated by a method widely known to those skilled in the art as Pharmacokinetics (PK) evaluation. In addition to the method of directly evaluating PK, it is also possible to predict the kinetic behavior from the amino acid sequence of the molecule by calculating the surface charge, isoelectric point, etc. of the molecule on software.
[0045] Safety evaluation The method for evaluating the safety of a molecule is not particularly limited, and examples thereof include immunogenicity prediction tools such as ISPRI Web-Based Immunogenicity Screening (EpiVax), HLA binding evaluation of fragment peptides of antigen-binding molecules, detection of T cell epitopes and evaluation of immunogenicity using MAPPs (MHC-Associated Peptide Proteomics) or T cell proliferation evaluation, etc. In addition, evaluation can be performed as long as it is measurable by methods known to those skilled in the art, such as binding to rheumatoid factor (RF), evaluation of immune responses using PBMC or whole blood, and platelet aggregation evaluation.
[0046] (Definitions of terms and methods used in machine learning) ·MBO(Model Based Optimization) MBO (Model Based Optimization) is an optimization method that does not directly optimize the characteristic values of a certain characteristic, but instead optimizes the predicted characteristic values of a certain characteristic, which are estimated using some kind of model.
[0047] ·TPE(Tree-structured Parzen Estimator) TPE (Tree-structured Parzen Estimator) is a type of Bayesian optimization. TPE involves the process of calculating the expected improvement of the output value for a certain input value based on the conditional probability of the output value for the input value and the probability for the output value for the function to be optimized. In other words, TPE is a method of optimizing a function using input values that maximize the expected improvement calculated in this way.
[0048] (Manufacturing method of the object) In one aspect of the present disclosure, methods for producing objects suitable for obtaining drug candidates can be carried out by methods known to those skilled in the art. When the subject is an antibody, the antibody can be produced using recombinant methods or constructs, for example, as described in U.S. Pat. No. 4,816,567. In one embodiment, a method of making an antibody includes culturing a host cell containing a nucleic acid encoding the antibody under conditions suitable for expression of the antibody candidate molecule compound described herein, and optionally recovering the antibody from the host cell (or host cell culture medium). The isolated nucleic acid encoding the antibody may encode an amino acid sequence comprising the VL and / or an amino acid sequence comprising the VH of the antibody (e.g., the light chain and / or the heavy chain of the antibody). The host cell containing such a nucleic acid includes (e.g., is transformed with) (1) a vector containing a nucleic acid encoding an amino acid sequence comprising the VL and an amino acid sequence comprising the VH of the antibody, or (2) a first vector containing a nucleic acid encoding an amino acid sequence comprising the VL of the antibody and a second vector containing a nucleic acid encoding an amino acid sequence comprising the VH of the antibody. In one embodiment, the host cell is eukaryotic (e.g., Chinese Hamster Ovary (CHO) cells) or lymphoid cells (e.g., Y0, NS0, Sp2 / 0 cells). Suitable host cells for cloning or expressing antibody-encoding vectors include prokaryotic or eukaryotic cells. For example, antibodies may be produced in bacteria, particularly if glycosylation and Fc effector functions are not required. For expression of antibody fragments and polypeptides in bacteria, see, for example, U.S. Patent Nos. 5,648,237, 5,789,199, and 5,840,523. (See also Charlton, Methods in Molecular Biology, Vol. 248 (BKC Lo, ed., Humana Press, Totowa, NJ, 2003), pp. 245-254, which describes expression of antibody fragments in E. coli.) After expression, the antibody may be isolated in a soluble fraction from the bacterial cell paste and can be further purified.
[0049] When the target substance is a peptide compound or a cyclic peptide compound, it can be produced by liquid phase synthesis, solid phase synthesis using Fmoc synthesis, Boc synthesis, or a combination thereof. Liquid phase synthesis and solid phase synthesis can be carried out by methods well known to those skilled in the art. Solid phase synthesis is a method in which a compound is bound to a solid, and the compound is chemically reacted with a reagent on the solid resin to synthesize the target compound. Solid phase peptide synthesis is a method in which a desired amino acid or peptide is bound to a solid resin, and further desired amino acids or peptides are sequentially linked to the amino acids or peptides bound to the solid resin to extend the peptide chain and synthesize the peptide. The peptide bound to this solid resin can be separated from the solid resin to obtain the target peptide.
[0050] First embodiment FIG. 1 is a block diagram showing an example of a drug discovery system 100 including a molecular design apparatus 1 according to the first embodiment. The drug discovery system 100 is a system for creating new objects suitable as drug candidates. The system provides a method for generating new objects having predetermined properties, such as a specific biological activity (e.g., binding to a specific protein). Drugs include, but are not limited to, small molecule drugs, medium molecule drugs, biological drugs, cells, nucleic acid drugs, biopharmaceuticals, or potential active agents such as other active agents. Objects include molecular structures that have a desired or defined biological activity (e.g., binding to a specific protein preferentially over other proteins). Molecules that are drug candidates include biomolecules and compounds, including various molecules such as nucleic acids, peptides, cyclic peptides, proteins, antibodies, target molecule binding molecules, polymeric compounds, medium molecule compounds, and small molecule compounds. Although not described in this specification, the drug discovery system 100 may include a selection device for a molecule that interacts with a drug target, a lead molecule creation device, etc. The drug discovery system 100 may be, for example, an information processing system configured including the disclosure of WO2020 / 246617.
[0051] The drug discovery system 100 includes a molecular design device 1. The molecular design device 1 searches for candidate molecules having desired properties and outputs information on the identified candidates. The output information is building block combination information of the candidate molecules. In other words, the molecular design device 1 identifies candidate molecules and outputs building block combination information of the identified candidate molecules. Here, a candidate molecule is a molecule that is expected to have desired properties. The candidate molecule building block combination information is information about the candidate molecule, and is building block combination information of some or all of the candidate molecule. The output candidate molecule building block combination information may include information indicating one candidate molecule, or may include information indicating multiple candidate molecules.
[0052] The drug discovery system 100 uses the candidate molecule building block combination information output from the molecular design device 1 to select a new object suitable for a drug candidate. For example, the drug discovery system 100 generates a candidate molecule based on the candidate molecule building block combination information output from the molecular design device 1, experimentally evaluates the properties of the candidate molecule, and selects a molecule having desired properties as a new object suitable for a drug candidate based on the results of the property evaluation. That is, the drug discovery system 100 can create a new object suitable for a drug candidate based on the results of the property evaluation performed by actually generating a candidate molecule using the candidate molecule building block combination information output from the molecular design device 1. In such a case, the candidate molecule can be said to be a molecule that can be a verification target for narrowing down candidates for the main component of a pharmaceutical.
[0053] The building block combination information of a molecule is information on a combination of some or all of the building blocks of the molecule. When the building block combination information is information on a combination of some of the building blocks of a molecule, the range of the sequence may be arbitrarily settable. A building block is a unit that constitutes a molecule. Building block combination information of a molecule relates to a combination of building blocks that constitute the molecule. In this application, an array containing individual components may be referred to as a "combination." The term "sequence" may also be used as an example of a building block combination.
[0054] The molecule to be designed is, for example, a protein. When the molecule is a protein, the building blocks are amino acids, and the building block combination information of the molecule is, for example, information of the amino acid sequence of the protein. The molecule to be designed is, for example, a nucleic acid. When the molecule is a nucleic acid, the building blocks are nucleotides. The sequence of the molecule is, for example, the nucleotide sequence of the nucleic acid. The molecular building block combination information is information about the nucleotide sequence. More specifically, when the designed molecule is an antibody, the molecular sequence is an amino acid sequence, and the building block is an amino acid. The building block combination information of the molecule is, for example, the full-length antibody sequence, for example, the amino acid sequence of VH or VL, or the sequence of a part of the antibody, such as CDR, FR, etc. For example, when the molecule to be designed is a cyclic peptide, the molecular sequence is an amino acid sequence containing unnatural amino acids, and the building blocks are natural amino acids and unnatural amino acids. The building block combination information of the molecule is information on the amino acid sequence containing the unnatural amino acids. When the molecule to be designed is a small molecule, the building block combination of the molecule is a combination of fragments, and the building blocks are fragments (fragment molecules that make up a small molecule). Also, for example, when the molecule to be designed is a nucleic acid, the molecular sequence is the base sequence and the building blocks are bases.
[0055] The desired characteristic is a characteristic required for a new object suitable for a drug candidate, and can be set arbitrarily. Non-limiting examples of the characteristic include binding ability to a predetermined in vivo target, binding ability, pharmacological activity, physical properties, kinetics, and safety, but are not limited thereto. For example, if the molecule to be designed is an antibody, the characteristic is, for example, the binding ability of the drug to a predetermined antigen. For example, if the molecule to be designed is mRNA (messenger-RNA), the characteristic is, for example, the translation ability of the protein.
[0056] The molecular design device 1 is an example of an inference device that infers a molecule having desired properties from prediction information described below. Examples of the desired properties include the ability to bind to a target molecule, efficacy, and drug-like properties including membrane permeability. Hereinafter, inferring a molecule having desired properties from prediction information may be rephrased as identifying a candidate molecule. In this embodiment, a molecule expected to have a desired property is one that exhibits good prediction value and low prediction uncertainty by a predictive model for the desired property.
[0057] The following describes the case where sequence information is processed as building block combination information. The molecular design apparatus 1 includes an inference unit 111 including, for example, a sequence information processing unit 111a, a property prediction unit 111b, a prediction information processing unit 111c, and a candidate molecule specifying unit 111d. The sequence information processing unit 111a prepares a sequence information set and outputs the prepared sequence information set to the property prediction unit 111b. The sequence information set is a set of sequence information of multiple molecules. The sequence information processing unit 111a may autonomously generate sequence information for each molecule, or may input sequence information for each molecule from another device. The sequence information set may be referred to as a building block combination information set, and the sequence information processing unit 111a may be referred to as a building block combination information processing unit. The property prediction unit 111b predicts the properties of a molecule for each element of the sequence information set of the sequence information input from the sequence information processing unit 111a, and outputs prediction information relating to the predicted properties to the prediction information processing unit 111c. The prediction information processing unit 111c acquires prediction information for each molecule from the property prediction unit 111b, and outputs the acquired prediction information to the candidate molecule identifying unit 111d. The candidate molecule specifying unit 111d specifies sequence information of at least one candidate molecule based on the prediction information. The candidate molecule specifying unit 111d may output output data indicating the specified sequence information to another device, or may store the output data in the storage unit 14 (described later).
[0058] In this embodiment, the sequence information set may be a set of sequence information indicating a virtual molecular sequence (sometimes called a "virtual sequence") generated by machine learning, or may be a set of sequence information of a real molecular sequence (sometimes called a "real sequence"), or may be a set including both sequence information indicating a virtual sequence and sequence information indicating a real sequence. For example, the sequence information set may have sequence information of a virtual sequence generated by a virtual sequence generation model. The sequence information set may also have sequence information indicating a real sequence obtained, for example, from an existing database or as an experimental result. It may also be sequence information extracted from all candidate combinations of building blocks by combinatorial optimization, which will be described later. The sequence information processing unit 111a performs a process of preparing a sequence information set. The sequence information processing unit 111a may further include a sequence information set acquisition unit (not shown) that executes an acquisition process for acquiring sequence information that is an output for certain input information using a machine learning model. In this case, specifically, the sequence information processing unit 111a acquires a sequence information set and outputs the acquired sequence information set to the characteristic prediction unit 111b.
[0059] In this embodiment, sequence information of a plurality of molecules is input to the property prediction unit 111b from the sequence information processing unit 111a. The property prediction unit 111b calculates a property prediction value and an estimate of the uncertainty of the prediction for each of the sequence information of a plurality of input molecules. In this embodiment, the property prediction unit 111b includes a predicted value calculation unit 111x and a prediction uncertainty estimation unit 111y. The predicted value calculation unit 111x inputs molecular sequence information and calculates a property predicted value using a prediction model. The property predicted value is a property value predicted using the prediction model. The prediction uncertainty estimation unit 111y estimates the uncertainty of the property predicted value calculated by the predicted value calculation unit 111x. A value indicating the estimated uncertainty is called an estimated value of prediction uncertainty. The predictive model is generated, for example, by learning based on training data having sequence information of individual molecules and multiple sets of characteristic evaluations for the molecules. The predicted value calculation unit 111x may calculate characteristic predicted values for a plurality of characteristics. When characteristic values are predicted for a plurality of characteristics, the prediction uncertainty estimation unit 111y may estimate the prediction uncertainty for the characteristic predicted value of each of the characteristics, or may estimate the prediction uncertainty for any one of the characteristic predicted values. The characteristic prediction unit 111b outputs, as prediction information for each molecule, a characteristic prediction value and an estimate of the prediction uncertainty of the characteristic prediction value to the prediction information processing unit 111c.
[0060] The candidate molecule identifying unit 111d performs a process of inferring molecules having desired properties based on the prediction information, and identifies candidate molecules. In this embodiment, the candidate molecule identifying unit 111d infers a molecule having a desired property based on the property prediction value and the estimated value of the prediction uncertainty for the sequence information of each molecule output from the property prediction unit 111b. The candidate molecule identifying unit 111d identifies at least one candidate molecule from a plurality of molecules whose sequence information is included in the sequence information set from the sequence information processing unit 111a based on the prediction information obtained from the prediction information processing unit 111c. In identifying the candidate molecule, the candidate molecule identifying unit 111d may use a property prediction value for at least one property, or may use property prediction values for multiple properties. In identifying the candidate molecule, the candidate molecule identifying unit 111d may use multiple property prediction values and an estimated value of the prediction uncertainty for at least one property prediction value as prediction information. That is, in identifying the candidate molecule, the constraint of the second property value (the prediction uncertainty for the property prediction value) may be taken into consideration when optimizing the first property value (the property prediction value). Furthermore, the candidate molecule identifying unit 111d may identify molecules having desired characteristics as candidate molecules based on characteristic quality values, which will be described later. The candidate molecule identifying unit 111d may select, for example, at least one molecule whose predicted property value and prediction uncertainty each satisfy a predetermined condition as a candidate molecule. The candidate molecule identifying unit 111d may select, for example, at least one molecule whose predicted property values and an estimated value of the prediction uncertainty for at least one predicted property value each satisfy a predetermined condition as a candidate molecule.
[0061] The candidate molecule identifying unit 111d may identify at least one candidate molecule based on a characteristic quality value calculated based on the characteristic prediction value and an estimate of the uncertainty of the prediction. In this case, the characteristic quality value may be calculated, for example, in the prediction information processing unit 111c or in the characteristic prediction unit 111b. The characteristic quality value is a value based on the characteristic prediction value and an estimate of the uncertainty of the prediction. The characteristic quality value can also be regarded as a response variable calculated using a predetermined function with the characteristic prediction value and the estimate of the uncertainty of the prediction as explanatory variables. For example, the characteristic quality value may be an index value that gives a smaller value as the estimated value of the uncertainty of the characteristic prediction becomes larger, and gives a larger value as the characteristic prediction value becomes larger. In other words, the characteristic quality value may be any value that increases as the characteristic prediction value increases and decreases as the estimated value of the uncertainty of the prediction increases. As an example, the characteristic quality value is the difference between the characteristic prediction value and a predetermined coefficient multiplied by the estimated value of the uncertainty of the prediction in a linear region, or the difference between the characteristic prediction value multiplied by the predetermined coefficient and the estimated value of the uncertainty of the prediction. The predetermined coefficient is a positive constant and indicates the degree of contribution of the characteristic value or the estimated value of the uncertainty of the prediction to the characteristic quality value. Therefore, when a property prediction value is defined that gives a larger value the more a desired property is predicted to be possessed or the degree of the property is higher, the property quality value is an index that shows a larger value as the property prediction value increases and the estimated value of the prediction uncertainty decreases. In this case, a larger property quality value means that the molecule actually generated is more likely to exhibit the expected property. In other words, it can be said that estimation processing using property quality values in molecular design realizes risk-averse estimation processing.
[0062] The characteristic quality value does not have to be the difference between the characteristic prediction value in the linear domain and a predetermined coefficient multiplied by the estimated value of the prediction uncertainty. The characteristic quality value may be a value obtained by raising either the characteristic prediction value or the estimated value of the prediction uncertainty to a predetermined number of times and calculating the difference or quotient between the other value. In addition, when a prediction value that gives a smaller value the higher the possibility of having a desired characteristic or the degree of the characteristic is defined, the characteristic prediction value may be determined by dividing the characteristic prediction value by the estimated value of the prediction uncertainty or by other types of calculations such as calculations in the logarithmic domain. The characteristic quality value may be calculated using both the prediction value and its uncertainty, and there are no limitations on the function, procedure, etc.
[0063] The candidate molecule identifying unit 111d identifies a molecule having a desired characteristic based on the calculated characteristic quality value. As an example, the candidate molecule identifying unit 111d increases the likelihood that a molecule having a desired characteristic will be identified as having a desired characteristic if the molecule has a better characteristic quality value among the molecules for which the characteristic prediction unit 111b has obtained a characteristic prediction value. Here, "good characteristic quality value" means that a characteristic quality value is obtained that indicates the possibility of being predicted to have a desired characteristic or that the degree of the characteristic is high. As another example, the candidate molecule identifying unit 111d may identify, as a molecule having a desired property, a molecule whose property quality value satisfies a predetermined condition among the molecules whose property predicted values have been calculated by the property prediction unit 111b. As yet another example, the candidate molecule identifying unit 111d may rank and rearrange a plurality of molecules whose property predicted values have been calculated by the property prediction unit 111b using the property quality value, and identify molecules within a predetermined range from the rearranged molecules as molecules having a desired property. Here, the candidate molecule identifying unit 111d may identify a certain number of a plurality of molecules as molecules having a desired property so that molecules having better property quality values are given priority. Furthermore, the candidate molecule identifying unit 111d may identify, as a molecule having a desired property, a molecule whose property predicted value and an estimated value of the uncertainty of the prediction each satisfy a predetermined condition.
[0064] The accuracy of the property prediction value calculated by the property prediction unit 111b affects the identification of the molecule. Depending on the training data of the property prediction model, the accuracy of property prediction based on sequence information varies. Here, the accuracy of prediction may be rephrased as the reliability of prediction or the certainty of prediction. Low prediction accuracy can be rephrased as low prediction reliability, low prediction certainty, or high prediction uncertainty. High prediction accuracy can be rephrased as high prediction reliability, high prediction certainty, or low prediction uncertainty. In the molecular design device 1 of this embodiment, not only the predicted property value but also the estimated value of the uncertainty of the prediction is used as an index in the inference process of a molecule having a desired property. That is, in the molecular design device 1 of this embodiment, when identifying a candidate molecule, the candidate molecule identifying unit uses a mathematical model that predicts the property value and estimates the uncertainty of the prediction, rather than a mathematical model that predicts the property value. This makes it possible to prevent a molecule with a high predicted property value but a low prediction probability from being identified as a candidate molecule by the inference process. That is, when a molecule is actually generated, it is possible to prevent the identification of a molecule that is highly likely to not have the properties as expected. As a result, this embodiment can further reduce the burden required for drug discovery.
[0065] Second Embodiment 2 is a block diagram showing an example of a drug discovery system 100 including a molecular design apparatus 1 according to the second embodiment. The following description will focus on the differences from the first embodiment. Unless otherwise specified, the description of the first embodiment will be used for the points in common.
[0066] The molecular design device 1 according to this embodiment is an example of an inference device that infers molecules expected to have desired properties from prediction information, which will be described later. The molecular design apparatus 1 includes an inference unit 111 including, for example, a sequence information processing unit 111a, a property prediction unit 111b, a prediction information processing unit 111c, and a candidate molecule specifying unit 111d.
[0067] The inference unit 111 according to this embodiment executes a combinatorial optimization algorithm (sometimes simply referred to as "combinatorial optimization" in this application) to infer molecules that are expected to have desired properties. The sequence information processing unit 111a executes combinatorial optimization to extract sequence information of an arbitrary number of molecules from a preset search space. That is, according to the sequence information processing unit 111a, a process of extracting sequence information of a part of molecules expected to have a desired property from candidates of sequences of building blocks in the search space by combinatorial optimization using updated extraction parameters is repeated. The search space can be set arbitrarily. For example, the search space may be set based on training data used for training a prediction model used for property prediction. The number of times of sequence information acquisition, that is, the number of times of repetition of sequence information acquisition, may be any natural number equal to or greater than 1. The number of times of sequence information acquisition may be set in advance in the inference unit 111. The inference unit 111 may count the number of times that the sequence information processing unit 111a actually extracts sequence information as the number of repetitions. In this embodiment, it means that the sequence information of the molecule included in the search space can be an element of the sequence information set.
[0068] The sequence information processing unit 111a outputs the extracted sequence information of the multiple molecules to the property prediction unit 111b. The property prediction unit 111b according to this embodiment uses a prediction model to generate prediction information indicating the properties of a molecule based on sequence information for each molecule input from the sequence information processing unit 111a, and outputs the prediction information generated for each molecule to the prediction information processing unit 111c. The prediction model is generated by learning based on training data configured to include, for example, multiple sets of sequence information for each molecule and property evaluation results for the molecule. The prediction information for each molecule includes, for example, a predicted characteristic value for each molecule. Therefore, the characteristic prediction unit 111b may include a predicted value calculation unit 111x that calculates a predicted characteristic value from the sequence information for each molecule. The prediction information for each molecule may include, for example, a predicted characteristic value for each molecule and an estimate of the uncertainty of the predicted characteristic value. The characteristic prediction unit 111b may be configured to include a prediction uncertainty estimation unit 111y that calculates an estimate of the uncertainty of the predicted characteristic value. Thus, the prediction information for each molecule output from the characteristic prediction unit 111b may include not only the predicted characteristic value for each molecule, but also the estimate of the uncertainty of the predicted characteristic value.
[0069] The predicted information processing unit 111c acquires predicted information for each molecule from the property prediction unit 111b. When the predicted information includes a property predicted value and an evaluation value of the uncertainty of the property predicted value, the predicted information processing unit 111c may calculate a property quality value for each molecule of sequence information of each molecule included in the sequence information set based on the property predicted value and the uncertainty of the property predicted value. In this embodiment, the sequence information processing unit 111a uses the property quality value calculated by the predicted information processing unit 111c as predicted information for updating the extraction parameters. Specific examples of the extraction parameters will be described later. If the counted number of repetitions is less than the preset number of times sequence information is acquired, the sequence information processing unit 111a updates the extraction parameters based on the acquired prediction information. Thus, in the combinatorial optimization, a series of processes including the extraction of sequence information based on the updated extraction parameters in the sequence information processing unit 111a, the acquisition of prediction information based on the extracted sequence information by the characteristic prediction unit 111b, and the updating of the extraction parameters by the sequence information processing unit 111a are repeated.
[0070] When the counted number of repetitions reaches a preset number of sequence information acquisitions, the sequence information processing unit 111a stops the processing of the combinatorial optimization algorithm, and the prediction information processing unit 111c outputs the prediction information for each molecule input from the property prediction unit 111b to the candidate molecule identifying unit 111d. The candidate molecule identifying unit 111d identifies at least one candidate molecule from the sequence information set in each round based on the prediction information for each molecule input from the prediction information processing unit 111c. When the number of sequence information acquisitions is N, the prediction information processing unit 111c identifies at least one candidate molecule from the sequence information of the molecules extracted N times. The candidate molecule identifying unit 111d may also identify at least one candidate molecule from the sequence information set of the molecules extracted N times. When the prediction information includes a property quality value, the candidate molecule identifying unit 111d can identify a candidate molecule based on the property quality value for each molecule. The candidate molecule identifying unit 111d outputs sequence information of the identified candidate molecule.
[0071] In this embodiment, in the combinatorial optimization, although not limited thereto, Bayesian optimization can be applied. As a Bayesian optimization method, for example, TPE (Tree-Structured Parzen Estimator) may be used, and as an evolutionary algorithm method, for example, NSGA II (Elitist Non-dominated Sorting Genetic Algorithm) may be used.
[0072] In this embodiment, by using combinatorial optimization, it is possible to infer molecules having desired properties without predicting the property values for each of the sequence information of all molecules included in the search space. If a huge number of molecules were all treated as processing targets, it would take an enormous amount of calculation time, but by performing combinatorial optimization on the search space as in this embodiment, it is possible to infer molecules having desired properties without treating all molecules. This can further reduce the burden required for drug discovery.
[0073] In this embodiment, the molecular design device 1 may use a mathematical model for estimating a characteristic value from sequence information as an objective function of the combinatorial optimization, or may use a mathematical model for estimating a characteristic quality value. Then, the output of the mathematical model is optimized by the combinatorial optimization. When using a mathematical model for estimating a characteristic quality value, it is possible to prevent a molecule having a low prediction probability even if the characteristic indicated by the characteristic prediction value is favorable from being specified as a candidate molecule. In other words, it is possible to reduce the possibility of specifying a molecule that is likely to not have the characteristics expected when the molecule is actually generated as a candidate molecule. This can further reduce the burden required for drug discovery.
[0074] FIG. 3 is an explanatory diagram for explaining an example of a sequence information set in the first and second embodiments. The "sequence ID" in FIG. 3 is an identifier for identifying the sequence information of each molecule included in the sequence information set. FIG. 3 shows sequence information for each molecule included in the sequence information set. The sequence information of each molecule shows multiple building blocks and the arrangement of each building block. The arrangement of each building block is shown between information D101 and information D102. For example, information D101 is information on a building block related to position H1 of the sequence of each molecule. VHL0001, VHL0002, VHL0003, VHL0004, etc. all indicate a molecule in which the building block at position H1 is M.
[0075] 4 is a diagram showing an example of a hardware configuration of the molecular design device 1 according to the first and second embodiments. The molecular design device 1 has a computer system that includes a control unit 11 having a processor 91 such as a CPU and a memory 92 connected to each other by a bus, and executes a predetermined program. The computer system can also be considered to function as the molecular design device 1 that includes the control unit 11, input unit 12, communication unit 13, storage unit 14, and output unit 15 by executing the program. In this application, "executing a program" or "running a program" includes the meaning of executing a process instructed by each of one or more commands written in the program.
[0076] More specifically, the processor 91 reads out a program stored in the storage unit 14 and stores the read out program in the memory 92. The processor 91 executes the program stored in the memory 92, thereby functioning as a molecular design device 1 including a control unit 11, an input unit 12, a communication unit 13, a storage unit 14 and an output unit 15. The program is stored in advance in the storage unit 14, for example.
[0077] The control unit 11 controls the operations of various functional units included in the molecular design apparatus 1. The control unit 11 includes, for example, a function of an inference unit 111 that executes an inference process for inferring a molecule having desired properties.
[0078] The input unit 12 includes input devices such as a mouse, a keyboard, and a touch panel. The input unit 12 may be configured as an interface that connects these input devices to the molecular design device 1. The input unit 12 accepts input of various information to the molecular design device 1. For example, training data for a prediction model is input to the input unit 12.
[0079] The communication unit 13 includes a communication interface for connecting the molecular design apparatus 1 to an external device. The communication unit 13 communicates with the external device via wired or wireless communication. The external device is, for example, a device that transmits training data for a prediction model. The external device may also be a device to which candidate molecule sequence information is transmitted. The molecular design apparatus 1 may be realized as a server device that is connected to the Internet and capable of mutual communication with one or more terminal devices.
[0080] The storage unit 14 is configured using a computer-readable storage medium device (non-transitory computer-readable recording medium) such as a magnetic hard disk device or a semiconductor storage device. The storage unit 14 stores various information related to the molecular design apparatus 1. The storage unit 14 stores information inputted via the input unit 12 or the communication unit 13, for example. The storage unit 14 stores various information generated by the processing executed by the control unit 11, for example. The storage unit 14 stores, for example, a prediction model, sequence information, various parameters used for combinatorial optimization (which may include the above-mentioned extracted parameters), and the like. The storage unit 14 stores, for example, the above-mentioned program.
[0081] The output unit 15 outputs various information. The output unit 15 includes a display device such as a CRT (Cathode Ray Tube) display, a liquid crystal display, or an organic EL (Electro-Luminescence) display. The output unit 15 may be configured as an interface that connects these display devices to the molecular design apparatus 1. The output unit 15 outputs information input to the input unit 12 or the communication unit 13, for example. The output unit 15 may output various information generated by the processing executed by the control unit 11, for example.
[0082] 5 is a diagram showing an example of the configuration of the control unit 11 according to each of the first and second embodiments. The control unit 11 includes an inference unit 111, an input control unit 112, a communication control unit 113, a memory control unit 114, and an output control unit 115. The inference unit 111 executes the above-mentioned inference process. The input control unit 112 controls the operation of the input unit 12. The communication control unit 113 controls the operation of the communication unit 13. The memory control unit 114 controls the operation of the memory unit 14. The output control unit 115 controls the operation of the output unit 15.
[0083] Next, an example of the inference process according to each embodiment will be described below. Fig. 6 is a flowchart showing an example of the flow of the process executed by the molecular design device 1 in the first embodiment. The sequence information processing unit 111a acquires a sequence information set including sequence information of a plurality of molecules (step S101). The property prediction unit 111b calculates a property predicted value for each molecule constituting the sequence information set, and calculates an estimate of the uncertainty of the property value (step S102). The candidate molecule identifying unit 111d identifies a candidate molecule based on the predicted characteristic value and the estimated value of the uncertainty of the prediction calculated for each molecule (step S103). As long as the estimated value of the uncertainty of the prediction by the characteristic predicting unit 111b can be quantified (Uncertainty Quantification), the method is not limited. For example, the standard deviation of the predicted characteristic value may be calculated as the estimated value of the uncertainty of the prediction. Furthermore, in the quantification of the uncertainty, a conformal prediction may be used. The characteristic predicting unit 111b may divide the learning data into a plurality of parts, calculate an error (prediction error) for the prediction for each part, and quantify the uncertainty based on the distribution of the errors. The candidate molecule specifying unit 111d outputs the sequence information of the specified candidate molecule as candidate molecule sequence information using the output unit 15 (step S104). After that, the control unit 11 ends the process of FIG.
[0084] The prediction information processing unit 111c may calculate a characteristic quality value based on the characteristic prediction value obtained in step S102 and the estimated value of the uncertainty of the prediction prior to step S103. The prediction information processing unit 111c may calculate a mean variance (MV) as an example of the characteristic quality value. The MV is obtained by subtracting a penalty function g(x) from the product of a risk tolerance parameter ρ and a characteristic prediction value f'(x) according to a predetermined prediction model (MV=ρf'(x)-g(x)). The risk tolerance parameter is a positive real value indicating the tolerance of the characteristic prediction value f'(x) against the uncertainty g(x), that is, the reliability of the prediction value f'(x). The prediction information processing unit 111c calculates a standard deviation as an example of the estimated value g(x) of the uncertainty. The MV may be applied to portfolio optimization in financial engineering. In financial engineering, the MV may be used as an index that integrates the expected reward as the average and the variance of the reward as the risk. MV is an index that can be used in various fields and for various purposes regardless of the optimization algorithm (see, for example, Q. Zhu and VYF Tan: Thompson Sampling Algorithms for Mean-Variance Bandits (2020 ICML), S. Takemori: Distributionally-Aware Kernelized Bandit Problems for Risk Aversion (2022 ICML)).
[0085] The prediction information processing unit 111c may execute calculations using uncertainty g(x) as a constraint, which is an example of a characteristic quality value. For example, the prediction information processing unit 111c may set a threshold value τ for a certain uncertainty g(x), and when the uncertainty g(x) is equal to or smaller than τ, set the characteristic quality value to the characteristic predicted value f'(x), and when the uncertainty g(x) is greater than τ, set the characteristic quality value to a sufficiently small value (for example, the minimum value that f'(x) can take). Such settings are expressed by equation (1).
number
[0086] The prediction information processing unit 111c may calculate the characteristic quality value based on a plurality of physical property values. For example, the prediction information processing unit 111c sets a threshold value τ1 of the first physical property value g1(x) and a threshold value τ2 of the second physical property value g2(x). When the first physical property value g1(x) is smaller than τ1 and the second physical property value g2(x) is smaller than τ2, the prediction information processing unit 111c may set the characteristic quality value to the characteristic predicted value f'(x). On the other hand, when the first physical property value g1(x) is equal to or larger than τ1 or the second physical property value g2(x) is equal to or larger than τ2, the prediction information processing unit 111c may set the characteristic quality value to a sufficiently small value (for example, the minimum value that f'(x) can take).
[0087] Alternatively, the prediction information processing unit 111c may calculate the characteristic quality value based on the predicted value Pr(x∈F|x) of the probability of satisfying the constraint. That is, the prediction information processing unit 111c may multiply the characteristic predicted value f'(x) by the prediction model by the predicted value of the probability of satisfying the constraint, and sample only the sequence information with a high predicted probability of satisfying the constraint. In this example, when the probability of satisfying the constraint is low, the characteristic quality value approaches 0, and the higher the probability of satisfying the constraint (the closer to 1), the closer the characteristic quality value approaches the value of the objective index itself. Such a characteristic quality value is expressed by equation (2).
number
[0088] The number of constraints may be one or more. For example, when calculating a characteristic quality value based on a plurality of physical property values, the prediction information processing unit 111c may set a constraint condition for each physical property value and multiply the characteristic predicted value f'(x) by the prediction model by the probability that all the conditions are satisfied. In this case, the probability that all the conditions are satisfied may be the product of the probabilities that the constraint conditions for each physical property value are satisfied.
[0089] FIG. 7 is a flowchart showing an example of the flow of processing executed by the molecular design device 1 in the second embodiment. The array information processing unit 111a acquires, as training data, a data set D including a plurality of data pairs that are pairs of a measurement value y with respect to an input value x (step S201). For example, the storage unit 14 stores in advance the data set D input from the input unit 12 or the communication unit 13. The array information processing unit 111a reads the data set D from the storage unit 14. Using the acquired training data, the array information processing unit 111a, for example, performs Gaussian process regression to learn a prediction model for predicting a measurement value with respect to an input value (step S202). The inference unit 111 performs combinatorial optimization to identify candidate molecules (step S203). The candidate molecule identification unit 111d outputs the sequence information of the identified candidate molecules using the output unit 15 (step S204). Thereafter, the process of FIG. 7 ends.
[0090] <Example of Application to MBO> Next, an example of application to MBO will be described. In this application example, offline MBO, which is a type of in-silico drug design method, and building block-based molecular design are combined and applied. Offline MBO is a method of searching for an optimal molecule within a "surrogate" prediction model generated from the acquired data. This method pertains to black-box optimization that treats a representative model as a black box and is also known as inverse analysis. Here, MBO is a technique in which a prediction model is learned using pre-accumulated experimental data as training data, drug candidate molecules are evaluated based on characteristic values obtained using the prediction model obtained by learning, and the prediction model is updated based on the evaluation results. In the following description, the prediction model may sometimes be referred to as a proxy model.
[0091] In this application example, discrete input values can be the optimization target of MBO. In MBO, a measurement value y for a discrete input value x∈X is given as an unknown function f(x), a data set D having n pairs of measurement values y for discrete input values x (n is an integer equal to or greater than 2) is given in advance, and it is assumed that measurement values y for additional input values x cannot be obtained. The measurement value y for the discrete input value x provides a correct answer (also called an oracle) for the function f(x). A general MBO aims to find an input value x that maximizes the function f(x) in the discrete value space X. The above method can be considered as a problem of finding an input value x that maximizes the characteristic quality value f(x)-λg(x) instead of the function f(x). Here, an array is applied as the input value x, and the function f(x) corresponds to a function for predicting a real value corresponding to the characteristic value. The penalty function g(x) is a function that gives a real value corresponding to the uncertainty of the value of the function f(x). λ corresponds to the coefficient α mentioned above. The input value x is represented using a vector representation.
[0092] In offline MBO, the array information processing unit 111a trains a prediction model in advance using a data set D having a plurality of pairs of measurement values y for discrete input values x as training data. In training the prediction model, the array information processing unit 111a sets a function value of an objective function for a discrete input value x as a predicted value y', and determines parameters of the prediction model so as to minimize an index value representing the magnitude of the difference between the predicted value y' and the measurement value y for the discrete input value x. In this application example, as described below, the inference unit 111 samples the discrete input value x, and updates the parameters of the combinatorial optimization based on the prediction result obtained for each sample using the prediction model.
[0093] The inference unit 111 can use, for example, the above-mentioned TPE in combinatorial optimization. TPE is a black-box optimization algorithm that combines Bayesian optimization with Parzen window density estimation. Since TPE can handle categorical parameters, it can be applied to building block-based molecular design.
[0094] In this application example, the inference unit 111 implements the function of TPE and is applied in molecular design as follows. TPE is a method aiming at maximizing the expected improvement (EI) of an objective function. First, the sequence information processing unit 111a applies a building block combination constituting a drug candidate molecule to a search space, and assigns an output from a prediction model to an objective score set Y corresponding to an input molecule set X sampled from the search space. For example, the sequence information processing unit 111a sets a property prediction value obtained by the prediction model as an objective score of TPE. The sequence information processing unit 111a samples the input molecule set X from the search space, that is, assigns a candidate building block for each modification position in the building block combination. For example, when the drug candidate molecule is a protein, a candidate amino acid is assigned for each modification position in the protein sequence. As an example, an Optuna implementation can be used to execute the TPE sampler.
[0095] Next, the characteristic prediction unit 111b calculates (evaluates) an estimated value of the objective function using the prediction model for the input values determined by sampling. The prediction information processing unit 111c outputs the selected input value and the calculated estimated value to the sequence information processing unit 111a, and the sequence information processing unit 111a updates (updates) the parameters of the optimization algorithm using the input value and the estimated value input from the prediction information processing unit 111c. Here, the sequence information processing unit 111a assigns the selected input value and the calculated estimated value to the input molecule set X sampled from the search space and the objective score set Y, respectively. The inference unit 111 repeats the processes of data division, sampling, evaluation, and update, so that the evaluation results of the input values (sample values) obtained by sampling are reflected in the prediction model, and a combination of building blocks with a larger estimated value of the objective function is probabilistically searched for.
[0096] FIG. 8 shows an example of the flow of a combinatorial optimization process executed by the molecular design apparatus 1 according to this application example. The inference unit 111 sets the initial value of the number of times array information is acquired to 0. The loop R20 includes the processes of steps S203a to S203c. The control unit 11 sets the execution condition of the loop R20 when the number of repetitions is equal to or less than a predetermined number of samplings N. The number of samplings corresponds to the number of times array information is acquired.
[0097] The sequence information processing unit 111a obtains sequence information of a plurality of molecules according to a combinatorial optimization algorithm (step S203a). For example, when TPE is used as the combinatorial optimization algorithm, the sequence information processing unit 111a divides a set of input values into two sets based on an estimated value (corresponding to the above-mentioned predicted characteristic value) obtained for each input value using a prediction model and a predetermined threshold γ. One set (called the "first set") is configured to include input values whose estimated value is equal to or greater than the threshold γ. The other set (called the "second set") is configured to include input values whose estimated value is less than the threshold γ. The sequence information processing unit 111a samples input values that maximize the expected improvement of the objective function. The expected improvement corresponds to the increase in the expected value of the objective function before and after the update, and is known to be proportional to p(x|y1) / p(x|y2) belonging to the first set for the input value x. p(x|y1) indicates the density distribution of the input value x for the first set. p(x|y2) indicates the density distribution of the input value x for the second set. That is, EI is an example of an extraction parameter. The array information processing unit 111a samples the input value that maximizes the calculated EI. The characteristic prediction unit 111b performs sequence evaluation using a prediction model. Here, the characteristic prediction unit 111b uses the prediction model to calculate an estimate of an objective function for the sampled input values (step S203b). The sequence information processing unit 111a assigns the selected input value and the calculated estimated value to an input molecule set X sampled from a search space and a set of objective scores Y. Thus, a new input value is added to the input molecule set X sampled from the search space in association with the objective function, thereby updating the extraction parameters of the combinatorial optimization algorithm related to the extraction of a sequence information set (step S203c).
[0098] Every time the process of steps S203a to S203c ends, the inference unit 111 updates (increments) the number of repetitions by adding 1. When the number of repetitions is equal to or less than the sampling count N, the inference unit 111 repeats the process of steps S203a to S203c. The inference unit 111 ends the processing of loop R20 when the number of repetitions exceeds the number of samplings N. The candidate molecule identification unit 111d may output a predetermined number of building block combinations in descending order from the one having the highest estimated value of the objective function as candidate molecule sequence information using the output unit 15. Then, the processing of FIG. 8 ends.
[0099] FIG. 9 shows another example of the flow of the combinatorial optimization process executed by the molecular design apparatus 1 according to this application example. The process of FIG. 9 is the same as the process of FIG. 8 in that the processes of steps S203a to S203c are repeated, but is different from the process of FIG. 8 in that steps S203d and S203e are provided instead of the loop R20. In step S203d, the inference unit 111 sets the sampling count N and the initial value 0 of the repetition count. After the process of step S203b, the process proceeds to step S203e. In step S203e, the inference unit 111 adds 1 to the repetition count at that time point, and judges whether the repetition count has reached N times. If it is judged that it has reached N times (step S203e YES), the process of FIG. 9 is terminated. If it is judged that it has not reached N times (step S203e NO), the process proceeds to step S203c. After the process of step S203c, the process proceeds to step S203a. 9 differs from the processing in FIG. 8 in that when the number of repetitions exceeds the number of samplings N, the inference unit 111 does not update the extraction parameters.
[0100] 8 and 9, it is assumed that the termination condition is set when the number of repetitions exceeds the number of times sequence information is acquired, but this is not necessarily limited to this. The termination condition may be, for example, a target value for the estimated value of the objective function that is set in advance, and the estimated value of the objective function reaches the target value.
[0101] In the verification example of this application, the TPE that uses the MV calculated from the predicted mean value and predicted variance value of Gaussian process regression as the objective function is sometimes called MV-TPE. The TPE that uses the predicted mean (Mean) of Gaussian process regression as the objective function is sometimes called Mean-TPE. In this application, the predicted mean of Gaussian process regression as the objective function refers to the characteristic value itself.
[0102] <First verification example> Next, a verification example of this embodiment will be described. In the first verification example, the safety of this application example was verified using a GFP dataset as training data for the prediction model. The GFP dataset has 56086 GFP (Green Fluorescent Protein) sequences and information on their brightness as a characteristic. GFP is a sample that is widely used in medical and biological research. The research goal is to generate bright GFP, and it is sometimes used as a benchmark for model-based optimization performance. The input value indicating each GFP sequence is expressed using a 756-dimensional vector representation. In this verification, when obtaining the vector representation, the Tasks Assessing Protein Embeddings method (TAPE) was used as a protein embedding model.
[0103] A prediction model was trained using a Gaussian process (GP) based on GFP sequences with a mutation number of 2 or less from the parent sequence avGFP (Aequorea victoria GFP) as training data. In the following description, the prediction model obtained by this training may be called a representative model (proxy model). In addition, the parent sequence avGFP may be simply called the parent sequence or template GFP. This procedure verifies sequences with few residue substitutions from the parent sequence, realizing a practical molecular optimization process. In this verification, a GFP sequence with a residue substitution (edit distance) of 2 or less was adopted, as illustrated in FIG. 10. As a pseudo-ground-truth model, LightGBM (Light Gradient Boosting Machine) was trained using all data in the GFP dataset. LightGBM may be used for ranking, classification, and other tasks based on decision tree algorithms. With this setting, the representative model covers GFP sequences around the parent sequence, and the pseudo-ground-truth model covers a wider GFP space as a search space.
[0104] The search space was defined as the mutations in the top 100 sequences of the training data. The search space includes 2 to 5 amino acid candidates for each of the 37 candidate mutation sites, as shown in Figure 17. The parameters of TPE were set as follows, as shown in Figure 18: the number of samplings was 3000, and multivariate "none" (i.e., the objective function has one variable). The top 10 sequences of the training data were used for warm-start initialization for each of MV-TPE and Mean-TPE.
[0105] Typically, in drug discovery, 10 to 100 sequences are evaluated in one batch. In this verification, as shown in Fig. 10, the top 96 sequences with the highest scores, i.e., the highest estimated values of the objective function, were selected as proposed sequences for evaluation using a pseudo-ground-truth model.
[0106] In the verification, the processes in Figures 7 and 8 were carried out using MV and the average as the objective function. The average was used as a comparative example. MV and the average each contain characteristic values, and GFP brightness was used as that characteristic value. Figure 11 shows the average as the estimated value of the objective function in the upper row as the optimization trajectory by TPE, and MV in the lower row. Both show the results of 10 optimization processes. The dashed lines show the estimated values for each sample in each run. Figure 12 shows the relationship between the mean (GP Mean) and standard deviation (GP Std) of the sample values in each round. By using MV as the objective function, the standard deviation is smaller than when the mean is used. This means that the uncertainty of the estimated value is reduced by using MV. Figure 13 shows the density distribution of the edit distance of each sample based on the template GFP. By using MV as the objective function, the edit distance is smaller overall than when the average is used. This means that by using MV, GFPs with fewer mutations from the template GFP are selected as samples.
[0107] Figure 14 shows the distribution of scores for the proposed sequences obtained using the pseudo-ground-truth model. The scores indicate estimates of the brightness of the fluorescence emitted by GFP. It is shown that using MV as the objective function results in higher brightness than using the mean. Figure 15 shows the edit distance of the proposed sequence based on the template GFP. Using MV as the objective function results in a smaller edit distance than using the average. This shows that using MV allows sequences with fewer mutations to be sampled, i.e., safe optimization can be achieved. Figure 16 shows the standard deviation (GP Std) of the proposed array. Using MV as the objective function results in a smaller standard deviation than using the average. This supports the idea that using MV allows for sampling of an array with less uncertainty in the brightness used as a characteristic value, thereby realizing safe optimization.
[0108] <Second verification example> Next, a second verification example will be described. In the second verification example, data on bispecific antibodies was used as training data, and a Gaussian process was executed to learn a prediction model. The training data used in this verification example is data showing the sequence and characteristic value of binding ability of a bispecific antibody whose antigens are MarvelD3 and CD3. Octet values were used as the characteristic value of binding ability. A vector expression obtained using TAPE was used as an input value for the prediction model. This vector expression expresses the protein sequence of the bispecific antibody sample.
[0109] MarvelD3 is a tight junction protein with four transmembrane domains. In this validation study, MarvelD3 was set as a candidate target for anticancer drugs. The development of bispecific antibodies that crosslink cancer antigens and antigens on T cells is expected to be applied to cancer treatment. In this validation study, the trained prediction model was used to select anti-MarvelD3 sequence candidates with better properties from the lead antibodies.
[0110] In this validation, the octet values of the antibody sequences were measured multiple times, and the antibody sequences and the octet values of the antibody sequences obtained by the measurements were used as inputs to the prediction model. By batch measurement, typically 100 or less octet values of the antibody sequences are obtained in one measurement. Here, the antibody for the measurement was obtained by the following procedure. First, a plasmid encoding a pre-designed heavy or light chain was prepared, and the recombinant antibody was transiently expressed using Expi293F cells. The antibody was captured from the culture supernatant using protein A and eluted into a buffer solution. The eluted buffer solution was mixed under reducing conditions to prepare a MarvelD3 / CD3 bispecific antibody. Here, selective heavy chain heterodimerization was performed by applying charge repulsion between the same heavy chains. The concentration of the antibody in the buffer solution was determined by the absorbance at 280 nm. Then, ion exchange chromatography was performed on the buffer solution containing the bispecific antibody to confirm that the desired MarvelD3 / CD3 antibody had been prepared.
[0111] Octet values were measured using the Octet HTX system. Extracellular vesicles bearing CD81 and human MarvelD3 proteins on their surfaces were captured on a sensor chip using anti-CD81 antibodies. After a baseline step of 600 seconds in D-PBS(-) containing 0.1% BSA, the association and dissociation responses were measured for 900 and 1500 seconds, respectively, in the same buffer containing 20 nM of antibody. The binding ability of the antibody is expressed as a shift in wavelength between the baseline step and the end of the association phase. Measurements were performed at a temperature of 30°C and an oscillation speed of 1000 times per minute during the baseline step, association, and dissociation phases.
[0112] In optimizing the prediction model, MV-TPE and Mean-TPE were performed as in the first validation example. In this validation, the top 48 sequences in terms of the estimated value of the objective function among the candidate sequences obtained by sampling were evaluated as proposed sequences.
[0113] Figure 19 shows the distribution of the mean and standard deviation of the estimated values obtained by sampling at each time for MV-TPE and Mean-TPE. The standard deviation indicates uncertainty, and the higher the mean of the estimated values, the better the octet value. According to Figure 19, when the mean is used as the objective function, the most frequent value of the standard deviation is 1.0, whereas when the MV is used as the objective function, the most frequent value of the standard deviation is 0.3. Furthermore, when the mean is used as the objective function, the most frequent value of the mean is 2.3, whereas when the MV is used, the most frequent value is 1.5. These results indicate that the uncertainty of the octet value as the characteristic prediction value is higher when the mean is used as the objective function than when the MV is used.
[0114] Figure 20 shows a t-distributed Stochastic Neighbor Embedding (t-SNE) visualization of sequences obtained by sampling. t-SNE is a method for compressing high-dimensional data into low dimensions and visualizing them. Figure 20 shows the distribution of vector representations by TAPE, which represents individual sequences, on a two-dimensional plane for each of the training data (Train), mean (Mean), and MV. According to Figure 20, when the mean is used as the objective function, the sampled sequence is far from the training data, whereas when MV is used, it is closer to the training data. This shows that MV-TPE can explore a safer area while suppressing uncertainty more than Mean-TPE, and also shows that pathological samples can be avoided.
[0115] Figure 21 shows the average value (GP Mean) of the estimated octet values of the 48 sequences to be evaluated when the TPE objective function is set to mean (Mean) and when it is set to MV. According to Figure 21, the average value of the octet values obtained using MV is lower than the average value of the octet values obtained using the mean as the objective function. Figure 22 shows the standard deviation of the estimated values (GP Std) for both the mean (Mean) and MV. According to Figure 22, the standard deviation of the octet values obtained using MV is lower than the standard deviation of the octet values obtained using the mean as the objective function. From this result, it can be inferred that the sequence sampled by MV has low uncertainty and is therefore less likely to lose stable expression and binding ability.
[0116] Figure 23 shows the distribution of expression levels of the 48 sequences to be evaluated for both mean and MV. According to Figure 23, when the mean is used as the objective function, the expression levels of the top 48 sequences are kept extremely low, and sufficient samples for measuring the octet value are not obtained. In contrast, when MV is used as the objective function, almost all of the top 48 sequences are significantly expressed. Figure 24 shows the distribution of octet values for each sequence for both MV and training data. The distribution trend of octet values is similar between MV and training data. This indicates that using MV as the objective function can obtain sequences with similar binding ability to the training data and, therefore, to the parent sequence. Furthermore, MV yielded several sequences with octet values higher than the maximum value in the training set.
[0117] <Summary> According to the above-mentioned demonstration example, it was shown that the offline model-based optimization (MBO) according to this application example can be applied to building block-based molecular design, which is a common method in drug design, and in-silico implementation (i.e., implementation by a computer) can be realized. Here, the tree-structured Parzen estimator (TPE), which is a method of Bayesian optimization, can be applied to building block-based molecular design related to the optimization of combinations of categorical variables. In addition, it was shown that MV used in financial engineering can avoid pathological behavior by applying offline MBO and contribute to the search for safe sequences. That is, by incorporating risk averse prediction in the MBO of molecular design, it was possible to prevent overestimation of the extrapolation region in the prediction of the properties of candidate molecules and prevent the proposal of molecules with low prediction reliability. In addition, the above-mentioned demonstration example showed that this embodiment is useful in searching for safe sequences in therapeutic antibody design.
[0118] The molecular design device 1 of the embodiment thus configured performs Bayesian optimization using a model that obtains a characteristic quality value based on drug candidate molecular sequence information as an objective function, thereby further reducing the burden required for drug discovery.
[0119] (Modification) The molecular design apparatus 1 may be implemented using a plurality of information processing devices communicatively connected via a network. In this case, each functional unit of the molecular design apparatus 1 may be distributed and implemented in a plurality of information processing devices. The drug discovery system 100 may also include a plant (not shown) that generates an optimal molecule estimated by the molecular design apparatus 1. The molecular design apparatus 1 outputs output information indicating the biological sequence of the optimal molecule estimated by the above method to the plant. The plant executes a process of generating a molecular compound having a biological sequence indicated by the output information input from the molecular design apparatus 1.
[0120] All or part of the functions of the molecular design apparatus 1 may be realized using hardware such as an ASIC (Application Specific integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. The program may be transmitted via an electric communication line.
[0121] The molecular design device 1 is an example of an estimation device. The above functions may be realized by other types of devices, such as a device whose main function is not molecular design, or an information device equipped with a general-purpose computer system.
[0122] For example, the present embodiment may be realized as a system including a processor and a memory. In the system, the memory is configured to store one or more instructions, which may cause the processor to execute the steps of: calculating a predicted property value of a molecule that is an element of a sequence information set, which is a set of sequence information of a plurality of different molecules, using a prediction model for predicting the property of the molecule from sequence information of the molecule, and estimating the uncertainty of the prediction; and searching for a candidate molecule having a more desired property based on the predicted property value and the uncertainty of the prediction. Furthermore, the instructions may be instructions to cause a processor to execute an acquisition step of acquiring a sequence information set, which is a collection of sequence information of a plurality of different molecules, in accordance with a combinatorial optimization algorithm; a prediction step of calculating a predicted property value of a molecule that is an element of the sequence information set, using a prediction model for predicting the properties of the molecule from its sequence information; an update step of updating extraction parameters of the combinatorial optimization algorithm, based on the predicted property value for each molecule, so that molecules having more desired properties are included; and a search step of searching for molecules having desired properties based on the predicted property values.
[0123] The present embodiment may be realized as a non-transitory computer-readable medium storing one or more instructions. The instructions may be instructions for causing a computer to execute a procedure of calculating a predicted property value of a molecule that is an element of a sequence information set, which is a set of sequence information of a plurality of different molecules, using a prediction model for predicting the property of the molecule from the sequence information of the molecule and estimating the uncertainty of the prediction, and a procedure of searching for a candidate molecule having a more desired property based on the predicted property value and the uncertainty of the prediction. The instructions may also be instructions for causing a computer to execute a procedure of acquiring a sequence information set, which is a set of sequence information of a plurality of different molecules, according to a combinatorial optimization algorithm, a prediction procedure of calculating a predicted property value of a molecule that is an element of the sequence information set, using a prediction model for predicting the property of the molecule from the sequence information of the molecule, an update procedure of updating an extraction parameter of the combinatorial optimization algorithm so that a molecule having a more desired property is included based on the predicted property value for each molecule, and a search procedure of searching for a molecule having a desired property based on the predicted property value.
[0124] This embodiment may be realized as a computer-implemented method using an artificial intelligence engine. The computer-implemented method includes a control step of performing a combinatorial optimization to estimate an optimal molecule, which is a drug candidate molecule that provides an optimal value of a desired property, among each of the drug candidate molecules represented by each element of the candidate set, and a step of outputting information on the optimal molecule, and the objective function of the combinatorial optimization is a model that performs a mathematical model that estimates a characteristic value of the desired property based on the drug candidate molecule sequence information and obtains the uncertainty of the result of the estimation, and obtains a characteristic quality value calculated based on the characteristic value and the predetermined uncertainty.
[0125] Although some embodiments have been described above, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are within the scope of the invention and its equivalents as well as the scope and spirit of the invention. The documents cited in this specification are incorporated herein by reference. [Explanation of symbols]
[0126] 1... molecular design device, 11... control unit, 12... input unit, 13... communication unit, 14... memory unit, 15... output unit, 100... drug discovery system, 111... inference unit, 112... input control unit, 113... communication control unit, 114... memory control unit, 115... output control unit, 91... processor, 92... memory.
Claims
1. An information processing system for identifying molecules suitable as drug candidates, For a molecule that is an element of a building block combination information set, which is a set of building block combination information of a plurality of different molecules, using a prediction model for predicting the characteristics of the molecule from the building block combination information of the molecule, calculating a predicted value of the characteristics of the molecule, and estimating the uncertainty of the prediction; a characteristic prediction unit; A candidate molecule identification unit that searches for candidates for molecules having desired characteristics based on the predicted characteristic value and the estimated value of the uncertainty of the prediction; An information processing system comprising:
2. Furthermore, it comprises a prediction information processing unit that calculates a characteristic quality value based on the predicted characteristic value and the estimated value of the uncertainty of the prediction, The candidate molecule identification unit identifies at least one candidate molecule from the elements of the building block combination information set based on the characteristic quality value. The information processing system according to claim 1.
3. The characteristic quality value is an output value of a predetermined function that takes two variables, the predicted characteristic value and the estimated value of the uncertainty of the prediction, as inputs. The information processing system according to claim 2.
4. The prediction information processing unit calculates the mean variance as an objective function for giving the characteristic quality value. The information processing system according to claim 2 or 3.
5. The characteristic quality value increases as the predicted characteristic value increases and decreases as the estimated value of the uncertainty of the prediction decreases. The information processing system according to claim 2 or 3.
6. The information processing system according to claim 2 or 3, comprising a building block combination information processing unit that uses a combinatorial optimization algorithm to obtain the building block combination information set based on the predicted characteristic value of the molecule or the characteristic quality value.
7. The characteristic quality value increases as the predicted characteristic value increases, the building block combination information set is obtained using a combinatorial optimization algorithm as the estimated value of the uncertainty of the prediction decreases, and the extraction parameters of the combinatorial optimization algorithm are updated based on the predicted characteristic value of the molecule or the characteristic quality value. The information processing system according to claim 2 or 3.
8. The combinatorial optimization algorithm is a tree-structured Parzen estimator. The information processing system according to claim 7.
9. The prediction model is a prediction model generated by learning based on the building block combination information of a plurality of molecules for training and the results of the property evaluation of the molecules. The information processing system according to any one of claims 1 to 3.
10. The building block combination information is the sequence information of the molecule. The information processing system according to any one of claims 1 to 3.
11. The molecule is at least one of nucleic acid, peptide, cyclic peptide, protein, antibody, and low molecular weight compound. The information processing system according to any one of claims 1 to 3.
12. The property is at least one property among binding ability, pharmacological activity, physical properties, kinetics, and safety. The information processing system according to any one of claims 1 to 3.
13. The molecule is a molecule that binds to a target molecule, and the property is the binding ability to the target molecule. The information processing system according to any one of claims 1 to 3.
14. The molecule is a protein, antibody, peptide, or cyclic peptide, and the building block combination information is information on the amino acid sequence. The information processing system according to any one of claims 1 to 3.
15. The estimated value of the uncertainty of the prediction is the standard deviation of the property prediction value. The information processing system according to any one of claims 1 to 3.
16. An information processing method in an information processing system for identifying a molecule suitable as a drug candidate, For a molecule that is an element of a building block combination information set, which is a set of building block combination information of a plurality of different molecules, using a prediction model for predicting the properties of the molecule from the building block combination information of the molecule, calculating a property prediction value of the molecule, and estimating the uncertainty of the prediction; Searching for candidates for molecules having more desired properties based on the property prediction value and the uncertainty of the prediction; An information processing method including.
17. To a computer, For a molecule that is an element of a building-block combination information set, which is a set of building-block combination information of a plurality of different molecules, a procedure for calculating a predicted value of the properties of the molecule using a prediction model for predicting the properties of the molecule from the building-block combination information of the molecule and estimating the uncertainty of the prediction, A procedure for searching for candidates for molecules having more desired properties based on the predicted property value and the uncertainty of the prediction, A program for causing the above to be executed.
18. A method for producing a molecular compound, comprising: Accessing a building-block combination information set, which is a set of building-block combination information of a plurality of different molecules, An input step of inputting the building-block combination information set into a prediction model, An inference step of searching for a molecule having more desired properties from the building-block combination information set based on the predicted property values and the estimated values of the uncertainty of the prediction for each molecule included in the building-block combination information set output from the prediction model, and specifying the molecule as a candidate molecule, An output step of outputting the building-block combination information related to the candidate molecule, A generation step of generating the molecular compound having the molecular sequence shown in the building-block combination information, A method having the above steps.
19. An information processing system for identifying a molecule suitable as a drug candidate, comprising: A building-block combination information processing unit that obtains a building-block combination information set, which is a set of building-block combination information of a plurality of different molecules, according to a combinatorial optimization algorithm, A property prediction unit that calculates a predicted property value of a molecule that is an element of the building-block combination information set using a prediction model for predicting the properties of the molecule from the building-block combination information of the molecule, A candidate molecule identification unit that searches for candidates for molecules having desired properties based on the predicted property value, The building-block combination information processing unit Updates the extraction parameters of the combinatorial optimization algorithm so that molecules having more desired properties are included based on the predicted property values of the molecules, An information processing system. **Claim 20**: The candidate molecule specifying unit searches for candidates of molecules having desired properties based on a first property prediction value calculated for a molecule that is an element of a first building block combination information set, which is a set of building block combination information acquired before the extraction parameters are updated, and a second property prediction value calculated for a molecule that is an element of a second building block combination information set, which is a set of building block combination information acquired after the extraction parameters are updated. The information processing system according to claim 19. **Claim 21**: The first building block combination information set and the second building block combination information set are sets each including building block combination information of different molecules. The information processing system according to claim 20. **Claim 22**: The building block combination information is the sequence information of the molecule. The information processing system according to any one of claims 19 to 21. **Claim 23**: The molecule is at least one of nucleic acid, peptide, cyclic peptide, protein, antibody, and low molecular compound. The information processing system according to any one of claims 19 to 21. **Claim 24**: The molecule is protein, antibody, peptide, or cyclic peptide, and the building block combination information is information on amino acid sequence. The information processing system according to any one of claims 19 to 21. **Claim 25**: The property is at least one property among binding ability, pharmacological activity, physical property, pharmacokinetics, and safety. The information processing system according to any one of claims 19 to 21. **Claim 26**: The molecule is a molecule that binds to a target molecule, and the property is the binding ability to the target molecule. The information processing system according to any one of claims 19 to 21. **Claim 27**: The property prediction unit calculates the property prediction value and calculates an estimated value of the uncertainty of the prediction. The information processing system according to any one of claims 19 to 21. **Claim 28**: The building block combination information processing unit updates the extraction parameters of the combinatorial optimization algorithm so that molecules having more desired properties are included based on the property prediction value of the molecule and the estimated value of the uncertainty of the prediction. The information processing system according to any one of claims 19 to 21.
29. Further comprising a prediction information processing unit that processes prediction information including the predicted characteristic value and / or the estimated value of the uncertainty of the prediction output from the characteristic prediction unit, The prediction information processing unit calculates a characteristic quality value based on the predicted characteristic value and the estimated value of the uncertainty of the prediction, The candidate molecule identification unit identifies at least one molecule from the building block combination information set based on the characteristic quality value, The information processing system according to claim 27.
30. The characteristic quality value increases as the predicted characteristic value increases and decreases as the estimated value of the uncertainty of the prediction decreases, The information processing system according to claim 29.
31. The characteristic quality value is an output value of a predetermined function that takes two variables of the predicted characteristic value and the estimated value of the uncertainty of the prediction as inputs, The information processing system according to claim 30.
32. The prediction information processing unit calculates the mean variance as an objective function that gives the characteristic quality value, The information processing system according to claim 29.
33. The estimated value of the uncertainty of the prediction is the standard deviation of the predicted characteristic value, The information processing system according to claim 28.
34. Using a tree-structured Parzen estimator as the combinatorial optimization algorithm, The information processing system according to any one of claims 19 to 21.
35. An information processing method in an information processing system for identifying a molecule suitable as a drug candidate, Obtaining a building block combination information set, which is a set of building block combination information of a plurality of different molecules, according to a combinatorial optimization algorithm; Calculating a predicted characteristic value of a molecule that is an element of the building block combination information set using a prediction model for predicting the characteristics of the molecule from the building block combination information of the molecule; Searching for a molecule having more desired characteristics based on the predicted characteristic value; Updating the extraction parameters of the combinatorial optimization algorithm so that the molecule having more desired characteristics is included based on the predicted characteristic value of the molecule. Information processing method.
36. In a computer, An acquisition procedure for obtaining a building block combination information set, which is a set of building block combination information of a plurality of different molecules, according to a combinatorial optimization algorithm, A prediction procedure for calculating a predicted value of the properties of a molecule that forms an element of the building block combination information set, using a prediction model for predicting the properties of the molecule from the building block combination information of the molecule, An update procedure for updating the extraction parameters of the combinatorial optimization algorithm so that the molecule having more desired properties is included, based on the predicted value of the properties for each molecule, A search procedure for searching for a molecule having desired properties based on the predicted value of the properties, A program for executing the above.
37. A method for producing a molecular compound, comprising: An acquisition step of obtaining a building block combination information set, which is a set of building block combination information of a plurality of different molecules, according to a combinatorial optimization algorithm, An input step of inputting the building block combination information set into a prediction model, An update step of updating the extraction parameters of the combinatorial optimization algorithm so that the molecule having more desired properties is included, based on the predicted value of the properties for each molecule included in the building block combination information set output from the prediction model, A step of searching for a molecule having more desired properties from the building block combination information set based on the predicted value of the properties and specifying it as a candidate molecule, An output step of outputting the building block combination information related to the candidate molecule, A generation step of generating the molecular compound having the molecular sequence shown in the building block combination information, A method comprising the above steps.
38. A molecular design apparatus comprising a control unit for inferring the properties of a molecule from the building block combination information of the molecule using a predetermined prediction model, The control unit: Obtains a building block combination information set, which is a set of building block combination information of a plurality of different molecules, Calculates a predicted value of the properties of a molecule that forms an element of the building block combination information set from the building block combination information of the molecule by the prediction model, and estimates the uncertainty of the prediction, Searching for candidates of molecules having desired properties based on the predicted property values and the estimated values of the uncertainty of the predictions. A molecular design device. **Claim 39**: A molecular design device comprising a control unit that infers the properties of a molecule from the building block combination information of the molecule using a predetermined prediction model, wherein the control unit obtains a building block combination information set, which is a set of building block combination information of a plurality of different molecules, according to a combinatorial optimization algorithm, calculates the predicted property values of the molecules that are elements of the building block combination information set using the prediction model, updates the extraction parameters of the combinatorial optimization algorithm based on the predicted property values of each molecule by the prediction model so that molecules having more desired properties are included, searches for candidates of molecules having desired properties based on the predicted property values, A molecular design device.