Peptide proposal system, peptide proposal method, peptide proposal program, and method for manufacturing peptide

The peptide suggestion system filters peptides using predicted values and uncertainty to set a Trust Region, improving the accuracy and reliability of peptide selection by focusing on reliable predictions.

WO2026115712A1PCT designated stage Publication Date: 2026-06-04CHUGAI PHARMA CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CHUGAI PHARMA CO LTD
Filing Date
2024-11-29
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing peptide prediction methods using machine learning may not necessarily lead to peptides with desired characteristics, necessitating improved accuracy in peptide suggestion systems.

Method used

A peptide suggestion system that filters candidate peptides based on predicted values, uncertainty, and distance from measured peptides, using a predictive model to suggest appropriate peptides by setting a Trust Region for local search.

Benefits of technology

Enables the proposal of peptides with desired properties by narrowing the search space to reliable predictions, enhancing accuracy and reliability in peptide selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024042347_04062026_PF_FP_ABST
    Figure JP2024042347_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A peptide proposal system 10 comprises: a candidate acquisition unit 11 that acquires candidate configuration information indicating the configurations of a plurality of candidate peptides that are candidates for characteristics measurement; a calculation unit 12 that uses a prediction model to calculate predicted values of the characteristics of the plurality of candidate peptides and uncertainties of the predicted values on the basis of the candidate configuration information; a measured configuration information acquisition unit 13 that acquires measured configuration information indicating the configurations of measured peptides of which the characteristics have been measured, the measured configuration information having been used to generate the prediction model; a filtering unit 14 that filters the candidate peptides on the basis of the candidate configuration information, the measured configuration information, and / or the uncertainties of the predicted values; and a determination unit 15 that determines, from the filtered candidate peptides, a candidate peptide of which the characteristics are to be measured, on the basis of the predicted values of the characteristics and / or the uncertainties of the predicted values.
Need to check novelty before this filing date? Find Prior Art

Description

Peptide Proposal System, Peptide Proposal Method, Peptide Proposal Program, and Method for Producing Peptide

[0001] The present invention relates to a peptide proposal system, a peptide proposal method, a peptide proposal program, and a method for producing a peptide.

[0002] Conventionally, Bayesian optimization is known as an optimization technique that can efficiently achieve a target value based on measurement. In Non-Patent Document 1, as a method for performing appropriate Bayesian optimization, it is proposed to limit the search in Bayesian optimization to a Trust Region, which is a reliable space based on measured data. In Non-Patent Document 2, it is proposed to define a Trust Region by Hamming distance for a space consisting of categorical variables.

[0003] Eriksson, David, et al. Scalable Global Optimization via Local Bayesian Optimization. Advances in Neural Information Processing Systems. 2019 Xingchen Wan, Vu Nguyen, Huong Ha, Binxin Ru, Cong Lu, Michael A. Osborne, Think Global and Act Local: Bayesian Optimisation over High-Dimensional Categorical and Mixed Search Spaces, arXiv:2102.07188, 2021

[0004] In machine learning in the pharmaceutical field, it is conceivable to use a trained learned prediction model to predict a peptide having characteristics that can be used for drug discovery and propose a peptide to be actually synthesized. However, simply predicting and proposing a peptide by machine learning may not necessarily lead to a peptide having desired characteristics. Therefore, it is desirable to improve the accuracy of prediction and proposal.

[0005] One embodiment of the present invention has been made in view of the above, and aims to provide a peptide suggestion system, a peptide suggestion method, and a peptide suggestion program that can appropriately suggest peptides, as well as a method for producing peptides.

[0006] To achieve the above objectives, the inventors have invented a peptide suggestion system, a peptide suggestion method, and a peptide suggestion program, as well as a method for producing peptides, which can suggest desired peptides by filtering using a predictive model. Specifically, the peptide suggestion system according to one embodiment of the present invention includes: a candidate acquisition means for acquiring candidate configuration information indicating the composition of a plurality of candidate peptides which are candidates for measuring characteristics; a calculation means for calculating predicted values ​​of the characteristics of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate configuration information acquired by the candidate acquisition means using a predictive model; a measured acquisition means for acquiring measured configuration information indicating the composition of measured peptides whose characteristics have been measured, which was used to generate the predictive model; i) the distance between the amino acid sequence of the candidate peptide indicated by the candidate configuration information acquired by the candidate acquisition means and the amino acid sequence of the measured peptide indicated by the measured configuration information acquired by the measured acquisition means; ii) the uncertainty of the predicted value related to the candidate peptide calculated by the calculation means; and iii) A filtering means for filtering candidate peptides based on a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, amino acid angle, and aromatic amino acid pattern between a candidate peptide and a measured peptide, based on candidate composition information obtained by a candidate acquisition means and measured composition information obtained by a measured acquisition means; and a determination means for determining a candidate peptide whose properties are to be measured based on at least one of the predicted value of the properties calculated by a calculation means and the uncertainty of the predicted value, from the candidate peptides after filtering by the filtering means.

[0007] In the peptide suggestion system according to one embodiment of the present invention, candidate peptides are filtered based on at least one selected from i) to iii) above. The candidate peptides after this filtering are appropriate. Therefore, according to the peptide suggestion system according to one embodiment of the present invention, peptides can be appropriately suggested.

[0008] The filtering means may further filter the candidate peptides based on the predicted values ​​of the characteristics of the candidate peptides calculated by the calculation means. This configuration allows for the proposal of even more appropriate peptides.

[0009] The filtering means may filter candidate peptides based on multiple different criteria, and the determination means may output information indicating the candidate peptides determined to have their characteristics measured, for each of the multiple criteria. With this configuration, it is possible to understand the proposed peptides for each criterion used to filter the candidate peptides.

[0010] The determination method may involve determining candidate peptides to measure properties using a Bayesian optimization acquisition function or the order of predicted property values. This configuration allows for the appropriate and reliable proposal of peptides.

[0011] The determination means may output information indicating the candidate peptides determined to measure the characteristics, for each pattern of amino acid species constituting the candidate peptide. With this configuration, it is possible to identify the proposed peptides for each pattern of amino acid species.

[0012] The amino acid species pattern may be as follows: i) NH-amino acids or N-alkyl amino acids ii) Cyclic amino acids or acyclic amino acids iii) At least one selected from those having an aromatic group or not having an aromatic group.

[0013] The predictive model may be a predictive model generated from measured configuration information obtained by a measured acquisition means. With this configuration, the predictive model can be used appropriately and reliably to calculate the predicted values ​​of the properties of multiple candidate peptides and the uncertainty of those predicted values. As a result, peptides can be proposed appropriately and reliably.

[0014] The properties may be at least one selected from molecular binding activity, molecular inhibitory activity, molecular and cellular activity, lipid solubility, solubility, membrane permeability, metabolic stability, pharmacokinetics, and toxicity. This configuration allows for the appropriate proposal of peptides relating to these properties.

[0015] Candidate peptides and measured peptides may be cyclic peptides. This configuration allows for the appropriate proposal of cyclic peptides.

[0016] A method for producing a peptide according to one embodiment of the present invention is a method for producing a candidate peptide determined using the peptide proposal system described above.

[0017] Incidentally, one embodiment of the present invention can be described as an invention of a peptide proposal system as described above, but it can also be described as an invention of a peptide proposal method and a peptide proposal program as follows. These are substantially the same invention, differing only in category, and produce similar functions and effects.

[0018] That is, a peptide proposal method according to one embodiment of the present invention is a peptide proposal method which is a method of operation of a peptide proposal system, comprising: a candidate acquisition step of acquiring candidate configuration information indicating the configuration of a plurality of candidate peptides which are candidates for measuring characteristics; a calculation step of using a prediction model to calculate predicted values ​​of the characteristics of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate configuration information acquired in the candidate acquisition step; a measured acquisition step of acquiring measured configuration information indicating the configuration of measured peptides whose characteristics have been measured, which was used to generate the prediction model; i) the distance between the amino acid sequence of the candidate peptide indicated by the candidate configuration information acquired in the candidate acquisition step and the amino acid sequence of the measured peptide indicated by the measured configuration information acquired in the measured acquisition step; ii) the uncertainty of the predicted value related to the candidate peptide calculated in the calculation step, and iii) A filtering step for filtering candidate peptides based on at least one selected from a comparison of at least one of the N-alkyl patterns, cyclic amino acid patterns, amino acid angles, and aromatic amino acid patterns between the candidate peptide and the measured peptide, based on candidate composition information obtained in the candidate acquisition step and measured composition information obtained in the measured acquisition step, or based on the uncertainty of the predicted values ​​for the candidate peptide calculated in the calculation step; and a determination step for determining candidate peptides to measure properties from the candidate peptides after filtering in the filtering step, based on at least one of the predicted values ​​of the properties calculated in the calculation step and the uncertainty of the predicted values.

[0019] Furthermore, a peptide proposal program according to one embodiment of the present invention includes a computer, a candidate acquisition means for acquiring candidate configuration information showing the composition of a plurality of candidate peptides which are candidates for measuring properties, a calculation means for calculating predicted values ​​of the properties of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate configuration information acquired by the candidate acquisition means using a prediction model, a measured acquisition means for acquiring measured configuration information showing the composition of measured peptides whose properties have been measured and used in generating the prediction model, and i) the distance between the amino acid sequence of the candidate peptide shown by the candidate configuration information acquired by the candidate acquisition means and the amino acid sequence of the measured peptide shown by the measured configuration information acquired by the measured acquisition means, ii) the uncertainty of the predicted value related to the candidate peptide calculated by the calculation means, and iii) A filtering means that filters candidate peptides based on a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, amino acid angle, and aromatic amino acid pattern between candidate peptides and measured peptides, based on candidate composition information obtained by candidate acquisition means and measured composition information obtained by measured acquisition means; and a determination means that determines candidate peptides to have their properties measured based on at least one of the predicted value of the properties calculated by calculation means and the uncertainty of the predicted value, from the candidate peptides after filtering by the filtering means.

[0020] According to one embodiment of the present invention, a peptide can be appropriately proposed.

[0021] This figure shows the configuration of a peptide proposal system according to an embodiment of the present invention. This figure schematically illustrates the conversion of peptides to vectors. This figure schematically illustrates the process of converting candidate configuration information to vectors. This figure schematically illustrates the calculation of predicted values ​​of characteristics using a prediction model. This is a graph schematically showing the uncertainty of predicted values ​​obtained using a prediction model. This flowchart shows the peptide proposal method, which is a process executed in the peptide proposal system according to an embodiment of the present invention. This figure shows the configuration of a peptide proposal program according to an embodiment of the present invention together with a recording medium.

[0022] The following describes in detail embodiments of the peptide proposal system, peptide proposal method, peptide proposal program, and peptide production method according to the present invention, with reference to the drawings. In the description of the drawings, the same elements are denoted by the same reference numerals, and redundant explanations are omitted.

[0023] Figure 1 shows the peptide suggestion system 10 according to this embodiment. The peptide suggestion system 10 is a system (device) used to suggest peptides (peptide compounds) having desired properties. Peptide suggestion refers to a process of finding peptides having desired properties under certain conditions and presenting the peptides found through this process to the user of the peptide suggestion system 10. For example, the peptide suggestion system 10 is used to suggest peptides having properties that can be used in drug discovery. Properties that can be used in drug discovery include, for example, binding affinity (molecular binding activity or molecular inhibitory activity), pharmacological activity (molecular and cellular activity), lipid solubility, solubility, membrane permeability, metabolic stability, pharmacokinetics, and toxicity. The peptide suggestion system 10 may also be used to suggest peptides for purposes other than drug discovery. Specifically, peptide suggestion is, for example, a suggestion of the sequence of amino acids (amino acid residues) that make up the peptide.

[0024] The peptides targeted for proposal using the peptide proposal system 10 may be cyclic peptides. For example, the cyclic peptides targeted for proposal using the peptide proposal system 10 may consist of 5 to 20 amino acid residues, 9 to 14 amino acid residues, or 11 amino acid residues. Cyclic peptides are used in drug discovery. The amino acids constituting the peptides targeted for proposal using the peptide proposal system 10 may include not only natural amino acids (20 types) but also unnatural amino acids. Furthermore, the peptide proposal system 10 may be used to propose peptides that can be used for purposes other than drug discovery.

[0025] Peptide proposals using the peptide proposal system 10 are carried out according to, for example, the following experimental design framework. The experimental design may be carried out using Bayesian optimization (BO). First, a pre-defined peptide is actually synthesized. Next, the properties of the synthesized peptide are measured (actually measured) by wet experiments. Subsequently, a predictive model is generated (constructed) that calculates predicted values ​​for the properties of an arbitrary peptide from the constituent information of the constituent of the arbitrary peptide, based on the constituent information of the constituent peptide and the measured peptide properties.

[0026] Next, the generated predictive model is used to calculate predicted values ​​for peptides that have not yet been synthesized or evaluated (whose properties have not been measured). Based on the calculation of predicted values, the peptide whose properties will be measured next is determined. Subsequently, the determined peptide is actually synthesized. Next, the properties of the synthesized peptide are measured (actually measured) by wet lab experiments. If the measured properties are the desired properties, then a peptide with the desired properties has been proposed. On the other hand, if the measured properties are not the desired properties, the predictive model is generated (reconstructed) again based on the constituent information showing the composition of the synthesized peptide and the properties of the measured peptide.

[0027] The next step in determining which peptide to synthesize and measure the properties of is performed from the perspective of obtaining a peptide with the desired properties and generating a predictive model that can calculate highly accurate predicted values ​​of those properties. By repeating the above process, a peptide with the desired properties is proposed.

[0028] The criteria for determining the peptides described above may differ with each iteration. For example, a structurally unique peptide (e.g., a peptide with three or more amino acid sequences modified from a peptide whose properties were measured, or a novel peptide) may be determined. Such a determination is called "exploration" or "exploratory proposal." Peptides determined by "exploratory proposals" may include those with uncertain predicted values, and by making "exploratory proposals," it is possible to quickly arrive at good peptides that are difficult to imagine from the current state. On the other hand, "exploratory proposals" are also likely to include peptides with reduced properties. Therefore, "exploratory proposals" greatly contribute to improving the accuracy of the prediction model when the initial prediction model of an iteration is considered to have difficulty making accurate predictions over a wide range. By making "exploratory proposals," it is possible to try modifications that change the skeletal structure of the peptides whose properties were measured (existing peptides), or to boldly optimize the combination of peptides whose properties were measured.

[0029] Instead of "exploration" or "exploratory proposals," it is also possible to determine peptides by making minor changes to the amino acid sequence of the peptides whose properties have been measured. Such a determination is called "utilization" or "utilization-oriented proposals." In "utilization-oriented proposals," the proposal is made within a range where the predicted value is likely to be accurate (for example, a range of peptides in which the amino acid sequence has been modified by one or two residues from the peptide whose properties have been measured), making it easier to find peptides with good properties. On the other hand, in "utilization-oriented proposals," it is difficult to find unexpectedly good peptides. Also, the diversity of peptides determined by "utilization-oriented proposals" tends to be low. Therefore, "utilization-oriented proposals" are effective when it is thought that the prediction model can make accurate predictions over a wide range by performing many iterations. By performing "utilization-oriented proposals," it is possible to identify chemists' oversights and to perform robust optimization of the combination of peptides whose properties have been measured (read optimization).

[0030] The peptide proposal system 10 is for appropriately proposing peptides within the framework described above. An appropriate peptide proposal is, for example, a proposal that more reliably obtains a peptide with desired properties. Alternatively, an appropriate peptide proposal is a proposal that obtains a peptide with more desirable properties. Alternatively, an appropriate peptide proposal is a proposal that obtains a peptide with desired properties through the synthesis and measurement of fewer peptides. Furthermore, an appropriate peptide proposal may achieve multiple of these.

[0031] The peptide proposal system 10 is specifically a computer including hardware such as a CPU (Central Processing Unit) and memory. The functions of the peptide proposal system 10 described later are performed by the operation of these components by programs, etc. The peptide proposal system 10 may be implemented by a single computer, or it may be implemented by a computer system consisting of multiple computers connected to each other by a network. The peptide proposal system 10 may have communication functions for acquiring information necessary for the processing of the peptide proposal system 10 as described below.

[0032] Next, the functions of the peptide suggestion system 10 according to this embodiment will be described. As shown in Figure 1, the peptide suggestion system 10 is configured to include a candidate acquisition unit 11, a calculation unit 12, a measured acquisition unit 13, a filtering unit 14, and a determination unit 15.

[0033] The candidate acquisition unit 11 is a candidate acquisition means that acquires candidate configuration information indicating the composition of multiple candidate peptides that are candidates for measuring characteristics. The candidate peptides are not particularly limited as long as they are peptides, but they may be linear peptides or cyclic peptides.

[0034] A candidate peptide is a peptide (hypothetical compound) that has not yet been actually synthesized in the peptide proposal described above, and is a candidate for synthesis and property measurement next. A candidate peptide may consist of a predetermined number of amino acids; for example, it may be a peptide composed of 5 to 20 amino acid residues, a peptide composed of 9 to 14 amino acid residues, or a peptide composed of 11 amino acid residues.

[0035] The properties to be measured are pre-defined properties and are the properties targeted for peptide proposal by the peptide proposal system 10. When proposing a peptide for drug discovery, the properties are related to drug discovery. The properties are, for example, either peptide activity (e.g., pharmacological activity) or drug-likeness indicators. Specifically, the properties may be at least one selected from binding affinity (molecular binding activity or molecular inhibitory activity), pharmacological activity (molecular and cellular activity), lipid solubility, solubility, membrane permeability, metabolic stability, pharmacokinetics, and toxicity. Molecular binding activity or molecular inhibitory activity may be binding activity or inhibitory activity to biomolecules, or binding activity or inhibitory activity to target molecules present in the body (e.g., target proteins). The target molecule is not particularly limited, but may be a molecule important for exhibiting the desired pharmacological effect, or a molecule involved in ADME (absorption, distribution, metabolism, and excretion). The properties may also be binding affinity to a pre-defined target.

[0036] More specifically, the activity among the characteristics is, for example, the IC50 value, the KD (equilibrium dissociation constant) value, and the coupling constant (k) of SPR (surface plasmon resonance). a , k d ), is the value of an assay evaluating binding to proteins. Activity may also be the α-screen value, which is a value that quantifies the interaction between target molecules (activity value). In addition, activity may be measured using methods other than those mentioned above, using SPR. SPR is a principle that allows for the label-free observation of intermolecular interactions (binding). By using this principle, information on the dynamic process of intermolecular interactions (binding constant and dissociation constant), binding affinity, and concentration of the analyte can be obtained.

[0037] Furthermore, drug-likeness indicators among the characteristics include test values ​​obtained from membrane permeability tests using artificial membranes (PAMPA (Parallel Artificial Membrane Permeability Assay)), membrane permeability tests using cells (Caco-2), metabolic stability tests, solubility tests, lipid solubility tests, metabolic stability tests, pharmacokinetic tests, or toxicity tests. Basically, the values ​​(predicted and measured values) that show the above characteristics are expressed as floating-point numbers. Depending on the distribution of the values, values ​​that have undergone logarithmic transformation, normalization, or box-cox transformation may be used. Note that the characteristics may be other than those listed above, as long as they are characteristics of a measurable peptide.

[0038] Candidate structural information includes, for example, information indicating the amino acids that make up the candidate peptide, and the position of each amino acid (the order (sequence) of the amino acids). Hereafter, the position of an amino acid will be referred to as the core. For an 11-amino acid peptide, there are cores 1 to 11. However, candidate structural information may be other than those mentioned above, as long as it indicates the composition of the candidate peptide and can be used in the following peptide proposals.

[0039] For example, multiple candidate peptides are generated by listing candidate amino acids for each core (position) and combining them (e.g., by exhaustive combinations). For example, the candidate amino acids for each core (position) may be the amino acids for each core of a measured peptide whose properties have been measured. In this case, instead of using all measured peptides, only a portion of the measured peptides may be used. The measured peptides used in this case may be those whose measured properties are in the top 1% (1st percentile), 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of all measured peptides. For example, the measured peptide used in this process may exhibit at least one of the following characteristics: high molecular binding activity, high molecular inhibitory activity, high cellular activity, high membrane permeability, high metabolic stability, good pharmacokinetics, and low toxicity.

[0040] Furthermore, candidate amino acids may be determined manually (by a chemist, etc.). For example, after the amino acids for each core of a measured peptide are designated as candidate amino acids, additional candidate amino acids may be added or removed manually, or both. In other words, the selection of candidate amino acids may be done manually. For example, amino acids that are not in stock and cannot be used for synthesis may be excluded. Alternatively, candidate amino acids for each core may be determined manually without relying on the amino acids for each core of a measured peptide. Furthermore, candidate amino acids for each core may be determined in ways other than those described above. In addition, multiple candidate peptides may be determined by methods other than determining candidate amino acids for each core. The amino acids are not particularly limited, but at least one of natural amino acids (20 types) and non-natural amino acids may be selected. In this specification, "amino acid" includes natural amino acids and non-natural amino acids. In this specification, "amino acid" may mean an amino acid residue. In this specification, "natural amino acids" refers to Gly (glycine), L-Ala (alanine), L-Ser (serine), L-Thr (threonine), L-Val (valine), L-Leu (leucine), L-Ile (isoleucine), L-Phe (phenylalanine), L-Tyr (tyrosine), L-Trp (tryptophan), L-His (histidine), L-Glu (glutamic acid), L-Asp (aspartic acid), L-Gln (glutamine), L-Asn (asparagine), L-Cys (cysteine), L-Met (methionine), L-Lys (lysine), L-Arg (arginine), and L-Pro (proline). In this specification, L-type amino acids included in natural amino acids may be written without the "L-". That is, L-Ala, which is a natural amino acid, may be written as Ala. In this specification, "non-natural amino acids" refers to amino acids other than natural amino acids. Examples of non-natural amino acids include β-amino acids, γ-amino acids, D-type amino acids, N-substituted amino acids other than Pro (for example, N-alkyl amino acids such as N-methylamino acids and N-ethylamino acids), α,α-disubstituted amino acids, and amino acids whose side chains differ from those of natural amino acids.

[0041] The candidate acquisition unit 11, for example, acquires information indicating candidate amino acids for each core, generates combinations of amino acids that become candidate peptides from the candidates, and acquires candidate composition information. The acquisition of information indicating candidate amino acids for each core is performed, for example, by reading out the information previously stored in the peptide proposal system 10 manually. Alternatively, the acquisition of the information may be performed by receiving information transmitted from another device. Further, the acquisition of the information may be performed by a method other than the above.

[0042] Alternatively, the candidate acquisition unit 11 may acquire candidate composition information generated manually or by another device. In this case, the acquisition of the information may be performed in the same manner as above. The number of candidate peptides related to the candidate composition information acquired by the candidate acquisition unit 11 is not particularly limited, but may be, for example, about 1 million to several hundred million. The candidate acquisition unit 11 outputs the candidate composition information for each of the acquired plurality of candidate peptides to the calculation unit 12, the filtering unit 14, and the determination unit 15.

[0043] The calculation unit 12 is a calculation means that calculates predicted values of characteristics of a plurality of candidate peptides and uncertainties of the predicted values based on the candidate composition information acquired by the candidate acquisition unit 11 using a prediction model.

[0044] The prediction model is a model used to calculate predicted values of peptide characteristics and uncertainties of the predicted values based on composition information indicating the composition of the peptide. The characteristics are the above-mentioned preset characteristics. The prediction model, for example, performs an input based on composition information indicating the composition of the peptide and outputs a predicted value. As the prediction model, a conventional prediction model may be used. For example, the prediction model may be a Gaussian process regression, a random forest, a support vector machine, or a neural network, a Deep ensemble, an ensemble model, bagging, or a Bayesian neural network. The prediction model may be generated by machine learning. Further, the prediction model may be other than the above.

[0045] The uncertainty of the predicted value is, for example, the variation of the predicted value. Specifically, for example, the uncertainty of the predicted value is the confidence interval of the predicted value of the prediction model. When using a Gaussian process regression as the prediction model, the confidence interval of the predicted value can be obtained by calculating the mean value and variance of the predicted value. The wider the confidence interval, the greater the uncertainty of the predicted value.

[0046] As will be described later, the prediction model is generated (constructed) using the measured peptides whose characteristics have been measured. The calculation unit 12 stores the prediction model until the time of calculating the predicted value of the characteristics of the candidate peptide and the uncertainty of the predicted value.

[0047] The calculation unit 12 inputs the candidate configuration information for each of the plurality of candidate peptides from the candidate acquisition unit 11. For each of the plurality of candidate peptides, the calculation unit 12 uses the prediction model to calculate the predicted value of the characteristics of the candidate peptide and the uncertainty of the predicted value based on the candidate configuration information. For example, the calculation of the predicted value of the characteristics and the uncertainty of the predicted value is performed as follows.

[0048] The calculation unit 12 converts the candidate configuration information for each of the plurality of candidate peptides into a format for input to the prediction model based on the conversion rules stored in advance. For example, as shown in FIG. 2, the calculation unit 12 converts the candidate configuration information into a vector having the number of dimensions preset as the format.

[0049] For example, the conversion is performed as follows: The amino acids constituting the peptide indicated by the candidate constituent information are converted into one-hot encoded vectors (e.g., vectors with a type dimension of amino acids), and these vectors are concatenated in the order of amino acids. Alternatively, the amino acids constituting the peptide indicated by the candidate constituent information are converted to SMILES (simplified molecular input line entry system) notation, and as shown in Figure 3, the amino acids in SMILES notation are converted into vectors using existing software (e.g., rdkit, mol2vec, or MolCLR), and these vectors are concatenated in the order of amino acids. The vector for each amino acid may be a fingerprint generated by rdkit, etc. Alternatively, the entire peptide may be represented as a graph, and the peptide may be converted into a vector using a graph network.

[0050] As shown in Figure 4, the calculation unit 12 inputs the vector for each candidate peptide into the prediction model to obtain predicted values ​​for the characteristics of the candidate peptide. In Figure 4, the vector to the left of the arrow is the vector for the candidate peptide, f(x) is the prediction model, and the value to the right of the arrow is the predicted value for the characteristic.

[0051] Furthermore, the calculation unit 12 calculates the uncertainty of the predicted value for each candidate peptide. The calculation of the uncertainty of the predicted value can be performed by a conventional method (for example, an existing package program) according to the prediction model.

[0052] Figure 5 shows a graph relating to predictions made by a prediction model. In the graph in Figure 5, the horizontal axis (x-axis) represents information based on the composition information of the peptide, which is the input to the prediction model, and the vertical axis (y-axis) represents the value of the peptide property (in the example in Figure 5, active membrane permeability). The solid line shows the predicted value of the peptide property, which is the output from the prediction model. The dashed line shows the true value (true function) of the peptide property. The measured value points relate to the measured peptide used to generate the prediction model. Since the prediction model is usually generated to output measured values ​​for measured value points, as shown in Figure 5, there is no (or little) uncertainty (variability) in the predicted value at the measured value points. Further away from the measured value points, accurate prediction becomes difficult, and the uncertainty (variability) in the predicted value increases.

[0053] The calculation unit 12 calculates predicted values ​​and uncertainties of predicted values ​​for pre-set characteristics. The calculation unit 12 may calculate predicted values ​​and uncertainties of predicted values ​​for multiple characteristics. A prediction model may be prepared for each characteristic. In addition, multiple prediction models (for example, 5 to 10 prediction models for each characteristic) may be generated for verification. Pipeline software may be used to generate prediction models.

[0054] The calculation unit 12 outputs the predicted values ​​of the characteristics calculated for each of the multiple candidate peptides, as well as information indicating the uncertainty of those predicted values, to the filtering unit 14 and the determination unit 15.

[0055] The measured data acquisition unit 13 is a measured data acquisition means that acquires measured configuration information indicating the composition of a measured peptide whose characteristics have been measured, which was used to generate the prediction model. The measured peptide may be a linear peptide or a cyclic peptide. The prediction model may be a prediction model generated from the measured configuration information acquired by the measured data acquisition unit 13.

[0056] The measured peptide is the peptide that, in the peptide proposal described above, is actually synthesized and its properties are measured, and the measured properties are used to generate the predictive model. The measured peptide may consist of a predetermined number of amino acids. For example, the measured peptide may consist of 5 to 20 amino acid residues, 9 to 14 amino acid residues, or 11 amino acid residues. The candidate peptide and the measured peptide may all consist of the same number of amino acids.

[0057] The measured peptides used to generate (construct) the initial predictive model can be set by any method (for example, a method following conventional experimental design). The measured peptides used to generate (reconstruct) the predictive model again shall be determined from candidate peptides in the peptide proposal system 10 as follows.

[0058] The prediction model is generated using training information (training data, teacher data) related to the measured peptide. The training information is, for example, information corresponding to the input to the prediction model and the output from the prediction model. The information corresponding to the input to the prediction model is information based on measured constituent information that shows the composition of the measured peptide. The measured constituent information for the measured peptide is the same as the candidate constituent information for the candidate peptide. The measured constituent information is, for example, information showing the amino acids that make up the measured peptide and the position of each amino acid (the order of amino acids). The input to the prediction model is, for example, a vector with a predetermined number of dimensions as described above. The vector of measured constituent information should be generated in the same way as the vector for candidate peptides. The information corresponding to the output from the prediction model is the measured values ​​of the properties of the measured peptide.

[0059] Furthermore, if the measured values ​​of a characteristic are outliers due to variability in wet experiments, the information related to those measurements may not be used in generating the predictive model. Even when measuring the same peptide, the measured values ​​can vary by 3 to 10 times. Variability in wet experiments can occur when experimental conditions (e.g., equipment or temperature conditions) are changed. Using outliers in generating the predictive model may lead to a significant decrease in prediction accuracy due to the influence of outliers. For example, outliers can be identified by the difference (experimental error) between the measured value and the predicted value calculated by the predictive model at that time. For example, values ​​where the difference from the predicted value exceeds a certain level can be considered outliers. Alternatively, outliers may be detected using machine learning techniques.

[0060] The generation of a predictive model based on the measured peptide information can be carried out in the same manner as before, depending on the generated predictive model (e.g., one based on Gaussian process regression, random forest, support vector machine, or neural network as described above). For example, the generation of the predictive model can be performed using machine learning.

[0061] The generation of the prediction model is performed up to the point when the calculation unit 12 calculates the predicted values ​​of the candidate peptide characteristics and the uncertainty of those predicted values. The generation of the prediction model may be performed by the peptide proposal system 10 or by an apparatus other than the peptide proposal system 10.

[0062] The measured data acquisition unit 13 acquires measured configuration information for each measured peptide used in generating the prediction model. The measured configuration information may pertain to all measured peptides used in generating the prediction model, or it may pertain to some of the measured peptides used in generating the prediction model. The measured configuration information is acquired, for example, by manually reading the information that has been previously stored in the peptide suggestion system 10. Alternatively, the information may be acquired by receiving information transmitted from another device. Furthermore, the information may be acquired by methods other than those described above.

[0063] The measured configuration information acquired by the measured data acquisition unit 13 pertains to the measured peptides used in generating the prediction model, and may include peptides not used in generating the prediction model, as long as they can be used in the processing described later. The measured data acquisition unit 13 outputs the measured configuration information for each acquired measured peptide to the filtering unit 14.

[0064] The filtering unit 14 is a filtering means that filters candidate peptides based on: i) the distance between the amino acid sequence of a candidate peptide indicated by candidate constituent information obtained by the candidate acquisition unit 11 and the amino acid sequence of a measured peptide indicated by measured constituent information obtained by the measured acquisition unit 13; ii) the uncertainty of the predicted value related to the candidate peptide calculated by the calculation unit 12; and iii) at least one selected from a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, amino acid angle, and aromatic amino acid pattern between the candidate peptide and the measured peptide, based on the candidate constituent information obtained by the candidate acquisition unit 11 and the measured constituent information obtained by the measured acquisition unit 13. The filtering unit 14 may further filter the candidate peptides based on the predicted value of the characteristics of the candidate peptide calculated by the calculation unit 12. The filtering unit 14 may also filter candidate peptides based on a plurality of different criteria.

[0065] The filtering of candidate peptides by the filtering unit 14 is intended to narrow down the candidate peptides so that the next candidate peptide to be actually synthesized and have its properties measured is appropriate, as described above. The filtering unit 14 uses any of the above information obtained by the candidate acquisition unit 11, the calculation unit 12, and the measured acquisition unit 13 to set a Trust Region, which is a local search space, and sets the candidate peptides contained in the Trust Region as the filtered candidate peptides, i.e., candidate peptides that may be determined to be actually synthesized and have their properties measured next.

[0066] A Trust Region is, for example, a filter that defines a compound space where predictions by a predictive model are reliable. The compound space for candidate peptides with properties usable for drug discovery increases exponentially with the number of amino acids and the number of amino acid combinations, becoming enormous. Freely proposing within this space risks becoming overly exploratory. By setting a Trust Region, the area where candidate peptides to be actually synthesized and have their properties measured can be proposed can be limited to the range where the predictive model is effective. As a result, excessive proposals of candidate peptides to be actually synthesized and have their properties measured can be prevented. Furthermore, as mentioned above, Trust Regions may be set with different rules for "exploration" and "utilization." It is best to efficiently propose peptides by balancing "exploration" and "utilization."

[0067] The filtering unit 14 performs filtering, for example, as follows: The filtering unit 14 receives candidate composition information for each of the multiple candidate peptides from the candidate acquisition unit 11. The filtering unit 14 receives predicted values ​​of the characteristics of each of the multiple candidate peptides and information indicating the uncertainty of said predicted values ​​from the calculation unit 12. The filtering unit 14 receives measured composition information for each measured peptide from the measured acquisition unit 13.

[0068] The filtering unit 14 stores rules for filtering in advance as filtering criteria, and performs filtering according to those rules. The filtering unit 14 may perform filtering using multiple rules. That is, one criterion may include multiple rules. The filtering unit 14 performs filtering using the input information that corresponds to the rules. Depending on the rule, the filtering unit 14 generates information for filtering and performs filtering based on the generated information. The filtering for each rule will be explained below.

[0069] i) The filtering method described below is based on the distance between the amino acid sequence of a candidate peptide, indicated by the candidate composition information obtained by the candidate acquisition unit 11, and the amino acid sequence of a measured peptide, indicated by the measured composition information obtained by the measured acquisition unit 13. For each candidate peptide, the filtering unit 14 calculates the distance (modification distance, edit distance) between the amino acid sequence of the candidate peptide, indicated by the candidate composition information, and the amino acid sequence of the measured peptide, indicated by the measured composition information, for all combinations with each measured peptide. For example, the filtering unit 14 considers each amino acid sequence as a string and calculates the distance between the strings. For example, the Hamming distance or the Levenshtein distance can be used as the distance between the strings. The calculation of this distance can be performed by a conventional method. Alternatively, the distance between amino acid sequences may be other than those mentioned above. For example, the distance between vectors based on the above-mentioned amino acid sequences may be calculated as the distance between amino acid sequences.

[0070] The filtering unit 14 identifies the smallest distance among all the calculated distances to each of the measured peptides for each candidate peptide, i.e., the distance to the nearest measured peptide. The identified distance becomes the distance for that candidate peptide.

[0071] The filtering unit 14 compares the identified distance with a preset threshold. The filtering unit 14 narrows down (retains) candidate peptides whose distance is below the threshold to be selected as candidate peptides to be actually synthesized and have their properties measured next. In other words, these candidate peptides are those that fall within the range of the Trust Region set by the amino acid sequence of the measured peptide and the threshold. The filtering unit 14 excludes (filters out) candidate peptides whose distance exceeds the threshold from being selected as candidate peptides to be actually synthesized and have their properties measured next.

[0072] The above filtering is based on the principle that the candidate peptides to be synthesized and then characterized should be within a certain range from the measured peptides in terms of the distance between amino acid sequences, thereby enabling the suggestion of appropriate peptides. To narrow the range of candidate peptides to be synthesized and then characterized (i.e., to focus on practical applications), the threshold should be reduced (i.e., limited to short distances). To broaden the range of candidate peptides (i.e., to focus on exploratory applications), the threshold should be increased (i.e., longer distances should be allowed).

[0073] ii) The filtering based on the uncertainty of the predicted value for the candidate peptide calculated by the calculation unit 12 is explained. The filtering unit 14 compares the magnitude of the uncertainty of the predicted value (for example, the value of the variance of the predicted value, which is the confidence interval of the predicted value) with a preset threshold. The filtering unit 14 narrows down (retains) candidate peptides whose magnitude of uncertainty of the predicted value is less than or equal to the threshold to be selected as candidate peptides to be actually synthesized and have their properties measured next. That is, these candidate peptides are candidate peptides within the range of the Trust Region set by the prediction model based on measured peptides and the threshold. The filtering unit 14 excludes candidate peptides whose magnitude of uncertainty of the predicted value exceeds the threshold from being selected as candidate peptides to be actually synthesized and have their properties measured next.

[0074] The filtering described above is based on the principle that selecting candidate peptides to be synthesized and then characterized should have a predictive uncertainty within a certain range, thereby enabling the proposal of appropriate peptides. To narrow the range of candidate peptides to be synthesized and characterized (i.e., to focus on practical applications), the threshold should be reduced (i.e., limited to low uncertainty). To broaden the range of candidate peptides (i.e., to focus on exploratory applications), the threshold should be increased (i.e., tolerating high uncertainty).

[0075] iii) The filtering process described here is based on a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, amino acid angles, and aromatic amino acid pattern between the candidate peptide and the measured peptide, using candidate composition information obtained by the candidate acquisition unit 11 and measured composition information obtained by the measured composition unit 13.

[0076] The N-alkyl pattern of a peptide is, for example, the pattern of the positions (order) of the N-alkylated amino acids that constitute the peptide. Alternatively, the N-alkyl pattern of a peptide may be, for example, the number of N-alkylated amino acids that constitute the peptide (N-alkyl count). "N-alkylated amino acid" refers to an amino acid in which the amino group of the amino acid main chain is alkylated to, for example, a methyl group or an ethyl group. The filtering unit 14 calculates the N-alkyl pattern of a candidate peptide from the candidate constituent information. The filtering unit 14 calculates the N-alkyl pattern of a measured peptide from the measured constituent information. In this case, the candidate constituent information and the measured constituent information are information that allows the calculation of the N-alkyl pattern of the peptide relating to the information. Alternatively, the candidate constituent information and the measured constituent information may include the N-alkyl pattern of the peptide relating to the information.

[0077] The filtering unit 14 compares the N-alkyl pattern of each of the multiple candidate peptides with the N-alkyl patterns of each of the measured peptides. Based on the results of this comparison, the filtering unit 14 narrows down (retains) the candidate peptides whose N-alkyl pattern matches any of the N-alkyl patterns of the measured peptides to be selected as candidate peptides to be actually synthesized and have their properties measured next. In other words, these candidate peptides are candidate peptides within the range of the Trust Region set by the N-alkyl patterns of the measured peptides. The filtering unit 14 excludes candidate peptides whose N-alkyl pattern does not match any of the N-alkyl patterns of the measured peptides from the selection of candidate peptides to be actually synthesized and have their properties measured next. In this way, the N-alkyl patterns of the measured peptides become the N-alkyl patterns of the acceptable candidate peptides.

[0078] The above filtering is based on the principle that the N-alkyl pattern of the candidate peptide to be synthesized and its properties measured next should correspond to the N-alkyl pattern of the measured peptide, thereby enabling the proposal of appropriate peptides. Furthermore, if the range of candidate peptides to be synthesized and their properties measured next is to be narrowed (i.e., set towards practical application), the measured peptides related to the measured constituent information should be limited. For example, the measured peptides related to the measured constituent information should be defined as the main measured peptides. If the range of candidate peptides is to be broadened (i.e., set towards exploration), the measured peptides related to the measured constituent information should be diverse.

[0079] The cyclic amino acid pattern of a peptide is, for example, the pattern of the positions (order) of cyclic amino acids, such as proline (Pro), that constitute the peptide. Alternatively, the cyclic amino acid pattern of a peptide may be, for example, the number of cyclic amino acids that constitute the peptide (number of cyclic amino acids). A "cyclic amino acid" refers to an amino acid that has a cyclic structure in which the nitrogen atom of the amino group of the amino acid and any atom of the side chain come together to form a ring. Examples of cyclic amino acids include proline (Pro), which has a five-membered ring cyclic structure, but unnatural amino acids such as Aze(2) (CAS number: 2133-34-8), which has a four-membered ring cyclic structure, and Pic(2) (CAS number: 3105-95-1), which has a six-membered ring cyclic structure, may also be used. The filtering unit 14 calculates the cyclic amino acid pattern of a candidate peptide from the candidate constituent information. The filtering unit 14 also calculates the cyclic amino acid pattern of a measured peptide from the measured constituent information. In this case, the candidate constituent information and the measured constituent information are information that allows for the calculation of the cyclic amino acid pattern of the peptide related to that information. Alternatively, the candidate and measured constituent information may include the cyclic amino acid pattern of the peptide related to that information.

[0080] The filtering unit 14 compares the cyclic amino acid pattern of each of the multiple candidate peptides with the cyclic amino acid patterns of each of the measured peptides. Based on the results of this comparison, the filtering unit 14 narrows down (retains) the candidate peptides whose cyclic amino acid pattern matches any of the measured peptides to be selected as candidate peptides to be synthesized and have their properties measured next. In other words, these candidate peptides are within the range of the Trust Region set by the cyclic amino acid patterns of the measured peptides. The filtering unit 14 excludes candidate peptides whose cyclic amino acid pattern does not match any of the measured peptides from being selected as candidate peptides to be synthesized and have their properties measured next. In this way, the cyclic amino acid patterns of the measured peptides become the cyclic amino acid patterns of the acceptable candidate peptides.

[0081] The above filtering is based on the principle that the cyclic amino acid pattern of the candidate peptide to be synthesized and its properties measured next should correspond to the cyclic amino acid pattern of the measured peptide, thereby enabling the proposal of appropriate peptides. Furthermore, if the range of candidate peptides to be synthesized and their properties measured next is to be narrowed (i.e., set towards practical application), the measured peptides related to the measured constituent information should be limited. For example, the measured peptides related to the measured constituent information should be defined as the main measured peptides. If the range of candidate peptides is to be broadened (i.e., set towards exploration), the measured peptides related to the measured constituent information should be diverse.

[0082] The angle related to amino acids is, for example, the dihedral angle of the peptide bond between amino acids (i.e., the main chain dihedral angle). This dihedral angle includes the dihedral angle ψ of the Ca-C bond and the dihedral angle φ of the Ca-N bond. Alternatively, the angle related to amino acids may be the dihedral angle of a β-amino acid.

[0083] The filtering unit 14 calculates the range of dihedral angles of candidate peptides from candidate constituent information. The filtering unit 14 also calculates the dihedral angles of measured peptides from measured constituent information. In this case, the candidate constituent information and measured constituent information are information that allows for the calculation of the dihedral angles and range of dihedral angles of the peptides related to the information. The calculation of the dihedral angles and range of dihedral angles can be performed by conventional methods (e.g., structural analysis, molecular dynamics simulation, or prediction). Alternatively, the candidate constituent information and measured constituent information may include the dihedral angles of the peptides related to the information. The dihedral angles used for calculation may be a portion of a pre-set set of dihedral angles, or all dihedral angles of the main chain. An example of software used in molecular dynamics simulation is Amber 18. The calculation conditions using Amber 18 are, for example, those described in Mikimasa Tanada et al. Development of Orally Bioavailable Peptides Targeting an Intracellular Protein: From a Hit to a Clinical KRAS Inhibitor, J. Am. Chem. Soc. The conditions described in 2023, 145, 16610-16620 may also be used.

[0084] The filtering unit 14 compares the range of dihedral angles of each of the multiple candidate peptides with the dihedral angles of each of the measured peptides. Specifically, the filtering unit 14 may check whether each binding dihedral angle within the measured peptide falls within the range of possible dihedral angles of the candidate peptide. As a result of this comparison, the filtering unit 14 narrows down (retains) the candidate peptides whose dihedral angle range includes any of the dihedral angles of the measured peptides to be selected as candidate peptides to be actually synthesized and have their properties measured next. In other words, these candidate peptides are candidate peptides within the range of Trust Region set by the dihedral angles of the measured peptides. The filtering unit 14 excludes candidate peptides whose dihedral angle range does not include any of the dihedral angles of the measured peptides from being selected as candidate peptides to be actually synthesized and have their properties measured next. In this way, the dihedral angles of the measured peptides become the acceptable dihedral angles of the candidate peptides.

[0085] The above filtering is based on the principle that the dihedral angle of the candidate peptide to be synthesized and its properties measured next should correspond to the dihedral angle of the already measured peptide, thereby enabling the suggestion of appropriate peptides. Furthermore, if the range of candidate peptides to be synthesized and their properties measured next is narrowed (i.e., set towards practical application), then, as described above, it is acceptable if the dihedral angle of the already measured peptide falls within the range of the candidate peptide's dihedral angle. If the range of candidate peptides is broadened (i.e., set towards exploration), then, as described above, it is acceptable if the dihedral angle of the already measured peptide falls within a range that is further expanded from the range of the candidate peptide's dihedral angle. In other words, even if the dihedral angle of the already measured peptide does not fall within the range of the candidate peptide's dihedral angle, it may be acceptable under certain circumstances.

[0086] The aromatic amino acid pattern (Aromatic Ring (AR) pattern) of a peptide is, for example, the number of amino acids in the peptide that have an aromatic ring, such as a phenyl group, in their side chains (ARC: Aromatic Ring Count). Alternatively, the aromatic amino acid pattern may be the pattern of the positions (order) of amino acids in the peptide that have an aromatic ring, such as a phenyl group, in their side chains. The filtering unit 14 calculates the aromatic amino acid pattern of a candidate peptide from the candidate constituent information. The filtering unit 14 also calculates the aromatic amino acid pattern of a measured peptide from the measured constituent information. In this case, the candidate constituent information and the measured constituent information are information that allows for the calculation of the aromatic amino acid pattern of the peptide related to that information. Alternatively, the candidate constituent information and the measured constituent information may include the aromatic amino acid pattern of the peptide related to that information. "Aromatic amino acids," "amino acids having aromatic rings," or "aromatic amino acids" include, for example, natural amino acids such as phenylalanine (Phe), tyrosine (Tyr), tryptophan (Trp), and histidine (His), as well as non-natural amino acids such as homophenylalanine (Hph).

[0087] The filtering unit 14 compares the aromatic amino acid pattern of each of the multiple candidate peptides with the aromatic amino acid patterns of each of the measured peptides. Based on the results of this comparison, the filtering unit 14 narrows down (retains) the candidate peptides whose aromatic amino acid pattern matches that of any of the measured peptides to be selected as candidate peptides to be synthesized and have their properties measured next. In other words, these candidate peptides are within the range of the Trust Region set by the aromatic amino acid patterns of the measured peptides. The filtering unit 14 excludes candidate peptides whose aromatic amino acid pattern does not match that of any of the measured peptides from being selected as candidate peptides to be synthesized and have their properties measured next. In this way, the aromatic amino acid patterns of the measured peptides become the aromatic amino acid patterns of the acceptable candidate peptides.

[0088] The above filtering is based on the principle that the aromatic amino acid pattern of the candidate peptide to be synthesized and its properties measured next should correspond to the aromatic amino acid pattern of the measured peptide, thereby enabling the proposal of appropriate peptides. Furthermore, if the range of candidate peptides to be synthesized and their properties measured next is to be narrowed (i.e., set towards practical application), the measured peptides related to the measured constituent information should be limited. For example, the measured peptides related to the measured constituent information should be defined as the main measured peptides. If the range of candidate peptides is to be broadened (i.e., set towards exploration), the measured peptides related to the measured constituent information should be diverse.

[0089] The filtering methods described in i) and iii) above are based on the peptide's backbone. It is known that the properties of cyclic peptides change significantly depending on the characteristics of their backbone. Therefore, filtering based on the peptide's backbone, as described above, is effective.

[0090] The filtering unit 14 performs at least one of the filtering methods i) to iii) described above. That is, the filtering unit 14 may perform all of the filtering methods i) to iii), or it may perform only one of the filtering methods i) to iii). In addition to at least one of the filtering methods i) to iii), the filtering unit 14 may also perform the following filtering methods.

[0091] The filtering unit 14 may filter candidate peptides using the N-alkyl pattern, cyclic amino acid pattern, angle related to amino acids, or aromatic amino acid pattern (as used in filtering iii above), and may also filter candidate peptides without using information related to measured peptides. The filtering unit 14 may store the filtering rules set in advance for such cases and use them to perform filtering.

[0092] The filtering unit 14 may further filter candidate peptides other than those in i) to iii). For example, the filtering unit 14 further filters candidate peptides based on predicted values ​​of the characteristics of the candidate peptides calculated by the calculation unit 12. The filtering unit 14 filters using predicted values ​​of characteristics that have been set in advance to be used for filtering. The characteristics used for filtering are, for example, pre-set binding affinity to a target, pharmacological activity, lipophilicity, solubility, membrane permeability, metabolic stability, pharmacokinetics, or toxicity.

[0093] The filtering unit 14 compares the predicted values ​​of the candidate peptide's properties with a preset threshold. Based on the results of this comparison, the filtering unit 14 narrows down (retains) the candidate peptides to be selected as candidates for actual synthesis and property measurement. For example, if a higher value of a property is desirable, the filtering unit 14 narrows down (retains) the candidate peptides whose predicted property values ​​are above the threshold to be selected as candidates for actual synthesis and property measurement. The filtering unit 14 excludes candidate peptides whose predicted property values ​​are below the threshold from being selected as candidates for actual synthesis and property measurement.

[0094] The filtering unit 14 may further filter candidate peptides using conventional rules. Conventional rules include, for example, the drug-likeness rule for cyclic peptides. The drug-likeness rule is shown, for example, in Atsushi Ohata et al. Validation of a New Methodology to Create Oral Drugs beyond the Rule of 5 for Intracellular Tough Targets, J. Am. Chem. Soc. 2023, 44, 24035–24051, International Publication No. 2013 / 100132 and International Publication No. 2018 / 225864.

[0095] Furthermore, the filtering unit 14 may further filter candidate peptides using information about the main chain of the candidate peptides, which are cyclic peptides, in a manner other than those described above. For example, the main chain information of the candidate peptides used may include: whether or not the amide group of the amino acids constituting the cyclic peptide is modified (NH, NMe / N-alkyl), the number, pattern and type; whether or not the amino acids constituting the cyclic peptide have a cyclic structure (cyclic amino acid, acyclic amino acid), the number, pattern and type; whether or not the amino acids constituting the cyclic peptide have an aromatic group (having an aromatic group, not having an aromatic group), the number, pattern and type; and the range of angles between each bond within the amino acids constituting the cyclic peptide (main chain dihedral angles ψ, φ, etc.).

[0096] The filtering unit 14 outputs information indicating the filtering results, specifically information indicating the candidate peptides remaining after filtering, to the determination unit 15. The candidate peptides remaining after filtering are at least those contained in Trust Region.

[0097] As described above, the filtering unit 14 may perform filtering based on different criteria for each iteration of peptide proposals. That is, the filtering unit 14 may perform filtering with a different degree of exploration (balance between "exploration" and "utilization") for each iteration of peptide proposals. For example, the filtering unit 14 may store the number of iterations in association with the criteria and perform filtering based on the criteria corresponding to the number of iterations. For example, the filtering unit 14 may perform filtering corresponding to the "exploration-oriented proposals" described above for a small number of iterations, and filtering corresponding to the "utilization-oriented proposals" described above for a large number of iterations.

[0098] The filtering unit 14 may perform filtering based on multiple different criteria and obtain multiple filtering results corresponding to those criteria. That is, the filtering unit 14 may perform filtering with multiple different levels of exploration and obtain multiple filtering results. Each of the multiple criteria consists of one or more of the filtering rules described above. For example, the multiple criteria may correspond to the "exploration" and "utilization" described above. The multiple criteria may include the same rule. The multiple criteria may be made different from each other by using different information (e.g., threshold values) for filtering that are set in advance in the same rule. In this case, the filtering unit 14 outputs to the determination unit 15, for each criterion, information indicating the filtering result, which is information indicating the candidate peptides that remain after filtering.

[0099] The determination unit 15 is a determination means that determines a candidate peptide for which characteristics are to be measured, based on at least one of the predicted value of the characteristics calculated by the calculation unit 12 and the uncertainty of the predicted value, from the candidate peptide after filtering by the filtering unit 14. The determination unit 15 may output information indicating the candidate peptide determined to be measured for characteristics, for each pattern of amino acid species constituting the candidate peptide. The patterns of amino acid species are the following i) to iii): i) NH-amino acids or N-alkyl amino acids ii) Cyclic amino acids or acyclic amino acids iii) At least one selected from having an aromatic group or not having an aromatic group. "NH-amino acids" are amino acid residues in which the amide group formed by the amino group of the main chain of the amino acid constituting the candidate peptide is not modified (unsubstituted). On the other hand, "N-alkyl amino acids" are amino acid residues in which the amide group formed by the amino group of the main chain of the amino acid constituting the candidate peptide is modified (N-alkylated). "Cyclic amino acids" mean amino acids that have a cyclic structure in which the nitrogen atom of the amino group of the amino acid and any atom of the side chain come together to form a ring. Examples of cyclic amino acids include proline (Pro). "Acyclic amino acids" are amino acids other than cyclic amino acids. The determination unit 15 may output information indicating candidate peptides determined to measure properties for each of several criteria. The determination unit 15 may determine candidate peptides to measure properties using a Bayesian optimization acquisition function or the order of predicted property values. For example, the determination unit 15 may determine candidate peptides using a ranking algorithm.

[0100] The determination unit 15, for example, determines the candidate peptide to be actually synthesized and have its properties measured next from the candidate peptides remaining after filtering, as follows: The determination unit 15 receives candidate composition information for each of the multiple candidate peptides from the candidate acquisition unit 11. The determination unit 15 receives information from the calculation unit 12 indicating the predicted values ​​of the properties for each of the multiple candidate peptides and the uncertainty of those predicted values. The determination unit 15 receives information from the filtering unit 14 indicating the candidate peptides remaining after filtering, which is information indicating the results of filtering.

[0101] The determination unit 15 selects the top N candidate peptides from the remaining candidate peptides after filtering, in order of their predicted values ​​for a predetermined characteristic, the target index, as the candidate peptides to be synthesized and have their characteristics measured next. N is a predetermined number. If a larger value for the target index is desirable, the determination unit 15 selects the top N candidate peptides in descending order of their predicted values ​​as the candidate peptides to be synthesized and have their characteristics measured next.

[0102] Alternatively, the determination unit 15 uses a Bayesian optimization acquisition function to determine the candidate peptide to be synthesized and have its properties measured next, based on the predicted values ​​of the properties of each candidate peptide and the uncertainty of those predicted values. The determination using the Bayesian optimization acquisition function can be performed using a conventional Bayesian optimization method (e.g., an existing package program). For example, conventional Bayesian optimization methods include Thompson sampling, UCB, EI, PI, qEI, and qUCB. An acquisition function with a parameter for adjusting the degree of exploration may also be used. One such parameter is the UCB hyperparameter β. In this case as well, the determination unit 15 may store a predetermined number of candidate peptides to be determined (e.g., 100) and determine that number of candidate peptides.

[0103] Furthermore, the determination unit 15 may determine a candidate peptide to be synthesized and have its properties measured next by a method other than those described above, provided that it is based on at least one of the predicted value of the properties and the uncertainty of the predicted value, and that it enables the suggestion of an appropriate peptide.

[0104] The determination unit 15 outputs information indicating candidate peptides that will be actually synthesized and have their properties measured after the determination. This output presents the peptide to the user of the peptide suggestion system 10. For example, the determination unit 15 outputs candidate composition information of the determined candidate peptide. The determination unit 15 may also output information related to the determined candidate peptide.

[0105] The information output may also include information used in the above processing (for example, predicted values ​​of properties (e.g., binding affinity (molecular binding activity or molecular inhibitory activity), pharmacological activity (molecular and cellular activity), lipophilicity, solubility, membrane permeability, metabolic stability, pharmacokinetics, and toxicity), uncertainty of said predicted values, distance from the amino acid sequence of the measured peptide, number of N-alkylated amino acids in the peptide, and number of amino acids having aromatic rings such as phenyl groups in the side chains in the peptide). The information output may also include information not used in the above processing (for example, molecular weight and cLog P). Of the above information, molecular weight, cLog P, number of N-alkylated amino acids in the peptide, and number of amino acids having aromatic rings such as phenyl groups in the side chains in the peptide are drug-likeness indices. The information not used in the above processing may be calculated by the determination unit 15 from the candidate composition information at the time of information output. The determination unit 15 may also set the order of the information indicating the determined candidate peptides to the order in which the candidate peptides were determined (for example, in order of predicted values ​​of the target indices). By referring to candidate structural information, the structure (e.g., amino acid sequence) of the selected candidate peptide can be determined.

[0106] cLog P is a computer-calculated distribution coefficient and can be determined according to the principles described in "CLOGP Reference Manual Dailylight Version 4.9 (Release Date: August 1, 2011, https: / / www.daylight.com / dayhtml / doc / clogp / )". One example of how to calculate cLog P is provided by Dailylight Chemical Information Systems, Inc. One possible method is to use Daylight Version 4.95 (release date: August 1, 2011, ClogP algorithm version 5.4, database version 28, https: / / www.daylight.com / dayhtml / doc / release_notes / index.html) for the calculation.

[0107] The principles described in "CLOGP Reference Manual Daily Version 4.9 (Release date: August 1, 2011, https: / / www.dailylight.com / dayhtml / doc / clogp / )" are as described in paragraph

[0101] of International Publication No. 2024 / 214825.

[0108] For example, the decision unit 15 transmits the information to the user's terminal of the peptide suggestion system 10. Alternatively, the decision unit 15 may display the information on a display device provided by the peptide suggestion system 10. The information may be output as, for example, spreadsheet data or as a web screen. The decision unit 15 may also output the information by methods and to destinations other than those described above.

[0109] The user of the peptide suggestion system 10 can refer to the information output from the peptide suggestion system 10 when actually synthesizing a peptide and measuring its properties. In the peptide suggestion process, the information is used as a reference, and a new peptide is actually synthesized and its properties are measured. For example, all of the peptides determined by the peptide suggestion system 10 as candidate peptides to be actually synthesized and their properties measured may be actually synthesized and their properties measured. Alternatively, some of the peptides determined by the peptide suggestion system 10 as candidate peptides to be actually synthesized and their properties measured may be actually synthesized and their properties measured. The above method of peptide synthesis is a method for producing candidate peptides determined using the peptide suggestion system 10.

[0110] If the measured characteristics match the desired characteristics, then a peptide possessing the desired characteristics has been proposed. On the other hand, if the measured characteristics do not match the desired characteristics, a predictive model is generated (reconstructed) again based on the constituent information showing the composition of the synthesized peptide and the characteristics of the measured peptide. The peptide whose characteristics have been measured is designated as the new measured peptide, and the above processing is performed again by the peptide proposal system 10 using the generated predictive model. This process is repeated until a peptide possessing the desired characteristics can be proposed. In other words, if one round is defined as the process from acquiring the information necessary to determine the candidate peptide whose characteristics to be measured (data collection) to determining the candidate peptide whose characteristics to be measured, then the rounds are repeated until a peptide possessing the desired characteristics can be proposed.

[0111] The determination unit 15 may output information indicating candidate peptides determined to measure characteristics, for each pattern of amino acid species constituting the candidate peptide. For example, the determination unit 15 may output information indicating candidate peptides for each NH-amino acid or N-alkyl amino acid. Alternatively, the determination unit 15 may output information indicating candidate peptides for each cyclic amino acid or acyclic amino acid. The determination unit 15 may also output information indicating candidate peptides for each amino acid species having an aromatic group or not having an aromatic group. Furthermore, the determination unit 15 may output information indicating candidate peptides for each N-alkyl pattern. Additionally, the determination unit 15 may output information indicating candidate peptides for each aromatic amino acid pattern.

[0112] When the filtering unit 14 performs filtering based on multiple different criteria, the determination unit 15 determines, for each criterion, a candidate peptide to be synthesized and have its properties measured next from the candidate peptides remaining after filtering. The determination unit 15 outputs information indicating the candidate peptide determined for each criterion. The number of candidate peptides to be determined may be set in advance for each criterion. For example, the number of candidate peptides to be determined may be 100 for the criteria related to "utilization" and "exploration," while the number of candidate peptides to be determined may be 50 for the criterion that leans towards "utilization" and uses a stricter filter based on predicted membrane permeability values.

[0113] As described above, by separating the output according to the pattern or criteria of the amino acid species that constitute the candidate peptide, it becomes easier for the user of the peptide suggestion system 10 (chemist, etc.) to select a peptide to actually synthesize from the candidate peptides indicated by the output information. The above describes the functions of the peptide suggestion system 10 according to this embodiment.

[0114] Next, the peptide proposal method, which is a process (operation method performed by the peptide proposal system 10) executed by the peptide proposal system 10 according to this embodiment, will be explained using the flowchart in Figure 6. In this process, the candidate acquisition unit 11 acquires candidate configuration information showing the composition of a plurality of candidate peptides that are candidates for measuring characteristics (S01, candidate acquisition step). Subsequently, the calculation unit 12 uses a prediction model to calculate predicted values ​​of the characteristics of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate configuration information acquired by the candidate acquisition unit 11 (S02, calculation step). In addition, the measured acquisition unit 13 acquires measured configuration information showing the composition of measured peptides whose characteristics have been measured and which were used to generate the prediction model (S03, measured acquisition step). Note that the processes of acquiring candidate configuration information by the candidate acquisition unit 11 and calculating predicted values ​​of characteristics and the uncertainty of said predicted values ​​by the calculation unit 12 (S01, S02) and the process of acquiring measured configuration information by the measured acquisition unit 13 (S03) can be performed independently of each other, so they do not necessarily have to be performed in the order described above.

[0115] Next, the filtering unit 14 filters the candidate peptides (S04, filtering step). This filtering is performed based on at least one selected from i) the distance between the amino acid sequence of the candidate peptide indicated by the candidate composition information obtained by the candidate acquisition unit 11 and the amino acid sequence of the measured peptide indicated by the measured composition information obtained by the measured composition information obtained by the measured composition information unit 13, ii) the uncertainty of the predicted value for the candidate peptide calculated by the calculation unit 12, and iii) a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, amino acid angle, and aromatic amino acid pattern between the candidate peptide and the measured peptide based on the candidate composition information obtained by the candidate acquisition unit 11 and the measured composition information obtained by the measured composition information unit 13. Furthermore, filtering by the filtering unit 14 may be performed based on factors other than those described above.

[0116] Next, the determination unit 15 determines a candidate peptide whose properties are to be measured from the candidate peptides filtered by the filtering unit 14, based on at least one of the predicted values ​​of the properties calculated by the calculation unit 12 and the uncertainty of the predicted values ​​(S05, determination step). Subsequently, information indicating the candidate peptide whose properties are to be measured, determined by the determination unit 15, is output (S06). This output is used as a reference to actually synthesize a new peptide, and the properties of the peptide are measured. If the measured properties are the desired properties, then a peptide with the desired properties has been proposed. On the other hand, if the measured properties are not the desired properties, a prediction model is generated (reconstructed) again based on the configuration information indicating the composition of the synthesized peptide and the properties of the measured peptide. The peptide whose properties have been measured is designated as a new measured peptide, and the peptide proposal method, which is the process performed by the peptide proposal system 10, is repeated until a peptide with the desired properties is proposed using the generated prediction model. The above is the peptide proposal method according to this embodiment.

[0117] In this embodiment, candidate peptides are filtered based on at least one selected from i) to iii) in the description of the filtering unit 14. That is, candidate peptides to be actually synthesized and have their properties measured are determined from those contained in the Trust Region corresponding to at least one selected from i) to iii) as described above. The candidate peptides after filtering, i.e., the candidate peptides contained in the Trust Region, are suitable for peptide proposal. Therefore, according to this embodiment, peptides can be appropriately proposed.

[0118] As in this embodiment, the candidate peptides and measured peptides may be cyclic peptides. With this configuration, cyclic peptides can be appropriately proposed. However, the peptides proposed by the peptide proposal system 10 do not necessarily have to be cyclic peptides.

[0119] As in this embodiment, the filtering unit 14 may further filter the candidate peptides based on the predicted values ​​of the characteristics of the candidate peptides calculated by the calculation unit 12. This configuration allows for the suggestion of even more appropriate peptides. However, further filtering based on the predicted values ​​of the characteristics of the candidate peptides is not necessarily required.

[0120] As in this embodiment, the properties may be at least one selected from binding affinity (molecular binding activity or molecular inhibitory activity), pharmacological activity (molecular and cellular activity), lipophilicity, solubility, membrane permeability, metabolic stability, pharmacokinetics, and toxicity. With this configuration, peptides relating to the properties can be appropriately proposed. However, the properties used may be other than those described above.

[0121] As in this embodiment, the determination unit 15 may output information indicating the candidate peptides determined to be measured for each pattern of amino acid species constituting the candidate peptide. With this configuration, it is possible to understand the proposed peptides for each pattern of amino acid species.

[0122] In this case, the amino acid species patterns may be the following i) to iii): i) NH-amino acids or N-alkyl amino acids ii) cyclic amino acids or acyclic amino acids iii) at least one selected from having an aromatic group or not having an aromatic group. With this configuration, the proposed peptides can be identified for each of the above patterns. However, it is not necessary to output information indicating candidate peptides for each pattern of amino acid species constituting the candidate peptide.

[0123] As in this embodiment, the filtering unit 14 may filter candidate peptides based on multiple different criteria, and the determination unit 15 may output information indicating the candidate peptides determined to be measured for each of the multiple criteria. With this configuration, it is possible to understand the peptides proposed for each criterion used to filter the candidate peptides. However, filtering candidate peptides based on multiple different criteria and outputting information indicating candidate peptides for each of the multiple criteria are not necessarily required.

[0124] As in this embodiment, the prediction model may be a prediction model generated from measured configuration information acquired by the measured information acquisition unit 13. With this configuration, the predicted values ​​of the characteristics of multiple candidate peptides and the uncertainty of said predicted values ​​can be calculated appropriately and reliably using the prediction model. As a result, peptides can be proposed appropriately and reliably. However, the prediction model may be generated from information other than measured configuration information.

[0125] As in this embodiment, the determination unit 15 may determine candidate peptides for which characteristics are to be measured using the Bayesian optimization acquisition function or the order of the predicted values ​​of the characteristics. With this configuration, peptides can be proposed appropriately and reliably. However, the determination of candidate peptides for which characteristics are to be measured may be performed by methods other than those described above, as long as it is based on at least one of the predicted values ​​of the characteristics and the uncertainty of said predicted values.

[0126] Next, a peptide proposal program for executing the processing by the series of peptide proposal systems 10 described above will be explained. As shown in Figure 7, the peptide proposal program 100 is stored in a program storage area 111 formed on a computer-readable recording medium 110 that is inserted into and accessed by a computer, or is provided by the computer. The recording medium 110 may be a non-temporary recording medium.

[0127] The peptide proposal program 100 comprises a candidate acquisition module 101, a calculation module 102, a measured acquisition module 103, a filtering module 104, and a determination module 105. The functions realized by executing the candidate acquisition module 101, the calculation module 102, the measured acquisition module 103, the filtering module 104, and the determination module 105 are the same as the functions of the candidate acquisition unit 11, the calculation unit 12, the measured acquisition unit 13, the filtering unit 14, and the determination unit 15 of the peptide proposal system 10 described above.

[0128] Furthermore, the peptide proposal program 100 may be configured such that part or all of it is transmitted via a transmission medium such as a communication line, received and recorded (including installed) by other equipment. Also, each module of the peptide proposal program 100 may be installed on multiple computers, not just one. In that case, the series of processes described above will be performed by a computer system consisting of these multiple computers.

[0129] [Note] As can be seen from the various examples above, this disclosure includes the following aspects: (Note 1) Candidate acquisition means for acquiring candidate configuration information showing the composition of a plurality of candidate peptides that are candidates for measuring characteristics; calculation means for calculating predicted values ​​of the characteristics of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate configuration information acquired by the candidate acquisition means using a prediction model; measured acquisition means for acquiring measured configuration information showing the composition of measured peptides whose characteristics have been measured, which was used to generate the prediction model; i) the distance between the amino acid sequence of the candidate peptide shown by the candidate configuration information acquired by the candidate acquisition means and the amino acid sequence of the measured peptide shown by the measured configuration information acquired by the measured acquisition means; ii) the uncertainty of the predicted value related to the candidate peptide calculated by the calculation means; and iii) a filtering means for filtering the candidate peptide based on at least one selected from a comparison of the N-alkyl pattern, cyclic amino acid pattern, angle related to amino acids, and aromatic amino acid pattern of the candidate peptide and the measured peptide based on the candidate configuration information acquired by the candidate acquisition means and the measured configuration information acquired by the measured acquisition means. A peptide proposal system comprising: a determination means for determining a candidate peptide whose properties are to be measured, based on at least one of the predicted values ​​of the properties calculated by the calculation means and the uncertainty of the predicted values, from the candidate peptides after filtering by the filtering means. (Note 2) The peptide proposal system according to Note 1, wherein the filtering means further filters the candidate peptides based on the predicted values ​​of the properties of the candidate peptides calculated by the calculation means. (Note 3) The peptide proposal system according to Note 1 or 2, wherein the filtering means filters the candidate peptides based on a plurality of different criteria, and the determination means outputs information indicating the candidate peptides determined to have their properties measured, for each of the plurality of criteria.(Note 4) The peptide proposal system according to any one of Notes 1 to 3, wherein the determination means determines a candidate peptide to measure the properties using a Bayesian optimization acquisition function or the order of predicted values ​​of the properties. (Note 5) The peptide proposal system according to any one of Notes 1 to 4, wherein the determination means outputs information indicating the candidate peptide determined to measure the properties, for each pattern of amino acid species constituting the candidate peptide. (Note 6) The peptide proposal system according to Note 5, wherein the amino acid species pattern is selected from the following i) to iii): i) NH-amino acids or N-alkyl amino acids ii) cyclic amino acids or acyclic amino acids iii) having an aromatic group or not having an aromatic group (Note 7) The peptide proposal system according to any one of Notes 1 to 6, wherein the prediction model is a prediction model generated from measured configuration information obtained by the measured acquisition means. (Note 8) The peptide proposal system according to any one of Notes 1 to 7, wherein the aforementioned characteristic is at least one selected from molecular binding activity, molecular inhibitory activity, molecular and cellular activity, lipid solubility, solubility, membrane permeability, metabolic stability, pharmacokinetics, and toxicity. (Note 9) The peptide proposal system according to any one of Notes 1 to 8, wherein the candidate peptide and the measured peptide are cyclic peptides. (Note 10) A method for producing a candidate peptide determined using the peptide proposal system according to any one of Notes 1 to 9.(Note 11) A peptide proposal method which is a method of operation of a peptide proposal system, comprising: a candidate acquisition step of acquiring candidate configuration information indicating the composition of a plurality of candidate peptides which are candidates for measuring characteristics; a calculation step of calculating predicted values ​​of the characteristics of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate configuration information acquired in the candidate acquisition step using a prediction model; a measured acquisition step of acquiring measured configuration information indicating the composition of measured peptides whose characteristics have been measured, which was used to generate the prediction model; i) the distance between the amino acid sequence of the candidate peptide indicated by the candidate configuration information acquired in the candidate acquisition step and the amino acid sequence of the measured peptide indicated by the measured configuration information acquired in the measured acquisition step; ii) the uncertainty of the predicted value related to the candidate peptide calculated in the calculation step; and iii) a filtering step of filtering the candidate peptide based on at least one selected from a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, angle related to amino acids, and aromatic amino acid pattern between the candidate peptide and the measured peptide, based on the candidate configuration information acquired in the candidate acquisition step and the measured configuration information acquired in the measured acquisition step. A peptide proposal method comprising: a determination step of determining a candidate peptide to measure properties from the candidate peptides after filtering in the filtering step, based on at least one of the predicted values ​​of the properties calculated in the calculation step and the uncertainty of said predicted values.(Note 12) The computer comprises: candidate acquisition means for acquiring candidate configuration information showing the composition of a plurality of candidate peptides that are candidates for measuring characteristics; calculation means for calculating predicted values ​​of the characteristics of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate configuration information acquired by the candidate acquisition means using a prediction model; measured acquisition means for acquiring measured configuration information showing the composition of measured peptides whose characteristics have been measured, which was used to generate the prediction model; i) the distance between the amino acid sequence of the candidate peptide shown by the candidate configuration information acquired by the candidate acquisition means and the amino acid sequence of the measured peptide shown by the measured configuration information acquired by the measured acquisition means; ii) the uncertainty of the predicted value related to the candidate peptide calculated by the calculation means; and iii) filtering means for filtering the candidate peptide based on at least one selected from a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, angle related to amino acids, and aromatic amino acid pattern between the candidate peptide and the measured peptide, based on the candidate configuration information acquired by the candidate acquisition means and the measured configuration information acquired by the measured acquisition means. A peptide proposal program that operates as a determination means for determining a candidate peptide whose properties are to be measured, based on at least one of the predicted values ​​of the properties calculated by the calculation means and the uncertainty of the predicted values, from candidate peptides after filtering by the filtering means.(Note 13) A peptide proposal system comprising: a filtering means for filtering candidate peptides based on candidate configuration information showing the configurations of a plurality of candidate peptides that are candidates for measuring properties, the configurations of which have been calculated using a prediction model, and measured configuration information showing the configurations of measured peptides whose properties have been measured and which were used to generate the prediction model, i) the distance between the amino acid sequence of the candidate peptide shown by the candidate configuration information and the amino acid sequence of the measured peptide shown by the measured configuration information, ii) the uncertainty of the prediction value relating to the candidate peptide, and iii) at least one selected from a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, angle relating to amino acids, and aromatic amino acid pattern between the candidate peptide and the measured peptide based on the candidate configuration information and the measured configuration information; and a determination means for determining a candidate peptide for measuring properties from the candidate peptides after filtering by the filtering means, based on at least one of the prediction value of properties and the uncertainty of the prediction value calculated by the calculation means. (Note 14) The peptide proposal system according to Note 13, further comprising calculation means for calculating predicted values ​​of the characteristics of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate configuration information using the prediction model. (Note 15) The peptide proposal system according to Note 13 or 14, further comprising candidate acquisition means for acquiring the candidate configuration information. (Note 16) The peptide proposal system according to any one of Notes 13 to 15, further comprising measured acquisition means for acquiring the measured configuration information.(Note 17) A peptide proposal method comprising: a filtering step of filtering candidate peptides based on candidate configuration information showing the configurations of a plurality of candidate peptides that are candidates for measuring properties, for which predicted values ​​of properties and the uncertainty of said predicted values ​​have been calculated using a prediction model, and measured configuration information showing the configurations of measured peptides whose properties have been measured and which were used to generate the prediction model, i) the distance between the amino acid sequence of the candidate peptide shown by the candidate configuration information and the amino acid sequence of the measured peptide shown by the measured configuration information, ii) the uncertainty of the predicted values ​​relating to the candidate peptides, and iii) at least one selected from a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, angle relating to amino acids, and aromatic amino acid pattern between the candidate peptide and the measured peptide based on the candidate configuration information and the measured configuration information; and a determination step of determining a candidate peptide for measuring properties from the candidate peptides after filtering in the filtering step, based on at least one of the predicted values ​​of properties and the uncertainty of said predicted values ​​calculated by the calculation means. (Note 18) The peptide proposal method according to Note 17, further comprising a calculation step of calculating predicted values ​​of the characteristics of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate constituent information using the prediction model. (Note 19) The peptide proposal method according to Note 17 or 18, further comprising a candidate acquisition step of acquiring the candidate constituent information. (Note 20) The peptide proposal method according to any one of Notes 17 to 19, further comprising a measured acquisition step of acquiring the measured constituent information.(Note 21) A step of generating a retrained predictive model based on the measured characteristics of the candidate peptide determined in the determination step; a second filtering step of filtering the second candidate peptide based on: i) the distance between the amino acid sequence of the second candidate peptide shown by the second candidate configuration information and the amino acid sequence of the measured peptide shown by the second measured configuration information, i) the uncertainty of the predicted value relating to the second candidate peptide, and iii) at least one selected from a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, amino acid angle, and aromatic amino acid pattern between the second candidate peptide and the measured peptide based on the second candidate configuration information and the second measured configuration information; and the second measured configuration information used to generate the retrained predictive model, i) the distance between the amino acid sequence of the second candidate peptide shown by the second candidate configuration information and the amino acid sequence of the measured peptide shown by the second measured configuration information; ii) the uncertainty of the predicted value relating to the second candidate peptide; and iii) at least one selected from a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, amino acid angle, and aromatic amino acid pattern between the second candidate peptide and the measured peptide based on the second candidate configuration information and the second measured configuration information. A peptide proposal method according to any one of Appendix 17 to 20, comprising: a second determination step of determining a candidate peptide to measure properties from the second candidate peptide after filtering in the second filtering step, based on at least one of the predicted value of the properties and the uncertainty of the predicted value. (Appendix 22) A peptide proposal method according to any one of Appendix 17 to 21, wherein the criterion for the second filtering step is, compared to the criterion for the filtering step, i) the distance between the amino acid sequence of the candidate peptide indicated by the second candidate constituent information and the amino acid sequence of the property-measured peptide indicated by the second measured constituent information is short, ii) the uncertainty of the predicted value relating to the second candidate peptide is low, and iii) the similarity between the N-alkyl pattern, cyclic amino acid pattern, angle relating to amino acids, and aromatic amino acid pattern of the second candidate peptide and the measured peptide is high, based on the second candidate constituent information and the second measured constituent information.With this configuration, candidate peptides to be characterized are determined based on two types of filtering steps: a first filtering step and a second filtering step. By considering these two types of filtering steps, it is possible to propose a wide range of candidate peptides, from identifying lead peptides for exploration purposes to optimizing lead peptides for practical applications. (Note 23) A peptide proposal program that causes a computer to operate as: a filtering means for filtering candidate peptides based on: candidate configuration information showing the configurations of a plurality of candidate peptides that are candidates for measuring characteristics, for which predicted values ​​of characteristics and uncertainty of said predicted values ​​have been calculated using a prediction model; measured configuration information showing the configurations of measured peptides whose characteristics have been measured and which were used to generate the prediction model; i) the distance between the amino acid sequence of the candidate peptide shown by the candidate configuration information and the amino acid sequence of the measured peptide shown by the measured configuration information; ii) uncertainty of the predicted value relating to the candidate peptide; and iii) at least one selected from a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, angle relating to amino acids, and aromatic amino acid pattern of the candidate peptide and the measured peptide based on the candidate configuration information and the measured configuration information; and a determination means for determining a candidate peptide for measuring characteristics from the candidate peptides after filtering by the filtering means, based on at least one of the predicted values ​​of characteristics and uncertainty of said predicted values ​​calculated by the calculation means.

[0130] 10...Peptide proposal system, 11...Candidate acquisition unit, 12...Calculation unit, 13...Measured acquisition unit, 14...Filtering unit, 15...Decision unit, 100...Peptide proposal program, 101...Candidate acquisition module, 102...Calculation module, 103...Measured acquisition module, 104...Filtering module, 105...Decision module, 110...Recording medium, 111...Program storage area.

Claims

1. Candidate acquisition means for acquiring candidate configuration information showing the composition of a plurality of candidate peptides that are candidates for measuring characteristics; calculation means for calculating predicted values ​​of the characteristics of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate configuration information acquired by the candidate acquisition means using a prediction model; measured acquisition means for acquiring measured configuration information showing the composition of measured peptides whose characteristics have been measured, which was used to generate the prediction model; i) the distance between the amino acid sequence of the candidate peptide shown by the candidate configuration information acquired by the candidate acquisition means and the amino acid sequence of the measured peptide shown by the measured configuration information acquired by the measured acquisition means; ii) the uncertainty of the predicted value related to the candidate peptide calculated by the calculation means; and iii) at least one selected from a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, angle related to amino acids, and aromatic amino acid pattern between the candidate peptide and the measured peptide, based on the candidate configuration information acquired by the candidate acquisition means and the measured configuration information acquired by the measured acquisition means. A peptide proposal system comprising: a determination means for determining a candidate peptide whose properties are to be measured from candidate peptides after filtering by the filtering means, based on at least one of the predicted values ​​of the properties calculated by the calculation means and the uncertainty of the predicted values.

2. The peptide proposal system according to claim 1, wherein the filtering means further filters the candidate peptide based on the predicted values ​​of the characteristics of the candidate peptide calculated by the calculation means.

3. The peptide proposal system according to claim 1 or 2, wherein the filtering means filters the candidate peptides based on a plurality of different criteria, and the determination means outputs information indicating the candidate peptides determined to be used for measuring characteristics, for each of the plurality of criteria.

4. The peptide proposal system according to any one of claims 1 to 3, wherein the determination means determines candidate peptides to measure properties using the acquisition function of Bayesian optimization or the order of predicted values ​​of properties.

5. The peptide proposal system according to any one of claims 1 to 4, wherein the determination means outputs information indicating candidate peptides determined to measure characteristics, for each pattern of amino acid species constituting the candidate peptide.

6. The peptide proposed system according to claim 5, wherein the pattern of amino acid species is selected from the following i) to iii): i) NH-amino acids or N-alkyl amino acids ii) cyclic amino acids or acyclic amino acids iii) having an aromatic group or not having an aromatic group.

7. The peptide proposal system according to any one of claims 1 to 6, wherein the predictive model is a predictive model generated from measured configuration information obtained by the measured acquisition means.

8. The peptide proposed system according to any one of claims 1 to 7, wherein the aforementioned property is at least one selected from molecular binding activity, molecular inhibitory activity, molecular and cellular activity, lipid solubility, solubility, membrane permeability, metabolic stability, pharmacokinetics, and toxicity.

9. The peptide proposal system according to any one of claims 1 to 8, wherein the candidate peptide and the measured peptide are cyclic peptides.

10. A method for producing a candidate peptide determined using the peptide proposal system described in any one of claims 1 to 9.

11. A peptide proposal method which is a method of operation for a peptide proposal system, comprising: a candidate acquisition step of acquiring candidate configuration information indicating the composition of a plurality of candidate peptides which are candidates for measuring properties; a calculation step of using a prediction model to calculate predicted values ​​of the properties of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate configuration information acquired in the candidate acquisition step; a measured acquisition step of acquiring measured configuration information indicating the composition of measured peptides whose properties have been measured, which was used to generate the prediction model; i) the distance between the amino acid sequence of the candidate peptide indicated by the candidate configuration information acquired in the candidate acquisition step and the amino acid sequence of the measured peptide indicated by the measured configuration information acquired in the measured acquisition step; ii) the uncertainty of the predicted value related to the candidate peptide calculated in the calculation step; and iii) a filtering step of filtering the candidate peptide based on at least one selected from a comparison of at least one of the N-alkyl pattern, cyclic amino acid pattern, angle related to amino acids, and aromatic amino acid pattern between the candidate peptide and the measured peptide, based on the candidate configuration information acquired in the candidate acquisition step and the measured configuration information acquired in the measured acquisition step. A peptide proposal method comprising: a determination step of determining a candidate peptide to measure properties from the candidate peptides after filtering in the filtering step, based on at least one of the predicted values ​​of the properties calculated in the calculation step and the uncertainty of said predicted values.

12. A computer comprising: candidate acquisition means for acquiring candidate configuration information showing the composition of a plurality of candidate peptides that are candidates for measuring characteristics; calculation means for calculating predicted values ​​of the characteristics of the plurality of candidate peptides and the uncertainty of said predicted values ​​based on the candidate configuration information acquired by the candidate acquisition means using a prediction model; measured acquisition means for acquiring measured configuration information showing the composition of measured peptides whose characteristics have been measured, which was used to generate the prediction model; i) the distance between the amino acid sequence of the candidate peptide shown by the candidate configuration information acquired by the candidate acquisition means and the amino acid sequence of the measured peptide shown by the measured configuration information acquired by the measured acquisition means; ii) the uncertainty of the predicted value related to the candidate peptide calculated by the calculation means; and iii) at least one selected from a comparison of the N-alkyl pattern, cyclic amino acid pattern, angle related to amino acids, and aromatic amino acid pattern between the candidate peptide and the measured peptide based on the candidate configuration information acquired by the candidate acquisition means and the measured configuration information acquired by the measured acquisition means. A peptide proposal program that operates as a determination means for determining a candidate peptide whose properties are to be measured, based on at least one of the predicted values ​​of the properties calculated by the calculation means and the uncertainty of the predicted values, from candidate peptides after filtering by the filtering means.