Prediction method, prediction device, prediction program, and prediction model

A computer-based method predicts immunogenicity and other properties of protein-based drugs using physicochemical analysis, addressing inefficiencies in conventional methods by providing rapid and resource-effective predictions.

WO2026004863A1PCT designated stage Publication Date: 2026-01-02CHUGAI PHARMA CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/022732
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-06-24
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Conventional methods for predicting the immunogenicity of protein-based pharmaceuticals are time-consuming and resource-intensive, requiring detection of T cell proliferation after several days, leading to inefficiencies in the drug development process.

Method used

A computer-implemented method for predicting immunogenicity, pharmacokinetics, and non-specific binding of test substances based on physicochemical properties such as charge, hydrophobicity, and dipole moment, using a trained neural network model to analyze molecular structures.

Benefits of technology

Enables rapid and efficient prediction of biological properties, reducing the need for extensive resources and time, and improving the selection of suitable drug candidates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025022732_02012026_PF_FP_ABST
    Figure JP2025022732_02012026_PF_FP_ABST
Patent Text Reader

Abstract

This prediction method to be executed by a computer includes a prediction step (S13) in which biological properties of a test substance are predicted on the basis of physicochemical properties based on at least the electrical charge or the hydrophobicity of the test substance. The biological properties can be at least one property selected from the group consisting of the immunogenicity, the pharmacokinetics, and the nonspecific binding properties of the test substance. The physicochemical properties can be further based on a dipole moment. The physicochemical properties can also be based on at least electrical charge distribution or hydrophobicity distribution. The physicochemical properties can also be based on at least deviation in electrical charge distribution or deviation in hydrophobicity distribution. The electrical charge distribution can also be based on electrical charge and a dipole moment, and the hydrophobicity distribution can also be based on hydrophobicity and a dipole moment.
Need to check novelty before this filing date? Find Prior Art

Description

Prediction method, prediction device, prediction program, and prediction model

[0001] One aspect of the present disclosure relates to a biological property prediction method, a biological property prediction device, a biological property prediction program, and a biological property prediction model for predicting the biological property of a test substance.

[0002] Today, many pharmaceuticals with complex structures, such as proteins, are on the market. In developing pharmaceuticals whose active ingredients are derived from proteins, it is necessary to predict the developability of a large number of test substances from an early stage of development.

[0003] Research is being conducted into methods for selecting test substances having desirable biological properties as candidate substances through screening. As an approach for predicting immunogenicity, which is one of the biological properties, an assay method is known in which cells derived from human tissues are stimulated with the test substance, and the number of cells producing cytokines such as IL-2 in the early reaction stage before T cell proliferation and cell proliferation due to the immune response become active is used as a predictive index (for example, Patent Document 1 listed below).

[0004] Special Publication No. 2009-528044

[0005] However, conventional methods use the immune response of T cells to predict the immunogenicity of a test substance, or to screen for drug candidate substances from among the test substances, and involve prediction or screening at the time when T cell proliferation is active (after 5 to 7 days after the start of culture in the presence of the test substance) or at the initial response stage before cell proliferation becomes active (after 2 to 3 days after the start of culture in the presence of the test substance).As a result, it takes a certain amount of time before T cell proliferation or initial response can be detected at the protein level, and there are issues with efficiency in the process of selecting candidate substances suitable for drugs, such as the need for resources such as equipment, manpower, and costs, and the need for testers to acquire technical skills.

[0006] Therefore, it is desirable to more easily predict the biological properties of test substances.

[0007] The biological property prediction method, biological property prediction device, biological property prediction program, and biological property prediction model according to one aspect of the present disclosure can be described as follows: [1] A prediction method implemented by a computer, comprising a prediction step of predicting the biological property of a test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance. [2] The prediction method according to [1], wherein the biological property is at least one property selected from the group consisting of immunogenicity, pharmacokinetics, and non-specific binding of the test substance. [3] A computer-implemented immunogenicity prediction method, comprising a prediction step of predicting the immunogenicity of the test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance. [4] The immunogenicity prediction method according to [3], wherein the physicochemical properties are further based on a dipole moment. [5] The immunogenicity prediction method according to [3] or [4], wherein the physicochemical properties are based on at least one of a charge distribution or a hydrophobicity distribution. [6] The immunogenicity prediction method according to any one of [3] to [5], wherein the physicochemical property is based on at least one of a bias in charge distribution or a bias in hydrophobicity distribution. [7] The immunogenicity prediction method according to [5] or [6], wherein the charge distribution is based on charge and dipole moment, and the hydrophobicity distribution is based on hydrophobicity and dipole moment. [8] The immunogenicity prediction method according to any one of [3] to [7], wherein the test substance is a protein or an antibody. [9] The immunogenicity prediction method according to any one of [3] to [8], wherein the test substance is an antibody, and the physicochemical property is a physicochemical property of an Fv (fragment variable) region of the test substance.

[10] The immunogenicity prediction method according to any one of [3] to [9], wherein the prediction step determines whether the immunogenicity of the test substance is higher or lower than a predetermined standard.

[11] The immunogenicity prediction method according to any one of [3] to

[10] , wherein the physicochemical property is based on at least one of a bias in charge distribution or a bias in hydrophobicity distribution, and the prediction step predicts that the ADA (anti-drug antibody) expression rate of the test substance will be lower than a predetermined ADA expression rate if the bias in charge distribution or the bias in hydrophobicity distribution is within a predetermined criterion.

[12] The immunogenicity prediction method according to any one of [3] to

[10] , wherein the prediction step predicts that the ADA (anti-drug antibody) expression rate of the test substance will be lower than a predetermined ADA expression rate if the overall charge in the fragment variable (Fv) region of the test substance is within a predetermined first criterion and the dipole moment of the test substance is within a predetermined second criterion.

[13] The immunogenicity prediction method according to any one of [3] to

[10] , wherein the prediction step predicts that the ADA (Anti-drug antibody) expression rate of the test substance is lower than a predetermined ADA expression rate when the degree of hydrophobicity of the test substance is within a predetermined first criterion and the dipole moment of the test substance is within a predetermined second criterion.

[14] The immunogenicity prediction method according to

[13] , wherein the degree of hydrophobicity of the test substance is a value obtained by normalizing a hydrophobic surface area with a total surface area.

[15] The immunogenicity prediction method according to any one of [3] to

[14] , wherein the prediction step predicts the immunogenicity of each of a plurality of test substances different from one another, and further includes an output step of outputting information about at least one test substance for which the prediction result from the prediction step satisfies a predetermined criterion.

[16] The immunogenicity prediction method according to any one of [3] to

[15] , wherein the prediction step further comprises making the prediction based on known physicochemical properties and known immunogenicity of each of a plurality of known substances.

[17] The immunogenicity prediction method according to any one of [3] to

[16] , wherein the prediction step comprises making the prediction by applying the physicochemical properties of the test substance to a trained model trained based on the known physicochemical properties and known immunogenicity of each of a plurality of known substances.

[18] The immunogenicity prediction method according to any one of [3] to

[17] , further comprising an evaluation step of evaluating the physicochemical properties of the test substance, wherein the prediction step makes the prediction based on the physicochemical properties evaluated in the evaluation step.

[19] The immunogenicity prediction method according to

[18] , wherein the evaluation step makes the evaluation based on a molecular model of the test substance generated on molecular modeling software.

[20] The immunogenicity prediction method according to

[18] or

[19] , further comprising a generation step of generating a molecular model of the test substance on molecular modeling software, wherein the evaluation step makes the evaluation based on the molecular model generated in the generation step.

[21] The immunogenicity prediction method according to

[19] or

[20] , wherein the molecular model is based on the amino acid sequence of the test substance.

[22] The immunogenicity prediction method according to any one of [3] to

[21] , further comprising an identifying step of identifying elements of the test substance that contribute to the variation in the physicochemical property, wherein the predicting step predicts the immunogenicity of the test substance by substituting the elements identified in the identifying step.

[23] The immunogenicity prediction method according to

[22] , wherein the test substance is a protein or an antibody, and the elements are amino acid residues.

[24] An immunogenicity prediction device comprising at least one processor, wherein the at least one processor predicts the immunogenicity of the test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance.

[25] An immunogenicity prediction program that causes a computer to function as a prediction unit that predicts the immunogenicity of the test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance.

[26] An immunogenicity prediction model, which is a trained model for causing a computer to function to output information regarding the immunogenicity of a test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance, the immunogenicity prediction model being configured by a neural network in which weighting coefficients have been trained using training data including information regarding the known physicochemical properties and information regarding the known immunogenicity of each of a plurality of known substances, and performing a calculation based on the trained weighting coefficients on the information regarding the physicochemical properties of the test substance input to the neural network, and outputting information regarding the immunogenicity of the test substance.

[27] A pharmacokinetic prediction method executed by a computer, comprising a prediction step of predicting the pharmacokinetics of the test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance and the dipole moment of the test substance.

[28] The pharmacokinetic prediction method according to

[27] , wherein the physicochemical properties are based on at least one of a charge distribution or a hydrophobicity distribution.

[29] The pharmacokinetic prediction method according to

[27] or

[28] , wherein the physicochemical property is based on at least one of a bias in charge distribution or a bias in hydrophobicity distribution.

[30] The pharmacokinetic prediction method according to

[28] or

[29] , wherein the charge distribution is based on charge and dipole moment, and the hydrophobicity distribution is based on hydrophobicity and dipole moment.

[31] The pharmacokinetic prediction method according to any one of

[27] to

[30] , wherein the test substance is a protein or an antibody.

[32] The pharmacokinetic prediction method according to any one of

[27] to

[31] , wherein the test substance is an antibody, and the physicochemical property is a physicochemical property of an Fv (Fragment variable) region of the test substance.

[33] The pharmacokinetic prediction method according to any one of

[27] to

[32] , further comprising an evaluation step of evaluating the physicochemical properties of the test substance, wherein the prediction step performs the prediction based on the physicochemical properties evaluated in the evaluation step.

[34] The pharmacokinetic prediction method according to

[33] , wherein the evaluation step performs the evaluation based on a molecular model of the test substance generated on molecular modeling software.

[35] The pharmacokinetic prediction method of

[33] or

[34] , further comprising a generation step of generating a molecular model of the test substance on molecular modeling software, wherein the evaluation step performs evaluation based on the molecular model generated in the generation step.

[36] The pharmacokinetic prediction method of

[34] or

[35] , wherein the molecular model is based on the amino acid sequence of the test substance.

[37] A pharmacokinetic prediction device comprising at least one processor, wherein the at least one processor predicts the pharmacokinetics of the test substance based on physicochemical properties based on at least one of the charge and hydrophobicity of the test substance and the dipole moment of the test substance.

[38] A pharmacokinetic prediction program for causing a computer to function as a prediction unit that predicts the pharmacokinetics of the test substance based on physicochemical properties based on at least one of the charge and hydrophobicity of the test substance and the dipole moment of the test substance.

[39] A computer-implemented method for predicting nonspecific binding, comprising a prediction step of predicting the nonspecific binding of a test substance based on physicochemical properties based on the charge distribution of the test substance.

[40] The method for predicting nonspecific binding according to

[39] , wherein the physicochemical properties are based on a bias in the charge distribution.

[41] The method for predicting nonspecific binding according to

[39] or

[40] , wherein the charge distribution is based on charge and dipole moment.

[42] The method for predicting nonspecific binding according to any one of

[39] to

[41] , wherein the test substance is a protein or an antibody.

[43] The method for predicting nonspecific binding according to any one of

[39] to

[42] , wherein the test substance is an antibody, and the physicochemical properties are those of an Fv (fragment variable) region of the test substance.

[44] The method for predicting nonspecific binding according to any one of

[39] to

[43] , further comprising an evaluation step of evaluating the physicochemical properties of the test substance, wherein the prediction step performs the prediction based on the physicochemical properties evaluated in the evaluation step.

[45] The method for predicting nonspecific binding according to

[44] , wherein the evaluation step performs the evaluation based on a molecular model of the test substance generated on molecular modeling software.

[46] The method for predicting non-specific binding according to

[44] or

[45] , further comprising a generation step of generating a molecular model of the test substance on molecular modeling software, wherein the evaluation step performs evaluation based on the molecular model generated in the generation step.

[47] The method for predicting non-specific binding according to

[45] or

[46] , wherein the molecular model is based on the amino acid sequence of the test substance.

[48] A non-specific binding prediction device comprising at least one processor, wherein the at least one processor predicts the non-specific binding of the test substance based on physicochemical properties based on the charge distribution of the test substance.

[49] A non-specific binding prediction program for causing a computer to function as a prediction unit that predicts the non-specific binding of the test substance based on the physicochemical properties based on the charge distribution of the test substance.

[0008] According to one aspect of the present disclosure, the biological properties of a test substance can be more easily predicted.

[0009] 1 is a diagram showing an example of the functional configuration of a prediction device according to an embodiment. FIG. 2 is a diagram showing an example of the hardware configuration of a computer used in the prediction device according to an embodiment. FIG. 3 is a diagram showing the configuration of a prediction program according to an embodiment together with an auxiliary storage device. FIG. 4 is a diagram showing an example of charge distribution on the surface of an antibody. FIG. 5 is a diagram showing an example of no bias towards positive charges in the charge distribution of the Fv region of a low immunogenic antibody. FIG. 6 is a diagram showing an example of bias towards positive charges in the charge distribution of the Fv region of a high immunogenic antibody. FIG. 7 is a diagram showing an example of no bias towards negative charges in the charge distribution of the Fv region of a low immunogenic antibody. FIG. 8 is a diagram showing an example of bias towards negative charges in the charge distribution of the Fv region of a high immunogenic antibody. FIG. 9 is a diagram showing an example of small or almost no charge / hydrophobic patches. FIG. 10 is a diagram showing an example of the presence of large charge / hydrophobic patches. FIG. 11 is a list (part 1) of 110 types of antibody pharmaceuticals for which ADA incidence rates have been reported. FIG. 12 is a list (part 2) of 110 types of antibody pharmaceuticals for which ADA incidence rates have been reported. FIG. 13 is a list (part 3) of 110 types of antibody pharmaceuticals for which ADA incidence rates have been reported. 1 is a list (part 4) of 110 types of antibody pharmaceuticals for which ADA incidence rates have been reported. 2 is a list (part 5) of 110 types of antibody pharmaceuticals for which ADA incidence rates have been reported. 3 is a list (part 6) of 110 types of antibody pharmaceuticals for which ADA incidence rates have been reported. 4 is a list (part 6) of 110 types of antibody pharmaceuticals for which ADA incidence rates exceed 30%. 5 is a scatter plot (part 1) showing the charge distribution of test samples. 6 is a scatter plot (part 1) showing the hydrophobicity distribution of test samples. 7 is a table (part 1) showing the prediction performance of prediction based on test samples. 8 is a diagram showing the ADA incidence rate according to the bias in charge distribution and / or the bias in hydrophobicity distribution. 9 is a flowchart showing an example of a learning process performed by the immunogenicity prediction device according to the embodiment. 10 is a flowchart showing an example of a prediction process performed by the immunogenicity prediction device according to the embodiment. 11 is a scatter plot showing the relationship between systemic clearance and pI based on molecular structure. 12 is a scatter plot showing the relationship between systemic clearance and hydrophobic surface area. 13 is a scatter plot (part 2) showing the charge distribution of test samples. 14 is a scatter plot (part 2) showing the hydrophobicity distribution of test samples. 1 is a table (part 2) showing the predictive performance of predictions based on test samples. FIG. 2 is a scatter plot showing the relationship between ECM binding score and pI based on molecular structure.FIG. 1 is a scatter plot showing the relationship between ECM binding score and hydrophobic surface area. FIG. 2 is a scatter plot (part 3) showing the charge distribution of test samples. FIG. 3 is a scatter plot (part 3) showing the hydrophobicity distribution of test samples. FIG. 4 is a table (part 3) showing the predictive performance of predictions based on test samples. FIG. 5 is a diagram showing the group to which the antibodies according to the examples belong. FIG. 6 is a scatter plot showing the charge distribution of the antibodies according to the examples.

[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the description of the drawings, the same elements are designated by the same reference numerals, and duplicate explanations will be omitted. Furthermore, the embodiments of the present disclosure in the following description are specific examples of the present invention, and the present invention is not limited to these embodiments unless otherwise specified to limit the present invention.

[0011] FIG. 1 is a diagram showing an example of the functional configuration of a prediction device 1 (biological property prediction device, immunogenicity prediction device, pharmacokinetics prediction device, or non-specific binding prediction device) according to an embodiment. The prediction device 1 is a computer device that predicts the biological properties of a test substance. The biological properties of the test substance include immunogenicity, non-specific binding, and pharmacokinetics. As an example, the biological property predicted by the prediction device 1 is at least one selected from the group consisting of immunogenicity, non-specific binding, and pharmacokinetics.

[0012] Examples of test substances include polypeptides or proteins (e.g., antibodies or recombinant proteins), protein-non-protein conjugates (e.g., antibody-drug conjugates and glycosylated antibodies), complexes of two or more polypeptides or proteins, peptide-non-protein conjugates, and mixtures of two or more of these. The combinations of two or more may belong to the same category or different categories, and are not limited thereto. Examples of non-proteins include small molecule drugs, glycans, nucleic acids, and lipids. Examples of test substances include antibodies and antibody fragments having partial antibody structures, recombinant proteins, and impurities such as host-derived proteins (HCPs; host cell proteins) that may be present in formulations. The test substance may be an antibody. That is, the test substance may be a protein or an antibody. In this embodiment, antibodies are primarily considered as test substances, but this is not limited thereto. When predicting immunogenicity using the prediction device 1, the test substance is not particularly limited as long as it is a substance that can cause immunogenicity in humans.

[0013] The term "antibody" is used in the broadest sense and encompasses various antibody structures, including, but not limited to, monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), antibody-drug conjugates, antibody-protein conjugates, and antibody fragments (e.g., single-chain antibody molecules such as Fv, Fab, scFv (single-chain Fv) and VHH (variable domain of heavy chain of heavy chain), and conjugates formed by binding these), as long as they exhibit the desired antigen-binding activity.

[0014] 1, the prediction device 1 includes a storage unit 10, an acquisition unit 11, a generation unit 12, an evaluation unit 13, an identification unit 14, a learning unit 15, and a prediction unit 16. Details of each functional block will be described later.

[0015] Each functional block of the prediction device 1 is assumed to function within the prediction device 1, but is not limited to this. For example, some of the functional blocks of the prediction device 1 may function in a computer device different from the prediction device 1 and connected to the prediction device 1 through a network, while appropriately transmitting and receiving information to and from the prediction device 1. Furthermore, some functional blocks of the prediction device 1 may be omitted, multiple functional blocks may be integrated into one functional block, or one functional block may be separated into multiple functional blocks.

[0016] FIG. 2 is a diagram illustrating an example of the hardware configuration of a computer used in the prediction device 1. As shown in FIG. 2, the prediction device 1 is physically configured as a computer system including (at least one) CPU (Central Processing Unit) 100, which is a central processing unit (processor), RAM (Random Access Memory) 101 and ROM (Read Only Memory) 102, which are main storage devices, input / output devices 103 such as a keyboard, microphone, and display, a communication module 104, which is a data transmission / reception device, and an auxiliary storage device 105 such as a hard disk and SSD (Solid State Drive). The CPU 100, RAM 101, ROM 102, input / output device 103, communication module 104, and auxiliary storage device 105 may each be configured with multiple components. The functions of each functional block illustrated in FIG. 1 are realized by loading specific computer software onto hardware such as the CPU 100 and RAM 101 illustrated in FIG. 2, operating the input / output device 103 and communication module 104 under the control of the CPU 100, and reading and writing data from and to the RAM 101 and auxiliary storage device 105.

[0017] Next, a description will be given of a prediction program P1 that causes a computer to execute a series of processes by the prediction device 1. As shown in Fig. 3, the prediction program P1 is inserted into a computer and accessed, or is stored in a program storage area formed in an auxiliary storage device 105 provided in the computer. More specifically, the prediction program P1 is stored in a program storage area formed in the auxiliary storage device 105 provided in the prediction device 1.

[0018] The prediction program P1 is configured to include a storage module P10, an acquisition module P11, a generation module P12, an evaluation module P13, a specification module P14, a learning module P15, and a prediction module P16. Functions realized by executing the storage module P10, the acquisition module P11, the generation module P12, the evaluation module P13, the specification module P14, the learning module P15, and the prediction module P16 are similar to the functions of the storage unit 10, the acquisition unit 11, the generation unit 12, the evaluation unit 13, the specification unit 14, the learning unit 15, and the prediction unit 16 of the prediction device 1 described above, respectively.

[0019] The prediction program P1 is a program that causes the prediction device 1, which is a computer (one or more CPUs), to function as a storage unit 10, an acquisition unit 11, a generation unit 12, an evaluation unit 13, an identification unit 14, a learning unit 15, and a prediction unit 16.

[0020] The prediction program P1 may be configured such that a part or all of it is transmitted via a transmission medium such as a communication line, and is received and stored (including installed) by another device. Furthermore, each module of the prediction program P1 may be installed on one of multiple computers, rather than on a single computer. In this case, the series of processes of the prediction program P1 described above are performed by a computer system consisting of the multiple computers.

[0021] Hereinafter, each function of the prediction device 1 shown in FIG. 1 will be described.

[0022] The storage unit 10 stores any information that is used or output in processing, etc., of the prediction device 1. The storage unit 10 may store information calculated by each functional block of the prediction device 1. The information stored by the storage unit 10 may be referenced by each functional block of the prediction device 1 as appropriate.

[0023] The acquisition unit 11 acquires any information to be used or output in processing or the like of the prediction device 1. More specifically, the acquisition unit 11 may acquire any information from another device or the like via a network, from a user or the like via the input / output device 103, or from the storage unit 10. The acquisition unit 11 may output the acquired information to each functional block of the prediction device 1.

[0024] The generation unit 12 generates a molecular model of the test substance on molecular modeling software. For example, a user may input information about the test substance to the prediction device 1, the acquisition unit 11 may acquire the information and output it to the generation unit 12, and the generation unit 12 may generate a molecular model of the test substance on molecular modeling software based on the information. The information about the test substance that the user inputs to the prediction device 1 includes, for example, the common name and sequence of the test substance.

[0025] Examples of molecular modeling software that can be used to generate molecular models include MOE (Molecular Operating Environment) and PyMOL. In this embodiment, it is primarily assumed that MOE is used as the molecular modeling software, but the present invention is not limited to this.

[0026] The molecular model may be based on the amino acid sequence of the test substance. For example, the generation unit 12 may generate a molecular model of the test substance on molecular modeling software based on amino acid sequence information including a heavy chain "EVQLVESGGG..." and a light chain "DIQMTQSPSSL..." as information about the test substance.

[0027] The generation unit 12 can use the amino acid sequence of the protein (test substance) to generate a molecular model, but various structural data (for example, crystal structure data of similar molecules) can also be used to calibrate the generated molecular model. Furthermore, when the test substance is an antibody, the generation unit 12 can use the full-length sequence of the antibody or the Fv region (the variable region of the test substance such as an antibody) to generate a molecular model.

[0028] The generating unit 12 may store information about the generated molecular model in the storage unit 10 or may output the information to other functional blocks of the prediction device 1 .

[0029] The evaluation unit 13 evaluates the physicochemical properties of the test substance. For example, a user may input information about the test substance to the prediction device 1, the acquisition unit 11 may acquire the information and output it to the evaluation unit 13, and the evaluation unit 13 may evaluate the physicochemical properties of the test substance based on the information.

[0030] The evaluation unit 13 may evaluate the physicochemical properties of the test substance based on a molecular model of the test substance generated on molecular modeling software. The evaluation unit 13 may evaluate the physicochemical properties of the test substance based on the molecular model generated by the generation unit 12 (information related to the molecular model input from the generation unit 12).

[0031] The evaluation unit 13 may evaluate the physicochemical properties based on structural data or a structural model of the test substance. For example, structural data may include crystal structure data obtained using X-ray crystallography or cryo-electron microscopy, or crystal structure data registered in the RCSB Protein Data Bank (PDB). For example, structural models may include structural models generated using protein structure prediction software (e.g., AlphaFold).

[0032] The physicochemical properties are, for example, the physicochemical properties of the Fv region of the test substance.

[0033] The physicochemical property may be based on at least one of charge and hydrophobicity, i.e., the physicochemical property may be based on charge, or on hydrophobicity, or on charge and hydrophobicity.

[0034] The physicochemical property may further be based on the dipole moment, i.e., the physicochemical property may be based on the charge and the dipole moment, or on the hydrophobicity and the dipole moment, or on the charge, hydrophobicity and the dipole moment.

[0035] The physicochemical properties may be based on at least one of charge distribution and hydrophobicity distribution. That is, the physicochemical properties may be based on charge distribution, hydrophobicity distribution, or both charge distribution and hydrophobicity distribution. The physicochemical properties may be charge distribution properties determined by a combination of "Protein Net Charge" and "Protein Dipole Moment," as described below, or hydrophobicity distribution properties determined by a combination of "Hydrophobicity Score" and "Protein Dipole Moment." The physicochemical properties may further include an index related to protein size. Examples of protein size indexes include molecular weight and surface area. In this case, for example, when comparing proteins of different sizes, the hydrophobicity distribution properties or charge distribution properties may be normalized by molecular weight or surface area. For example, when comparing proteins other than antibodies, host cell-derived proteins (HCPs), VHH antibodies (variable domain of heavy chain of heavy chain antibodies), etc., the hydrophobicity distribution properties or charge distribution properties may be normalized using an index related to protein size.

[0036] The physicochemical property may be based on at least one of a bias in charge distribution and a bias in hydrophobicity distribution, i.e., the physicochemical property may be based on a bias in charge distribution, a bias in hydrophobicity distribution, or a bias in charge distribution and a bias in hydrophobicity distribution.

[0037] The charge distribution may be based on charge and dipole moment. The hydrophobicity distribution may be based on hydrophobicity and dipole moment.

[0038] Figure 4 is a diagram showing an example of charge distribution on the surface of an antibody. More specifically, Figure 4 shows a molecular model of an antibody (test substance) generated using molecular modeling software, and the charge distribution (positive and negative charges) on the antibody surface is shown by the molecular modeling software. As shown in Figure 4, the Fv region of the molecular model is included in the upper part of Figure 4.

[0039] Below, details will be given of a case where the predicted biological property is immunogenicity and the prediction device 1 is an immunogenicity prediction device.

[0040] The present inventors have found that highly immunogenic substances (such as antibody drugs) have a biased charge distribution in the Fv region, as will be explained in detail below with reference to Figures 5 to 8.

[0041] Figure 5 shows an example of the charge distribution of the Fv region of a low-immunogenicity antibody in which no bias toward positive charges is observed. More specifically, Figure 5 shows the charge distribution of the Fv region of ustekinumab, which has an ADA expression rate of 5.0%, in which no bias toward positive charges is observed.

[0042] Figure 6 shows an example of a highly immunogenic antibody in which a bias toward positive charges is observed in the charge distribution of the Fv region. More specifically, Figure 6 shows the charge distribution of the Fv region of Briakinumab, which has an ADA expression rate of 52.3%, and a region with a large bias toward positive charges is observed on the right side of Figure 6.

[0043] Figure 7 shows an example of the charge distribution of the Fv region of a low-immunogenic antibody in which no bias toward negative charges is observed. More specifically, Figure 7 shows the charge distribution of the Fv region of evolocumab, which has an ADA expression rate of 0.3%, and no bias toward negative charges is observed.

[0044] Figure 8 shows an example of a negative charge bias observed in the charge distribution of the Fv region of a highly immunogenic antibody. More specifically, Figure 8 shows the charge distribution of the Fv region of Blosozumab, which has an ADA expression rate of 40.0%, and a region with a large negative charge bias can be seen on the right side of Figure 8.

[0045] "Low immunogenicity" refers to, for example, an ADA expression rate below a predetermined level. "High immunogenicity" refers to, for example, an ADA expression rate equal to or higher than a predetermined level. The predetermined ADA expression rate is, for example, 30%, 25%, 20%, 15%, or 10%. In this embodiment, "low immunogenicity" refers to an ADA expression rate of less than 30%, and "high immunogenicity" refers to an ADA expression rate of 30% or higher, but is not limited thereto.

[0046] Here, since only qualitative information can be obtained from images such as those shown in Figures 5 to 8, it is difficult to formulate the relationship. The prediction device 1 quantitatively analyzes the charge distribution on the surface of the test substance using a quantitative analysis method. The inventors have also found that highly immunogenic substances have a biased hydrophobicity distribution in the Fv region, and the prediction device 1 also quantitatively analyzes the hydrophobicity distribution on the surface of the test substance using a quantitative analysis method. The prediction device 1 quantifies the charge distribution and hydrophobicity distribution on the surface of the molecule (antibody) using a molecular (antibody) model.

[0047] For example, the evaluation unit 13 evaluates the physicochemical properties using the physicochemical parameters "Protein Net Charge" (total protein charge), "Protein Dipole Moment", "Hydrophobicity Surface Area", and "van der Waals Surface Area", which are calculated by MOE (using existing technology) based on a molecular model of the test substance generated on MOE. The evaluation unit 13 may evaluate the physicochemical properties using a "Hydrophobicity Score", which is a value obtained by normalizing the "Hydrophobicity Surface Area" by the "van der Waals (VdW) Surface Area" (total surface area). The "Hydrophobicity Score" may be calculated using the following formula: "Hydrophobicity Score" = "Hydrophobicity Surface Area" / "van der Waals (VdW) Surface Area".

[0048] The combination of the above-mentioned physicochemical parameters is not particularly limited, and the characteristics of the charge distribution may be defined by a combination of "Protein Net Charge" and "Protein Dipole Moment," or the characteristics of the hydrophobicity distribution may be defined by a combination of "Hydrophobicity Score" and "Protein Dipole Moment." For example, the evaluation unit 13 evaluates the charge distribution and hydrophobicity distribution of the test substance using "Protein Net Charge," "Protein Dipole Moment," and "Hydrophobicity Score" calculated from the molecular model of the test substance generated on MOE.

[0049] The method for calculating the charge distribution and hydrophobicity distribution will be briefly explained using the images shown in Figures 9 and 10. Figure 9 shows an example in which there are few or no charge / hydrophobic patches. Note that a patch is a region where charged or hydrophobic amino acids gather to form a cluster. Figure 10 shows an example in which there are large charge / hydrophobic patches. "Protein Net Charge" indicates, for example, the total charge of the entire molecule in the molecular model. "Protein Dipole Moment" indicates, for example, the degree of polar amino acid side chain distribution. "Hydrophobicity Score" indicates, for example, the total normalized hydrophobic surface area.

[0050] In this embodiment, "Protein Net Charge" is referred to as "total protein charge" or "total charge" as appropriate. In this embodiment, "total protein charge" or "total charge" may be replaced with "Protein Net Charge" as appropriate, or vice versa.

[0051] In this embodiment, "Protein Dipole Moment" is referred to as "Protein Dipole Moment" or "Dipole Moment" as appropriate. In this embodiment, "Protein Dipole Moment" or "Dipole Moment" may be substituted with "Protein Dipole Moment" as appropriate, or vice versa.

[0052] In addition, in this embodiment, "Hydrophobicity Score" will be referred to as "Hydrophobicity Score" as appropriate. In this embodiment, "Hydrophobicity Score" may be replaced with "Hydrophobicity Score" as appropriate, and vice versa.

[0053] The evaluation unit 13 may store the evaluation results of the physicochemical properties in the storage unit 10 or output them to other functional blocks of the prediction device 1 .

[0054] The identification unit 14 identifies (estimates) elements of the test substance that contribute to variations (such as increases or decreases) in physicochemical properties (physicochemical parameters). The identification unit 14 may identify elements of the test substance that contribute to variations in physicochemical properties based on a molecular model of the test substance generated using molecular modeling software. The elements may be amino acid residues or translational modifications (such as glycosylation). For example, MOE's "patch analyzer" function can be used to automatically detect and output amino acid residues that form positively charged / negatively charged / hydrophobic clusters on the antibody surface. The identification unit 14 uses this function to automatically identify amino acid residues that contribute to variations in physicochemical properties.

[0055] The identification unit 14 may cause the storage unit 10 to store the identification results regarding the elements of the test substance, or may output the results to other functional blocks of the prediction device 1.

[0056] The learning unit 15 generates a prediction model, which is a trained model (prediction model) trained based on the known physicochemical properties and known biological properties of each of a plurality of known substances. The prediction model is used to set a threshold for prediction by the prediction unit 16, which will be described later.

[0057] When the biological property relates to immunogenicity, the learning unit 15 generates an immunogenicity prediction model, which is a trained model (prediction model) trained based on the known physicochemical properties and known immunogenicity of each of a plurality of known substances (known substances). The immunogenicity prediction model is used to set a threshold for prediction by the prediction unit 16, which will be described later.

[0058] The prediction model is an immunogenicity prediction model, which is a trained model for causing a computer to function to output information regarding the biological properties of a test substance based on the physicochemical properties of the test substance, and is composed of a neural network in which weighting coefficients have been trained using training data including information regarding the known physicochemical properties of each of a plurality of known substances and information regarding known development potential.The neural network performs calculations based on the trained weighting coefficients on the information regarding the physicochemical properties of the test substance inputted into the neural network, and outputs information regarding the biological properties of the test substance.

[0059] For example, in the case of an immunogenicity prediction model, the training data may include information on the "Protein Net Charge" of a known substance, information on the "Protein Dipole Moment" of the known substance, information on the "Hydrophobicity Score" of the known substance, and the ADA expression rate of the known substance. Note that when the prediction unit 16 described below makes a prediction based only on charge distribution, information on the "Hydrophobicity Score" may be omitted from the training data described above, and when a prediction is made based only on hydrophobic distribution, information on the "Protein Net Charge" may be omitted from the training data described above.

[0060] For example, the amino acid sequences and ADA incidence rates of antibody drugs that have been subjected to clinical trials and for which the incidence rates of ADA in humans have been reported may be used as training data. Specifically, the 110 types of antibody drugs (test samples) shown in Figures 11 to 16 may be used.

[0061] Figure 11 is a list (part 1) of 110 antibody pharmaceuticals for which ADA incidence rates have been reported. The list in Figure 11 includes the antibody names of antibody pharmaceuticals Nos. 1 to 20. Figure 12 is a list (part 2) of 110 antibody pharmaceuticals for which ADA incidence rates have been reported. The list in Figure 12 includes the antibody names of antibody pharmaceuticals Nos. 21 to 40. Figure 13 is a list (part 3) of 110 antibody pharmaceuticals for which ADA incidence rates have been reported. The list in Figure 13 includes the antibody names of antibody pharmaceuticals Nos. 41 to 60. Figure 14 is a list (part 4) of 110 antibody pharmaceuticals for which ADA incidence rates have been reported. The list in Figure 14 includes the antibody names of antibody pharmaceuticals Nos. 61 to 80. Figure 15 is a list (part 5) of 110 antibody pharmaceuticals for which ADA incidence rates have been reported. The list in Figure 15 includes the antibody names of antibody pharmaceuticals Nos. 81 to 100. Figure 16 is a list (part 6) of 110 antibody drugs for which the incidence of ADA has been reported. The list in Figure 16 includes the antibody names of antibody drugs 101 to 110.

[0062] Methods for obtaining the amino acid sequence of antibody drugs are well known to those skilled in the art, and can utilize, for example, patent information for each antibody drug or information from public databases. Furthermore, the ADA incidence rate can be calculated using labels issued by the Food and Drug Administration (FDA) or interview forms and published research information issued by the Pharmaceuticals and Medical Devices Agency (PMDA). Specifically, information from the following databases, which list approved drugs, can be used: FDA Label Database: https: / / www.accessdata.fda.gov / scripts / cder / daf / index.cfm PMDA Review Report Database: https: / / www.pmda.go.jp / PmdaSearch / iyakuSearch /

[0063] Information from individual papers may be used to calculate the ADA incidence rate for some unapproved drugs (mainly those whose development was discontinued during clinical trials).

[0064] For example, the following antibody drugs in Figure 11 (to the left of the first colon on each line) may use information from the following paper (to the right of the colon): AMG212: Hummel, HD et al. Pasotuxizumab, a BiTE((R)) immune therapy for castration-resistant prostate cancer: Phase I, dose-escalation study findings. Immunotherapy 13, 125-141, doi:10.2217 / imt-2020-0256 (2021). ATR-107: Hua, F. et al. Anti-IL21 receptormonoclonal antibody (ATR-107): Safety, pharmacokinetics, and pharmacodynamic evaluation in healthy volunteers: a phase I, first-in-human study. J ClinPharmacol 54, 14-22, doi:10.1002 / jcph.158 (2014). AMG211: Moek, KL et al. Phase I study of AMG 211 / MEDI-565 administered ascontinuous intravenous infusion (cIV) for relapsed / refractory gastrointestinal(GI) adenocarcinoma. Annals of Oncology 29, viii139-viii140,doi:10.1093 / annonc / mdy279.414 (2018). AMG211: Pishvaian, M. et al. Phase 1 Dose Escalation Study of MEDI-565, aBispecific T-Cell Engager that Targets Human Carcinoembryonic Antigen, inPatients With Advanced Gastrointestinal Adenocarcinomas.Clin Colorectal Cancer15, 345-351, doi:10.1016 / j.clcc.2016.07.009 (2016). LY3321367:Harding, J. J. et al. Blocking TIM-3 inTreatment-refractory Advanced Solid Tumors: A Phase Ia / b Study of LY3321367with or without an Anti-PD-L1 Antibody. Clin Cancer Res 27, 2168-2178,doi:10.1158 / 1078-0432.CCR-20-4405 (2021). Alemtuzumab:Li, Z., Richards, S., Surks, H. K.,Jacobs, A. & Panzara, M. A. Clinical pharmacology of alemtuzumab, ananti-CD52 immunomodulator, in multiple sclerosis. Clin Exp Immunol 194,295-314, doi:10.1111 / cei.13208 (2018). hA33 and Infliximab:Ito, S. et al. In vitro human helper T-cell assay to screen antibodydrug candidates for immunogenicity. J Immunotoxicol 16, 125-132,doi:10.1080 / 1547691X.2019.1604586 (2019). Briakinumab:Wu, J.J. et al. Briakinumab versusMethotrexate for Psoriasis. New England Journal of Medicine 366, 379-380,doi:10.1056 / NEJMc1113864 (2012). Briakinumab:Reich, K. et al. A 52-week trialcomparing briakinumab with methotrexate in patients with psoriasis.N Engl JMed 365, 1586-1596, doi:10.1056 / NEJMoa1010858 (2011). Cibisatamab:Melero, I. et al. Pharmacokinetics (PK)and Pharmacodynamics (PD) of a Novel Carcinoembryonic Antigen (CEA)T-cellBispecific Antibody (CEA-TCB) for the Treatment of CEA-Expressing Solid Tumors.ESMO 2017 poster (2017) Bococizumab:Ridker, P. M. et al. Lipid-ReductionVariability and Antidrug-Antibody Formation with Bococizumab. N Engl J Med 376,1517-1526, doi:10.1056 / NEJMoa1614062 (2017). Marstacimab:Cardinal, M. et al. A first-in-humanstudy of the safety, tolerability, pharmacokinetics and pharmacodynamics ofPF-06741086, an anti-tissue factor pathway inhibitor mAb, in healthyvolunteers. J Thromb Haemost 16, 1722-1731, doi:10.1111 / jth.14207 (2018). TYRP1-TCB:Spreafico, A. et al. Phase 1,first-in-human study of TYRP1-TCB (RO7293583), a novel TYRP1-targeting CD3T-cell engager, in metastatic melanoma: active drug monitoring to assess theimpact of immune response on drug exposure. Front Oncol 14, 1346502,doi:10.3389 / fonc.2024.1346502 (2024). Utomilumab: Segal, NH et al. Phase I Study ofSingle-Agent Utomilumab (PF-05082566), a 4-1BB / CD137 Agonist, in Patients with Advanced Cancer. Clin Cancer Res 24, 1816-1823,doi:10.1158 / 1078-0432.CCR-17-1922 (2018). Blosozumab: Recker, RR et al. A randomized, double-blind phase 2 clinical trial of blosozumab, a sclerostin antibody, inpostmenopausal women with low bone mineral density. J Bone Miner Res 30,216-224, doi:10.1002 / jbmr.2351 (2015). Atezolizumab: Wu, B. et al. Evaluation of atezolizumab immunogenicity: Clinical pharmacology (part 1). Clin Transl Sci15, 130-140, doi:10.1111 / cts.13127 (2022).

[0065] For example, the following antibody drug (to the left of the first colon) in Figure 12 may use information from the following paper (to the right of the colon): AMG110: Kebenko, M. et al. A multicenter phase 1 study of solitomab (MT110, AMG 110), a bispecific EpCAM / CD3 T-cell engager (BiTE®) antibody construct, in patients with refractory solid tumors. Oncoimmunology 7, e1450710, doi:10.1080 / 2162402X.2018.1450710 (2018).

[0066] For example, the following antibody drug (to the left of the first colon) in Figure 14 may use information from the following paper (to the right of the colon): Lebrikizumab: Simpson, EL et al. Efficacy and Safety of Lebrikizumab in Combination With Topical Corticosteroids in Adolescents and Adults With Moderate-to-Severe Atopic Dermatitis: A Randomized Clinical Trial (ADhere). JAMA Dermatol 159, 182-191, doi:10.1001 / jamadermatol.2022.5534(2023).

[0067] For example, the following antibody drug (to the left of the first colon) in Figure 16 may use information from the following paper (to the right of the colon): Gemtuzumab Ozogamicin: Montesinos, P. et al. A phase IV study evaluating QT interval, pharmacokinetics, and safety following fractionated dosing of gemtuzumab ozogamicin in patients with relapsed / refractory CD33-positive acute myeloid leukemia. Cancer ChemotherPharmacol 91, 441-446, doi:10.1007 / s00280-023-04516-9 (2023).

[0068] In this embodiment, learning was performed using the ADA expression rates of each antibody drug confirmed as of May 17, 2024 in the above database and the above paper.

[0069] An antibody drug with an ADA expression rate of more than 30% may be used as a positive control substance for immunogenicity prediction (although not limited to this). Specifically, the threshold of the immunogenicity prediction model may be calculated using the antibody drug shown in Figure 17 as a positive control substance. Figure 17 is a list of antibody drugs with an ADA expression rate of more than 30% (confirmed as of May 17, 2024). The criteria for the positive control substance used to set the threshold can be set appropriately depending on the purpose.

[0070] When a test sample is used as learning data, information regarding the "Protein Net Charge" of the test sample, information regarding the "Protein Dipole moment" of the test sample, and information regarding the "Hydrophobicity Score" of the test sample may be calculated using the functions of the generation unit 12 and evaluation unit 13 described above.

[0071] The following describes the verification results of the immunogenicity prediction model generated by the learning unit 15 based on the test sample. FIG. 18 is a scatter diagram (part 1) showing the charge distribution of the test sample. In FIG. 18, the unit of "dipole moment" on the horizontal axis is D (Debye), and the "total charge in the Fv region" on the vertical axis is unitless. FIG. 19 is a scatter diagram (part 1) showing the hydrophobicity distribution of the test sample. In FIG. 19, the unit of "dipole moment" on the horizontal axis is D (Debye), and the "hydrophobicity score" on the vertical axis is unitless. FIG. 20 is a table (part 1) showing the prediction performance of prediction based on the test sample. The table shown in FIG. 20 verifies the prediction performance for test samples with an ADA incidence rate of 30% or more. In the table shown in Figure 20, the accuracy rate indicates the overall hit / miss ratio, the specificity indicates whether "negatives" are overly suspected, the recall rate indicates whether "positives" are overlooked, and the precision rate indicates how accurately "positives" can be determined. Figure 21 shows the ADA occurrence rate according to the bias in charge distribution and / or the bias in hydrophobicity distribution. From the above, it can be seen that the ADA occurrence rate correlates with the bias in charge distribution and / or the bias in hydrophobicity distribution on the antibody surface. It can also be seen that the prediction method using charge distribution and hydrophobicity distribution by the prediction device 1 can accurately predict the immunogenicity of antibody pharmaceuticals.

[0072] 22 is a flowchart showing an example of the learning process executed by the prediction device 1. First, the generation unit 12 generates a molecular model of a known substance (step S1). Next, the evaluation unit 13 evaluates the physicochemical properties of the known substance based on the molecular model generated in step S1 (step S2). Next, the learning unit 15 generates training data based on the physicochemical properties of the known substance evaluated in step S2 and the known immunogenicity of the known substance (step S3). Steps S1 to S3 are repeated for each of the multiple known substances. Once training data has been generated for each of the multiple known substances, the learning unit 15 then performs training based on the training data to generate a trained model (immunogenicity prediction model) (step S4, learning step).

[0073] The learning unit 15 may store the generated immunogenicity prediction model in the storage unit 10 or output it to another functional block of the prediction device 1.

[0074] The prediction unit 16 predicts the immunogenicity of the test substance based on the physicochemical properties of the test substance. For example, a user may input information regarding the physicochemical properties of the test substance into the prediction device 1, the acquisition unit 11 may acquire the information and output it to the prediction unit 16, and the prediction unit 16 may predict the immunogenicity of the test substance based on the information. Furthermore, for example, the prediction unit 16 may predict the immunogenicity of the test substance based on the evaluation results of the physicochemical properties input from the evaluation unit 13. The prediction unit 16 may determine that the immunogenicity of the test substance is higher or lower than a predetermined standard. For example, the prediction unit 16 may determine that the immunogenicity of the test substance is an ADA expression rate of 30% or more, or an ADA expression rate of less than 30%.

[0075] For example, when setting the thresholds described above, parameter thresholds corresponding to 30%, 20%, and 10% are set, and learning is performed for each antibody drug, so that the prediction unit 16 can make a prediction such as "Antibody drug A has an ADA expression rate of 20% to 30%."

[0076] The "standard" in this embodiment may be a standard set based on an immunogenicity prediction model. The "standard" in this embodiment may not be an absolute numerical value such as "439 (Top 27.5%)", "3.21 (Top 25%)", and "-3.04 (Bottom 7.5%)" shown in FIG. 18, or "313 (Top 57.55%)" and "0.601 (Top 5%)" shown in FIG. 19, but may be a numerical value that can be changed when there is a change in the learning data.

[0077] The prediction unit 16 may predict that the ADA expression rate of the test substance is lower than a predetermined ADA expression rate when the total charge in the Fv region of the test substance is within a predetermined first criterion and the dipole moment of the test substance is within a predetermined second criterion. For example, based on the scatter diagram shown in Figure 18, the prediction unit 16 may predict that the ADA expression rate of the test substance is lower than "30%" when the total charge in the Fv region of the test substance is "3.21" or more and the dipole moment of the test substance is "439" or less. Also, based on the scatter diagram shown in Figure 18, the prediction unit 16 may predict that the ADA expression rate of the test substance is lower than "30%" when the total charge in the Fv region of the test substance is "-3.04" or more and "3.21" or less and the dipole moment of the test substance is "439" or less. Furthermore, for example, based on the scatter diagram shown in Figure 18, the prediction unit 16 may predict that the ADA expression rate of the test substance is lower than "30%" when the total charge in the Fv region of the test substance is "-3.04" or more and "3.21" or less, and the dipole moment of the test substance is "439" or more. Furthermore, for example, based on the scatter diagram shown in Figure 18, the prediction unit 16 may predict that the ADA expression rate of the test substance is lower than "30%" when the total charge in the Fv region of the test substance is "-3.04" or less, and the dipole moment of the test substance is "439" or less.

[0078] The prediction unit 16 may predict that the ADA expression rate of the test substance is lower than a predetermined ADA expression rate when the hydrophobicity degree (hydrophobicity score) of the test substance is within a predetermined first standard and the dipole moment of the test substance is within a predetermined second standard. The hydrophobicity degree of the test substance may be a value obtained by normalizing the hydrophobic surface area by the total surface area. For example, based on the scatter diagram shown in FIG. 19, the prediction unit 16 may predict that the ADA expression rate of the test substance is lower than "30%" when the hydrophobicity score of the test substance is "0.601" or higher and the dipole moment of the test substance is "313" or lower. Furthermore, based on the scatter diagram shown in FIG. 19, the prediction unit 16 may predict that the ADA expression rate of the test substance is lower than "30%" when the hydrophobicity score of the test substance is "0.601" or lower and the dipole moment of the test substance is "313" or lower. Furthermore, for example, based on the scatter plot shown in Figure 19, the prediction unit 16 may predict that the ADA expression rate of the test substance is lower than "30%" when the hydrophobicity score of the test substance is "0.601" or less and the dipole moment of the test substance is "313" or more.

[0079] The prediction unit 16 may predict that the ADA expression rate of the test substance is lower than a predetermined ADA expression rate when the charge distribution bias or hydrophobicity distribution bias is within a predetermined standard. For example, based on the diagram shown in FIG. 21, the prediction unit 16 may predict that the ADA expression rate of the test substance is lower than "30%" when the charge distribution bias is equal to or less than a predetermined first standard value and the hydrophobicity distribution bias is equal to or less than a predetermined second standard value. For example, the prediction unit 16 may predict that the ADA expression rate of the test substance is lower than "30%" when the charge distribution bias is equal to or less than a predetermined standard value. For example, the prediction unit 16 may predict that the ADA expression rate of the test substance is lower than "30%" when the hydrophobicity distribution bias is equal to or less than a predetermined standard value.

[0080] The prediction unit 16 may output values ​​of ADAs that have a particular clinical impact (specifically, ADAs that affect pharmacokinetics (PK), drug efficacy, and toxicity, and neutralizing antibodies).

[0081] The prediction unit 16 may transmit (output) the prediction result or the evaluation result to another device via a network, or may display (output) it to the user via the input / output device 103 .

[0082] The prediction unit 16 may predict the biological properties of each of a plurality of mutually different test substances, and output information about at least one test substance for which the prediction result satisfies a predetermined standard (e.g., the ADA expression rate is equal to or greater than a predetermined standard value). The output information about the test substance may include, for example, the general name and sequence of the test substance. The output may be performed by an output unit (not shown) included in the prediction device 1. For example, the prediction unit 16 may predict the biological properties of each of a plurality of mutually different test substances, and the output unit may output information about at least one test substance for which the prediction result by the prediction unit 16 satisfies a predetermined standard.

[0083] For example, by outputting information about test substances whose ADA expression rate is equal to or higher than a predetermined standard value, test substances predicted to be highly immunogenic can be excluded from drug candidates. Information about the test substances may also be output according to the prediction results. For example, information about the test substances may be output in ascending or descending order of ADA expression rate.

[0084] The prediction unit 16 may make a prediction further based on the known physicochemical properties and known immunogenicity of each of the multiple known substances. The prediction unit 16 may make a prediction by applying the physicochemical properties of the test substance to a trained model trained based on the known physicochemical properties and known immunogenicity of each of the multiple known substances. For example, the prediction unit 16 may predict the immunogenicity of the test substance (output information about immunogenicity) by inputting information about the physicochemical properties of the test substance into an immunogenicity prediction model stored by the storage unit 10. The data that the prediction unit 16 inputs into the immunogenicity prediction model during prediction may include information about the "Protein Net Charge" of the test substance, information about the "Protein Dipole Moment" of the test substance, and information about the "Hydrophobicity Score" of the test substance. When the prediction unit 16 makes a prediction based only on the charge distribution, information regarding the "Hydrophobicity Score" may be omitted from the above data, and when the prediction unit 16 makes a prediction based only on the hydrophobicity distribution, information regarding the "Protein Net Charge" may be omitted from the above data.

[0085] The prediction unit 16 may make a prediction based on the physicochemical properties evaluated by the evaluation unit 13 (the evaluation results input from the evaluation unit 13). The prediction unit 16 may predict the immunogenicity of a test substance in which the element identified by the identification unit 14 (the element indicated by the identification results input from the identification unit 14) is substituted (with any element).

[0086] 23 is a flowchart showing an example of the prediction process (immunogenicity prediction method) performed by the prediction device 1. First, the generation unit 12 generates a molecular model of the test substance (step S10, generation step). Next, the evaluation unit 13 evaluates the physicochemical properties of the test substance based on the molecular model generated in step S10 (step S11, evaluation step). Next, the identification unit 14 identifies factors that contribute to the variation in physicochemical properties (step S12, identification step), and (if any) performs S10 and S11 again for the test substance in which the identified factors have been replaced. Next, the immunogenicity of the test substance and the replaced test substance (if any) is predicted based on the physicochemical properties evaluated in step S11 (step S13, prediction step). Note that any one or more of steps S10, S11, and S12 may be omitted.

[0087] The prediction device 1 may also predict PK or non-specific binding. When the predicted biological property is pharmacokinetics, the prediction device 1 may also be referred to as a pharmacokinetic prediction device. When the predicted biological property is non-specific binding, the prediction device 1 may also be referred to as a non-specific binding prediction device. In either case, the processing other than that of the prediction unit is common to that of the immunogenicity prediction device described above, and therefore will not be described below.

[0088] The prediction unit 16 may predict antibodies with poor PK (or PK parameters) by quantifying the charge / hydrophobicity distribution of the antibody. Specifically, the prediction unit 16 may predict the pharmacokinetics of the test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance and the dipole moment of the test substance. The physicochemical properties may be based on at least one of the charge distribution or the hydrophobicity distribution. The physicochemical properties may be based on at least one of the bias in the charge distribution or the bias in the hydrophobicity distribution. The charge distribution may be based on the charge and the dipole moment, and the hydrophobicity distribution may be based on the hydrophobicity and the dipole moment.

[0089] Generally, the parameters for PK prediction are systemic clearance (systemic CL) or half-life (t 1/2) is used. Systemic clearance and half-life are correlated parameters. An "antibody with poor PK" can be rephrased as an "antibody with high systemic clearance." For example, an antibody with a CL of 0.5 L / day or more can generally be considered to have high systemic clearance. Therefore, a verification was performed using a dataset of antibodies for which systemic clearance information is available, and antibodies with a CL above a certain standard were selected as the prediction targets. As a result, the prediction unit 16 was able to predict antibodies with poor PK (or PK parameters) by quantifying the charge / hydrophobicity distribution of the antibody. When predicting PK for conventional IgG, ADCs (Antibody Drug Conjugates) whose payload results in higher systemic clearance than conventional antibodies, and scFv and Fab antibodies whose systemic clearance is higher than conventional antibodies due to the absence of Fc, may be excluded from the dataset. This results in higher prediction accuracy. Furthermore, when PK prediction is performed for each of the ADC, scFv, and Fab antibody, the data set may consist of only the data for each of the ADC, scFv, and Fab antibody.

[0090] The following describes the prediction of antibodies with high systemic clearance using a single parameter related to charge or hydrophobicity in conventional technology. FIG. 24 is a scatter plot showing the relationship between systemic clearance and pI based on molecular structure. pI (isoelectric point) is the isoelectric point. The scatter plot shown in FIG. 24 shows that it is difficult to accurately predict antibodies with high systemic clearance using only a single parameter, pI. FIG. 25 is a scatter plot showing the relationship between systemic clearance and hydrophobic surface area. The scatter plot shown in FIG. 25 shows that it is difficult to accurately predict antibodies with high systemic clearance using only a single parameter, hydrophobic surface area. As described above, it is not possible to accurately predict antibodies with high systemic clearance using a single parameter related to charge or hydrophobicity as in conventional technology.

[0091] The following describes the verification results of the pharmacokinetic prediction model generated by the learning unit 15 based on the test sample. Figure 26 is a scatter plot (part 2) showing the charge distribution of the test sample. In Figure 26, the unit of "protein dipole moment" on the horizontal axis is D (Debye), and the unit of "total protein charge" on the vertical axis is unitless. Figure 27 is a scatter plot (part 2) showing the hydrophobicity distribution of the test sample. In Figure 27, the unit of "protein dipole moment" on the horizontal axis is D (Debye), and the unit of "hydrophobicity score" on the vertical axis is unitless. Figure 28 is a table (part 2) showing the prediction performance of predictions based on the test sample. The table shown in Figure 28 verifies the prediction performance for test samples with a systemic clearance of "0.5" or higher. From the above, it can be seen that the prediction device 1 can accurately predict antibodies with high (poor) systemic clearance using the charge / hydrophobicity distribution.

[0092] For example, the prediction unit 16 may predict that the systemic clearance of the test substance will be greater than a predetermined systemic clearance when the overall protein charge of the test substance is within a predetermined first criterion and the protein dipole moment of the test substance is within a predetermined second criterion. For example, based on the scatter plot shown in Figure 26, the prediction unit 16 may predict that the systemic clearance of the test substance will be greater than "0.5" when the overall protein charge of the test substance is "3" or greater and the protein dipole moment of the test substance is "440" or greater.

[0093] For example, the prediction unit 16 may predict that the systemic clearance of the test substance will be greater than a predetermined systemic clearance when the hydrophobicity score of the test substance is within a predetermined first criterion and the protein dipole moment of the test substance is within a predetermined second criterion. For example, based on the scatter plot shown in Figure 27, the prediction unit 16 may predict that the systemic clearance of the test substance will be greater than "0.5" when the hydrophobicity score of the test substance is "0.575" or greater and the protein dipole moment of the test substance is "450" ​​or greater.

[0094] The prediction unit 16 may predict antibodies with high non-specific binding properties by quantifying the charge / hydrophobicity distribution of the antibody. Specifically, the prediction unit 16 may predict the non-specific binding properties of the test substance based on physicochemical properties based on the charge distribution of the test substance. The physicochemical properties may be based on the bias of the charge distribution. The charge distribution may be based on the charge and dipole moment.

[0095] As an evaluation index for nonspecific binding, adsorption assessment to columns (e.g., hydrophobic, hydrophilic, ion exchange, size exclusion columns), binding assays to cells (e.g., LSEC: liver sinusoidal endothelial cells), and extracellular matrix (ECM) binding assays are used. An "antibody with high nonspecific binding" can be rephrased as an "antibody with high ECM binding." Therefore, verification was performed using a dataset of antibodies for which ECM binding data had already been acquired. As described above, the prediction unit 16 was able to predict antibodies with high nonspecific binding by quantifying the charge / hydrophobicity distribution of the antibody. Note that antibodies with an ECM binding score of a certain value or higher may be used as the prediction target. As an example, antibodies with an ECM binding score of 5 or higher (normalized to the negative control tocilizumab), which are known to affect PK, were used as the prediction target.

[0096] The following describes the prediction of antibodies with high non-specific binding using a single parameter related to charge or hydrophobicity in conventional technology. FIG. 29 is a scatter plot showing the relationship between ECM binding score and pI based on molecular structure. The scatter plot shown in FIG. 29 shows that it is difficult to accurately predict antibodies with high ECM binding score (i.e., non-specific binding) using only a single parameter, pI. FIG. 30 is a scatter plot showing the relationship between ECM binding score and hydrophobic surface area. The scatter plot shown in FIG. 30 shows that it is difficult to accurately predict antibodies with high ECM binding score (i.e., non-specific binding) using only a single parameter, hydrophobic surface area. As described above, it is not possible to accurately predict antibodies with high non-specific binding (ECM binding) using a single parameter related to charge or hydrophobicity, as in conventional technology.

[0097] The following describes the verification results of the nonspecific binding prediction model generated by the learning unit 15 based on the test sample. FIG. 31 is a scatter plot (part 3) showing the charge distribution of the test sample. In FIG. 31, the unit of "protein dipole moment" on the horizontal axis is D (Debye), and the unit of "total protein charge" on the vertical axis is unitless. FIG. 32 is a scatter plot (part 3) showing the hydrophobicity distribution of the test sample. In FIG. 32, the unit of "protein dipole moment" on the horizontal axis is D (Debye), and the unit of "hydrophobicity score" on the vertical axis is unitless. FIG. 33 is a table (part 3) showing the prediction performance of predictions based on the test sample. The table shown in FIG. 33 verifies the prediction performance for test samples with an ECM binding score of "5.0" or higher. From the above, it can be seen that the prediction device 1 can accurately predict antibodies with high nonspecific binding (ECM binding) using charge distribution.

[0098] For example, the prediction unit 16 may predict that the ECM binding score of the test substance will be higher than a predetermined ECM binding score if the overall protein charge of the test substance is within a predetermined first criterion and the protein dipole moment of the test substance is within a predetermined second criterion. For example, based on the scatter plot shown in Figure 31, the prediction unit 16 may predict that the ECM binding score of the test substance will be higher than "5.0" if the overall protein charge of the test substance is "3.1" or higher and the protein dipole moment of the test substance is "350" or higher.

[0099] As described above, by quantifying the charge / hydrophobicity distribution of an antibody using the prediction device 1, it is possible to predict with high accuracy antibodies with poor pharmacokinetics (PK) and antibodies with high non-specific binding. For example, the prediction device 1 can make highly accurate predictions for antibodies with high systemic clearance, which is one of the PK parameters, by using both the charge distribution and the hydrophobicity distribution. Furthermore, for example, the prediction device 1 can make highly accurate predictions for antibodies with high ECM binding, which is one of the indicators of non-specific binding, by using the charge distribution. It is difficult to predict with high accuracy both systemic clearance and non-specific binding using the parameters used in conventional technology.

[0100] [Example] An example relating to the prediction of the immunogenicity of modified antibodies using the prediction device 1 will be described. The immunogenicity of multiple different modified antibodies with the same binding target was predicted using the indices "Protein Net Charge," "Protein Dipole Moment," and "Hydrophobicity Score." Specifically, prediction was performed using the amino acid sequences of the anti-C1s antibody group shown in the table in Figure 34. Figure 34 is a diagram showing the group to which the antibodies according to the example belong. Note that the amino acid sequences of the anti-C1s antibody group for which prediction was performed were determined by reference to the description in WO 2021 / 075479.

[0101] The prediction results using "Protein Net Charge" and "Protein Dipole Moment" are shown in Figure 35. Figure 35 is a scatter plot showing the charge distribution of an antibody according to an example.

[0102] As shown in Figure 35, even within a series of modified antibodies with the same binding target, the "Protein Net Charge" and "Protein Dipole Moment" used as indicators showed different values. As shown in Figure 35, in the predictions based on "Protein Net Charge" and "Protein Dipole Moment," AH0813-SG1, COS0637pHv1-SG1, and COS0637pHv2-SG1 belonging to the Prototype group all belonged to the high immunogenicity risk region, while COS637pHv4-TT91R, COS637pHv5-TT91R, COS637pHv6-TT91R, COS637pHv7-TT91R, COS637pHv8-TT91R, and COS637pHv9-TT91R belonging to the Improved group all belonged to the low immunogenicity risk region. Therefore, for these anti-C1s antibody groups, AH0813-SG1, COS0637pHv1-SG1, and COS0637pHv2-SG1 belonging to the Prototype group have a high immunogenicity risk, and COS637pHv4-TT91R, COS637pHv5-TT91R, COS637pHv6-TT91R, COS637pHv7-TT91R, COS637pHv8-TT91R, and COS637pHv9-TT91R belonging to the Improved group were determined to have a low immunogenicity risk. This result was consistent with the prediction results using the percentage of IL-2-secreting CD4 + T cells described in WO 2021 / 075479 (prediction method described in WO 2018 / 124005) as an index. Based on these results, predictions using "Protein Net Charge" and "Protein Dipole Moment" as indicators can be made to select drug candidate substances with low immunogenicity based solely on antibody amino acid sequence information. Therefore, by using such a prediction system, screening based on immunogenicity in the optimization process of antibody drug candidates can be carried out at lower cost and in a shorter time span than conventional methods.

[0103] In this example, the prediction results obtained by the method of the prediction device 1 were consistent with the results of the T cell assay, a conventional standard immunogenicity prediction method (confirmation of equivalence), demonstrating the usefulness of this method, which can achieve lower costs and higher throughput than conventional methods in screening drug candidate substances.

[0104] Next, the effects of the prediction device 1 that predicts biological characteristics will be described.

[0105] The prediction device 1 includes a prediction unit 16 that predicts the biological properties of a test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance. The biological properties may be at least one property selected from the group consisting of immunogenicity, pharmacokinetics, and nonspecific binding of the test substance. In this prediction device 1, the biological properties of the test substance are predicted based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance. This makes it easier to predict the biological properties of the test substance. Furthermore, because physicochemical properties based on at least one of the charge or hydrophobicity are correlated with the biological properties, the biological properties can be predicted with high accuracy.

[0106] The prediction device 1 includes a prediction unit 16 that predicts the immunogenicity of a test substance based on physicochemical properties based on at least one of the charge and hydrophobicity of the test substance. In this prediction device 1, the immunogenicity of the test substance is predicted based on physicochemical properties based on at least one of the charge and hydrophobicity of the test substance. This makes it possible to more easily predict the immunogenicity of the test substance. Furthermore, because there is a correlation between physicochemical properties based on at least one of the charge and hydrophobicity and immunogenicity, immunogenicity can be predicted with high accuracy.

[0107] The physicochemical properties may be further based on dipole moment. In this aspect, the immunogenicity of the test substance is predicted based on the physicochemical properties further based on dipole moment. Since the physicochemical properties further based on dipole moment and immunogenicity are correlated, immunogenicity can be predicted with high accuracy by this aspect.

[0108] The physicochemical properties may be based on at least one of charge distribution or hydrophobicity distribution. In this aspect, the immunogenicity of a test substance is predicted based on physicochemical properties based on at least one of charge distribution or hydrophobicity distribution. Since there is a correlation between physicochemical properties based on at least one of charge distribution or hydrophobicity distribution and immunogenicity, this aspect allows for accurate prediction of immunogenicity.

[0109] The physicochemical properties may be based on at least one of a bias in charge distribution or a bias in hydrophobic distribution. In this aspect, the immunogenicity of a test substance is predicted based on physicochemical properties based on at least one of a bias in charge distribution or a bias in hydrophobic distribution. Since there is a correlation between physicochemical properties based on at least one of a bias in charge distribution or a bias in hydrophobic distribution and immunogenicity, this aspect allows for accurate prediction of immunogenicity.

[0110] The charge distribution may be based on charge and dipole moment, and the hydrophobicity distribution may be based on hydrophobicity and dipole moment. In this aspect, the charge distribution and the hydrophobicity distribution can be obtained more easily and reliably.

[0111] The test substance may be a protein or an antibody. In this aspect, the immunogenicity of the protein or antibody can be predicted.

[0112] The test substance may be an antibody, and the physicochemical properties may be those of the Fv region of the test substance. In this aspect, the immunogenicity of an antibody is predicted based on the physicochemical properties of the Fv region of the antibody. Since the physicochemical properties of the Fv region and immunogenicity are correlated, this aspect allows for accurate prediction of immunogenicity.

[0113] The prediction unit 16 may determine whether the immunogenicity of the test substance is higher or lower than a predetermined standard. In this aspect, since the immunogenicity of the test substance is determined to be higher or lower than a predetermined standard, it is possible to easily determine, for example, test substances whose immunogenicity is higher or lower than the predetermined standard.

[0114] The physicochemical property may be based on at least one of a bias in charge distribution or a bias in hydrophobicity distribution, and the prediction unit 16 may predict that the ADA expression rate of the test substance is lower than a predetermined ADA expression rate when the bias in charge distribution or the bias in hydrophobicity distribution is within a predetermined standard. In this aspect, the ADA expression rate of the test substance can be predicted more reliably and easily.

[0115] The prediction unit 16 may predict that the ADA expression rate of the test substance is lower than a predetermined ADA expression rate when the overall charge in the Fv region of the test substance is within a predetermined first criterion and the dipole moment of the test substance is within a predetermined second criterion. In this aspect, the ADA expression rate of the test substance can be predicted more reliably and easily.

[0116] The prediction unit 16 may predict that the ADA expression rate of the test substance is lower than a predetermined ADA expression rate when the degree of hydrophobicity of the test substance is within a predetermined first criterion and the dipole moment of the test substance is within a predetermined second criterion. In this aspect, the ADA expression rate of the test substance can be predicted more reliably and easily.

[0117] The degree of hydrophobicity of a test substance may be calculated by normalizing the hydrophobic surface area to the total surface area, which allows the degree of hydrophobicity to be calculated more reliably and easily.

[0118] The prediction unit 16 may predict the immunogenicity of each of a plurality of mutually different test substances and output information regarding at least one test substance whose prediction result satisfies a predetermined criterion, or an output unit (not shown) may output information regarding at least one test substance whose prediction result satisfies a predetermined criterion. In such an aspect, information regarding at least one test substance whose predicted immunogenicity result satisfies a predetermined criterion among a plurality of mutually different test substances is output, making it possible to easily identify, for example, test substances whose predicted immunogenicity result satisfies a predetermined criterion.

[0119] The prediction unit 16 may make the prediction further based on the known physicochemical properties and known immunogenicity of each of the plurality of known substances. In this aspect, the prediction is made further based on the known physicochemical properties and known immunogenicity of each of the plurality of known substances, so that the immunogenicity can be predicted with higher accuracy.

[0120] The prediction unit 16 may predict the physicochemical properties of the test substance by applying them to a trained model (immunogenicity prediction model) that has been trained based on the known physicochemical properties and known immunogenicity of each of a plurality of known substances. In this aspect, since the prediction is made by applying them to a trained model that has been trained based on the known physicochemical properties and known immunogenicity of each of a plurality of known substances, immunogenicity can be predicted with higher accuracy.

[0121] The apparatus may further include an evaluation unit 13 that evaluates the physicochemical properties of the test substance, and the prediction unit 16 may make predictions based on the physicochemical properties evaluated by the evaluation unit 13. In such an aspect, the physicochemical properties of the test substance are evaluated and predictions are made based on the evaluated physicochemical properties, so that the physicochemical properties can be evaluated more accurately, and therefore immunogenicity can be predicted more accurately.

[0122] The evaluation unit 13 may perform the evaluation based on a molecular model of the test substance generated on molecular modeling software. In this aspect, the physicochemical properties are evaluated based on the molecular model of the test substance generated on molecular modeling software, so that the physicochemical properties can be evaluated more accurately, and therefore immunogenicity can be predicted more accurately.

[0123] The system may further include a generation unit 12 that generates a molecular model of the test substance on molecular modeling software, and the evaluation unit 13 may perform evaluation based on the molecular model generated by the generation unit 12. In this aspect, since the physicochemical properties are evaluated based on the molecular model of the test substance generated on molecular modeling software, the physicochemical properties can be evaluated more accurately, and therefore immunogenicity can be predicted more accurately.

[0124] The molecular model may be based on the amino acid sequence of the test substance. In such an aspect, immunogenicity is predicted based on a molecular model based on the amino acid sequence of the test substance. This makes it easier to predict the immunogenicity of a test substance having an amino acid sequence.

[0125] The system may further include an identification unit 14 that identifies elements of the test substance that contribute to the variation in physicochemical properties, and the prediction unit 16 may predict the immunogenicity of a test substance that has been substituted for the elements identified by the identification unit 14. In this aspect, elements of the test substance that contribute to the variation in physicochemical properties are identified, and the immunogenicity of a test substance that has been substituted for the identified elements is predicted. This makes it possible to easily determine, for example, substances with different immunogenicity.

[0126] The test substance may be a protein or an antibody, and the element may be an amino acid residue. In this aspect, the amino acid residues of the protein or antibody that contribute to the variation in physicochemical properties are identified, and the immunogenicity of proteins or antibodies in which the identified amino acid residues are substituted is predicted. This makes it possible to easily determine, for example, proteins or antibodies with different immunogenicity.

[0127] The prediction device 1 includes a prediction unit 16 that predicts the pharmacokinetics of a test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance and the dipole moment of the test substance. In this aspect, the pharmacokinetics of the test substance is predicted based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance and the dipole moment of the test substance. This makes it easier to predict the pharmacokinetics of the test substance. Furthermore, because there is a correlation between physicochemical properties based on at least one of the charge or hydrophobicity and the dipole moment and the pharmacokinetics, the pharmacokinetics can be predicted with high accuracy.

[0128] The physicochemical properties may be based on at least one of charge distribution or hydrophobicity distribution. In this aspect, the pharmacokinetics of the test substance is predicted based on the physicochemical properties based on at least one of charge distribution or hydrophobicity distribution. Since the physicochemical properties based on at least one of charge distribution or hydrophobicity distribution and the pharmacokinetics are correlated, this aspect allows for accurate prediction of the pharmacokinetics.

[0129] The physicochemical properties may be based on at least one of a bias in charge distribution or a bias in hydrophobicity distribution. In this aspect, the pharmacokinetics of the test substance is predicted based on the physicochemical properties based on at least one of a bias in charge distribution or a bias in hydrophobicity distribution. Since the physicochemical properties based on at least one of a bias in charge distribution or a bias in hydrophobicity distribution and the pharmacokinetics are correlated, this aspect allows for accurate prediction of the pharmacokinetics.

[0130] The charge distribution may be based on charge and dipole moment, and the hydrophobicity distribution may be based on hydrophobicity and dipole moment. In this aspect, the charge distribution and the hydrophobicity distribution can be obtained more easily and reliably.

[0131] The test substance may be a protein or an antibody. In this aspect, the pharmacokinetics of the protein or antibody can be predicted.

[0132] The test substance may be an antibody, and the physicochemical properties may be those of the Fv region of the test substance. In this aspect, the pharmacokinetics of the antibody is predicted based on the physicochemical properties of the Fv region of the antibody. Because the physicochemical properties of the Fv region correlate with the pharmacokinetics, this aspect allows for accurate prediction of the pharmacokinetics.

[0133] The system may further include an evaluation unit 13 that evaluates the physicochemical properties of the test substance, and the prediction unit 16 may make a prediction based on the physicochemical properties evaluated by the evaluation unit 13. In this aspect, the physicochemical properties of the test substance are evaluated and prediction is made based on the evaluated physicochemical properties, so that the physicochemical properties can be evaluated with higher accuracy, and therefore the pharmacokinetics can be predicted with higher accuracy.

[0134] The evaluation unit 13 may perform the evaluation based on a molecular model of the test substance generated on molecular modeling software. In this aspect, the physicochemical properties are evaluated based on the molecular model of the test substance generated on molecular modeling software, so that the physicochemical properties can be evaluated with higher accuracy, and therefore the pharmacokinetics can be predicted with higher accuracy.

[0135] The system may further include a generation unit 12 that generates a molecular model of the test substance on molecular modeling software, and the evaluation unit 13 may perform evaluation based on the molecular model generated by the generation unit 12. In this aspect, since the physicochemical properties are evaluated based on the molecular model of the test substance generated on molecular modeling software, the physicochemical properties can be evaluated with higher accuracy, and therefore the pharmacokinetics can be predicted with higher accuracy.

[0136] The molecular model may be based on the amino acid sequence of the test substance. In such an aspect, the pharmacokinetics is predicted based on a molecular model based on the amino acid sequence of the test substance. This makes it easier to predict the pharmacokinetics of a test substance having an amino acid sequence.

[0137] The prediction device 1 includes a prediction unit 16 that predicts the nonspecific binding of a test substance based on physicochemical properties based on the charge distribution of the test substance. In this aspect, the nonspecific binding of a test substance is predicted based on the physicochemical properties based on the charge distribution of the test substance. This makes it possible to more easily predict the nonspecific binding of a test substance. Furthermore, because the physicochemical properties based on the charge distribution and the nonspecific binding are correlated, the nonspecific binding can be predicted with high accuracy.

[0138] The physicochemical property may be based on a bias in charge distribution. In this aspect, the nonspecific binding of a test substance is predicted based on the physicochemical property based on the bias in charge distribution. Since the physicochemical property based on the bias in charge distribution and the nonspecific binding are correlated, this aspect allows for accurate prediction of the nonspecific binding.

[0139] The charge distribution may be based on the charge and dipole moment. In this aspect, the charge distribution can be obtained more easily and reliably.

[0140] The test substance may be a protein or an antibody. In this aspect, the non-specific binding property of the protein or antibody can be predicted.

[0141] The test substance may be an antibody, and the physicochemical properties may be those of the Fv region of the test substance. In this aspect, the non-specific binding of an antibody is predicted based on the physicochemical properties of the Fv region of the antibody. Because the physicochemical properties of the Fv region correlate with the non-specific binding, this aspect allows for accurate prediction of the non-specific binding.

[0142] The system may further include an evaluation unit 13 that evaluates the physicochemical properties of the test substance, and the prediction unit 16 may make a prediction based on the physicochemical properties evaluated by the evaluation unit 13. In this aspect, the physicochemical properties of the test substance are evaluated and prediction is made based on the evaluated physicochemical properties, so that the physicochemical properties can be evaluated with higher accuracy, and therefore the nonspecific binding property can be predicted with higher accuracy.

[0143] The evaluation unit 13 may perform the evaluation based on a molecular model of the test substance generated on molecular modeling software. In this aspect, the physicochemical properties are evaluated based on the molecular model of the test substance generated on molecular modeling software, so that the physicochemical properties can be evaluated with higher accuracy, and therefore the nonspecific binding property can be predicted with higher accuracy.

[0144] The system may further include a generation unit 12 that generates a molecular model of the test substance on molecular modeling software, and the evaluation unit 13 may perform evaluation based on the molecular model generated by the generation unit 12. In this aspect, since the physicochemical properties are evaluated based on the molecular model of the test substance generated on molecular modeling software, the physicochemical properties can be evaluated with higher accuracy, and therefore the nonspecific binding properties can be predicted with higher accuracy.

[0145] The molecular model may be based on the amino acid sequence of the test substance. In such an aspect, the non-specific binding property is predicted based on a molecular model based on the amino acid sequence of the test substance. This makes it easier to predict the non-specific binding property of a test substance having an amino acid sequence.

[0146] The immunogenicity prediction method executed by the prediction device 1 relates to a method for predicting immunogenicity based on various physicochemical parameters on a molecular model of a protein (e.g., an antibody). More specifically, the present invention is a method for generating a molecular model based on the amino acid sequence of a protein and predicting immunogenicity using physical property parameters calculated based on the generated model. Furthermore, the present invention can identify amino acid residues that contribute to parameter fluctuations and predict whether the immunogenicity of a test substance will be increased or decreased based on parameter fluctuations due to amino acid residue substitution.

[0147] The immunogenicity prediction method executed by the prediction device 1 predicts the immunogenicity of a test substance by combining physicochemical parameters on a molecular model. This immunogenicity prediction method predicts the immunogenicity of a test substance by utilizing characteristics on charge distribution and characteristics on hydrophobicity distribution. Furthermore, it is possible to identify amino acid residues that contribute to an increase or decrease in parameters, and predict an increase or decrease in the immunogenicity of the test substance from the parameters of a generated model in which the amino acid residues have been substituted. Specifically, for test substances requiring improved immunogenicity, such as vaccines, or test substances requiring reduced immunogenicity, such as antibody pharmaceuticals for specific diseases, it is possible to generate models in which amino acid residues have been substituted, and predict an increase or decrease in immunogenicity from the increase or decrease in each parameter.

[0148] The immunogenicity prediction method executed by the prediction device 1 specifically includes the following steps: (a) generating a molecular model using molecular modeling software based on the amino acid sequence of the target protein, (b) calculating physicochemical parameters based on the molecular model, and (c) determining the level of immunogenicity based on the calculated physicochemical parameters.

[0149] In the immunogenicity prediction method executed by the prediction device 1, step (a) can be omitted if crystal structure data or predicted structure data is available. Furthermore, step (b) can be used regardless of the type of software, as long as similar physicochemical parameters are used. Furthermore, by estimating amino acid residues that contribute to an increase or decrease in parameters from the generated model, and performing steps (a), (b), and (c) using an amino acid sequence in which specific amino acid residues have been substituted, it is possible to predict whether immunogenicity will be improved or reduced by the substitution of amino acid residues.

[0150] The immunogenicity prediction method executed by the prediction device 1 relates to in silico immunogenicity prediction.

[0151] The background to this is that only a limited number of antibody drugs currently on the market are highly immunogenic, that highly immunogenic antibody drugs tend to be difficult to develop, and that immunogenicity plays an important role in the developability of antibody drugs. Many factors affect the immunogenicity of antibody drugs. As for product-related factors, both the potential epitopes contained in the antibody drug and the physicochemical properties of the antibody drug itself affect ADA production.

[0152] According to this embodiment, it has been shown that highly immunogenic antibodies tend to have biases in the charge and hydrophobicity distributions on their surfaces. It has also been shown that quantitative prediction of charge and hydrophobicity distributions is useful for predicting the immunogenicity of antibodies. Here, charge distribution can be expressed by "Net Charge" and "Dipole Moment." Furthermore, hydrophobic distribution can be expressed by "Hydrophobic Score" and "Dipole Moment." According to this embodiment, verification results using anti-C1s antibodies have shown that results equivalent to those of T cell assays, a conventional immunogenicity prediction method, have been obtained. Conventional T cell assays are expensive (several million yen per assay) and have low throughput (two weeks per assay and can only predict a few samples simultaneously), and require actual antibodies. The method using the prediction device 1 can generate models and calculate parameters from only the amino acid sequence of the antibody, achieving lower costs and higher throughput than conventional methods. Furthermore, actual antibodies are not required.

[0153] 1...prediction device, 10...storage unit, 11...acquisition unit, 12...generation unit, 13...evaluation unit, 14...identification unit, 15...learning unit, 16...prediction unit, 100...CPU, 101...RAM, 102...ROM, 103...input / output device, 104...communication module, 105...auxiliary storage device, P1...prediction program, P10...storage module, P11...acquisition module, P12...generation module, P13...evaluation module, P14...identification module, P15...learning module, P16...prediction module.

Claims

1. A computer-implemented prediction method comprising a prediction step of predicting a biological property of a test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance.

2. The prediction method according to claim 1, wherein the biological property is at least one property selected from the group consisting of immunogenicity, pharmacokinetics, and non-specific binding of the test substance.

3. The prediction method according to claim 1 or 2, wherein the physicochemical property is further based on a dipole moment.

4. The prediction method according to any one of claims 1 to 3, wherein the physicochemical properties are based on at least one of charge distribution and hydrophobicity distribution.

5. The prediction method according to any one of claims 1 to 4, wherein the physicochemical properties are based on at least one of bias in charge distribution and bias in hydrophobicity distribution.

6. The prediction method according to claim 4 or 5, wherein the charge distribution is based on charge and dipole moment, and the hydrophobicity distribution is based on hydrophobicity and dipole moment.

7. The prediction method according to any one of claims 1 to 6, wherein the test substance is an antibody, and the physicochemical properties are those of an Fv (fragment variable) region of the test substance.

8. A prediction method according to any one of claims 1 to 7, wherein the physicochemical property is based on at least one of bias in charge distribution or bias in hydrophobicity distribution, and the prediction step predicts that the ADA (Anti-drug antibody) expression rate of the test substance is lower than a predetermined ADA expression rate when the bias in charge distribution or the bias in hydrophobicity distribution is within a predetermined standard.

9. The prediction method according to any one of claims 1 to 8, further comprising an evaluation step of evaluating the physicochemical properties of the test substance, and wherein the prediction step makes a prediction based on the physicochemical properties evaluated in the evaluation step.

10. The prediction method according to claim 9, further comprising a generation step of generating a molecular model of the test substance on molecular modeling software, and wherein the evaluation step performs evaluation based on the molecular model generated in the generation step.

11. A prediction method according to any one of claims 1 to 10, comprising at least one of the following (a) to (c): (a) a prediction step of predicting the immunogenicity of the test substance based on the physicochemical properties based on at least one of the charge and hydrophobicity of the test substance and the dipole moment of the test substance; (b) a prediction step of predicting the pharmacokinetics of the test substance based on the physicochemical properties based on at least one of the charge and hydrophobicity of the test substance and the dipole moment of the test substance; and (c) a prediction step of predicting the non-specific binding of the test substance based on the physicochemical properties based on the charge distribution of the test substance.

12. A prediction device comprising at least one processor, wherein the at least one processor predicts a biological property of a test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance.

13. A prediction program for causing a computer to function as a prediction unit that predicts the biological properties of a test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance.

14. A predictive model that is a trained model for causing a computer to function to output information regarding the biological properties of a test substance based on physicochemical properties based on at least one of the charge or hydrophobicity of the test substance, the predictive model being composed of a neural network in which weighting coefficients have been trained using training data including information regarding the known physicochemical properties and information regarding the known biological properties of each of a plurality of known substances, and performing calculations based on the trained weighting coefficients on the information regarding the physicochemical properties of the test substance input to the neural network, and outputting information regarding the biological properties of the test substance.

Citation Information

Patent Citations

  • Methods and systems for biopharmaceutical development

    JP2023548364A

  • Method for predicting immunogenic epitope and apparatus using same

    WO2024101854A1

  • Method for predicting immunogenicity of neoantigen epitope, and device using same

    WO2024101856A1

  • Prediction device, prediction method, and prediction program

    WO2024116360A1