Estimation system, method for estimating molecular conformation, and estimation program
Patent Information
- Application Number
- PCT/JP2024/038563
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-29
- Publication Date
- 2025-05-08
AI Technical Summary
The prior art is difficult to accurately estimate the formation of molecules capable of binding to the target molecule.
By an estimation system, the system includes at least one processor for obtaining the target molecular model, generating a molecular model capable of binding to the target molecule, calculating the interaction energy between the molecules and the interaction energy within the molecule, and performing a formation estimation based on these energy.
The accuracy of structural form estimation of molecules that can bind to the target molecule is improved, and the ability to estimate molecular binding patterns is enhanced.
Smart Images

Figure JP2024038563_08052025_PF_FP_ABST
Abstract
Description
Prediction system, molecular conformation prediction method, and prediction program
[0001] One aspect of the present disclosure relates to a prediction system, a method for predicting the conformation of a molecule, and a prediction program.
[0002] Patent Document 1 describes a method for searching for a structure of a cyclic molecule using a computer to search for a stable structure of a cyclic molecule in which n compound groups are cyclically linked. Patent Document 2 describes an apparatus for searching for a stable bonding structure of a molecule. Patent Document 3 describes a structure searching apparatus for searching for a stable structure of multiple molecules that interact with each other. Non-Patent Document 1 describes a method for estimating the conformation of a polymer chain having N monomers on a lattice using a quantum variational algorithm.
[0003] JP 2020-91518 A JP 2020-173643 A JP 2020-194487 A
[0004] Robert, A., Barkoutsos, P. K., Woerner, S. & Tavernelli, I. Resource-efficient quantum algorithm for protein folding. npj Quantum Inf. 7, 38 (2021).
[0005] It is desirable to more accurately predict the conformation of a molecule that can bind to a target molecule. One aspect of the present disclosure aims to provide a prediction system, a method for predicting the conformation of a molecule, and a prediction program.
[0006] A prediction system according to one aspect of the present disclosure includes at least one processor, which acquires a target molecule model representing the conformation of a target molecule, generates a target molecule model representing the conformation of a target molecule that is a molecule capable of binding to the target molecule, calculates an interaction energy between the target molecule and the target molecule as intermolecular interaction energy based on a positional relationship between the target molecule model and the target molecule model, calculates an interaction energy within the target molecule as intramolecular interaction energy, and predicts the conformation of the target molecule based on the intermolecular interaction energy and the intramolecular interaction energy.
[0007] A method for predicting a molecular conformation according to one aspect of the present disclosure is executed by a prediction system including at least one processor. The prediction method includes the steps of acquiring a target molecule model representing the conformation of a target molecule, generating a target molecule model representing the conformation of a target molecule that can bind to the target molecule, calculating an interaction energy between the target molecule and the target molecule as an intermolecular interaction energy based on a positional relationship between the target molecule model and the target molecule model, calculating an interaction energy within the target molecule as an intramolecular interaction energy, and predicting the conformation of the target molecule based on the intermolecular interaction energy and the intramolecular interaction energy.
[0008] An estimation program according to one aspect of the present disclosure causes a computer to execute the steps of: acquiring a target molecule model showing the conformation of the target molecule; generating a target molecule model showing the conformation of a target molecule that is a molecule capable of binding to the target molecule; calculating the interaction energy between the target molecule and the target molecule as intermolecular interaction energy based on the positional relationship between the target molecule model and the target molecule model; calculating the interaction energy within the target molecule as intramolecular interaction energy; and estimating the conformation of the target molecule based on the intermolecular interaction energy and the intramolecular interaction energy.
[0009] In this aspect, the interaction energy between the target molecule and the target molecule that bind to each other is calculated based on the positional relationship between these two molecules. The interaction energy within the target molecule is also calculated. The conformation of the target molecule is then estimated based on these two types of interaction energies. By taking into account these two types of interaction energies, the conformation of the target molecule that can bind to the target molecule can be estimated with higher accuracy.
[0010] According to one aspect of the present disclosure, the conformation of a molecule capable of binding to a target molecule can be predicted with greater accuracy.
[0011] 1 is a diagram illustrating an example of the functional configuration of an estimation system; FIG. 2 is a diagram illustrating an example of the hardware configuration of a computer that functions as an estimation system; FIG. 3 is a flowchart illustrating an example of processing executed by the estimation system; FIG. 4 is a diagram illustrating an example of a lattice system; FIG. 5 is a diagram illustrating an example of generating a target molecular model under constraint conditions; FIG. 6 is a diagram illustrating another example of generating a target molecular model under constraint conditions; FIG. 7 is a diagram illustrating the results of an example; FIG. 8 is a diagram illustrating the results of an example; FIG. 9 is a diagram illustrating the results of a comparative example.
[0012] Various examples of the present disclosure will be described in detail below with reference to the accompanying drawings. In the description of the drawings, the same or equivalent elements are designated by the same reference numerals, and redundant description will be omitted.
[0013] [System Overview] The prediction system according to the present disclosure is a computer system that predicts the conformation of a target molecule, which is a molecule that can bind to a target molecule. The prediction system processes a conformation model, which is electronic data that represents the conformation of a molecule. In the present disclosure, the conformation model of the target molecule is referred to as a "target molecular model," and the conformation model of the target molecule is referred to as a "target molecular model." The prediction system calculates the interaction energy between the target molecule and the target molecule as intermolecular interaction energy based on the target molecular model and the target molecular model. The prediction system also calculates the interaction energy within the target molecule as intramolecular interaction energy based on the target molecular model. The prediction system then predicts the conformation of the target molecule based on the intermolecular interaction energy and the intramolecular interaction energy.
[0014] In one example, the prediction system generates a target molecular model for each of a plurality of conformations of the target molecule, calculates intermolecular interaction energy and intramolecular interaction energy for the target molecular model, selects at least one conformation from the plurality of conformations of the target molecule based on the intermolecular interaction energy and intramolecular interaction energy obtained for each of the plurality of target molecular models, and outputs the selected conformation as a prediction result.
[0015] [System Configuration] FIG. 1 is a diagram showing the functional configuration of an estimation system 10 according to an example. In this example, the estimation system 10 includes functional modules: an acquisition unit 11, a generation unit 12, a calculation unit 13, and an estimation unit 14. The acquisition unit 11 is a functional module that acquires data necessary to estimate the conformation of a target molecule. In one example, the acquisition unit 11 acquires a target molecular model and target molecular data used to generate the target molecular model. The generation unit 12 is a functional module that generates at least one target molecular model. The calculation unit 13 is a functional module that calculates, for each of at least one target molecular model, an intermolecular interaction energy based on the positional relationship between the target molecular model and the target molecular model, and calculates an intramolecular interaction energy based on the target molecular model. The estimation unit 14 is a functional module that estimates the conformation of the target molecule based on the intermolecular interaction energy and intramolecular interaction energy of at least one target molecular model.
[0016] FIG. 2 is a diagram showing an example of the hardware configuration of a computer 100 functioning as the estimation system 10. For example, the computer 100 includes a processor 101, a main memory unit 102, an auxiliary memory unit 103, a communication control unit 104, an input device 105, and an output device 106. The processor 101 executes an operating system and application programs. The main memory unit 102 is composed of, for example, ROM and RAM. The auxiliary memory unit 103 is composed of, for example, a hard disk or flash memory, and generally stores larger amounts of data than the main memory unit 102. The communication control unit 104 is composed of, for example, a network card or a wireless communication module. The input device 105 is composed of, for example, a keyboard, a mouse, a touch panel, etc. The output device 106 is composed of, for example, a monitor and a speaker.
[0017] Each functional module of the estimation system 10 is realized by an estimation program 110 pre-stored in the auxiliary storage unit 103. Specifically, each functional module is realized by loading the estimation program 110 onto the processor 101 or the main storage unit 102 and causing the processor 101 to execute the estimation program 110. The processor 101 operates the communication control unit 104, the input device 105, or the output device 106 in accordance with the estimation program 110, and reads and writes data from and to the main storage unit 102 or the auxiliary storage unit 103. Data or a database required for processing may be stored in the main storage unit 102 or the auxiliary storage unit 103.
[0018] The estimation program 110 may be provided by being recorded on a non-transitory computer-readable storage medium such as a CD-ROM, a DVD-ROM, a semiconductor memory, etc. Alternatively, the estimation program 110 may be provided via a communication network as a data signal superimposed on a carrier wave.
[0019] The estimation system 10 may be configured with one computer 100 or with multiple computers 100. When multiple computers 100 are used, these computers 100 are connected via a communication network such as the Internet or an intranet, thereby logically constructing a single estimation system 10.
[0020] 3 is a flowchart showing, as a processing flow S1, an example of processing executed by the estimation system 10. The processing flow S1 is an example of a molecular conformation estimation method according to the present disclosure.
[0021] In step S11, the acquisition unit 11 acquires a target molecule model. The target molecule model may be a conformation model designated by an input operation or a selection operation by a user of the estimation system 10. The acquisition unit 11 may read the target molecule model from a predetermined storage device such as a database, or may receive the target molecule model from another computer such as a user terminal.
[0022] In one example, the acquisition unit 11 acquires a target molecular model in which at least one of a plurality of structural units (atomic groups) constituting a target molecule is represented by coarse-grained particles. For example, the target molecular model represents each of the plurality of structural units by a coarse-grained particle. Coarse-graining refers to a technique in which each structural unit (atomic group) constituting a system consisting of a large number of atoms is approximated as one or more particles called coarse-grained particles. This coarse-graining simplifies the structural units. Each structural unit constituting a target molecule may be represented by one or more coarse-grained particles. When one structural unit is represented by two coarse-grained particles, one coarse-grained particle may represent the main chain and the other coarse-grained particle may represent the side chain. The target molecular model may include both a structural unit represented by one coarse-grained particle and a structural unit represented by two or more coarse-grained particles.
[0023] In one example, the acquisition unit 11 may acquire a target molecule model showing the conformation of a part of the target molecule, for example, a target molecule model showing the conformation of an active site in the target molecule, or a target molecule model showing the conformation of the entire target molecule including the active site.
[0024] In step S12, the acquisition unit 11 acquires target molecule data to be used to generate a target molecule model. For example, the target molecule data is electronic data indicating the sequence of multiple building blocks that make up the target molecule. In one example, the target molecule data is specified by an input operation or a selection operation by a user. The acquisition unit 11 may read the target molecule data from a predetermined storage device such as a database, or may receive the target molecule data from another computer such as a user terminal.
[0025] In step S13, the generation unit 12 sets a lattice system based on the target molecular model. A lattice system is a two-dimensional or three-dimensional space in which multiple lattice points are connected to each other, and is used to set the conformation of the target molecule. The generation unit 12 sets the outer edge of the lattice system based on the target molecular model. In one example, the generation unit 12 may calculate the center of gravity of multiple building blocks that constitute the active site in the target molecule based on the target molecular model, and set the outer edge of the lattice system based on the center of gravity. For example, the generation unit 12 may set the outer edge of the lattice system at a position a predetermined distance away from the calculated center of gravity. The generation unit 12 generates the lattice system by arranging multiple unit cells within the outer edge. The generation unit 12 may use one of a regular tetrahedral lattice, a cubic lattice, and a square lattice as the unit cell. If the coordinate values of each lattice point in the lattice system are expressed as integers, the generation unit 12 may rearrange each coarse-grained particle of the target molecular model by a technique such as scaling to use such a lattice system. In one example, before setting the grid system, the generation unit 12 may perform at least one of translation and rotation on at least one of the target molecular model and the target molecular model represented by the target molecular data to reposition the at least one model.
[0026] 4 is a diagram showing an example of a lattice system. In this example, a target molecule model 200 represents each of a plurality of structural units that make up the active site of the target molecule using coarse-grained particles 201. The generation unit 12 calculates the center of gravity 210 of the plurality of coarse-grained particles (structural units) 201, and sets the outer edge of the lattice system 220 at a position a predetermined distance away from this center of gravity. The generation unit 12 generates the lattice system 220 by arranging unit cells within the outer edge.
[0027] Returning to FIG. 3 , in step S14, the generation unit 12 generates a target molecular model in a lattice system based on the target molecular data so that the target molecular model satisfies the constraints in the lattice system. The generation unit 12 generates the target molecular model to reflect the arrangement of multiple structural units constituting the target molecule. The generation unit 12 generates a target molecular model that reflects the arrangement by arranging the jth structural unit and the (j+1)th structural unit constituting the target molecule adjacent to each other, where j is an integer greater than or equal to 1. In the present disclosure, "adjacent" two structural units (or two coarse-grained particles) means that one structural unit (or coarse-grained particle) and the other structural unit (or coarse-grained particle) are close to each other within a predetermined distance. The predetermined distance may be 3.8 Å.
[0028] In one example, the generator 12 generates a target molecular model in which at least one of multiple structural units (atomic groups) constituting the target molecule is represented by coarse-grained particles, similar to the target molecular model. For example, the generator 12 generates a target molecular model in which each of multiple structural units is represented by coarse-grained particles. The generator 12 generates a target molecular model in which the jth structural unit and the (j+1)th structural unit constituting the target molecule are represented by coarse-grained particles, where j is an integer greater than or equal to 1. The generator 12 represents (approximates) a certain structural unit by one or more coarse-grained particles. The generator 12 may represent a structural unit by two coarse-grained particles, for example, by a coarse-grained particle representing a main chain and a coarse-grained particle representing a side chain. In this example, the generator 12 arranges the coarse-grained particle representing the main chain and the coarse-grained particle representing the side chain adjacent to each other. The target molecular model may include both a structural unit coarse-grained by one particle and a structural unit coarse-grained by two or more particles.
[0029] The generator 12 may generate a target molecular model in which at least half or all of the multiple building blocks constituting the target molecule are represented by multiple coarse-grained particles. For example, the generator 12 may generate a target molecular model in which at least half or all of the multiple building blocks constituting the target molecule are represented by two coarse-grained particles. The generator 12 may generate a target molecular model in which at least half or all of the multiple building blocks constituting the target molecule are represented by a coarse-grained particle representing a main chain and a coarse-grained particle representing a side chain. In one example, the generator 12 generates a target molecular model in a lattice system by arranging a coarse-grained particle representing a main chain and a coarse-grained particle representing a side chain adjacent to each other for at least half or all of the multiple building blocks constituting the target molecule in the lattice system.
[0030] The generation unit 12 places each coarse-grained particle of the target molecular model at a lattice point so that only one coarse-grained particle is placed at one lattice point. The generation unit 12 places two coarse-grained particles corresponding to the jth and (j+1)th constituent units of the target molecule adjacent to each other in the lattice system, where j is an integer of 1 or greater, to generate the target molecular model in the lattice system. In this case, the lattice point L at which the coarse-grained particle representing the jth constituent unit is placed is a The adjacent lattice point L b A coarse-grained particle representing the (j+1)th structural unit is placed at the lattice point L where the coarse-grained particle representing the main chain is placed. When a structural unit is represented by a coarse-grained particle representing a main chain and a coarse-grained particle representing a side chain, the generation unit 12 places the coarse-grained particle representing the main chain and the coarse-grained particle representing the side chain adjacent to each other, and generates a target molecular model in the lattice system. In this case, the lattice point L where the coarse-grained particle representing the main chain is placed is c The adjacent lattice point L d Coarse-grained particles representing side chains are placed on the
[0031] Constraints refer to conditions that must be satisfied when generating a target molecular model. In one example, constraints include intermolecular constraints regarding the positional relationship between the target molecular model and the target molecular model. For example, the intermolecular constraint indicates a requirement to avoid steric hindrance between the target molecule and the target molecule, i.e., a requirement that the coarse-grained particles of the target molecular model maintain a predetermined distance or more from the coarse-grained particles of the target molecular model. Constraints may also include intramolecular constraints regarding the positional relationship between multiple building blocks that make up the target molecular model. For example, the intramolecular constraint indicates a requirement to avoid steric hindrance between multiple building blocks that make up the target molecule, i.e., a requirement that the coarse-grained particles do not collide with each other within the target molecular model. The intramolecular constraint may be a requirement that multiple coarse-grained particles of the target molecule maintain a predetermined distance or more from each other, or may be a requirement that multiple coarse-grained particles of the target molecule must not be positioned at a single grid point. Constraints may also include a cyclization constraint indicating a requirement that the target molecule includes a cyclic structure, i.e., a requirement that three or more coarse-grained particles in the target molecular model are cyclically bonded. In one example, the constraints include an intermolecular constraint and / or an intramolecular constraint and / or a cyclization constraint.
[0032] FIG. 5 shows an example of generating a target molecular model under constraint conditions. For convenience, FIG. 5 shows a lattice system 220 in a simple two-dimensional shape with a square unit cell. In this example, it is assumed that each constituent unit of the target molecular model is represented by a single coarse-grained particle. The constraint conditions 310 shown in this example include an intermolecular constraint 311, an intramolecular constraint 312, and a cyclization constraint 313. For convenience, these three constraints are explained individually in this example, but it should be noted that the generation unit 12 generates a target molecular model so as to satisfy all of the constraint conditions 310 within the lattice system.
[0033] The generation unit 12 generates the target molecular model 230 in the lattice system 220 based on the intermolecular constraint 311 so that no steric hindrance occurs between the coarse-grained particle 231 of the target molecular model 230 and the coarse-grained particle of the target molecular model 200. After placing the fourth coarse-grained particle 231a, the generation unit 12 places the fifth coarse-grained particle to the right of the coarse-grained particle 231a so that no steric hindrance occurs between the coarse-grained particle of the target molecular model 200 and the coarse-grained particle of the target molecular model 200.
[0034] The generation unit 12 generates the target molecular model 240 in the lattice system 220 based on the intramolecular constraint 312 so that no steric hindrance occurs between the coarse-grained particles 241 of the target molecular model 240. After placing the fourth coarse-grained particle 241a, the generation unit 12 places the fifth coarse-grained particle to the right or below the coarse-grained particle 241a so that no steric hindrance occurs between the fifth coarse-grained particle and the other coarse-grained particles 241.
[0035] The generator 12 generates the target molecular model 250 in the lattice system 220 based on the cyclization constraint 313 so that a cyclic structure 259 is formed by three or more coarse-grained particles 251 .
[0036] FIG. 6 shows another example of generating a target molecular model under constraint conditions. FIG. 6 also shows a lattice system 220 in a simple two-dimensional shape. In this example, it is assumed that each structural unit of the target molecular model is represented by two coarse-grained particles. Hatched circles represent main chains, and open circles represent side chains. Hereinafter, the coarse-grained particles representing main chains will be referred to as "main chain particles," and the coarse-grained particles representing side chains will be referred to as "side chain particles." The constraint conditions 320 shown in this example include an intermolecular constraint 321, an intramolecular constraint 322, and a cyclization constraint 323. For convenience, these three constraints will be described separately in this example, but it should be noted that the generation unit 12 generates a target molecular model so as to satisfy all of the constraint conditions 320.
[0037] The generation unit 12 generates the target molecular model 260 in the lattice system 220 based on the intermolecular constraint 311 so that the main chain particles 261 and side chain particles 262 of the target molecular model 260 do not cause steric hindrance between them and the coarse-grained particles of the target molecular model 200. After arranging the main chain particle 261a and side chain particle 262a representing the third structural unit, the generation unit 12 arranges the main chain particle and side chain particle representing the fourth structural unit so that no steric hindrance occurs between them and the coarse-grained particles of the target molecular model 200. The generation unit 12 arranges the fourth main chain particle to the right of the main chain particle 261a, and arranges the fourth side chain particle to the right of or above the fourth main chain particle.
[0038] Based on the intramolecular constraints 322, the generation unit 12 generates the target molecular model 270 within the lattice system 220 so that no steric hindrance occurs between the coarse-grained particles of the target molecular model 270. After arranging the main chain particle 271a and side chain particle 272a representing the third structural unit, the generation unit 12 arranges the main chain particle and side chain particle representing the fourth structural unit so that no steric hindrance occurs between the main chain particle 271 and the side chain particle 272. The generation unit 12 arranges the fourth main chain particle to the right of the main chain particle 271a, and arranges the fourth side chain particle to the right, above, or below the fourth main chain particle.
[0039] The generation unit 12 generates the target molecular model 280 in the lattice system 220 so that a cyclic structure 289 is formed by three or more main chain particles 281 based on the cyclization constraint 323. The generation unit 12 places side chain particles 282 adjacent to each main chain particle 281.
[0040] 3 , in step S15, the calculation unit 13 calculates the interaction energy of the conformation of the target molecule based on the target molecule model and the object molecule model. In one example, the calculation unit 13 calculates the intramolecular interaction energy and the intermolecular interaction energy, and calculates the sum of these two interaction energies as the total interaction energy of the conformation of the target molecule.
[0041] The calculation of intramolecular interaction energy will now be described. In one example, the calculation unit 13 calculates multiple interaction energies between multiple structural units that make up the target molecule as multiple inter-unit interaction energies. The calculation unit 13 calculates the intramolecular interaction energy based on the multiple inter-unit interaction energies. For example, the calculation unit 13 calculates the sum of the multiple inter-unit interaction energies as the intramolecular interaction energy. Such calculation is represented by formula (1). Variable E intra indicates the intramolecular interaction energy. The variables i and j each indicate a number for identifying the structural unit of the target molecule. ij is a coefficient set based on the distance between structural unit i and structural unit j. ijis "1" if the distance is less than a predetermined threshold Td, and is "0" otherwise. However, when two structural units are directly covalently bonded, the variable a ij is "0". An example of this is when the jth constitutional unit and the (j+1)th constitutional unit are covalently bonded. ij denotes the inter-unit interaction energy occurring between constituent units i and j.
[0042] The calculation of intermolecular interaction energy will now be described. In one example, the calculation unit 13 calculates multiple interaction energies between multiple structural units constituting the target molecule and multiple structural units constituting the target molecule as multiple inter-unit interaction energies. The calculation unit 13 calculates the intermolecular interaction energy based on the multiple inter-unit interaction energies. For example, the calculation unit 13 calculates the sum of the multiple inter-unit interaction energies as the intermolecular interaction energy. Such calculation is represented by equation (2). Variable E inter represents the intermolecular interaction energy. The variable i represents a number for identifying a constituent unit of the target molecule, and the variable k represents a number for identifying a constituent unit of the target molecule. The variable b ik is a coefficient set based on the distance between structural unit i and structural unit k. ik The variable e is set to "1" if the distance is less than a predetermined threshold Td, and is set to "0" otherwise. ik denotes the inter-unit interaction energy occurring between two structural units i and k.
[0043] Both formulas (1) and (2) represent the calculation of the sum of the interaction energies between two structural units whose distance from each other is less than a threshold value Td, i.e., between two structural units that are located close to each other. When the structural units of the subject molecule and the target molecule are amino acids, the calculation unit 13 may determine the inter-unit interaction energy and the threshold value Td in both formulas (1) and (2) based on the Miyazawa-Jernigan matrix (MJ matrix) (J. Mol. Biol. (1996) 256, 623-644, Table 3). The MJ matrix defines the interaction energy between two amino acids. The variable a in formulas (1) and (2) ij , b ik The distance referenced to determine the value of may be the Euclidean distance. The threshold Td is 6.5 Å (6.5×10 -10 m).
[0044] As shown in formula (3), the calculation unit 13 calculates the intramolecular interaction energy E intra and intermolecular mutual energy E inter The sum of these is the total interaction energy E total It is calculated as follows.
[0045] As shown in step S16, in one example, the generation unit 12 and the calculation unit 13 cooperate to perform a search that repeats the generation of a target molecular model and the calculation of interaction energy for multiple conformations of the target molecule. If processing is to be performed for another conformation (NO in step S16), the process returns to step S14. The generation unit 12 generates a new target molecular model in an existing lattice system (e.g., lattice system 220) based on the target molecular data so that the target molecular model satisfies the constraints within the lattice system. In the repeated step S15, the calculation unit 13 calculates the interaction energy (e.g., total interaction energy) of the new conformation of the target molecule based on the target molecular model and the new target molecular model.
[0046] If the search is to be terminated (YES in step S16), the process proceeds to step S17. In step S17, the estimation unit 14 estimates a conformation of the target molecule (target molecule model). In one example, the estimation unit 14 selects at least one conformation from a plurality of conformations of the target molecule based on the interaction energies of each of the conformations, and sets the selected conformation as the estimation result. The estimation unit 14 may select a conformation based on the total interaction energy. For example, the estimation unit 14 may select a conformation with the smallest total interaction energy, or may select at least one conformation with a total interaction energy less than a predetermined threshold Te. The estimation unit 14 may select a conformation based on an intermolecular interaction energy. For example, the estimation unit 14 may select a conformation with the smallest intermolecular interaction energy, or may select at least one conformation with an intermolecular interaction energy less than a predetermined threshold Ti.
[0047] In step S18, the estimation unit 14 outputs an estimation result indicating at least one estimated conformation (target molecular model). For example, the estimation unit 14 may output, as the estimation result, electronic data representing each conformation in a format that allows the conformation to be rendered by computer graphics (CG). The estimation unit 14 may output an estimation result that further indicates the interaction energy of each estimated conformation, for example, an estimation result that further indicates at least one of the total interaction energy, the intermolecular interaction energy, and the intramolecular interaction energy. The estimation unit 14 may store the estimation result in a storage device such as the auxiliary storage unit 103, display the prediction result on a monitor, or transmit the prediction result to another computer such as a user terminal.
[0048] As shown in process flow S1, the generation unit 12 generates multiple target molecular models (conformations) of the target molecule while changing the orientation or position of the lattice system relative to the target molecular model. The calculation unit 13 calculates the intermolecular interaction energy and intramolecular interaction energy for each generated target molecular model (conformation) and calculates the sum of these two interaction energies as the total interaction energy. The estimation unit 14 selects at least one conformation from the multiple conformations of the target molecule based on the interaction energy (e.g., the total interaction energy or the intermolecular interaction energy) of each of the multiple conformations of the target molecule.
[0049] As described above, in one example, the estimation unit 14 outputs the conformation with the smallest total interaction energy as the estimated result. This process can be considered as mathematical optimization. In this case, Equation (3) can be considered as the objective function, and the total interaction energy E total can be said to be the objective variable. The total interaction energy E total The conformation for which is smallest is the conformation that is likely to be the most stable (i.e., the most stable structure).
[0050] [Molecules] The molecules are not particularly limited and can be appropriately selected depending on the purpose. For example, the molecules may be low molecules (low molecular weight compounds), medium molecules (medium molecular weight compounds), or polymers (polymer compounds). In the present disclosure, low molecules or low molecular weight compounds refer to compounds with a molecular weight of less than 500 g / mol. In the present disclosure, medium molecules or medium molecular weight compounds refer to compounds with a molecular weight of 500 g / mol or more and less than 10,000 g / mol. In the present disclosure, polymers or polymer compounds refer to compounds with a molecular weight of 10,000 g / mol or more.
[0051] The molecule may be a biomolecule or a non-biomolecule. The molecule may be an antigen-binding molecule such as a nucleic acid, peptide, protein, or antibody, or may be a molecule capable of binding to a target molecule (a target molecule-binding molecule). The molecule may be a drug candidate molecule. The peptide may be a linear peptide or a cyclic peptide. When the molecule is a peptide, protein, or antibody, the building blocks of the molecule are amino acids. When the molecule is a nucleic acid, the building blocks of the molecule are nucleosides or nucleotides. In one example, the target molecule may be a protein such as a receptor or enzyme, and the target molecule may be a peptide. In the present disclosure, a protein serving as a target molecule is also referred to as a "target protein," and a peptide serving as a target molecule is also referred to as a "target peptide." The number of amino acid residues contained in the peptide may be, for example, 5 to 30, 7 to 20, 8 to 18, or 9 to 15. The peptide may be linear, branched, or cyclic. When the peptide is a linear peptide or a cyclic peptide, the linear peptide as the target molecule is also referred to as a "target linear peptide," and the cyclic peptide as the target molecule is also referred to as a "target cyclic peptide."
[0052] The term "amino acid" includes natural amino acids and unnatural amino acids. A "natural amino acid" is any L-amino acid selected from Gly (glycine), L-Ala (alanine), L-Ser (serine), L-Thr (threonine), L-Val (valine), L-Leu (leucine), L-Ile (isoleucine), L-Phe (phenylalanine), L-Tyr (tyrosine), L-Trp (tryptophan), L-His (histidine), L-Glu (glutamic acid), L-Asp (aspartic acid), L-Gln (glutamine), L-Asn (asparagine), L-Cys (cysteine), L-Met (methionine), L-Lys (lysine), L-Arg (arginine), and L-Pro (proline). An "unnatural amino acid" is an amino acid other than natural amino acids. Examples of unnatural amino acids include β-amino acids, γ-amino acids, D-amino acids, N-substituted amino acids other than Pro, α,α-disubstituted amino acids, and amino acids whose side chains differ from those of natural amino acids. In the case of α-amino acids, the "main chain of an amino acid" refers to the linear portion composed of an amino group, an α-carbon, and a carboxyl group. In the case of β-amino acids, the "main chain of an amino acid" refers to the linear portion composed of an amino group, a β-carbon, an α-carbon, and a carboxyl group. In the case of γ-amino acids, the "main chain of an amino acid" refers to the linear portion composed of an amino group, a γ-carbon, a β-carbon, an α-carbon, and a carboxyl group. In the case of α-amino acids, the "side chain of an amino acid" refers to the group and / or atom bonded to the carbon (α-carbon) to which the amino group and carboxyl group are bonded. For example, the methyl group of Ala is the side chain of an amino acid. On the other hand, Gly (glycine) does not have a side chain. In the case of β-amino acids, the group and / or atom attached to at least one of the α-carbon and β-carbon can be the side chain of the amino acid. In the case of γ-amino acids, the group and / or atom attached to at least one of the α-carbon, β-carbon, and γ-carbon can be the side chain of the amino acid.
[0053] [Modifications] The technology of the present disclosure has been described in detail above based on various examples. However, the technology of the present disclosure is not limited to the above examples. Various modifications are possible within the scope of the gist of the present disclosure.
[0054] In the above process flow S1, the estimation system 10 generates multiple conformations of the target molecule (target molecular model) and selects at least one conformation from the multiple conformations. However, the estimation system 10 may calculate intermolecular interaction energy and intramolecular interaction energy for one conformation of the target molecule (target molecular model) and estimate the conformation of the target molecule based on these interaction energies. In other words, the repetition shown in step S16 of the process flow S1 may be omitted.
[0055] The estimation system may be constructed as a server in a client-server system, or may be implemented in a stand-alone computer. Alternatively, the estimation system may be implemented in a user terminal that can access a predetermined database via a communication network such as the Internet.
[0056] The processing steps of the method executed by at least one processor are not limited to the examples in the above embodiments. For example, some of the steps or processes described above may be omitted, or the steps may be executed in a different order. Furthermore, any two or more of the steps described above may be combined, or some of the steps may be modified or deleted. Alternatively, other steps may be executed in addition to the steps described above.
[0057] In the present disclosure, when comparing the magnitude of two numerical values, either of the two criteria "greater than or equal to" and "greater than" may be used, or either of the two criteria "less than or equal to" and "less than" may be used.
[0058] In the present disclosure, the expression "at least one processor executes a first process, executes a second process, ... executes an nth process" or an expression corresponding thereto indicates a concept including a case where the processor that executes n processes from the first process to the nth process changes midway. In other words, this expression indicates a concept including both a case where all n processes are executed by the same processor and a case where the processor changes among the n processes according to an arbitrary policy.
[0059] In this disclosure, the term "to" indicating a range is an inclusive expression. For example, "A to B" means a range equal to or greater than A and equal to or less than B.
[0060] In this disclosure, the term "about" when used in conjunction with a numerical value means a range of plus and minus 10% of that numerical value.
[0061] The term "and / or" is used herein to refer to each of the objects listed before and after "and / or" or any combination thereof. For example, "A, B and / or C" includes each of the objects "A," "B," and "C," as well as the combinations "A and B," "A and C," "B and C," and "A and B and C."
[0062] [Supplementary Notes] As can be seen from the various examples above, the present disclosure includes the following aspects. (Supplementary Note 1) An estimation system comprising at least one processor, wherein the at least one processor: acquires a target molecule model representing the conformation of a target molecule; generates a target molecule model representing the conformation of a target molecule that is a molecule capable of binding to the target molecule; calculates an interaction energy between the target molecule and the target molecule as an intermolecular interaction energy based on a positional relationship between the target molecule model and the target molecule model; calculates an interaction energy within the target molecule as an intramolecular interaction energy; and estimates the conformation of the target molecule based on the intermolecular interaction energy and the intramolecular interaction energy. (Supplementary Note 2) The estimation system according to Supplementary Note 1, wherein the at least one processor: calculates a plurality of interaction energies between a plurality of structural units constituting the target molecule and a plurality of structural units constituting the target molecule as a plurality of inter-unit interaction energies; and calculates the intermolecular interaction energy based on the plurality of inter-unit interaction energies. (Supplementary Note 3) The estimation system according to Supplementary Note 1 or 2, wherein the at least one processor calculates a plurality of interaction energies between a plurality of structural units constituting the target molecule as a plurality of inter-unit interaction energies, and calculates the intramolecular interaction energy based on the plurality of inter-unit interaction energies. (Supplementary Note 4) The estimation system according to Supplementary Note 2 or 3, wherein the at least one processor obtains the target molecular model representing the plurality of structural units constituting the target molecule using coarse-grained particles, and generates the target molecular model representing the plurality of structural units constituting the target molecule using the coarse-grained particles. (Supplementary Note 5) The estimation system according to Supplementary Note 4, wherein the at least one processor arranges the j-th structural unit and the (j+1)-th structural unit constituting the target molecule adjacent to each other, where j is an integer greater than or equal to 1, and generates the target molecular model representing the j-th structural unit and the (j+1)-th structural unit using the coarse-grained particles.(Supplementary Note 6) The estimation system according to Supplementary Note 4 or 5, wherein the at least one processor generates the target molecular model, in which at least one of the plurality of structural units constituting the target molecule is represented by a plurality of the coarse-grained particles. (Supplementary Note 7) The estimation system according to Supplementary Note 6, wherein the at least one processor generates the target molecular model, in which at least one of the plurality of structural units constituting the target molecule is represented by two of the coarse-grained particles. (Supplementary Note 8) The estimation system according to any one of Supplements 4 to 7, wherein the at least one processor generates the target molecular model, in which at least one of the plurality of structural units constituting the target molecule is represented by the coarse-grained particle representing a main chain and the coarse-grained particle representing a side chain. (Supplementary Note 9) The estimation system according to Supplementary Note 8, wherein the at least one processor generates the target molecular model by arranging the coarse-grained particle representing the main chain and the coarse-grained particle representing the side chain adjacent to each other. (Supplementary Note 10) The estimation system according to any one of Supplementary Notes 1 to 9, wherein the at least one processor: generates the target molecular model for each of the plurality of conformations of the target molecule, calculates the intermolecular interaction energy based on the positional relationship between the target molecular model and the generated target molecular model, and selects at least one conformation from the plurality of conformations of the target molecule based on the intermolecular interaction energy for each of the plurality of conformations of the target molecule. (Supplementary Note 11) The estimation system according to Supplementary Note 10, wherein the at least one processor: generates the target molecular model for each of the plurality of conformations of the target molecule, calculates an interaction energy within the target molecule as an intramolecular interaction energy based on the generated target molecular model, and calculates the sum of the intramolecular interaction energy and the intermolecular interaction energy as a total interaction energy, and selects the at least one conformation from the plurality of conformations of the target molecule based on the total interaction energy for each of the plurality of conformations of the target molecule.(Supplementary Note 12) The estimation system according to any one of Supplements 1 to 11, wherein the at least one processor generates the target molecular model so as to satisfy constraints including an intermolecular constraint on the positional relationship between the target molecular model and the target molecular model. (Supplementary Note 13) The estimation system according to Supplementary Note 12, wherein the intermolecular constraint indicates avoidance of steric hindrance between the target molecule and the target molecule. (Supplementary Note 14) The estimation system according to Supplementary Note 12 or 13, wherein the at least one processor sets an outer boundary of a lattice system for arranging the conformation of the target molecule based on the target molecular model, and generates the target molecular model in the lattice system so that the target molecular model satisfies the constraints in the lattice system. (Supplementary Note 15) The estimation system according to Supplementary Note 14, wherein the at least one processor obtains the target molecular model indicating the conformation of an active site in the target molecule, and sets the outer boundary of the lattice system based on the center of gravity of multiple building blocks that constitute the active site. (Supplementary Note 16) The estimation system according to Supplementary Note 14 or 15, wherein a unit cell of the lattice system is any one selected from a regular tetrahedral lattice, a cubic lattice, and a square lattice. (Supplementary Note 17) The estimation system according to any one of Supplementary Notes 14 to 16, wherein the at least one processor generates the target molecule model in the lattice system by arranging two coarse-grained particles corresponding to the jth and (j+1)th constitutional units constituting the target molecule adjacent to each other in the lattice system, where j is an integer greater than or equal to 1. (Supplementary Note 18) The estimation system according to any one of Supplementary Notes 14 to 17, wherein the at least one processor generates the target molecule model in the lattice system by arranging a coarse-grained particle representing a main chain and a coarse-grained particle representing a side chain adjacent to each other in the lattice system for at least one of a plurality of constitutional units constituting the target molecule. (Supplementary Note 19) The estimation system according to any one of Supplementary Notes 12 to 18, wherein the constraint conditions further include at least one of an intramolecular constraint indicating that steric hindrance is avoided between multiple building blocks constituting the target molecule, and a cyclization constraint indicating that the target molecule contains a cyclic structure.(Supplementary Note 20) The prediction system according to any one of Supplementary Notes 1 to 19, wherein the target molecule is a protein, and the object molecule is a peptide. (Supplementary Note 21) A method for predicting a molecular conformation, executed by a prediction system having at least one processor, comprising: a step of acquiring a target molecule model indicating the conformation of the target molecule, a step of generating a target molecule model indicating the conformation of a object molecule that is a molecule capable of binding to the target molecule, a step of calculating an interaction energy between the object molecule and the target molecule as an intermolecular interaction energy based on a positional relationship between the target molecule model and the object molecule model, a step of calculating an interaction energy within the object molecule as an intramolecular interaction energy, and a step of predicting the conformation of the object molecule based on the intermolecular interaction energy and the intramolecular interaction energy. (Supplementary Note 22) An estimation program that causes a computer to execute the steps of: obtaining a target molecule model that shows the conformation of a target molecule; generating a target molecule model that shows the conformation of a target molecule that is a molecule that can bind to the target molecule; calculating an interaction energy between the target molecule and the target molecule as intermolecular interaction energy based on a positional relationship between the target molecule model and the target molecule model; calculating an interaction energy within the target molecule as intramolecular interaction energy; and estimating the conformation of the target molecule based on the intermolecular interaction energy and the intramolecular interaction energy. (Supplementary Note 23) The estimation system according to Supplementary Note 4 or 5, wherein the at least one processor generates the target molecule model that represents at least half of the plurality of building blocks that constitute the target molecule with a plurality of the coarse-grained particles. (Supplementary Note 24) The estimation system according to Supplementary Note 4 or 5, wherein the at least one processor generates the target molecule model that represents all of the plurality of building blocks that constitute the target molecule with a plurality of the coarse-grained particles.(Supplementary Note 25) The estimation system according to Supplementary Note 6, wherein the at least one processor generates the target molecular model in which at least half of the plurality of structural units constituting the target molecule are represented by two of the coarse-grained particles. (Supplementary Note 26) The estimation system according to Supplementary Note 6, wherein the at least one processor generates the target molecular model in which all of the plurality of structural units constituting the target molecule are represented by two of the coarse-grained particles. (Supplementary Note 27) The estimation system according to any one of Supplements 4 to 7, wherein the at least one processor generates the target molecular model in which at least half of the plurality of structural units constituting the target molecule are represented by the coarse-grained particle representing a main chain and the coarse-grained particle representing a side chain. (Supplementary Note 28) The estimation system according to any one of Supplements 4 to 7, wherein the at least one processor generates the target molecular model in which all of the plurality of structural units constituting the target molecule are represented by the coarse-grained particle representing a main chain and the coarse-grained particle representing a side chain. (Supplementary Note 29) The estimation system according to any one of Supplementary Notes 14 to 17, wherein the at least one processor generates the target molecule model in the lattice system by arranging, in the lattice system, coarse-grained particles representing main chains and coarse-grained particles representing side chains adjacent to each other for at least half of the plurality of structural units constituting the target molecule. (Supplementary Note 30) The estimation system according to any one of Supplementary Notes 14 to 17, wherein the at least one processor generates the target molecule model in the lattice system by arranging, in the lattice system, coarse-grained particles representing main chains and coarse-grained particles representing side chains adjacent to each other for all of the plurality of structural units constituting the target molecule.
[0063] According to Supplementary Notes 1, 21, and 22, the interaction energy between the target molecule and the target molecule that bind to each other is calculated based on the positional relationship between these two molecules. The interaction energy within the target molecule is also calculated. The conformation of the target molecule is then estimated based on these two types of interaction energies. By considering these two types of interaction energies, the conformation of the target molecule that can bind to the target molecule can be estimated with greater accuracy. By accurately estimating this conformation, the binding mode of the target molecule to the target molecule can be estimated.
[0064] According to Supplementary Note 2, the intermolecular interaction energy is calculated based on the interaction energy between the structural units of the target molecule and the structural units of the target molecule, so that a more accurate intermolecular interaction energy can be obtained. This accurate calculation can further increase the accuracy of estimating the conformation of the target molecule.
[0065] According to Supplementary Note 3, the intramolecular interaction energy is calculated based on the interaction energy between the structural units of the target molecule, so that a more accurate intramolecular interaction energy can be obtained. This accurate calculation can further increase the accuracy of estimating the conformation of the target molecule.
[0066] According to Supplementary Note 4, the constituent units of both the target molecule and the target molecule are represented by coarse-grained particles. This coarse-graining simplifies the target molecule model and the target molecule model, making it easier to calculate the interaction energy and reducing the calculation time. As a result, the conformation of the target molecule can be estimated more quickly.
[0067] According to Supplementary Note 5, a target molecule model using coarse-grained particles is generated so as to reflect the arrangement of the constituent units of the target molecule. By using a target molecule model that more accurately reflects the structure of the target molecule, the accuracy of estimating the conformation of the target molecule can be further improved.
[0068] According to Supplementary Note 6, since at least a portion of the plurality of structural units of the target molecule is represented by a plurality of coarse-grained particles, a target molecular model that shows the structure of the target molecule in more detail is generated. By using this target molecular model, the conformation of the target molecule can be estimated in more detail.
[0069] According to Supplementary Note 7, at least a part of the multiple structural units of the target molecule is represented by two coarse-grained particles. By using this target molecule model, it is possible to achieve a balance between shortening the calculation time for the interaction energy and detailed estimation of the conformation.
[0070] According to Supplementary Note 8, at least a portion of the multiple structural units of a target molecule is represented by two coarse-grained particles representing the main chain and the side chain, thereby generating a target molecular model that represents the structure or function of the structural units. By using this target molecular model, the conformation of the target molecule can be predicted in more detail and more accurately.
[0071] According to Supplementary Note 9, a target molecule model using coarse-grained particles is generated so as to reflect the arrangement of the main chain and side chains. By using a target molecule model that more accurately reflects the structure of the target molecule, the accuracy of estimating the conformation of the target molecule can be further improved.
[0072] According to Supplementary Note 10, a plurality of target molecular models (conformations) are generated for a certain target molecule, and at least one conformation is selected based on the intermolecular interaction energy of each target molecular model. This search taking into account the intermolecular interaction energy makes it possible to more accurately estimate a highly plausible conformation from among multiple candidate conformations for the target molecule.
[0073] According to Supplementary Note 11, a plurality of target molecular models (conformations) are generated for a given target molecule, and at least one conformation is selected based on the total interaction energy of each target molecular model. This search, which takes into account both intermolecular interaction energy and intramolecular interaction energy, makes it possible to more accurately estimate a highly plausible conformation from among multiple candidate conformations for the target molecule.
[0074] According to Supplementary Note 12, a target molecular model is generated so as to satisfy the intermolecular constraints regarding the positional relationship between the target molecular model and the target molecular model, thereby obtaining a target molecular model that matches the behavior of the target molecule and the target molecule in the real world. By using this target molecular model, the conformation of the target molecule can be more accurately estimated.
[0075] According to Supplementary Note 13, a target molecule model is generated so as to avoid steric hindrance between the target molecule and the target molecule, thereby obtaining an appropriate target molecule model that does not cause molecular distortion or collisions between molecules. By using this target molecule model, the conformation of the target molecule can be more accurately predicted.
[0076] According to Supplementary Note 14, a target molecular model is generated in consideration of constraints within a lattice system set based on a target molecular model. The introduction of the lattice system simplifies the generation of the target molecular model, allowing the target molecular model to be generated at high speed.
[0077] According to Supplementary Note 15, the outer boundary of the lattice system is set based on the center of gravity of the active site of the target molecule. By setting the range of the lattice system in this manner, it is possible to efficiently generate a target molecule model that matches the target molecule and its behavior in the real world.
[0078] According to Supplementary Note 16, by setting the unit cell of the lattice system to a regular tetrahedral lattice, a cubic lattice, or a square lattice, it is possible to generate a target molecular model with high accuracy while simplifying the generation of the target molecular model.
[0079] According to Supplementary Note 17, a target molecule model is generated in a lattice system so as to reflect the arrangement of the constituent units of the target molecule. By using a target molecule model that more accurately reflects the structure of the target molecule, the accuracy of estimating the conformation of the target molecule can be further improved.
[0080] According to Supplementary Note 18, a target molecule model using coarse-grained particles is generated so as to reflect the arrangement of the main chain and side chains. By using a target molecule model that more accurately reflects the structure of the target molecule, the accuracy of estimating the conformation of the target molecule can be further improved.
[0081] According to Supplementary Note 19, a target molecular model is generated so as to satisfy intramolecular constraints or cyclization constraints regarding the positional relationships of multiple structural units within the target molecular model, thereby obtaining a target molecular model that matches the behavior of the target molecule in the real world. By using this target molecular model, the conformation of the target molecule can be more accurately estimated.
[0082] According to Supplementary Note 20, it is possible to more accurately predict the conformation of a peptide (target peptide) that can bind to a target protein. By accurately predicting the conformation, it is possible to predict the binding mode of the peptide to the target protein. Such prediction can contribute to the design of peptides that are expected to have improved drug efficacy.
[0083] According to Supplementary Note 23, at least half of the plurality of structural units of the target molecule are represented by a plurality of coarse-grained particles, so that a target molecular model showing the structure of the target molecule in more detail is generated. By using this target molecular model, the conformation of the target molecule can be estimated in more detail.
[0084] According to Supplementary Note 24, since all of the multiple structural units of the target molecule are represented by multiple coarse-grained particles, a target molecule model that shows the structure of the target molecule in more detail is generated. By using this target molecule model, the conformation of the target molecule can be estimated in more detail.
[0085] According to Supplementary Note 25, at least half of the multiple structural units of the target molecule are represented by two coarse-grained particles. By using this target molecule model, it is possible to achieve a balance between shortening the calculation time for the interaction energy and detailed estimation of the conformation.
[0086] According to Supplementary Note 26, all of the multiple structural units of a target molecule are represented by two coarse-grained particles. By using this target molecule model, it is possible to achieve a balance between shortening the calculation time for interaction energy and detailed estimation of the conformation.
[0087] According to Supplementary Note 27, at least half of the multiple structural units of a target molecule are represented by two coarse-grained particles representing the main chain and side chain, thereby generating a target molecular model that represents the structure or function of the structural units. By using this target molecular model, the conformation of the target molecule can be predicted in more detail and more accurately.
[0088] According to Supplementary Note 28, all of the multiple structural units of a target molecule are represented by two coarse-grained particles representing the main chain and side chain, thereby generating a target molecular model that shows the structure or function of the structural units. By using this target molecular model, the conformation of the target molecule can be predicted in more detail and more accurately.
[0089] According to Supplementary Note 29, a target molecule model using coarse-grained particles is generated so as to reflect the arrangement of the main chain and side chains. By using a target molecule model that more accurately reflects the structure of the target molecule, the accuracy of estimating the conformation of the target molecule can be further improved.
[0090] According to Supplementary Note 30, a target molecule model using coarse-grained particles is generated so as to reflect the arrangement of the main chain and side chains. By using a target molecule model that more accurately reflects the structure of the target molecule, the accuracy of estimating the conformation of the target molecule can be further improved.
[0091] Although examples of the technology disclosed herein will be described, the technology disclosed herein is not limited to these examples.
[0092] [Example 1] (Calculation target) Coordinate data of the crystal structure of a complex in which a cyclic peptide having splicing inhibitory activity is bound to the UHM (U2AF homology motif) domain of SPF45 (Splicing factor 45) protein is registered in the Protein Data Bank (PDB) (https: / / www.rcsb.org / ). The identifier (PDB ID) of the crystal structure in PDB is "5LSO". The crystal structure "5LSO" is composed of four polymers, with the A chain and B chain corresponding to the UHM domain of SPF45, and the C chain and D chain corresponding to the cyclic peptide. The cyclic peptide of the C chain is bound to the UHM domain of the A chain of the SPF45 protein, and the cyclic peptide of the D chain is bound to the UHM domain of the B chain of the SPF45 protein. The coordinate data of the complex model of A and C chains of the crystal structure "5LSO" was used for verification.
[0093] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the C chain of the crystal structure "5LSO" were extracted. The extracted amino acid residues were Leu309, Met312, Val313, Asp319, Asp321, Leu322, Glu325, Thr326, Glu329, Cys330, Leu372, Arg375, Tyr376, Phe377, Gly378, Gly379, Arg380, and Val382, a total of 18 amino acid residues. Peptide docking simulations were performed on these amino acid residues.
[0094] (Coarse-graining of target protein) Of the extracted amino acid residues, Val313, Thr326, Cys330, Gly378, Gly379, and Val382 were approximated by one coarse-grained particle, and the remaining amino acid residues were approximated by two coarse-grained particles. Finally, the target protein was approximated by a total of 30 coarse-grained particles.
[0095] When an amino acid residue is approximated by one coarse-grained particle, the coarse-grained particle is placed at the center of gravity of all heavy atoms constituting the amino acid residue. When an amino acid residue is approximated by two coarse-grained particles, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain is placed at the center of gravity of the heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms constituting the amino acid residue. The coarse-grained particle corresponding to the side chain is placed at the center of gravity of the heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms constituting the amino acid residue.
[0096] (Target peptide) The target peptide had an amino acid sequence of Lys-Ser-Arg-Trp-Asp-Glu. The peptide was cyclized by forming an amide bond between the amine in the side chain of Lys1 and the carboxyl group in the side chain of Glu6.
[0097] (Coarse-graining of target peptide) Of the amino acid residues constituting the target peptide, Ser2 was approximated by one coarse-grained particle, and the remaining amino acid residues were approximated by two coarse-grained particles. Finally, the target peptide was approximated by a total of 11 coarse-grained particles.
[0098] For Ser2, coarse-grained particles were placed at the positions of the centers of gravity of all the heavy atoms that constitute the amino acid residues. For amino acid residues other than Ser2, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms that constitute the amino acid residues. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms that constitute the amino acid residues.
[0099] (Scaling and Relocation of Coarse-Grained Model) In the list of Cartesian coordinates indicating the positions of all coarse-grained particles constituting the target protein, the maximum and minimum values in the x-axis direction are called max_X and min_X, respectively. Similarly, in the list, the maximum and minimum values in the y-axis direction are called max_Y and min_Y, respectively, and the maximum and minimum values in the z-axis direction are called max_Z and min_Z, respectively. Assuming that the Cartesian coordinates indicating the position of an arbitrary coarse-grained particle are represented as (X, Y, Z), this coarse-grained particle was relocated to the position represented by the following coordinates (4).
[0100] This operation was applied to all of the coarse-grained particles constituting the target protein and all of the coarse-grained particles constituting the target peptide, and the coarse-grained models were rearranged for each of the target protein and the target peptide.
[0101] In the list of Cartesian coordinates indicating the positions of the coarse-grained particles constituting the target protein after rearrangement, the value obtained by rounding up the decimal point of the minimum value in the x-axis direction in the negative direction is called min_X'. Similarly, in the list, the value obtained by rounding up the decimal point of the minimum value in the y-axis direction in the negative direction is called min_Y', and the value obtained by rounding up the decimal point of the minimum value in the z-axis direction in the negative direction is called min_Z'. If the Cartesian coordinates indicating the position of any coarse-grained particle after rearrangement are represented as (X', Y', Z'), the coarse-grained particle was further rearranged to the position (X'-min_X', Y'-min_Y', Z'-min_Z'). This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0102] (Translation and rotation of the coarse-grained model) The Cartesian coordinates indicating the position of the i-th coarse-grained particle constituting the target protein and the target peptide are (x i , y i , z i The Cartesian coordinates indicating the position of any coarse-grained particle are represented as (X, Y, Z), and the total number of coarse-grained particles constituting the target protein and the target peptide is represented as N. The coarse-grained particles were placed at the position indicated by the following coordinates (5).
[0103] The coarse-grained particle located at coordinate (5) was rotated by 1.571 rad around the z-axis, and then the coarse-grained particle was translated by +2.533 in the x-axis direction, +2.533 in the y-axis direction, and +2.533 in the z-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide. This process performed translation and rotation of the coarse-grained model for each of the target protein and target peptide.
[0104] The Cartesian coordinates indicating the position of any coarse-grained particle that was translated and rotated were expressed as (X', Y', Z'), and the coarse-grained particle was relocated to the position of the following coordinates (6). This operation was applied to all coarse-grained particles that make up the target protein and all coarse-grained particles that make up the target peptide, and the coarse-grained models of the target protein and the target peptide were relocated.
[0105] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (5,5,4), (6,6,5), (6,4,3), (4,6,3), (4,4,5), (5,7,6), (7,5,6), (7,7,4), (5,3,2), (7,3,4), (7,5,2), (3,5,2), (3,7,4), (5,7,2), (3,3,4), (3,5,6), (5,3,6), (6,8,7), (4,8,5), (4,6,7), (5,9,8), (7,7,8), (7,9,6), (3,9,6), (5,9,4), (3,7,8), (5,5,8), (8,6,7), (8,4,5), (6,4,7), (9,5,8), (9,7,6), (9,3,6), (9,54), (7,3,8), (8,8,5), (8,6,3), (6,8,3), (9,9,4), (9,7,2), (7,9,2), (6,2,1), (4,4,1), (4,2,3), (5,1,0), (7,1,2), (7,3,0), (3,3,0), (5,5,0), (3,1,2), (5,1,4), (8,2,3), (6,2,5), (9,1,4), (9,3,2), (7,1,6), (8,4,1), (6,6,1), (9,5,0), (7,7,0), (2,6,1), (2,4,3), (1,5,0), (1,7,2), (3,7,0), (1,3,2), (1,5,4), (2,8,3), (2,6,5), (1,9,4), (3,9,2), (1,7,6), (4,8,1), (5,9,0), (2,2,5), (1,1,4), (1,3,6), (3,1,6), (2,4,7), (1,5,8), (3,3,8), (4,2,7), (5,1,8), (6,10,9), (4,10,7), (4,8,9), (5,11,10), (7,9,10), (7,11,8), (3,11,8), (5,11,6), (3,9,10), (5,7,10), (8,8,9), (6,6,9), (9,7,10), (9,9,8), (7,510), (8,10,7), (6,10,5), (9,11,6), (7,11,4), (2,10,5), (2,8,7), (1,11,6), (3,11,4), (1,9,8), (4,10,3), (5,11,2), (2,6,9), (1,7,10), (3,5,10), (4,4,9), (5,3,10), (10,6,9), (10,4,7), (8,4,9), (11,5,10), (11,7,8), (11,3,8), (11,5,6), (9,3,10), (10,8,7), (10,6,5), (11,9,6), (11,7,4), (10,2,5), (8,2,7), (11,1,6), (11,3,4), (9,1,8), (10,4,3), (11,5,2), (6,2,9), (7,1,10), (10,10,5), (10,8,3), (8,10,3), (11,11,4), (11,9,2), (9,11,2), (10,6,1), (8,8,1), (11,7,0), (9,9,0), (6,10,1), (7,11,0), (6,0,-1), (4,2,-1), (4,0,1), (5,-1,-2), (7,-1,0), (7,1,-2), (3,1,-2), (5,3,-2), (3,-1,0), (5,-1,2), (8,0,1), (6,0,3), (9,-1,2), (9,1,0), (7,-1,4), (8,2,-1), (6,4,-1), (9,3,-2), (7,5,-2), (2,4,-1), (2,2,1), (1,3,-2), (3,5,-2), (1,1,0), (4,6,-1), (5,7,-2), (2,0,3), (1,-1,2), (3,-1,4), (4,0,5), (5,-1,6), (10,0,3), (8,0,5), (11,-1,4), (11,1,2), (9,-1,6), (10,2,1), (11,3,0), (6,0,7), (7,-1,8), (10,4,-1), (8,6,-1), (11,5,-2), (9,7,-2), (6,8,-1), (7,9,-2), (0,6,-1), (0,4,1), (-1,5,-2), (-1,7,0), (1,7,-2), (-1,3,0), (-1,5,2), (0,8,1), (0,6,3), (-1,9,2), (1,9,0), (-1,7,4),(2,8,-1), (3,9,-2), (0,2,3), (-1,1,2), (-1,3,4), (0,4,5), (-1,5,6), (0,10,3), (0,8,5), (-1,11,4), (1,11,2), (-1,9,6), (2,10,1), (3,11,0), (0,6,7), (-1,7,8), (4,10,-1), (5,11,-2), (0,0,5), (-1,-1,4), (-1,1,6), (1,-1,6), (0,2,7), (-1,3,8), (1,1,8), (2,0,7), (3,-1,8), (0,4,9), (-1,5,10), (1,3,10), (2,2,9), (3,1,10), (4,0,9), (5,-1,10),
[0106] (Objective Function) The total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. In formulas (1) and (2), the value indicated by the MJ matrix was used as the inter-unit interaction energy, and the variable a ij , b ik The threshold value Td for determining the above is set as follows:
[0107] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the side chain of Lys1 and the coarse-grained particle corresponding to the side chain of Glu6 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0108] (Mathematical Optimization) To perform mathematical optimization including calculation of interaction energy, OR-Tools version 9.5.2237 (https: / / pypi.org / project / ortools / 9.5.2237 / ), a constraint programming method, was used. Using OR-Tools, a coarse-grained model in which the target protein was rearranged by translation and rotation was input. Then, the objective variable E was calculated within the range that satisfied the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was finally obtained as the docking model. total The minimum value of was −193.44.
[0109] (Verification) The degree of agreement between the obtained docking model and a coarse-grained model of the C chain of the crystal structure "5LSO" was verified. As an index of the verification, the root mean square deviation (RMSD) shown in the following formula (7) was calculated. Here, the variable d irepresents the Euclidean distance between the position of the i-th coarse-grained particle in the docking model and the position of the i-th coarse-grained particle in the crystal structure rearranged by translation and rotation. N represents the total number of coarse-grained particles constituting the target peptide. As a result, a docking model with an RMSD of 2.87 Å from the rearranged crystal structure was obtained.
[0110] The docking model and rearranged crystal structure in Example 1 are shown in Figure 7. The shaded circles represent coarse-grained particles of the target protein. The black circles represent coarse-grained particles corresponding to the main chain of the target peptide, and the white circles represent coarse-grained particles corresponding to the side chain of the target peptide.
[0111] Example 2 (Calculation Object) Coordinate data for the crystal structure of a complex between HIV-1 (human immunodeficiency virus type 1) integrase protein and a cyclic peptide has been registered in the PDB. The identifier for this crystal structure in the PDB is "3WNE." The crystal structure "3WNE" is composed of four polymers, with chains A and B corresponding to the HIV-1 integrase protein and chains C and D corresponding to the cyclic peptide. Both the cyclic peptide of chain C and the cyclic peptide of chain D are bound to an HIV-1 integrase protein dimer composed of chains A and B. The coordinate data for the complex model of chains A, B, and C of the crystal structure "3WNE" was used for verification.
[0112] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the C chain of the crystal structure "3WNE" were extracted. The extracted amino acid residues were Asp167, Gln168, Ala169, Glu170, His171, Thr174, and Met178 for the A chain, and Gln95, Leu102, Thr125, Ala128, Ala129, and Trp132 for the B chain, for a total of 13 amino acid residues. A docking simulation of the cyclic peptide was performed for these amino acid residues.
[0113] (Coarse-graining of target protein) Each extracted amino acid residue was approximated by two coarse-grained particles. One coarse-grained particle corresponded to the main chain, and the other corresponded to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned atomic names such as CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned atomic names other than CA, C, N, and O. Finally, the target protein was approximated by a total of 26 coarse-grained particles.
[0114] (Target peptide) The amino acid sequence of the target peptide was Pro-Lys-Ile-Asp-Asn-Gly. The peptide was cyclized by forming an amide bond between the amine in the main chain of Pro1 and the carboxyl group in the main chain of Gly6.
[0115] (Coarse-graining of target peptide) Of the amino acid residues constituting the target peptide, Gly6 was approximated by one coarse-grained particle, and the remaining amino acid residues were approximated by two coarse-grained particles. Finally, the target peptide was approximated by a total of 11 coarse-grained particles.
[0116] For Gly6, a coarse-grained particle model was placed at the center of gravity of all the heavy atoms that constitute the amino acid residue. For amino acid residues other than Gly6, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain was placed at the center of gravity of the heavy atoms that constitute the amino acid residue and are assigned atomic names such as CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the center of gravity of the heavy atoms that constitute the amino acid residue and are assigned atomic names other than CA, C, N, and O.
[0117] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0118] (Translation and rotation of the coarse-grained model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. Then, the coarse-grained particle placed at coordinate (5) was rotated by 3.142 rad around the x-axis and by 4.712 rad around the z-axis. The coarse-grained particle was then translated by +1.267 in the y-axis direction and +2.533 in the z-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0119] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 147 lattice points located at the following orthogonal coordinates. (4,5,3), (5,6,4), (5,4,2), (3,6,2), (3,4,4), (4,7,5), (6,5,5), (6,7,3), (5,8,6), (3,8,4), (3,6,6), (7,6,6), (7,4,4), (5,4,6), (7,8,4), (7,6,2), (5,8,2), (4,3,1), (6,3,3), (6,5,1), (5,2,0), (3,4,0), (3,2,2), (7,2,2), (5,2,4), (7,4,0), (5,6,0), (2,5,1), (2,7,3), (4,7,1), (1,6,0), (1,4,2), (1,8,2), (1,6,4), (3,8,0), (2,3,3), (2,5,5), (4,3,5), (1,2,4), (1,4,6), (3,2,6), (4,9,7), (6,7,7), (6,9,5), (5,10,8), (3,10,6), (3,8,8), (7,8,8), (5,6,8), (7,10,6), (5,10,4), (2,9,5), (4,9,3), (1,10,4), (1,8,6), (3,10,2), (2,7,7), (4,5,7), (1,6,8), (3,4,8), (8,5,7), (8,7,5), (9,6,8), (9,4,6), (7,4,8), (9,8,6), (9,6,4), (8,3,5), (8,5,3), (9,2,4), (7,2,6), (9,4,2), (6,3,7), (5,2,8), (8,9,3), (9,10,4), (9,8,2), (7,10,2), (8,7,1), (9,6,0), (7,8,0), (6,9,1), (5,10,0), (4,1,-1), (6,1,1), (6,3,-1), (5,0,-2), (3,2,-2), (3,0,0), (7,0,0), (5,0,2), (7,2,-2), (5,4,-2), (2,3,-1), (4,5,-1), (1,4,-2), (1,2,0), (3,6,-2), (2,1,1), (4,1,3), (1,0,2), (3,0,4), (8,1,3), (8,3,1), (9,0,2), (7,0,4), (9,2,0), (6,1,5), (5,0,6), (8,5,-1), (9,4,-2), (7,6,-2), (6,7,-1), (5,8,-2), (0,5,-1), (0,7,1), (2,7,-1), (-1,6,-2), (-1,4,0), (-1,8,0), (-1,6,2), (1,8,-2), (0,3,1), (0,5,3), (-1,2,2), (-1,4,4), (0,9,3), (2,9,1), (-1,10,2), (-1,8,4), (1,10,0), (0,7,5), (-1,6,6), (4,9,-1), (3,10,-2), (0,1,3), (0,3,5), (2,1,5), (-1,0,4), (-1,2,6), (1,0,6), (0,5,7), (2,3,7), (-1,4,8), (1,2,8), (4,1,7), (3,0,8),
[0120] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0121] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set. (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedral lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the main chain of Pro1 and the coarse-grained particle corresponding to Gly6 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0122] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value of was −166.61.
[0123] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the C chain of the crystal structure "3WNE" was verified. As a result, a docking model with an RMSD of 3.95 Å from the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 2 are shown in Figure 7. The representation of each circle is the same as in Example 1.
[0124] [Example 3] (Calculation target) Coordinate data for the crystal structure of a complex between an angiotensin II peptide and the Fab (fragment antigen-binding) region of an antibody that recognizes it is registered in the PDB. The identifier for this crystal structure in the PDB is "2CK0." The crystal structure "2CK0" is composed of three polymers, with the L chain and H chain corresponding to the Fab region of the antibody and the P chain corresponding to the angiotensin II peptide. The coordinate data for the complex model of the L chain, H chain, and P chain of the crystal structure "2CK0" was used for verification.
[0125] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the P chain of the crystal structure "2CK0" were extracted. The extracted amino acid residues were His31, Tyr32, Ser91, Tyr92, Asn93, Leu94, and Tyr95 for the L chain, and Asn35, Trp47, Arg50, Arg52, Gly52C, Phe53, Asn54, Ala56, Tyr58, Asp100, and Gly101 for the H chain, for a total of 18 amino acid residues. Peptide docking simulations were performed on these amino acid residues.
[0126] (Coarse-graining of target protein) Of the extracted amino acid residues, Ser91 of the L chain and Gly52C, Ala56, and Gly101 of the H chain were approximated by one coarse-grained particle, and the remaining amino acid residues were approximated by two coarse-grained particles. Finally, the target protein was approximated by a total of 32 coarse-grained particles.
[0127] When an amino acid residue is approximated by one coarse-grained particle, the coarse-grained particle is placed at the center of gravity of all heavy atoms constituting the amino acid residue. When an amino acid residue is approximated by two coarse-grained particles, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain is placed at the center of gravity of the heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms constituting the amino acid residue. The coarse-grained particle corresponding to the side chain is placed at the center of gravity of the heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms constituting the amino acid residue.
[0128] (Target peptide) The amino acid sequence of the target peptide was Cys-Lys-Glu-Trp-Leu-Ser-Thr-Ala-Pro-Cys-Gly. The peptide was cyclized by forming a disulfide bond between the side chains of Cys1 and Cys10.
[0129] (Coarse-graining of target peptide) Of the amino acid residues constituting the target peptide, Cys1, Ser6, Thr7, Ala8, Pro9, Cys10, and Gly11 were approximated by one coarse-grained particle, and the remaining amino acid residues were approximated by two coarse-grained particles. Finally, the target peptide was approximated by a total of 15 coarse-grained particles.
[0130] For Cys1, Ser6, Thr7, Ala8, Pro9, Cys10, and Gly11, coarse-grained particle models were placed at the center of gravity of all constituent heavy atoms. For amino acid residues other than Cys1, Ser6, Thr7, Ala8, Pro9, Cys10, and Gly11, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain was placed at the center of gravity of the heavy atoms that make up the amino acid residue and are assigned atomic names such as CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the center of gravity of the heavy atoms that make up the amino acid residue and are assigned atomic names other than CA, C, N, and O.
[0131] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0132] (Translation and Rotation of Coarse-Grained Model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. The coarse-grained particle placed at coordinate (5) was then rotated by 3.142 rad around the x-axis and 3.142 rad around the z-axis. The coarse-grained particle was then translated by +2.533 in the x-axis direction and +1.267 in the y-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0133] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (5,5,5), (6,6,6), (6,4,4), (4,6,4), (4,4,6), (5,7,7), (7,5,7), (7,7,5), (5,3,3), (7,3,5), (7,5,3), (3,5,3), (3,7,5), (5,7,3), (3,3,5), (3,5,7), (5,3,7), (6,8,8), (4,8,6), (4,6,8), (5,9,9), (7,7,9), (7,9,7), (3,9,7), (5,9,5), (3,7,9), (5,5,9), (8,6,8), (8,4,6), (6,4,8), (9,5,9), (9,7,7), (9,3,7), (9,55), (7,3,9), (8,8,6), (8,6,4), (6,8,4), (9,9,5), (9,7,3), (7,9,3), (6,2,2), (4,4,2), (4,2,4), (5,1,1), (7,1,3), (7,3,1), (3,3,1), (5,5,1), (3,1,3), (5,1,5), (8,2,4), (6,2,6), (9,1,5), (9,3,3), (7,1,7), (8,4,2), (6,6,2), (9,5,1), (7,7,1), (2,6,2), (2,4,4), (1,5,1), (1,7,3), (3,7,1), (1,3,3), (1,5,5), (2,8,4), (2,6,6), (19,5), (3,9,3), (1,7,7), (4,8,2), (5,9,1), (2,2,6), (1,1,5), (1,3,7), (3,1,7), (2,4,8), (1,5,9), (3,3,9), (4,2,8), (5,1,9), (6,10,10), (4,10,8), (4,8,10), (5,11,11), (7,9,11), (7,11,9), (3,11,9), (5,11,7), (3,9,11), (5,7,11), (8,8,10), (6,6,10), (9,7,11), (9,9,9), (7,5,11), (8,10,8), (6,10,6), (9,11,7), (7,11,5), (2,10,6),(2,8,8), (1,11,7), (3,11,5), (1,9,9), (4,10,4), (5,11,3), (2,6,10), (1,7,11), (3,5,11), (4,4,10), (5,3,11), (10,6,10), (10,4,8), (8,4,10), (11,5,11), (11,7,9), (11,3,9), (11,5,7), (9,3,11), (10,8,8), (10,6,6), (11,9,7), (11,7,5), (10,2,6), (8,2,8), (11,1,7), (11,3,5), (9,1,9), (10,4,4), (11,5,3), (6,2,10), (7,1,11), (10,10,6), (10,8,4), (8,10,4), (11,11,5), (11,9,3), (9,11,3), (10,6,2), (8,8,2), (11,7,1), (9,9,1), (6,10,2), (7,11,1), (6,0,0), (4,2,0), (4,0,2), (5,-1,-1), (7,-1,1), (7,1,-1), (3,1,-1), (5,3,-1), (3,-1,1), (5,-1,3), (8,0,2), (6,0,4), (9,-1,3), (9,1,1), (7,-1,5), (8,2,0), (6,4,0), (9,3,-1), (7,5,-1), (2,4,0), (2,2,2), (1,3,-1), (3,5,-1), (1,1,1), (4,6,0), (5,7,-1), (2,0,4), (1,-1,3), (3,-1,5), (4,0,6), (5,-1,7), (10,0,4), (8,0,6), (11,-1,5), (11,1,3), (9,-1,7), (10,2,2), (11,3,1), (6,0,8), (7,-1,9), (10,4,0), (8,6,0), (11,5,-1), (9,7,-1), (6,8,0), (7,9,-1), (0,6,0), (0,4,2), (-1,5,-1), (-1,7,1), (1,7,-1), (-1,3,1), (-1,5,3), (0,8,2), (0,6,4), (-1,9,3), (1,9,1), (-1,7,5),(2,8,0), (3,9,-1), (0,2,4), (-1,1,3), (-1,3,5), (0,4,6), (-1,5,7), (0,10,4), (0,8,6), (-1,11,5), (1,11,3), (-1,9,7), (2,10,2), (3,11,1), (0,6,8), (-1,7,9), (4,10,0), (5,11,-1), (0,0,6), (-1,-1,5), (-1,1,7), (1,-1,7), (0,2,8), (-1,3,9), (1,1,9), (2,0,8), (3,-1,9), (0,4,10), (-1,5,11), (1,3,11), (2,2,10), (3,1,11), (4,0,10), (5,-1,11),
[0134] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0135] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to Cys1 and the coarse-grained particle corresponding to Cys10 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0136] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value was −307.34.
[0137] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the P chain of the crystal structure "2CK0" was verified. As a result, a docking model with an RMSD of 5.45 Å from the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 3 are shown in Figure 7. The representation of each circle is the same as in Example 1.
[0138] [Example 4] (Calculation target) Coordinate data of the crystal structure of a complex in which an oxytocin peptide is bound to a neurophysin protein is registered in the PDB. The identifier of this crystal structure in the PDB is "1NPO." The crystal structure "1NPO" is composed of four polymers, with the A and C chains corresponding to neurophysin, and the B and D chains corresponding to the oxytocin peptide. The oxytocin peptide of the B chain is bound to the neurophysin of the A chain, and the oxytocin of the D chain is bound to the neurophysin of the C chain. The coordinate data of the complex model of the A and B chains of the crystal structure "1NPO" was used for verification.
[0139] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the B chain of the crystal structure "1NPO" were extracted. The extracted amino acid residues were Leu5, Leu7, Arg8, Cys10, Cys21, Phe22, Gly23, Pro24, Cys44, Gln45, Glu47, Asn48, Leu50, Pro51, Ser52, Pro53, Cys54, Gln55, Gly65, Asn75, and Asp76, a total of 21 amino acid residues. Peptide docking simulations were performed on these amino acid residues.
[0140] (Coarse-graining of target protein) Of the extracted amino acid residues, Cys10, Cys21, Gly23, Pro24, Cys44, Pro51, Ser52, Pro53, Cys54, and Gly65 were approximated with one coarse-grained particle, and the remaining amino acid residues were approximated with two coarse-grained particles. Finally, the target protein was approximated with a total of 32 coarse-grained particles.
[0141] When an amino acid residue is approximated by one coarse-grained particle, the coarse-grained particle is arranged at the position of the center of gravity of all heavy atoms constituting the amino acid residue. When an amino acid residue is approximated by two coarse-grained particles, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain is arranged at the position of the center of gravity of a heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms constituting the amino acid residue. The coarse-grained particle corresponding to the side chain is arranged at the position of the center of gravity of a heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms constituting the amino acid residue.
[0142] (Target peptide) The amino acid sequence of the target peptide was Cys-Tyr-Ile-Gln-Asn-Cys-Pro-Leu-Gly. The peptide was cyclized by the formation of a disulfide bond between Cys1 and Cys6.
[0143] (Coarse-graining of target peptide) Of the amino acid residues constituting the target peptide, Cys1, Cys6, Pro7, and Gly9 were approximated by one coarse-grained particle, and the remaining amino acid residues were approximated by two coarse-grained particles. Finally, the target peptide was approximated by a total of 14 coarse-grained particles.
[0144] When an amino acid residue is approximated by one coarse-grained particle, the coarse-grained particle model is placed at the center of gravity of all the heavy atoms that constitute the amino acid residue. When an amino acid residue is approximated by two coarse-grained particles, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain is placed at the center of gravity of the heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms that constitute the amino acid residue. The coarse-grained particle corresponding to the side chain is placed at the center of gravity of the heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms that constitute the amino acid residue.
[0145] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0146] (Translation and rotation of the coarse-grained model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. Then, the coarse-grained particle placed at coordinate (5) was rotated by 1.047 rad around the x-axis and by 4.189 rad around the z-axis. The coarse-grained particle was then translated by +1.267 in the x-axis direction, +2.533 in the y-axis direction, and +1.267 in the z-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0147] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (6,4,5), (7,5,6), (7,3,4), (5,5,4), (5,3,6), (6,6,7), (8,4,7), (8,6,5), (6,2,3), (8,2,5), (8,4,3), (4,4,3), (4,6,5), (6,6,3), (4,2,5), (4,4,7), (6,2,7), (7,7,8), (5,7,6), (5,5,8), (6,8,9), (8,6,9), (8,8,7), (4,8,7), (6,8,5), (4,6,9), (6,4,9), (9,5,8), (9,3,6), (7,3,8), (10,4,9), (10,6,7), (10,2,7), (10,4,5), (8,2,9), (9,7,6), (9,5,4), (7,7,4), (10,8,5), (10,6,3), (8,8,3), (7,1,2), (5,3,2), (5,1,4), (6,0,1), (8,0,3), (8,2,1), (4,2,1), (6,4,1), (4,0,3), (6,0,5), (9,1,4), (7,1,6), (10,0,5), (10,2,3), (8,0,7), (9,3,2), (7,5,2), (10,4,1), (8,6,1), (3,5,2), (3,3,4), (2,4,1), (2,6,3), (4,6,1), (2,2,3), (2,4,5), (3,7,4), (3,5,6), (2,8,5), (4,8,3), (2,6,7), (5,7,2), (6,8,1), (3,1,6), (2,0,5), (2,2,7), (4,0,7), (3,3,8), (2,4,9), (4,2,9), (5,1,8), (6,0,9), (7,9,10), (5,9,8), (5,7,10), (6,10,11), (8,8,11), (8,10,9), (4,10,9), (6,10,7), (4,8,11), (6,6,11), (9,7,10), (7,5,10), (10,6,11), (10,8,9), (8,4,11), (9,9,8), (7,9,6), (10,10,7), (8,10,5), (3,9,6), (3,7,8), (2,10,7), (4,10,5), (2,8,9), (5,9,4), (6,10,3), (3,5,10), (2,6,11), (4,4,11), (5,3,10), (6,2,11), (11,5,10), (11,3,8), (9,3,10), (12,4,11), (12,6,9), (12,2,9), (12,4,7), (10,2,11), (11,7,8), (11,5,6), (12,8,7), (12,6,5), (11,1,6), (9,1,8), (12,0,7), (12,2,5), (10,0,9), (11,3,4), (12,4,3), (7,1,10), (8,0,11), (11,9,6), (11,7,4), (9,9,4), (12,10,5), (12,8,3), (10,10,3), (11,5,2), (9,7,2), (12,6,1), (10,8,1), (7,9,2), (8,10,1), (7,-1,0), (5,1,0), (5,-1,2), (6,-2,-1), (8,-2,1), (8,0,-1), (4,0,-1), (6,2,-1), (4,-2,1), (6,-2,3), (9,-1,2), (7,-1,4), (10,-2,3), (10,0,1), (8,-2,5), (9,1,0), (7,3,0), (10,2,-1), (8,4,-1), (3,3,0), (3,1,2), (2,2,-1), (4,4,-1), (2,0,1), (5,5,0), (6,6,-1), (3,-1,4), (2,-2,3), (4,-2,5), (5,-1,6), (6,-2,7), (11,-1,4), (9,-1,6), (12,-2,5), (12,0,3), (10,-2,7), (11,1,2), (12,2,1), (7,-1,8), (8,-2,9), (11,3,0), (9,5,0), (12,4,-1), (10,6,-1), (7,7,0), (8,8,-1), (1,5,0), (1,3,2), (0,4,-1), (0,6,1), (2,6,-1), (0,2,1), (0,4,3), (1,7,2), (1,5,4), (0,8,3), (2,8,1), (0,6,5), (3,7,0), (4,8,-1), (1,1,4), (0,0,3), (0,2,5), (1,3,6), (0,4,7), (1,9,4), (1,7,6), (0,10,5), (2,10,3), (0,8,7), (3,9,2), (4,10,1), (1,5,8), (0,6,9), (5,9,0), (6,10,-1), (1,-1,6), (0,-2,5), (0,0,7), (2,-2,7), (1,1,8), (0,2,9), (2,0,9), (3,-1,8), (4,-2,9), (1,3,10), (0,4,11), (2,2,11), (3,1,10), (4,0,11), (5,-1,10), (6,-2,11),
[0148] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0149] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the main chain of Cys1 and the coarse-grained particle corresponding to the main chain of Cys6 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0150] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value of was −313.45.
[0151] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the B chain of the crystal structure "1NPO" was verified. As a result, a docking model with an RMSD of 4.40 Å with the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 4 are shown in Figure 7. The representation of each circle is the same as in Example 1.
[0152] Example 5 (Calculation Object) Coordinate data for the crystal structure of a complex between HIV-1 (human immunodeficiency virus type 1) integrase protein and a cyclic peptide has been registered in the PDB. The identifier for this crystal structure in the PDB is "3AV9." The crystal structure "3AV9" is composed of four polymers, with the A and B chains corresponding to the HIV-1 integrase protein and the X and Y chains corresponding to the cyclic peptide. The X chain cyclic peptide is bound to an HIV-1 integrase protein dimer composed of the A and B chains, and the Y chain cyclic peptide is also bound to an HIV-1 integrase protein dimer composed of the A and B chains. The coordinate data for the complex model of the A, B, and Y chains of the crystal structure "3AV9" was used for verification.
[0153] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the Y chain of the crystal structure "3AV9" were extracted. The extracted amino acid residues were Gln95, Thr124, Thr125, Ala128, Ala129, Trp131, and Trp132 for the A chain, and Asp167, Gln168, Ala169, Glu170, His171, Thr174, and Met178 for the B chain, for a total of 14 amino acid residues. A docking simulation of the cyclic peptide was performed for these amino acid residues.
[0154] (Coarse-graining of target protein) The extracted amino acid residues were approximated with two coarse-grained particles. One coarse-grained particle corresponded to the main chain, and the other corresponded to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned the atomic names CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned the atomic names other than CA, C, N, and O. Finally, the target protein was approximated with a total of 28 coarse-grained particles.
[0155] (Target peptide) The amino acid sequence of the target peptide was Ser-Ala-Lys-Ile-Asp-Asn-Leu-Asp. The peptide was cyclized by forming an amide bond between the amine in the main chain of Ser1 and the carboxyl group in the main chain of Asp8.
[0156] (Coarse-graining of target peptide) The amino acid residues constituting the target peptide were approximated by two coarse-grained particles. One coarse-grained particle corresponded to the main chain, and the other corresponded to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atoms constituting the amino acid residue that were assigned the atomic names CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atoms constituting the amino acid residue that were assigned the atomic names other than CA, C, N, and O. Finally, the target peptide was approximated by a total of 16 coarse-grained particles.
[0157] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0158] (Translation and rotation of the coarse-grained model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. Then, the coarse-grained particle placed at coordinate (5) was rotated by 4.712 rad around the x-axis and by 3.142 rad around the z-axis. The coarse-grained particle was then translated by +1.267 in the x-axis direction and +2.533 in the y-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0159] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (5,5,5), (6,6,6), (6,4,4), (4,6,4), (4,4,6), (5,7,7), (7,5,7), (7,7,5), (5,3,3), (7,3,5), (7,5,3), (3,5,3), (3,7,5), (5,7,3), (3,3,5), (3,5,7), (5,3,7), (6,8,8), (4,8,6), (4,6,8), (5,9,9), (7,7,9), (7,9,7), (3,9,7), (5,9,5), (3,7,9), (5,5,9), (8,6,8), (8,4,6), (6,4,8), (9,5,9), (9,7,7), (9,3,7), (9,55), (7,3,9), (8,8,6), (8,6,4), (6,8,4), (9,9,5), (9,7,3), (7,9,3), (6,2,2), (4,4,2), (4,2,4), (5,1,1), (7,1,3), (7,3,1), (3,3,1), (5,5,1), (3,1,3), (5,1,5), (8,2,4), (6,2,6), (9,1,5), (9,3,3), (7,1,7), (8,4,2), (6,6,2), (9,5,1), (7,7,1), (2,6,2), (2,4,4), (1,5,1), (1,7,3), (3,7,1), (1,3,3), (1,5,5), (2,8,4), (2,6,6), (19,5), (3,9,3), (1,7,7), (4,8,2), (5,9,1), (2,2,6), (1,1,5), (1,3,7), (3,1,7), (2,4,8), (1,5,9), (3,3,9), (4,2,8), (5,1,9), (6,10,10), (4,10,8), (4,8,10), (5,11,11), (7,9,11), (7,11,9), (3,11,9), (5,11,7), (3,9,11), (5,7,11), (8,8,10), (6,6,10), (9,7,11), (9,9,9), (7,5,11), (8,10,8), (6,10,6), (9,11,7), (7,11,5), (2,10,6),(2,8,8), (1,11,7), (3,11,5), (1,9,9), (4,10,4), (5,11,3), (2,6,10), (1,7,11), (3,5,11), (4,4,10), (5,3,11), (10,6,10), (10,4,8), (8,4,10), (11,5,11), (11,7,9), (11,3,9), (11,5,7), (9,3,11), (10,8,8), (10,6,6), (11,9,7), (11,7,5), (10,2,6), (8,2,8), (11,1,7), (11,3,5), (9,1,9), (10,4,4), (11,5,3), (6,2,10), (7,1,11), (10,10,6), (10,8,4), (8,10,4), (11,11,5), (11,9,3), (9,11,3), (10,6,2), (8,8,2), (11,7,1), (9,9,1), (6,10,2), (7,11,1), (6,0,0), (4,2,0), (4,0,2), (5,-1,-1), (7,-1,1), (7,1,-1), (3,1,-1), (5,3,-1), (3,-1,1), (5,-1,3), (8,0,2), (6,0,4), (9,-1,3), (9,1,1), (7,-1,5), (8,2,0), (6,4,0), (9,3,-1), (7,5,-1), (2,4,0), (2,2,2), (1,3,-1), (3,5,-1), (1,1,1), (4,6,0), (5,7,-1), (2,0,4), (1,-1,3), (3,-1,5), (4,0,6), (5,-1,7), (10,0,4), (8,0,6), (11,-1,5), (11,1,3), (9,-1,7), (10,2,2), (11,3,1), (6,0,8), (7,-1,9), (10,4,0), (8,6,0), (11,5,-1), (9,7,-1), (6,8,0), (7,9,-1), (0,6,0), (0,4,2), (-1,5,-1), (-1,7,1), (1,7,-1), (-1,3,1), (-1,5,3), (0,8,2), (0,6,4), (-1,9,3), (1,9,1), (-1,7,5),(2,8,0), (3,9,-1), (0,2,4), (-1,1,3), (-1,3,5), (0,4,6), (-1,5,7), (0,10,4), (0,8,6), (-1,11,5), (1,11,3), (-1,9,7), (2,10,2), (3,11,1), (0,6,8), (-1,7,9), (4,10,0), (5,11,-1), (0,0,6), (-1,-1,5), (-1,1,7), (1,-1,7), (0,2,8), (-1,3,9), (1,1,9), (2,0,8), (3,-1,9), (0,4,10), (-1,5,11), (1,3,11), (2,2,10), (3,1,11), (4,0,10), (5,-1,11),
[0160] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0161] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the main chain of Ser1 and the coarse-grained particle corresponding to the main chain of Asp8 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0162] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value of was −219.25.
[0163] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the Y chain of the crystal structure "3AV9" was verified. As a result, a docking model with an RMSD of 4.08 Å with the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 5 are shown in Figure 8. The representation of each circle is the same as in Example 1.
[0164] Example 6 (Calculation Object) Coordinate data for the crystal structure of a complex between HIV-1 (human immunodeficiency virus type 1) integrase protein and a cyclic peptide has been registered in the PDB. The identifier for this crystal structure in the PDB is "3AVA." The crystal structure "3AVA" is composed of four polymers, with chains A and B corresponding to the HIV-1 integrase protein and chains X and Y corresponding to the cyclic peptide. The cyclic peptide of the X chain is bound to an HIV-1 integrase protein dimer composed of chains A and B, and the cyclic peptide of the Y chain is also bound to an HIV-1 integrase protein dimer composed of chains A and B. The coordinate data for the complex model of chains A, B, and Y of the crystal structure "3AVA" was used for verification.
[0165] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the Y chain of the crystal structure "3AVA" were extracted. The extracted amino acid residues were Gln95, Thr124, Thr125, Ala128, Ala129, Trp131, and Trp132 for the A chain, and Asp167, Gln168, Ala169, Glu170, His171, Thr174, and Met178 for the B chain, for a total of 14 amino acid residues. A docking simulation of the cyclic peptide was performed for these amino acid residues.
[0166] (Coarse-graining of target protein) The extracted amino acid residues were approximated with two coarse-grained particles. One coarse-grained particle corresponded to the main chain, and the other corresponded to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned atomic names such as CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned atomic names other than CA, C, N, and O. Finally, the target protein was approximated with a total of 28 coarse-grained particles.
[0167] (Target peptide) The amino acid sequence of the target peptide was Ala-Leu-Lys-Ile-Asp-Asn-Leu-Asp. The peptide was cyclized by forming an amide bond between the main chain amine of Ala1 and the main chain carboxyl group of Asp8.
[0168] (Coarse-graining of target peptide) The amino acid residues constituting the target peptide were approximated by two coarse-grained particles. One coarse-grained particle corresponded to the main chain, and the other corresponded to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atoms constituting the amino acid residue that were assigned the atomic names CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atoms constituting the amino acid residue that were assigned the atomic names other than CA, C, N, and O. Finally, the target peptide was approximated by a total of 16 coarse-grained particles.
[0169] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0170] (Translation and rotation of the coarse-grained model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. Then, the coarse-grained particle placed at coordinate (5) was rotated by 4.712 rad around the x-axis and by 3.142 rad around the z-axis. The coarse-grained particle was then translated by +2.533 in the x-axis direction, +2.533 in the y-axis direction, and +2.533 in the z-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0171] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (5,5,5), (6,6,6), (6,4,4), (4,6,4), (4,4,6), (5,7,7), (7,5,7), (7,7,5), (5,3,3), (7,3,5), (7,5,3), (3,5,3), (3,7,5), (5,7,3), (3,3,5), (3,5,7), (5,3,7), (6,8,8), (4,8,6), (4,6,8), (5,9,9), (7,7,9), (7,9,7), (3,9,7), (5,9,5), (3,7,9), (5,5,9), (8,6,8), (8,4,6), (6,4,8), (9,5,9), (9,7,7), (9,3,7), (9,55), (7,3,9), (8,8,6), (8,6,4), (6,8,4), (9,9,5), (9,7,3), (7,9,3), (6,2,2), (4,4,2), (4,2,4), (5,1,1), (7,1,3), (7,3,1), (3,3,1), (5,5,1), (3,1,3), (5,1,5), (8,2,4), (6,2,6), (9,1,5), (9,3,3), (7,1,7), (8,4,2), (6,6,2), (9,5,1), (7,7,1), (2,6,2), (2,4,4), (1,5,1), (1,7,3), (3,7,1), (1,3,3), (1,5,5), (2,8,4), (2,6,6), (19,5), (3,9,3), (1,7,7), (4,8,2), (5,9,1), (2,2,6), (1,1,5), (1,3,7), (3,1,7), (2,4,8), (1,5,9), (3,3,9), (4,2,8), (5,1,9), (6,10,10), (4,10,8), (4,8,10), (5,11,11), (7,9,11), (7,11,9), (3,11,9), (5,11,7), (3,9,11), (5,7,11), (8,8,10), (6,6,10), (9,7,11), (9,9,9), (7,5,11), (8,10,8), (6,10,6), (9,11,7), (7,11,5), (2,10,6),(2,8,8), (1,11,7), (3,11,5), (1,9,9), (4,10,4), (5,11,3), (2,6,10), (1,7,11), (3,5,11), (4,4,10), (5,3,11), (10,6,10), (10,4,8), (8,4,10), (11,5,11), (11,7,9), (11,3,9), (11,5,7), (9,3,11), (10,8,8), (10,6,6), (11,9,7), (11,7,5), (10,2,6), (8,2,8), (11,1,7), (11,3,5), (9,1,9), (10,4,4), (11,5,3), (6,2,10), (7,1,11), (10,10,6), (10,8,4), (8,10,4), (11,11,5), (11,9,3), (9,11,3), (10,6,2), (8,8,2), (11,7,1), (9,9,1), (6,10,2), (7,11,1), (6,0,0), (4,2,0), (4,0,2), (5,-1,-1), (7,-1,1), (7,1,-1), (3,1,-1), (5,3,-1), (3,-1,1), (5,-1,3), (8,0,2), (6,0,4), (9,-1,3), (9,1,1), (7,-1,5), (8,2,0), (6,4,0), (9,3,-1), (7,5,-1), (2,4,0), (2,2,2), (1,3,-1), (3,5,-1), (1,1,1), (4,6,0), (5,7,-1), (2,0,4), (1,-1,3), (3,-1,5), (4,0,6), (5,-1,7), (10,0,4), (8,0,6), (11,-1,5), (11,1,3), (9,-1,7), (10,2,2), (11,3,1), (6,0,8), (7,-1,9), (10,4,0), (8,6,0), (11,5,-1), (9,7,-1), (6,8,0), (7,9,-1), (0,6,0), (0,4,2), (-1,5,-1), (-1,7,1), (1,7,-1), (-1,3,1), (-1,5,3), (0,8,2), (0,6,4), (-1,9,3), (1,9,1), (-1,7,5),(2,8,0), (3,9,-1), (0,2,4), (-1,1,3), (-1,3,5), (0,4,6), (-1,5,7), (0,10,4), (0,8,6), (-1,11,5), (1,11,3), (-1,9,7), (2,10,2), (3,11,1), (0,6,8), (-1,7,9), (4,10,0), (5,11,-1), (0,0,6), (-1,-1,5), (-1,1,7), (1,-1,7), (0,2,8), (-1,3,9), (1,1,9), (2,0,8), (3,-1,9), (0,4,10), (-1,5,11), (1,3,11), (2,2,10), (3,1,11), (4,0,10), (5,-1,11),
[0172] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0173] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the main chain of Ala1 and the coarse-grained particle corresponding to the main chain of Asp8 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0174] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value was −298.81.
[0175] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the Y chain of the crystal structure "3AVA" was verified. As a result, a docking model with an RMSD of 3.79 Å from the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 6 are shown in Figure 8. The representation of each circle is the same as in Example 1.
[0176] Example 7 (Calculation Object) Coordinate data for the crystal structure of a complex between HIV-1 (human immunodeficiency virus type 1) integrase protein and a cyclic peptide is registered in the PDB. The identifier for this crystal structure in the PDB is "3AVB." The crystal structure "3AVB" is composed of four polymers, with the A and B chains corresponding to the HIV-1 integrase protein and the X and Y chains corresponding to the cyclic peptide. The X chain cyclic peptide is bound to an HIV-1 integrase protein dimer composed of the A and B chains, and the Y chain cyclic peptide is also bound to an HIV-1 integrase protein dimer composed of the A and B chains. The coordinate data for the complex model of the A, B, and Y chains of the crystal structure "3AVB" were used for verification.
[0177] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the Y chain of the crystal structure "3AVB" were extracted. The extracted amino acid residues were Gln95, Thr124, Thr125, Ala128, Ala129, Trp131, and Trp132 for the A chain, and Asp167, Gln168, Ala169, Glu170, His171, Thr174, and Met178 for the B chain, for a total of 14 amino acid residues. A docking simulation of the cyclic peptide was performed on these amino acid residues.
[0178] (Coarse-graining of target protein) The extracted amino acid residues were approximated with two coarse-grained particles. One coarse-grained particle corresponded to the main chain, and the other corresponded to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned the atomic names CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned the atomic names other than CA, C, N, and O. Finally, the target protein was approximated with a total of 28 coarse-grained particles.
[0179] (Target peptide) The amino acid sequence of the target peptide was Ser-Leu-Lys-Ile-Asp-Asn-Leu-Asp. The peptide was cyclized by forming an amide bond between the amine in the main chain of Ser1 and the carboxyl group in the main chain of Asp8.
[0180] (Coarse-graining of target peptide) The amino acid residues constituting the target peptide were approximated by two coarse-grained particles. One coarse-grained particle corresponded to the main chain, and the other corresponded to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atoms constituting the amino acid residue that were assigned the atomic names CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atoms constituting the amino acid residue that were assigned the atomic names other than CA, C, N, and O. Finally, the target peptide was approximated by a total of 16 coarse-grained particles.
[0181] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0182] (Translation and rotation of the coarse-grained model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. Then, the coarse-grained particle placed at coordinate (5) was rotated by 4.712 rad around the x-axis and by 3.142 rad around the z-axis. The coarse-grained particle was then translated by +2.533 in the x-axis direction, +2.533 in the y-axis direction, and +2.533 in the z-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0183] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (5,5,5), (6,6,6), (6,4,4), (4,6,4), (4,4,6), (5,7,7), (7,5,7), (7,7,5), (5,3,3), (7,3,5), (7,5,3), (3,5,3), (3,7,5), (5,7,3), (3,3,5), (3,5,7), (5,3,7), (6,8,8), (4,8,6), (4,6,8), (5,9,9), (7,7,9), (7,9,7), (3,9,7), (5,9,5), (3,7,9), (5,5,9), (8,6,8), (8,4,6), (6,4,8), (9,5,9), (9,7,7), (9,3,7), (9,55), (7,3,9), (8,8,6), (8,6,4), (6,8,4), (9,9,5), (9,7,3), (7,9,3), (6,2,2), (4,4,2), (4,2,4), (5,1,1), (7,1,3), (7,3,1), (3,3,1), (5,5,1), (3,1,3), (5,1,5), (8,2,4), (6,2,6), (9,1,5), (9,3,3), (7,1,7), (8,4,2), (6,6,2), (9,5,1), (7,7,1), (2,6,2), (2,4,4), (1,5,1), (1,7,3), (3,7,1), (1,3,3), (1,5,5), (2,8,4), (2,6,6), (19,5), (3,9,3), (1,7,7), (4,8,2), (5,9,1), (2,2,6), (1,1,5), (1,3,7), (3,1,7), (2,4,8), (1,5,9), (3,3,9), (4,2,8), (5,1,9), (6,10,10), (4,10,8), (4,8,10), (5,11,11), (7,9,11), (7,11,9), (3,11,9), (5,11,7), (3,9,11), (5,7,11), (8,8,10), (6,6,10), (9,7,11), (9,9,9), (7,5,11), (8,10,8), (6,10,6), (9,11,7), (7,11,5), (2,10,6),(2,8,8), (1,11,7), (3,11,5), (1,9,9), (4,10,4), (5,11,3), (2,6,10), (1,7,11), (3,5,11), (4,4,10), (5,3,11), (10,6,10), (10,4,8), (8,4,10), (11,5,11), (11,7,9), (11,3,9), (11,5,7), (9,3,11), (10,8,8), (10,6,6), (11,9,7), (11,7,5), (10,2,6), (8,2,8), (11,1,7), (11,3,5), (9,1,9), (10,4,4), (11,5,3), (6,2,10), (7,1,11), (10,10,6), (10,8,4), (8,10,4), (11,11,5), (11,9,3), (9,11,3), (10,6,2), (8,8,2), (11,7,1), (9,9,1), (6,10,2), (7,11,1), (6,0,0), (4,2,0), (4,0,2), (5,-1,-1), (7,-1,1), (7,1,-1), (3,1,-1), (5,3,-1), (3,-1,1), (5,-1,3), (8,0,2), (6,0,4), (9,-1,3), (9,1,1), (7,-1,5), (8,2,0), (6,4,0), (9,3,-1), (7,5,-1), (2,4,0), (2,2,2), (1,3,-1), (3,5,-1), (1,1,1), (4,6,0), (5,7,-1), (2,0,4), (1,-1,3), (3,-1,5), (4,0,6), (5,-1,7), (10,0,4), (8,0,6), (11,-1,5), (11,1,3), (9,-1,7), (10,2,2), (11,3,1), (6,0,8), (7,-1,9), (10,4,0), (8,6,0), (11,5,-1), (9,7,-1), (6,8,0), (7,9,-1), (0,6,0), (0,4,2), (-1,5,-1), (-1,7,1), (1,7,-1), (-1,3,1), (-1,5,3), (0,8,2), (0,6,4), (-1,9,3), (1,9,1), (-1,7,5),(2,8,0), (3,9,-1), (0,2,4), (-1,1,3), (-1,3,5), (0,4,6), (-1,5,7), (0,10,4), (0,8,6), (-1,11,5), (1,11,3), (-1,9,7), (2,10,2), (3,11,1), (0,6,8), (-1,7,9), (4,10,0), (5,11,-1), (0,0,6), (-1,-1,5), (-1,1,7), (1,-1,7), (0,2,8), (-1,3,9), (1,1,9), (2,0,8), (3,-1,9), (0,4,10), (-1,5,11), (1,3,11), (2,2,10), (3,1,11), (4,0,10), (5,-1,11),
[0184] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0185] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the main chain of Ser1 and the coarse-grained particle corresponding to the main chain of Asp8 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0186] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value was −298.28.
[0187] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the Y chain of the crystal structure "3AVB" was verified. As a result, a docking model with an RMSD of 3.83 Å from the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 7 are shown in Figure 8. The representation of each circle is the same as in Example 1.
[0188] Example 8 (Calculation Object) Coordinate data for the crystal structure of a complex between HIV-1 (human immunodeficiency virus type 1) integrase protein and a cyclic peptide has been registered in the PDB. The identifier for this crystal structure in the PDB is "3AVI." The crystal structure "3AVI" is composed of four polymers, with the A and B chains corresponding to the HIV-1 integrase protein and the D and F chains corresponding to the cyclic peptide. The D chain cyclic peptide is bound to an HIV-1 integrase protein dimer composed of the A and B chains, and the F chain cyclic peptide is also bound to an HIV-1 integrase protein dimer composed of the A and B chains. The coordinate data for the complex model of the A, B, and F chains of the crystal structure "3AVI" was used for verification.
[0189] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the F chain of the crystal structure "3AVI" were extracted. The extracted amino acid residues were Asp167, Gln168, Ala169, Glu170, His171, Thr174, and Met178 for the A chain, and Gln95, Ala98, Leu102, Thr124, Thr125, Ala128, Ala129, Trp131, and Trp132 for the B chain, for a total of 16 amino acid residues. A docking simulation of the cyclic peptide was performed for these amino acid residues.
[0190] (Coarse-graining of target protein) Of the extracted amino acid residues, Ala169 and Thr174 of chain A and Ala98, Thr124, Thr125, Ala128, and Ala129 of chain B were approximated with one coarse-grained particle, and the remaining amino acid residues were approximated with two coarse-grained particles. Finally, the target protein was approximated with a total of 25 coarse-grained particles.
[0191] When an amino acid residue is approximated by one coarse-grained particle, the coarse-grained particle is arranged at the position of the center of gravity of all heavy atoms constituting the amino acid residue. When an amino acid residue is approximated by two coarse-grained particles, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain is arranged at the position of the center of gravity of a heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms constituting the amino acid residue. The coarse-grained particle corresponding to the side chain is arranged at the position of the center of gravity of a heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms constituting the amino acid residue.
[0192] (Target peptide) The amino acid sequence of the target peptide was Ser-Leu-Lys-Ile-Asp-Asn-Met-Asp. The peptide was cyclized by forming an amide bond between the amine in the main chain of Ser1 and the carboxyl group in the main chain of Asp8.
[0193] (Coarse-graining of target peptide) Of the amino acid residues constituting the target peptide, Ser1 was approximated by one coarse-grained particle, and the remaining amino acid residues were approximated by two coarse-grained particles. Finally, the target peptide was approximated by a total of 15 coarse-grained particles.
[0194] For Ser1, a coarse-grained particle model was placed at the center of gravity of all the heavy atoms that constitute the amino acid residue. For amino acid residues other than Ser1, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain was placed at the center of gravity of the heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms that constitute the amino acid residue. The coarse-grained particle corresponding to the side chain was placed at the center of gravity of the heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms that constitute the amino acid residue.
[0195] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0196] (Translation and rotation of the coarse-grained model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. Then, the coarse-grained particle placed at coordinate (5) was rotated by 1.571 rad around the x-axis and by 3.142 rad around the z-axis. The coarse-grained particle was then translated by +2.533 in the x-axis direction and +2.533 in the z-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0197] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (4,5,5), (5,6,6), (5,4,4), (3,6,4), (3,4,6), (4,7,7), (6,5,7), (6,7,5), (4,3,3), (6,3,5), (6,5,3), (2,5,3), (2,7,5), (4,7,3), (2,3,5), (2,5,7), (4,3,7), (5,8,8), (3,8,6), (3,6,8), (4,9,9), (6,7,9), (6,9,7), (2,9,7), (4,9,5), (2,7,9), (4,5,9), (7,6,8), (7,4,6), (5,4,8), (8,5,9), (8,7,7), (8,3,7), (8,5,5), (6,3,9), (7,8,6), (7,6,4), (5,8,4), (8,9,5), (8,7,3), (6,9,3), (5,2,2), (3,4,2), (3,2,4), (4,1,1), (6,1,3), (6,3,1), (2,3,1), (4,5,1), (2,1,3), (4,1,5), (7,2,4), (5,2,6), (8,1,5), (8,3,3), (6,1,7), (7,4,2), (5,6,2), (8,5,1), (6,7,1), (1,6,2), (1,4,4), (0,5,1), (0,7,3), (2,7,1), (0,3,3), (0,5,5), (1,8,4), (1,6,6), (0,9,5), (2,9,3), (0,7,7), (3,8,2), (4,9,1), (1,2,6), (0,1,5), (0,3,7), (2,1,7), (1,4,8), (0,5,9), (2,3,9), (3,2,8), (4,1,9), (5,10,10), (3,10,8), (3,8,10), (4,11,11), (6,9,11), (6,11,9), (2,11,9), (4,11,7), (2,9,11), (4,7,11), (7,8,10), (5,6,10), (8,7,11), (8,9,9), (6,5,11), (7,10,8), (5,10,6), (8,11,7), (6,11,5), (1,10,6),(1,8,8), (0,11,7), (2,11,5), (0,9,9), (3,10,4), (4,11,3), (1,6,10), (0,7,11), (2,5,11), (3,4,10), (4,3,11), (9,6,10), (9,4,8), (7,4,10), (10,5,11), (10,7,9), (10,3,9), (10,5,7), (8,3,11), (9,8,8), (9,6,6), (10,9,7), (10,7,5), (9,2,6), (7,2,8), (10,1,7), (10,3,5), (8,1,9), (9,4,4), (10,5,3), (5,2,10), (6,1,11), (9,10,6), (9,8,4), (7,10,4), (10,11,5), (10,9,3), (8,11,3), (9,6,2), (7,8,2), (10,7,1), (8,9,1), (5,10,2), (6,11,1), (5,0,0), (3,2,0), (3,0,2), (4,-1,-1), (6,-1,1), (6,1,-1), (2,1,-1), (4,3,-1), (2,-1,1), (4,-1,3), (7,0,2), (5,0,4), (8,-1,3), (8,1,1), (6,-1,5), (7,2,0), (5,4,0), (8,3,-1), (6,5,-1), (1,4,0), (1,2,2), (0,3,-1), (2,5,-1), (0,1,1), (3,6,0), (4,7,-1), (1,0,4), (0,-1,3), (2,-1,5), (3,0,6), (4,-1,7), (9,0,4), (7,0,6), (10,-1,5), (10,1,3), (8,-1,7), (9,2,2), (10,3,1), (5,0,8), (6,-1,9), (9,4,0), (7,6,0), (10,5,-1), (8,7,-1), (5,8,0), (6,9,-1), (-1,6,0), (-1,4,2), (-2,5,-1), (-2,7,1), (0,7,-1), (-2,3,1), (-2,5,3), (-1,8,2), (-1,6,4), (-2,9,3), (0,9,1), (-2,7,5), (1,8,0),(2,9,-1), (-1,2,4), (-2,1,3), (-2,3,5), (-1,4,6), (-2,5,7), (-1,10,4), (-1,8,6), (-2,11,5), (0,11,3), (-2,9,7), (1,10,2), (2,11,1), (-1,6,8), (-2,7,9), (3,10,0), (4,11,-1), (-1,0,6), (-2,-1,5), (-2,1,7), (0,-1,7), (-1,2,8), (-2,3,9), (0,1,9), (1,0,8), (2,-1,9), (-1,4,10), (-2,5,11), (0,3,11), (1,2,10), (2,1,11), (3,0,10), (4,-1,11),
[0198] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0199] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the main chain of Ser1 and the coarse-grained particle corresponding to the main chain of Asp8 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0200] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value was −229.84.
[0201] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the F chain of the crystal structure "3AVI" was verified. As a result, a docking model with an RMSD of 4.34 Å from the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 8 are shown in Figure 8. The representation of each circle is the same as in Example 1.
[0202] Example 9 (Calculation Object) Coordinate data for the crystal structure of a complex between HIV-1 (human immunodeficiency virus type 1) integrase protein and a cyclic peptide is registered in the PDB. The identifier for this crystal structure in the PDB is "3AVJ." The crystal structure "3AVJ" is composed of four polymers, with the A and B chains corresponding to the HIV-1 integrase protein and the D and F chains corresponding to the cyclic peptide. The D chain cyclic peptide is bound to an HIV-1 integrase protein dimer composed of the A and B chains, and the F chain cyclic peptide is also bound to an HIV-1 integrase protein dimer composed of the A and B chains. The coordinate data for the complex model of the A, B, and F chains of the crystal structure "3AVJ" were used for verification.
[0203] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the F chain of the crystal structure "3AVJ" were extracted. The extracted amino acid residues were Asp167, Gln168, Ala169, Glu170, His171, Thr174, and Met178 for the A chain, and Gln95, Ala98, Leu102, Thr124, Thr125, Ala128, Ala129, Trp131, and Trp132 for the B chain, for a total of 16 amino acid residues. A docking simulation of the cyclic peptide was performed for these amino acid residues.
[0204] (Coarse-graining of target protein) Of the extracted amino acid residues, Ala169 and Thr174 of chain A and Ala98, Thr124, Thr125, Ala128, and Ala129 of chain B were approximated with one coarse-grained particle, and the remaining amino acid residues were approximated with two coarse-grained particles. Finally, the target protein was approximated with a total of 25 coarse-grained particles.
[0205] When an amino acid residue is approximated by one coarse-grained particle, the coarse-grained particle is arranged at the position of the center of gravity of all heavy atoms constituting the amino acid residue. When an amino acid residue is approximated by two coarse-grained particles, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain is arranged at the position of the center of gravity of a heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms constituting the amino acid residue. The coarse-grained particle corresponding to the side chain is arranged at the position of the center of gravity of a heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms constituting the amino acid residue.
[0206] (Target peptide) The amino acid sequence of the target peptide was Ala-Leu-Lys-Ile-Asp-Asn-Met-Asp. The peptide was cyclized by forming an amide bond between the main chain amine of Ala1 and the main chain carboxyl group of Asp8.
[0207] (Coarse-graining of target peptide) Among the amino acid residues constituting the target peptide, Ala1 was approximated by one coarse-grained particle, and the remaining amino acid residues were approximated by two coarse-grained particles. Finally, the target peptide was approximated by a total of 15 coarse-grained particles.
[0208] For Ala1, a coarse-grained particle model was placed at the center of gravity of all the heavy atoms that constitute the amino acid residue. For amino acid residues other than Ala1, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain was placed at the center of gravity of the heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms that constitute the amino acid residue. The coarse-grained particle corresponding to the side chain was placed at the center of gravity of the heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms that constitute the amino acid residue.
[0209] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0210] (Translation and rotation of the coarse-grained model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. Then, the coarse-grained particle placed at coordinate (5) was rotated by 1.571 rad around the x-axis and by 3.142 rad around the z-axis. The coarse-grained particle was then translated by +2.533 in the x-axis direction and +2.533 in the z-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0211] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (4,5,5), (5,6,6), (5,4,4), (3,6,4), (3,4,6), (4,7,7), (6,5,7), (6,7,5), (4,3,3), (6,3,5), (6,5,3), (2,5,3), (2,7,5), (4,7,3), (2,3,5), (2,5,7), (4,3,7), (5,8,8), (3,8,6), (3,6,8), (4,9,9), (6,7,9), (6,9,7), (2,9,7), (4,9,5), (2,7,9), (4,5,9), (7,6,8), (7,4,6), (5,4,8), (8,5,9), (8,7,7), (8,3,7), (8,5,5), (6,3,9), (7,8,6), (7,6,4), (5,8,4), (8,9,5), (8,7,3), (6,9,3), (5,2,2), (3,4,2), (3,2,4), (4,1,1), (6,1,3), (6,3,1), (2,3,1), (4,5,1), (2,1,3), (4,1,5), (7,2,4), (5,2,6), (8,1,5), (8,3,3), (6,1,7), (7,4,2), (5,6,2), (8,5,1), (6,7,1), (1,6,2), (1,4,4), (0,5,1), (0,7,3), (2,7,1), (0,3,3), (0,5,5), (1,8,4), (1,6,6), (0,9,5), (2,9,3), (0,7,7), (3,8,2), (4,9,1), (1,2,6), (0,1,5), (0,3,7), (2,1,7), (1,4,8), (0,5,9), (2,3,9), (3,2,8), (4,1,9), (5,10,10), (3,10,8), (3,8,10), (4,11,11), (6,9,11), (6,11,9), (2,11,9), (4,11,7), (2,9,11), (4,7,11), (7,8,10), (5,6,10), (8,7,11), (8,9,9), (6,5,11), (7,10,8), (5,10,6), (8,11,7), (6,11,5), (1,10,6),(1,8,8), (0,11,7), (2,11,5), (0,9,9), (3,10,4), (4,11,3), (1,6,10), (0,7,11), (2,5,11), (3,4,10), (4,3,11), (9,6,10), (9,4,8), (7,4,10), (10,5,11), (10,7,9), (10,3,9), (10,5,7), (8,3,11), (9,8,8), (9,6,6), (10,9,7), (10,7,5), (9,2,6), (7,2,8), (10,1,7), (10,3,5), (8,1,9), (9,4,4), (10,5,3), (5,2,10), (6,1,11), (9,10,6), (9,8,4), (7,10,4), (10,11,5), (10,9,3), (8,11,3), (9,6,2), (7,8,2), (10,7,1), (8,9,1), (5,10,2), (6,11,1), (5,0,0), (3,2,0), (3,0,2), (4,-1,-1), (6,-1,1), (6,1,-1), (2,1,-1), (4,3,-1), (2,-1,1), (4,-1,3), (7,0,2), (5,0,4), (8,-1,3), (8,1,1), (6,-1,5), (7,2,0), (5,4,0), (8,3,-1), (6,5,-1), (1,4,0), (1,2,2), (0,3,-1), (2,5,-1), (0,1,1), (3,6,0), (4,7,-1), (1,0,4), (0,-1,3), (2,-1,5), (3,0,6), (4,-1,7), (9,0,4), (7,0,6), (10,-1,5), (10,1,3), (8,-1,7), (9,2,2), (10,3,1), (5,0,8), (6,-1,9), (9,4,0), (7,6,0), (10,5,-1), (8,7,-1), (5,8,0), (6,9,-1), (-1,6,0), (-1,4,2), (-2,5,-1), (-2,7,1), (0,7,-1), (-2,3,1), (-2,5,3), (-1,8,2), (-1,6,4), (-2,9,3), (0,9,1), (-2,7,5), (1,8,0),(2,9,-1), (-1,2,4), (-2,1,3), (-2,3,5), (-1,4,6), (-2,5,7), (-1,10,4), (-1,8,6), (-2,11,5), (0,11,3), (-2,9,7), (1,10,2), (2,11,1), (-1,6,8), (-2,7,9), (3,10,0), (4,11,-1), (-1,0,6), (-2,-1,5), (-2,1,7), (0,-1,7), (-1,2,8), (-2,3,9), (0,1,9), (1,0,8), (2,-1,9), (-1,4,10), (-2,5,11), (0,3,11), (1,2,10), (2,1,11), (3,0,10), (4,-1,11),
[0212] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0213] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the main chain of Ala1 and the coarse-grained particle corresponding to the main chain of Asp8 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0214] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value was −235.20.
[0215] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the F chain of the crystal structure "3AVJ" was verified. As a result, a docking model with an RMSD of 4.33 Å from the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 9 are shown in Figure 9. The representation of each circle is the same as in Example 1.
[0216] Example 10 (Calculation Object) Coordinate data for the crystal structure of a complex between HIV-1 (human immunodeficiency virus type 1) integrase protein and a cyclic peptide has been registered in the PDB. The identifier for this crystal structure in the PDB is "3AVK." The crystal structure "3AVK" is composed of four polymers, with the A and B chains corresponding to the HIV-1 integrase protein and the D and F chains corresponding to the cyclic peptide. The D chain cyclic peptide is bound to an HIV-1 integrase protein dimer composed of the A and B chains, and the F chain cyclic peptide is also bound to an HIV-1 integrase protein dimer composed of the A and B chains. The coordinate data for the complex model of the A, B, and F chains of the crystal structure "3AVK" was used for verification.
[0217] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the F chain of the crystal structure "3AVK" were extracted. The extracted amino acid residues were Asp167, Gln168, Ala169, Glu170, His171, Thr174, and Met178 for the A chain, and Gln95, Ala98, Leu102, Thr124, Thr125, Ala128, Ala129, Trp131, and Trp132 for the B chain, for a total of 16 amino acid residues. A docking simulation of the cyclic peptide was performed on these amino acid residues.
[0218] (Coarse-graining of target protein) Of the extracted amino acid residues, Ala169 and Thr174 of chain A and Ala98, Thr124, Thr125, Ala128, and Ala129 of chain B were approximated into one coarse-grained particle, and the remaining amino acid residues were approximated into two coarse-grained particles. Finally, the target protein was approximated into a total of 25 coarse-grained particles.
[0219] When an amino acid residue is approximated by one coarse-grained particle, the coarse-grained particle is arranged at the position of the center of gravity of all heavy atoms constituting the amino acid residue. When an amino acid residue is approximated by two coarse-grained particles, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain is arranged at the position of the center of gravity of a heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms constituting the amino acid residue. The coarse-grained particle corresponding to the side chain is arranged at the position of the center of gravity of a heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms constituting the amino acid residue.
[0220] (Target peptide) The amino acid sequence of the target peptide was Ser-Leu-Lys-Ile-Asp-Asn-Glu-Asp. The peptide was cyclized by forming an amide bond between the amine in the main chain of Ser1 and the carboxyl group in the main chain of Asp8.
[0221] (Coarse-graining of target peptide) Among the amino acid residues constituting the target peptide, Ala1 was approximated by one coarse-grained particle, and the remaining amino acid residues were approximated by two coarse-grained particles. Finally, the target peptide was approximated by a total of 15 coarse-grained particles.
[0222] For Ser1, a coarse-grained particle model was placed at the center of gravity of all the heavy atoms that constitute the amino acid residue. For amino acid residues other than Ser1, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain was placed at the center of gravity of the heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms that constitute the amino acid residue. The coarse-grained particle corresponding to the side chain was placed at the center of gravity of the heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms that constitute the amino acid residue.
[0223] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0224] (Translation and rotation of the coarse-grained model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. Then, the coarse-grained particle placed at coordinate (5) was rotated by 3.142 rad around the x-axis and by 1.571 rad around the z-axis. The coarse-grained particle was then translated by +1.267 in the x-axis direction and +1.267 in the z-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0225] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (4,4,5), (5,5,6), (5,3,4), (3,5,4), (3,3,6), (4,6,7), (6,4,7), (6,6,5), (4,2,3), (6,2,5), (6,4,3), (2,4,3), (2,6,5), (4,6,3), (2,2,5), (2,4,7), (4,2,7), (5,7,8), (3,7,6), (3,5,8), (4,8,9), (6,6,9), (6,8,7), (2,8,7), (4,8,5), (2,6,9), (4,4,9), (7,5,8), (7,3,6), (5,3,8), (8,4,9), (8,6,7), (8,2,7), (8,4,5), (6,2,9), (7,7,6), (7,5,4), (5,7,4), (8,8,5), (8,6,3), (6,8,3), (5,1,2), (3,3,2), (3,1,4), (4,0,1), (6,0,3), (6,2,1), (2,2,1), (4,4,1), (2,0,3), (4,0,5), (7,1,4), (5,1,6), (8,0,5), (8,2,3), (6,0,7), (7,3,2), (5,5,2), (8,4,1), (6,6,1), (1,5,2), (1,3,4), (0,4,1), (0,6,3), (2,6,1), (0,2,3), (0,4,5), (1,7,4), (1,5,6), (0,8,5), (2,8,3), (0,6,7), (3,7,2), (4,8,1), (1,1,6), (0,0,5), (0,2,7), (2,0,7), (1,3,8), (0,4,9), (2,2,9), (3,1,8), (4,0,9), (5,9,10), (3,9,8), (3,7,10), (4,10,11), (6,8,11), (6,10,9), (2,10,9), (4,10,7), (2,8,11), (4,6,11), (7,7,10), (5,5,10), (8,6,11), (8,8,9), (6,4,11), (7,9,8), (5,9,6), (8,10,7), (6,10,5), (1,9,6), (1,7,8), (0,10,7), (2,10,5), (0,8,9), (3,9,4), (4,10,3), (1,5,10), (0,6,11), (2,4,11), (3,3,10), (4,2,11), (9,5,10), (9,3,8), (7,3,10), (10,4,11), (10,6,9), (10,2,9), (10,4,7), (8,2,11), (9,7,8), (9,5,6), (10,8,7), (10,6,5), (9,1,6), (7,1,8), (10,0,7), (10,2,5), (8,0,9), (9,3,4), (10,4,3), (5,1,10), (6,0,11), (9,9,6), (9,7,4), (7,9,4), (10,10,5), (10,8,3), (8,10,3), (9,5,2), (7,7,2), (10,6,1), (8,8,1), (5,9,2), (6,10,1), (5,-1,0), (3,1,0), (3,-1,2), (4,-2,-1), (6,-2,1), (6,0,-1), (2,0,-1), (4,2,-1), (2,-2,1), (4,-2,3), (7,-1,2), (5,-1,4), (8,-2,3), (8,0,1), (6,-2,5), (7,1,0), (5,3,0), (8,2,-1), (6,4,-1), (1,3,0), (1,1,2), (0,2,-1), (2,4,-1), (0,0,1), (3,5,0), (4,6,-1), (1,-1,4), (0,-2,3), (2,-2,5), (3,-1,6), (4,-2,7), (9,-1,4), (7,-1,6), (10,-2,5), (10,0,3), (8,-2,7), (9,1,2), (10,2,1), (5,-1,8), (6,-2,9), (9,3,0), (7,5,0), (10,4,-1), (8,6,-1), (5,7,0), (6,8,-1), (-1,5,0), (-1,3,2), (-2,4,-1), (-2,6,1), (0,6,-1), (-2,2,1), (-2,4,3), (-1,7,2), (-1,5,4), (-2,8,3), (0,8,1), (-2,6,5), (1,7,0),(2,8,-1), (-1,1,4), (-2,0,3), (-2,2,5), (-1,3,6), (-2,4,7), (-1,9,4), (-1,7,6), (-2,10,5), (0,10,3), (-2,8,7), (1,9,2), (2,10,1), (-1,5,8), (-2,6,9), (3,9,0), (4,10,-1), (-1,-1,6), (-2,-2,5), (-2,0,7), (0,-2,7), (-1,1,8), (-2,2,9), (0,0,9), (1,-1,8), (2,-2,9), (-1,3,10), (-2,4,11), (0,2,11), (1,1,10), (2,0,11), (3,-1,10), (4,-2,11),
[0226] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0227] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the main chain of Ser1 and the coarse-grained particle corresponding to the main chain of Asp8 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0228] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value of was −194.35.
[0229] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the F chain of the crystal structure "3AVK" was verified. As a result, a docking model with an RMSD of 3.92 Å with the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 10 are shown in Figure 9. The representation of each circle is the same as in Example 1.
[0230] Example 11 (Calculation Object) Coordinate data for the crystal structure of a complex between HIV-1 (human immunodeficiency virus type 1) integrase protein and a cyclic peptide is registered in the PDB. The identifier for this crystal structure in the PDB is "3AVL." The crystal structure "3AVL" is composed of four polymers, with chains A and B corresponding to the HIV-1 integrase protein and chains E and F corresponding to the cyclic peptide. The cyclic peptide of chain E is bound to an HIV-1 integrase protein dimer composed of chains A and B, and the cyclic peptide of chain F is also bound to an HIV-1 integrase protein dimer composed of chains A and B. The coordinate data for the complex model of chains A, B, and F of the crystal structure "3AVL" was used for verification.
[0231] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the F chain of the crystal structure "3AVL" were extracted. The extracted amino acid residues were Asp167, Gln168, Ala169, Glu170, His171, Thr174, and Met178 for the A chain, and Gln95, Thr124, Thr125, Ala128, Ala129, Trp131, and Trp132 for the B chain, for a total of 14 amino acid residues. A docking simulation of the cyclic peptide was performed for these amino acid residues.
[0232] (Coarse-graining of target protein) Of the extracted amino acid residues, Ala169 and Thr174 of chain A and Thr124, Thr125, Ala128, and Ala129 of chain B were approximated with one coarse-grained particle, and the remaining amino acid residues were approximated with two coarse-grained particles. Finally, the target protein was approximated with a total of 22 coarse-grained particles.
[0233] When an amino acid residue is approximated by one coarse-grained particle, the coarse-grained particle is arranged at the position of the center of gravity of all heavy atoms constituting the amino acid residue. When an amino acid residue is approximated by two coarse-grained particles, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain is arranged at the position of the center of gravity of a heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms constituting the amino acid residue. The coarse-grained particle corresponding to the side chain is arranged at the position of the center of gravity of a heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms constituting the amino acid residue.
[0234] (Target peptide) The amino acid sequence of the target peptide was Ala-Thr-Lys-Ile-Asp-Asn-Leu-Asp. The peptide was cyclized by forming an amide bond between the main chain amine of Ala1 and the main chain carboxyl group of Asp8.
[0235] (Coarse-graining of target peptide) Of the amino acid residues constituting the target peptide, Ala1 and Thr2 were approximated by one coarse-grained particle, and the remaining amino acid residues were approximated by two coarse-grained particles. Finally, the target peptide was approximated by a total of 14 coarse-grained particles.
[0236] For Ala1 and Thr2, coarse-grained particle models were placed at the center of gravity of all heavy atoms that constitute the amino acid residues. For amino acid residues other than Ala1 and Thr2, one coarse-grained particle corresponds to the main chain, and the other coarse-grained particle corresponds to the side chain. The coarse-grained particle corresponding to the main chain was placed at the center of gravity of the heavy atom that is assigned an atomic name such as CA, C, N, or O among the heavy atoms that constitute the amino acid residue. The coarse-grained particle corresponding to the side chain was placed at the center of gravity of the heavy atom that is assigned an atomic name other than CA, C, N, and O among the heavy atoms that constitute the amino acid residue.
[0237] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0238] (Translation and rotation of the coarse-grained model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. Then, the coarse-grained particle placed at coordinate (5) was rotated by 4.712 rad around the x-axis and by 1.571 rad around the z-axis. The coarse-grained particle was then translated by +2.533 in the y-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0239] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (4,5,5), (5,6,6), (5,4,4), (3,6,4), (3,4,6), (4,7,7), (6,5,7), (6,7,5), (4,3,3), (6,3,5), (6,5,3), (2,5,3), (2,7,5), (4,7,3), (2,3,5), (2,5,7), (4,3,7), (5,8,8), (3,8,6), (3,6,8), (4,9,9), (6,7,9), (6,9,7), (2,9,7), (4,9,5), (2,7,9), (4,5,9), (7,6,8), (7,4,6), (5,4,8), (8,5,9), (8,7,7), (8,3,7), (8,5,5), (6,3,9), (7,8,6), (7,6,4), (5,8,4), (8,9,5), (8,7,3), (6,9,3), (5,2,2), (3,4,2), (3,2,4), (4,1,1), (6,1,3), (6,3,1), (2,3,1), (4,5,1), (2,1,3), (4,1,5), (7,2,4), (5,2,6), (8,1,5), (8,3,3), (6,1,7), (7,4,2), (5,6,2), (8,5,1), (6,7,1), (1,6,2), (1,4,4), (0,5,1), (0,7,3), (2,7,1), (0,3,3), (0,5,5), (1,8,4), (1,6,6), (0,9,5), (2,9,3), (0,7,7), (3,8,2), (4,9,1), (1,2,6), (0,1,5), (0,3,7), (2,1,7), (1,4,8), (0,5,9), (2,3,9), (3,2,8), (4,1,9), (5,10,10), (3,10,8), (3,8,10), (4,11,11), (6,9,11), (6,11,9), (2,11,9), (4,11,7), (2,9,11), (4,7,11), (7,8,10), (5,6,10), (8,7,11), (8,9,9), (6,5,11), (7,10,8), (5,10,6), (8,11,7), (6,11,5), (1,10,6),(1,8,8), (0,11,7), (2,11,5), (0,9,9), (3,10,4), (4,11,3), (1,6,10), (0,7,11), (2,5,11), (3,4,10), (4,3,11), (9,6,10), (9,4,8), (7,4,10), (10,5,11), (10,7,9), (10,3,9), (10,5,7), (8,3,11), (9,8,8), (9,6,6), (10,9,7), (10,7,5), (9,2,6), (7,2,8), (10,1,7), (10,3,5), (8,1,9), (9,4,4), (10,5,3), (5,2,10), (6,1,11), (9,10,6), (9,8,4), (7,10,4), (10,11,5), (10,9,3), (8,11,3), (9,6,2), (7,8,2), (10,7,1), (8,9,1), (5,10,2), (6,11,1), (5,0,0), (3,2,0), (3,0,2), (4,-1,-1), (6,-1,1), (6,1,-1), (2,1,-1), (4,3,-1), (2,-1,1), (4,-1,3), (7,0,2), (5,0,4), (8,-1,3), (8,1,1), (6,-1,5), (7,2,0), (5,4,0), (8,3,-1), (6,5,-1), (1,4,0), (1,2,2), (0,3,-1), (2,5,-1), (0,1,1), (3,6,0), (4,7,-1), (1,0,4), (0,-1,3), (2,-1,5), (3,0,6), (4,-1,7), (9,0,4), (7,0,6), (10,-1,5), (10,1,3), (8,-1,7), (9,2,2), (10,3,1), (5,0,8), (6,-1,9), (9,4,0), (7,6,0), (10,5,-1), (8,7,-1), (5,8,0), (6,9,-1), (-1,6,0), (-1,4,2), (-2,5,-1), (-2,7,1), (0,7,-1), (-2,3,1), (-2,5,3), (-1,8,2), (-1,6,4), (-2,9,3), (0,9,1), (-2,7,5), (1,8,0),(2,9,-1), (-1,2,4), (-2,1,3), (-2,3,5), (-1,4,6), (-2,5,7), (-1,10,4), (-1,8,6), (-2,11,5), (0,11,3), (-2,9,7), (1,10,2), (2,11,1), (-1,6,8), (-2,7,9), (3,10,0), (4,11,-1), (-1,0,6), (-2,-1,5), (-2,1,7), (0,-1,7), (-1,2,8), (-2,3,9), (0,1,9), (1,0,8), (2,-1,9), (-1,4,10), (-2,5,11), (0,3,11), (1,2,10), (2,1,11), (3,0,10), (4,-1,11),
[0240] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0241] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the main chain of Ala1 and the coarse-grained particle corresponding to the main chain of Asp8 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0242] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value was −193.68.
[0243] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the F chain of the crystal structure "3AVL" was verified. As a result, a docking model with an RMSD of 4.76 Å with the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 11 are shown in Figure 9. The representation of each circle is the same as in Example 1.
[0244] Example 12 (Calculation Object) Coordinate data for the crystal structure of a complex between HIV-1 (human immunodeficiency virus type 1) integrase protein and a cyclic peptide has been registered in the PDB. The identifier for this crystal structure in the PDB is "3AVM." The crystal structure "3AVM" is composed of four polymers, with the A and B chains corresponding to the HIV-1 integrase protein and the D and F chains corresponding to the cyclic peptide. The D chain cyclic peptide is bound to an HIV-1 integrase protein dimer composed of the A and B chains, and the F chain cyclic peptide is also bound to an HIV-1 integrase protein dimer composed of the A and B chains. The coordinate data for the complex model of the A, B, and F chains of the crystal structure "3AVM" was used for verification.
[0245] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the F chain of the crystal structure "3AVM" were extracted. The extracted amino acid residues were Asp167, Gln168, Ala169, Glu170, His171, Thr174, and Met178 for the A chain, and Gln95, Thr124, Thr125, Ala128, Ala129, Trp131, and Trp132 for the B chain, for a total of 14 amino acid residues. A docking simulation of the cyclic peptide was performed on these amino acid residues.
[0246] (Coarse-graining of target protein) The extracted amino acid residues were approximated with two coarse-grained particles. One coarse-grained particle corresponded to the main chain, and the other corresponded to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned the atomic names CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned the atomic names other than CA, C, N, and O. Finally, the target protein was approximated with a total of 28 coarse-grained particles.
[0247] (Target peptide) The amino acid sequence of the target peptide was Ser-Arg-Lys-Ile-Asp-Asn-Leu-Asp. The peptide was cyclized by forming an amide bond between the amine in the main chain of Ser1 and the carboxyl group in the main chain of Asp8.
[0248] (Coarse-graining of target peptide) The amino acid residues constituting the target peptide were approximated by two coarse-grained particles. One coarse-grained particle corresponded to the main chain, and the other corresponded to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atoms constituting the amino acid residue that were assigned the atomic names CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atoms constituting the amino acid residue that were assigned the atomic names other than CA, C, N, and O. Finally, the target peptide was approximated by a total of 16 coarse-grained particles.
[0249] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0250] (Translation and rotation of the coarse-grained model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. Then, the coarse-grained particle placed at coordinate (5) was rotated by 3.142 rad around the x-axis and 3.142 rad around the z-axis. The coarse-grained particle was then translated by +2.533 in the z-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0251] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (5,5,5), (6,6,6), (6,4,4), (4,6,4), (4,4,6), (5,7,7), (7,5,7), (7,7,5), (5,3,3), (7,3,5), (7,5,3), (3,5,3), (3,7,5), (5,7,3), (3,3,5), (3,5,7), (5,3,7), (6,8,8), (4,8,6), (4,6,8), (5,9,9), (7,7,9), (7,9,7), (3,9,7), (5,9,5), (3,7,9), (5,5,9), (8,6,8), (8,4,6), (6,4,8), (9,5,9), (9,7,7), (9,3,7), (9,55), (7,3,9), (8,8,6), (8,6,4), (6,8,4), (9,9,5), (9,7,3), (7,9,3), (6,2,2), (4,4,2), (4,2,4), (5,1,1), (7,1,3), (7,3,1), (3,3,1), (5,5,1), (3,1,3), (5,1,5), (8,2,4), (6,2,6), (9,1,5), (9,3,3), (7,1,7), (8,4,2), (6,6,2), (9,5,1), (7,7,1), (2,6,2), (2,4,4), (1,5,1), (1,7,3), (3,7,1), (1,3,3), (1,5,5), (2,8,4), (2,6,6), (19,5), (3,9,3), (1,7,7), (4,8,2), (5,9,1), (2,2,6), (1,1,5), (1,3,7), (3,1,7), (2,4,8), (1,5,9), (3,3,9), (4,2,8), (5,1,9), (6,10,10), (4,10,8), (4,8,10), (5,11,11), (7,9,11), (7,11,9), (3,11,9), (5,11,7), (3,9,11), (5,7,11), (8,8,10), (6,6,10), (9,7,11), (9,9,9), (7,5,11), (8,10,8), (6,10,6), (9,11,7), (7,11,5), (2,10,6),(2,8,8), (1,11,7), (3,11,5), (1,9,9), (4,10,4), (5,11,3), (2,6,10), (1,7,11), (3,5,11), (4,4,10), (5,3,11), (10,6,10), (10,4,8), (8,4,10), (11,5,11), (11,7,9), (11,3,9), (11,5,7), (9,3,11), (10,8,8), (10,6,6), (11,9,7), (11,7,5), (10,2,6), (8,2,8), (11,1,7), (11,3,5), (9,1,9), (10,4,4), (11,5,3), (6,2,10), (7,1,11), (10,10,6), (10,8,4), (8,10,4), (11,11,5), (11,9,3), (9,11,3), (10,6,2), (8,8,2), (11,7,1), (9,9,1), (6,10,2), (7,11,1), (6,0,0), (4,2,0), (4,0,2), (5,-1,-1), (7,-1,1), (7,1,-1), (3,1,-1), (5,3,-1), (3,-1,1), (5,-1,3), (8,0,2), (6,0,4), (9,-1,3), (9,1,1), (7,-1,5), (8,2,0), (6,4,0), (9,3,-1), (7,5,-1), (2,4,0), (2,2,2), (1,3,-1), (3,5,-1), (1,1,1), (4,6,0), (5,7,-1), (2,0,4), (1,-1,3), (3,-1,5), (4,0,6), (5,-1,7), (10,0,4), (8,0,6), (11,-1,5), (11,1,3), (9,-1,7), (10,2,2), (11,3,1), (6,0,8), (7,-1,9), (10,4,0), (8,6,0), (11,5,-1), (9,7,-1), (6,8,0), (7,9,-1), (0,6,0), (0,4,2), (-1,5,-1), (-1,7,1), (1,7,-1), (-1,3,1), (-1,5,3), (0,8,2), (0,6,4), (-1,9,3), (1,9,1), (-1,7,5),(2,8,0), (3,9,-1), (0,2,4), (-1,1,3), (-1,3,5), (0,4,6), (-1,5,7), (0,10,4), (0,8,6), (-1,11,5), (1,11,3), (-1,9,7), (2,10,2), (3,11,1), (0,6,8), (-1,7,9), (4,10,0), (5,11,-1), (0,0,6), (-1,-1,5), (-1,1,7), (1,-1,7), (0,2,8), (-1,3,9), (1,1,9), (2,0,8), (3,-1,9), (0,4,10), (-1,5,11), (1,3,11), (2,2,10), (3,1,11), (4,0,10), (5,-1,11),
[0252] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0253] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the main chain of Ser1 and the coarse-grained particle corresponding to the main chain of Asp8 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0254] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value was −226.76.
[0255] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the F chain of the crystal structure "3AVM" was verified. As a result, a docking model with an RMSD of 4.17 Å with the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 12 are shown in Figure 9. The representation of each circle is the same as in Example 1.
[0256] Example 13 (Calculation Subject) Coordinate data for the crystal structure of a complex between HIV-1 (human immunodeficiency virus type 1) integrase protein and a cyclic peptide has been registered in the PDB. The identifier for this crystal structure in the PDB was "3AVN." The crystal structure "3AVN" is composed of four polymers, with the A and B chains corresponding to the HIV-1 integrase protein and the G and H chains corresponding to the cyclic peptide. The cyclic peptide of the G chain is bound to an HIV-1 integrase protein dimer composed of the A and B chains, and the cyclic peptide of the H chain is also bound to an HIV-1 integrase protein dimer composed of the A and B chains. The coordinate data for the complex model of the A, B, and H chains of the crystal structure "3AVN" were used for verification.
[0257] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the H chain of the crystal structure "3AVN" were extracted. The extracted amino acid residues were Asp167, Gln168, Ala169, Glu170, His171, Thr174, and Met178 for the A chain, and Gln95, Thr124, Thr125, Ala128, Ala129, Trp131, and Trp132 for the B chain, for a total of 14 amino acid residues. A docking simulation of the cyclic peptide was performed for these amino acid residues.
[0258] (Coarse-graining of target protein) The extracted amino acid residues were approximated with two coarse-grained particles. One coarse-grained particle corresponded to the main chain, and the other corresponded to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned the atomic names CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atoms that make up the amino acid residue and that are assigned the atomic names other than CA, C, N, and O. Finally, the target protein was approximated with a total of 28 coarse-grained particles.
[0259] (Target peptide) The amino acid sequence of the target peptide was Ser-His-Lys-Ile-Asp-Asn-Leu-Asp. The peptide was cyclized by forming an amide bond between the amine in the main chain of Ser1 and the carboxyl group in the main chain of Asp8.
[0260] (Coarse-graining of target peptide) The amino acid residues constituting the target peptide were approximated by two coarse-grained particles. One coarse-grained particle corresponded to the main chain, and the other corresponded to the side chain. The coarse-grained particle corresponding to the main chain was placed at the position of the center of gravity of the heavy atoms constituting the amino acid residue that were assigned the atomic names CA, C, N, or O. The coarse-grained particle corresponding to the side chain was placed at the position of the center of gravity of the heavy atoms constituting the amino acid residue that were assigned the atomic names other than CA, C, N, and O. Finally, the target peptide was approximated by a total of 16 coarse-grained particles.
[0261] (Scaling and rearrangement of coarse-grained model) Scaling and rearrangement of the coarse-grained model were performed in the same manner as in Example 1. That is, rearrangement to the position represented by the above coordinate (4) was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby rearranging the coarse-grained models for each of the target protein and the target peptide. Then, an operation of further rearranging the rearranged coarse-grained particles to positions (X'-min_X', Y'-min_Y', Z'-min_Z') was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0262] (Translation and rotation of the coarse-grained model) As in Example 1, the i-th coarse-grained particle of the target protein was placed at the position indicated by the coordinate (5) above. Then, the coarse-grained particle placed at coordinate (5) was rotated by 3.142 rad around the x-axis and 3.142 rad around the z-axis. The coarse-grained particle was then translated by +2.533 in the z-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby performing translation and rotation of the coarse-grained model for each of the target protein and the target peptide. Next, as in Example 1, all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide were relocated to the coordinate (6) above, thereby relocating the coarse-grained models of the target protein and the target peptide, respectively.
[0263] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 239 lattice points located at the following orthogonal coordinates. (5,5,5), (6,6,6), (6,4,4), (4,6,4), (4,4,6), (5,7,7), (7,5,7), (7,7,5), (5,3,3), (7,3,5), (7,5,3), (3,5,3), (3,7,5), (5,7,3), (3,3,5), (3,5,7), (5,3,7), (6,8,8), (4,8,6), (4,6,8), (5,9,9), (7,7,9), (7,9,7), (3,9,7), (5,9,5), (3,7,9), (5,5,9), (8,6,8), (8,4,6), (6,4,8), (9,5,9), (9,7,7), (9,3,7), (9,55), (7,3,9), (8,8,6), (8,6,4), (6,8,4), (9,9,5), (9,7,3), (7,9,3), (6,2,2), (4,4,2), (4,2,4), (5,1,1), (7,1,3), (7,3,1), (3,3,1), (5,5,1), (3,1,3), (5,1,5), (8,2,4), (6,2,6), (9,1,5), (9,3,3), (7,1,7), (8,4,2), (6,6,2), (9,5,1), (7,7,1), (2,6,2), (2,4,4), (1,5,1), (1,7,3), (3,7,1), (1,3,3), (1,5,5), (2,8,4), (2,6,6), (19,5), (3,9,3), (1,7,7), (4,8,2), (5,9,1), (2,2,6), (1,1,5), (1,3,7), (3,1,7), (2,4,8), (1,5,9), (3,3,9), (4,2,8), (5,1,9), (6,10,10), (4,10,8), (4,8,10), (5,11,11), (7,9,11), (7,11,9), (3,11,9), (5,11,7), (3,9,11), (5,7,11), (8,8,10), (6,6,10), (9,7,11), (9,9,9), (7,5,11), (8,10,8), (6,10,6), (9,11,7), (7,11,5), (2,10,6),(2,8,8), (1,11,7), (3,11,5), (1,9,9), (4,10,4), (5,11,3), (2,6,10), (1,7,11), (3,5,11), (4,4,10), (5,3,11), (10,6,10), (10,4,8), (8,4,10), (11,5,11), (11,7,9), (11,3,9), (11,5,7), (9,3,11), (10,8,8), (10,6,6), (11,9,7), (11,7,5), (10,2,6), (8,2,8), (11,1,7), (11,3,5), (9,1,9), (10,4,4), (11,5,3), (6,2,10), (7,1,11), (10,10,6), (10,8,4), (8,10,4), (11,11,5), (11,9,3), (9,11,3), (10,6,2), (8,8,2), (11,7,1), (9,9,1), (6,10,2), (7,11,1), (6,0,0), (4,2,0), (4,0,2), (5,-1,-1), (7,-1,1), (7,1,-1), (3,1,-1), (5,3,-1), (3,-1,1), (5,-1,3), (8,0,2), (6,0,4), (9,-1,3), (9,1,1), (7,-1,5), (8,2,0), (6,4,0), (9,3,-1), (7,5,-1), (2,4,0), (2,2,2), (1,3,-1), (3,5,-1), (1,1,1), (4,6,0), (5,7,-1), (2,0,4), (1,-1,3), (3,-1,5), (4,0,6), (5,-1,7), (10,0,4), (8,0,6), (11,-1,5), (11,1,3), (9,-1,7), (10,2,2), (11,3,1), (6,0,8), (7,-1,9), (10,4,0), (8,6,0), (11,5,-1), (9,7,-1), (6,8,0), (7,9,-1), (0,6,0), (0,4,2), (-1,5,-1), (-1,7,1), (1,7,-1), (-1,3,1), (-1,5,3), (0,8,2), (0,6,4), (-1,9,3), (1,9,1), (-1,7,5),(2,8,0), (3,9,-1), (0,2,4), (-1,1,3), (-1,3,5), (0,4,6), (-1,5,7), (0,10,4), (0,8,6), (-1,11,5), (1,11,3), (-1,9,7), (2,10,2), (3,11,1), (0,6,8), (-1,7,9), (4,10,0), (5,11,-1), (0,0,6), (-1,-1,5), (-1,1,7), (1,-1,7), (0,2,8), (-1,3,9), (1,1,9), (2,0,8), (3,-1,9), (0,4,10), (-1,5,11), (1,3,11), (2,2,10), (3,1,11), (4,0,10), (5,-1,11),
[0264] (Objective Function) As in Example 1, the total interaction energy of the target peptide was calculated using the above formulas (1) to (3), and formula (3) was used as the objective function for mathematical optimization. The use of the MJ matrix and the setting of the threshold Td were also the same as in Example 1.
[0265] (Constraints) To generate a conformational model of the target peptide, the following five constraints were set: (1) All coarse-grained particles constituting the target peptide must be placed at the lattice points of a regular tetrahedron lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the main chain of the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the main chain of the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) When the i-th amino acid residue of the target peptide is approximated by two coarse-grained particles, the coarse-grained particle corresponding to the main chain of the amino acid residue and the coarse-grained particle corresponding to the side chain of the amino acid residue must be placed at two adjacent lattice points. (4) For the target peptide, the coarse-grained particle corresponding to the main chain of Ser1 and the coarse-grained particle corresponding to the main chain of Asp8 must be placed at two adjacent lattice points. (5) The coarse-grained particle of the target peptide must not be placed at a grid point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0266] (Mathematical Optimization) By the same process as in Example 1, the objective variable E is calculated within the range that satisfies the above constraints. total The arrangement pattern of the coarse-grained particles of the target peptide that gives the minimum value of E was obtained as the docking model. total The minimum value was −226.35.
[0267] (Verification) Using the same method as in Example 1, the degree of agreement between the obtained docking model and a coarse-grained model of the H chain of the crystal structure "3AVN" was verified. As a result, a docking model with an RMSD of 4.04 Å from the rearranged crystal structure was obtained. The docking model and rearranged crystal structure in Example 13 are shown in Figure 9. The representation of each circle is the same as in Example 1.
[0268] [Comparative Example 1] (Calculation Object) Coordinate data for the crystal structure of a complex between HIV-1 (human immunodeficiency virus type 1) integrase protein and a cyclic peptide is registered in the PDB. The identifier for this crystal structure in the PDB is "3WNE." The crystal structure "3WNE" is composed of four polymers, with chains A and B corresponding to the HIV-1 integrase protein and chains C and D corresponding to the cyclic peptide. Both the cyclic peptide of chain C and the cyclic peptide of chain D are bound to an HIV-1 integrase protein dimer composed of chains A and B. The coordinate data for the complex model of chains A, B, and C of the crystal structure "3WNE" was used for verification.
[0269] (Target Protein) Among the amino acid residues constituting the target protein, amino acid residues containing atoms located within 5 Å of any atom constituting the C chain of the crystal structure "3WNE" were extracted. The extracted amino acid residues were Asp167, Gln168, Ala169, Glu170, His171, Thr174, and Met178 for the A chain, and Gln95, Leu102, Thr125, Ala128, Ala129, and Trp132 for the B chain, for a total of 13 amino acid residues. A docking simulation of the cyclic peptide was performed for these amino acid residues.
[0270] (Coarse-graining of target protein) Each extracted amino acid residue was approximated by one coarse-grained particle. For each amino acid residue, a coarse-grained particle model was placed at the position of the center of gravity of all heavy atoms constituting the amino acid residue, and thus the target protein was finally approximated by 13 coarse-grained particles.
[0271] (Target peptide) The amino acid sequence of the target peptide was Pro-Lys-Ile-Asp-Asn-Gly. The peptide was cyclized by forming an amide bond between the amine in the main chain of Pro1 and the carboxyl group in the main chain of Gly6.
[0272] (Coarse-graining of target peptide) Each amino acid residue constituting the target peptide was approximated by one coarse-grained particle. Ultimately, the target peptide was approximated by a total of six coarse-grained particles. Coarse-grained particle models were placed at the center of gravity of all heavy atoms constituting the target peptide.
[0273] (Scaling and rearrangement of coarse-grained model) In the list of Cartesian coordinates showing the positions of all coarse-grained particles that make up the target protein, the maximum value on the x-axis is called max_X, the minimum value on the x-axis is called min_X, the maximum value on the y-axis is called max_Y, the minimum value on the y-axis is called min_Y, the maximum value on the z-axis is called max_Z, and the minimum value on the z-axis is called min_Z. The Cartesian coordinates showing the position of any coarse-grained particle are (X, Y, Z), and this coarse-grained particle was rearranged to the position represented by the following coordinates (8).
[0274] This operation was applied to all of the coarse-grained particles constituting the target protein and all of the coarse-grained particles constituting the target peptide, and the coarse-grained models were rearranged for each of the target protein and the target peptide.
[0275] In the list of Cartesian coordinates indicating the positions of the coarse-grained particles constituting the target protein after rearrangement, the value obtained by rounding up the decimal point of the minimum value in the x-axis direction in the negative direction is called min_X'. Similarly, in the list, the value obtained by rounding up the decimal point of the minimum value in the y-axis direction in the negative direction is called min_Y', and the value obtained by rounding up the decimal point of the minimum value in the z-axis direction in the negative direction is called min_Z'. If the Cartesian coordinates indicating the position of any coarse-grained particle after rearrangement are represented as (X', Y', Z'), the coarse-grained particle was further rearranged to the position (X'-min_X', Y'-min_Y', Z'-min_Z'). This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide, thereby further rearranging the coarse-grained models of the target protein and the target peptide.
[0276] (Translation and rotation of the coarse-grained model) The Cartesian coordinates indicating the position of the i-th coarse-grained particle constituting the target protein and the target peptide are (x i , y i , z i The Cartesian coordinates indicating the position of any coarse-grained particle are represented as (X, Y, Z), and the total number of coarse-grained particles constituting the target protein and the target peptide is represented as N. The coarse-grained particles were placed at the positions indicated by the following coordinates (9).
[0277] The coarse-grained particle located at coordinate (9) was rotated counterclockwise around the x-axis by 3.142 rad and counterclockwise around the z-axis by 4.172 rad. The coarse-grained particle was then translated by +1.267 in the y-axis direction and +2.533 in the z-axis direction. This operation was applied to all coarse-grained particles constituting the target protein and all coarse-grained particles constituting the target peptide. This process performed translation and rotation of the coarse-grained model for each of the target protein and target peptide.
[0278] The Cartesian coordinates indicating the position of any coarse-grained particle that was translated and rotated were expressed as (X', Y', Z'), and the coarse-grained particle was relocated to the position of the following coordinates (10). This operation was applied to all coarse-grained particles that make up the target protein and all coarse-grained particles that make up the target peptide, and the coarse-grained models of the target protein and target peptide were relocated.
[0279] (Generation of Lattice Structure) A regular tetrahedron lattice was generated by setting 147 lattice points located at the following orthogonal coordinates. (4,4,3), (5,5,4), (5,3,2), (3,5,2), (3,3,4), (4,6,5), (6,4,5), (6,6,3), (5,7,6), (3,7,4), (3,5,6), (7,5,6), (7,3,4), (5,3,6), (7,7,4), (7,5,2), (5,7,2), (4,2,1), (6,2,3), (6,4,1), (5,1,0), (3,3,0), (3,1,2), (7,1,2), (5,1,4), (7,3,0), (5,5,0), (2,4,1), (2,6,3), (4,6,1), (1,5,0), (1,3,2), (1,7,2), (1,5,4), (3,7,0), (2,2,3), (2,4,5), (4,2,5), (1,1,4), (1,3,6), (3,1,6), (4,8,7), (6,6,7), (6,8,5), (5,9,8), (3,9,6), (3,7,8), (7,7,8), (5,5,8), (7,9,6), (5,9,4), (2,8,5), (4,8,3), (1,9,4), (1,7,6), (3,9,2), (2,6,7), (4,4,7), (1,5,8), (3,3,8), (8,4,7), (8,6,5), (9,5,8), (9,3,6), (7,3,8), (9,7,6), (9,5,4), (8,2,5), (8,4,3), (9,1,4), (7,1,6), (9,3,2), (6,2,7), (5,1,8), (8,8,3), (9,9,4), (9,7,2), (7,9,2), (8,6,1), (9,5,0), (7,7,0), (6,8,1), (5,9,0), (4,0,-1), (6,0,1), (6,2,-1), (5,-1,-2), (3,1,-2), (3,-1,0), (7,-1,0), (5,-1,2), (7,1,-2), (5,3,-2), (2,2,-1), (4,4,-1), (1,3,-2), (1,1,0), (3,5,-2), (2,0,1), (4,0,3), (1,-1,2), (3,-1,4), (8,0,3), (8,2,1), (9,-1,2), (7,-1,4), (9,1,0), (6,0,5), (5,-1,6), (8,4,-1), (9,3,-2), (7,5,-2), (6,6,-1), (5,7,-2), (0,4,-1), (0,6,1), (2,6,-1), (-1,5,-2), (-1,3,0), (-1,7,0), (-1,5,2), (1,7,-2), (0,2,1), (0,4,3), (-1,1,2), (-1,3,4), (0,8,3), (2,8,1), (-1,9,2), (-1,7,4), (1,9,0), (0,6,5), (-1,5,6), (4,8,-1), (3,9,-2), (0,0,3), (0,2,5), (2,0,5), (-1,-1,4), (-1,1,6), (1,-1,6), (0,4,7), (2,2,7), (-1,3,8), (1,1,8), (4,0,7), (3,-1,8),
[0280] (Objective Function) The following equation (11) was used as the objective function for mathematical optimization. The first term on the right side represents the intramolecular interaction energy of the target peptide, and is expressed as the following equation (12). Variable a ij is the shortest Euclidean distance between a coarse-grained particle belonging to the i-th amino acid residue of the target peptide and a coarse-grained particle belonging to the j-th amino acid residue of the target peptide. If the shortest Euclidean distance is less than If it is equal to or greater than this, the variable e is 0. ij is the energy value assigned to the pair of the type of the i-th amino acid residue in the target peptide and the type of the j-th amino acid residue in the target peptide in the Miyazawa-Jernigan matrix. N is the total number of coarse-grained particles that make up the target peptide. The second term on the right side of equation (11) represents the intermolecular interaction energy between the target protein and the target peptide, and is expressed as shown in the following equation (13). Variable b ij is the shortest Euclidean distance between the coarse-grained particle belonging to the i-th amino acid residue of the target peptide and the coarse-grained particle belonging to the j-th amino acid residue of the target protein. If the shortest Euclidean distance is less than If it is equal to or greater than this, the variable e is 0. ij is the energy value assigned to the pair of the type of the i-th amino acid residue and the type of the j-th amino acid residue in the target peptide in the Miyazawa-Jernigan matrix. N is the total number of coarse-grained particles that make up the target peptide. L is the total number of coarse-grained particles that make up the target protein.
[0281] (Constraints) The following four constraints were set when performing mathematical optimization calculations. (1) All coarse-grained particles constituting the target peptide are placed at lattice points of a regular tetrahedral lattice, and multiple coarse-grained particles must not be placed at one lattice point. (2) The coarse-grained particle corresponding to the i-th amino acid residue of the target peptide and the coarse-grained particle corresponding to the (i+1)-th amino acid residue of the target peptide must be placed at two adjacent lattice points. (3) The coarse-grained particle corresponding to Pro1 of the target peptide and the coarse-grained particle corresponding to the main chain region of Gly6 must be placed at two adjacent lattice points. (4) The coarse-grained particle of the target peptide must not be placed at a lattice point whose Euclidean distance from the coarse-grained particle constituting the target protein is within the following value:
[0282] (Mathematical Optimization) To perform mathematical optimization, including calculation of interaction energy, OR-Tools version 9.5.2237 (https: / / pypi.org / project / ortools / 9.5.2237 / ), a constraint programming method, was used. Using OR-Tools, a coarse-grained model in which the target protein was rearranged by translation and rotation was input. Then, a docking model was obtained, which was an arrangement pattern of coarse-grained particles constituting the target peptide that gave the minimum value of the objective function E within the range satisfying the aforementioned constraints. The minimum value of the objective function E was −72.51.
[0283] (Verification) The degree of agreement between the obtained docking model and a coarse-grained model of the C chain of the crystal structure "3WNE" was verified. As an index of the verification, the root mean square deviation (RMSD) on the actual scale, shown in the following formula (14), was calculated. Here, the variable d i represents the Euclidean distance between the position of the i-th coarse-grained particle in the docking model and the position of the i-th coarse-grained particle in the coarse-grained model of the target peptide after rearrangement by translation and rotation operations. N is the total number of coarse-grained particles constituting the target peptide. As a result, it was confirmed that a docking model with a root mean square deviation value of 8.25 Å was obtained from the coarse-grained model of the target peptide after translation and rotation operations. The docking model and rearranged crystal structure in Comparative Example 1 are shown in Figure 10. The shaded circles represent coarse-grained particles of the target protein. The black circles represent coarse-grained particles corresponding to the amino acid residues of the target peptide. As shown in Figure 10, it was found that when an amino acid residue is approximated by a single coarse-grained particle, the docking model and the crystal structure diverge from each other.
[0284] 10...estimation system, 11...acquisition unit, 12...generation unit, 13...calculation unit, 14...estimation unit, 110...estimation program, 200...target molecular model, 201...coarse-grained particles, 210...center of gravity, 220...lattice system, 310, 320...constraint conditions, 311, 321...intermolecular constraints, 312, 322...intramolecular constraints, 313, 323...cyclization constraints, 230, 240, 250, 260, 270, 280...target molecular model, 231, 241, 251, 261, 262, 271, 272, 281, 282...coarse-grained particles.
Claims
1. A prediction system comprising at least one processor, wherein the at least one processor obtains a target molecule model showing a conformation of a target molecule, generates a target molecule model showing a conformation of a target molecule that is a molecule capable of binding to the target molecule, calculates an interaction energy between the target molecule and the target molecule as intermolecular interaction energy based on a positional relationship between the target molecule model and the target molecule model, calculates an interaction energy within the target molecule as intramolecular interaction energy, and predicts the conformation of the target molecule based on the intermolecular interaction energy and the intramolecular interaction energy.
2. The estimation system described in claim 1, wherein the at least one processor calculates multiple interaction energies between multiple constituent units constituting the target molecule and multiple constituent units constituting the target molecule as multiple inter-unit interaction energies, and calculates the intermolecular interaction energy based on the multiple inter-unit interaction energies.
3. The estimation system described in claim 1 or 2, wherein the at least one processor calculates multiple interaction energies between multiple structural units that constitute the target molecule as multiple inter-unit interaction energies, and calculates the intramolecular interaction energy based on the multiple inter-unit interaction energies.
4. The estimation system described in claim 2 or 3, wherein the at least one processor obtains the target molecule model that represents the multiple constituent units that constitute the target molecule using coarse-grained particles, and generates the target molecule model that represents the multiple constituent units that constitute the target molecule using the coarse-grained particles.
5. The estimation system according to claim 4, wherein the at least one processor arranges the jth structural unit and the (j+1)th structural unit constituting the target molecule adjacent to each other, where j is an integer of 1 or more, and generates the target molecule model in which the jth structural unit and the (j+1)th structural unit are represented by the coarse-grained particles.
6. The estimation system according to claim 4 or 5, wherein the at least one processor generates the target molecule model in which at least one of the plurality of building blocks constituting the target molecule is represented by a plurality of the coarse-grained particles.
7. The estimation system according to claim 6, wherein the at least one processor generates the target molecule model in which at least one of the plurality of building blocks constituting the target molecule is represented by two of the coarse-grained particles.
8. The estimation system according to any one of claims 4 to 7, wherein the at least one processor generates the target molecule model in which at least one of the plurality of structural units constituting the target molecule is represented by the coarse-grained particles representing a main chain and the coarse-grained particles representing a side chain.
9. The estimation system according to claim 8, wherein the at least one processor generates the target molecular model by arranging the coarse-grained particle representing the main chain and the coarse-grained particle representing the side chain adjacent to each other.
10. The estimation system according to any one of claims 1 to 9, wherein the at least one processor: generates a target molecular model for each of a plurality of conformations of the target molecule; calculates the intermolecular interaction energy based on the positional relationship between the target molecular model and the generated target molecular model; and selects at least one conformation from the plurality of conformations of the target molecule based on the intermolecular interaction energy for each of the plurality of conformations of the target molecule.
11. The prediction system according to claim 10, wherein the at least one processor generates a target molecular model for each of the multiple conformations of the target molecule, calculates an interaction energy within the target molecule as an intramolecular interaction energy based on the generated target molecular model, calculates the sum of the intramolecular interaction energy and the intermolecular interaction energy as a total interaction energy, and selects the at least one conformation from the multiple conformations of the target molecule based on the total interaction energy for each of the multiple conformations of the target molecule.
12. The estimation system according to any one of claims 1 to 11, wherein the at least one processor generates the target molecular model so as to satisfy constraint conditions including an intermolecular constraint regarding the positional relationship between the target molecular model and the target molecular model.
13. The prediction system of claim 12, wherein the intermolecular constraints indicate avoidance of steric hindrance between the subject molecule and the target molecule.
14. The estimation system described in claim 12 or 13, wherein the at least one processor sets an outer boundary of a lattice system for arranging the conformation of the target molecule based on the target molecular model, and generates the target molecular model within the lattice system so that the target molecular model satisfies the constraints within the lattice system.
15. The prediction system according to claim 14, wherein the at least one processor obtains the target molecule model representing the conformation of an active site in the target molecule, and sets the outer boundary of the lattice system based on the center of gravity of multiple building blocks that constitute the active site.
16. The estimation system according to claim 14 or 15, wherein a unit cell of the lattice system is any one selected from a regular tetrahedral lattice, a cubic lattice, and a square lattice.
17. The estimation system according to any one of claims 14 to 16, wherein the at least one processor generates the target molecule model in the lattice system by arranging two coarse-grained particles corresponding to the jth and (j+1)th building blocks constituting the target molecule adjacent to each other in the lattice system, where j is an integer equal to or greater than 1.
18. The estimation system according to any one of claims 14 to 17, wherein the at least one processor generates the target molecule model within the lattice system by arranging a coarse-grained particle representing a main chain and a coarse-grained particle representing a side chain adjacent to each other within the lattice system for at least one of a plurality of building blocks constituting the target molecule.
19. The prediction system according to any one of claims 12 to 18, wherein the constraint conditions further include at least one of an intramolecular constraint indicating that steric hindrance is avoided between multiple building blocks constituting the target molecule, and a cyclization constraint indicating that the target molecule contains a cyclic structure.
20. The estimation system according to any one of claims 1 to 19, wherein the target molecule is a protein, and the target molecule is a peptide.
21. A method for predicting a molecular conformation, executed by a prediction system having at least one processor, comprising: a step of obtaining a target molecular model showing the conformation of a target molecule; a step of generating a target molecular model showing the conformation of a target molecule that is a molecule capable of binding to the target molecule; a step of calculating an interaction energy between the target molecule and the target molecule as intermolecular interaction energy based on a positional relationship between the target molecular model and the target molecular model; a step of calculating an interaction energy within the target molecule as intramolecular interaction energy; and a step of predicting the conformation of the target molecule based on the intermolecular interaction energy and the intramolecular interaction energy.
22. An estimation program that causes a computer to execute the steps of: obtaining a target molecule model that indicates the conformation of a target molecule; generating a target molecule model that indicates the conformation of a target molecule that can bind to the target molecule; calculating the interaction energy between the target molecule and the target molecule as intermolecular interaction energy based on the positional relationship between the target molecule model and the target molecule model; calculating the interaction energy within the target molecule as intramolecular interaction energy; and estimating the conformation of the target molecule based on the intermolecular interaction energy and the intramolecular interaction energy.
Citation Information
Patent Citations
Structure search method of cyclic molecule and structure search device as well as program
JP2020091518A
Bond structure search device, bond structure search method, and bond structure search program
JP2020173643A
Structure search device, structure search method, and structure search program
JP2020194487A
Method of searching the structure of stable biopolymer-ligand molecule composite
WO1993020525A1