Structure search program, structure search method, and information processing device
By generating and switching coarse-grained models with linear springs and back-mapping to atomic structures, the method expands the search range of protein structural space, overcoming limitations of conventional techniques.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- FUJITSU LTD
- Filing Date
- 2024-11-08
- Publication Date
- 2026-05-20
AI Technical Summary
Existing methods like MD, CGMD, and conventional CG-ENM struggle to explore the structural space of proteins beyond the vicinity of the reference structure, limiting the sampling of structures far from the initial configuration.
The method involves generating coarse-grained models by replacing protein atoms with particles connected by linear springs, switching reference structures through CG-ENM calculations, and back-mapping to atomic structures, thereby expanding the search range of structural space.
This approach allows for the sampling of a broader range of protein structures, including those with significant deviations from the initial structure, enhancing the exploration of structural space.
Smart Images

Figure 2026083982000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a structure exploration program and the like.
Background Art
[0002] Proteins are known to exhibit various functions due to structural changes, and it is required to widely explore the structural space.
[0003] For example, as conventional techniques for exploring the structure of proteins, there are classical molecular dynamics (Molecular Dynamics, hereinafter referred to as MD), coarse-grained MD (Coarse-Grained MD, hereinafter referred to as CGMD), and coarse-grained elastic network model (Coarse-Grained Elastic Network Model, hereinafter referred to as CG-ENM).
[0004] From the perspective of computational cost, it is difficult for MD and CGMD to widely explore the structural space. On the other hand, CG-ENM (Adaptive CG-ENM) is a method that can explore the structural space of proteins more widely compared to MD and CGMD. Although Tirion-type CG-ENM is widely known as CG-ENM (Non-Patent Document 1), in recent years, Adaptive CG-ENM (Non-Patent Document 2) has been reported by Kaneda, one of the inventors.
[0005] FIG. 7 is a diagram showing the results of exploring the structural space by the prior art. The horizontal axis of graph G1 in FIG. 7 corresponds to the first principal component (PC1) in principal component analysis. The vertical axis of graph G1 corresponds to the second principal component (PC2) in principal component analysis. Also, the color intensity of each point in graph G1 corresponds to Q-score: contact fraction of native contact pairs (hereinafter sometimes referred to as Q-score), and the darker the color, the larger the Q-score. The Q-score is a score indicating how appropriate the structure of the protein is, and the more appropriate the structure, the larger the score.
[0006] Point p1 corresponds to the reference structure (initial structure). For example, the reference structure is the structure of an existing protein. Examples of existing protein structures include the Holo structure and the Apo structure.
[0007] MD, CGMD, and CG-ENM perform a structural space search (sampling) on a reference structure. Point clouds p2 and p3 are the sampling results from MD. For example, the MD computation time required to sample point cloud p2 was "50 ns". The MD computation time required to sample point cloud p3 was "1 μs". Point cloud p4 is the sampling result from a conventional Tirion-type CG-ENM. Point cloud p5 is the sampling result from Adaptive CG-ENM. [Prior art documents] [Patent Documents]
[0008] [Patent Document 1] Japanese Patent Publication No. 2010-113473 [Patent Document 2] Japanese Patent Publication No. 2003-272980 [Patent Document 3] U.S. Patent Application Publication No. 2004 / 0102941 [Patent Document 4] U.S. Patent Application Publication No. 2008 / 0082305 [Non-patent literature]
[0009] [Non-Patent Document 1] MM Tirion, Phys. Rev. Lett. 77, 1905-1908 (1996). [Non-Patent Document 2] R. Kanada et al. -J. Chem. Theory Comput. 18, 2062-2074 (2022). [Overview of the project] [Problems that the invention aims to solve]
[0010] However, even when using CG-ENM to search the structure space, the search space is limited to the vicinity of the reference structure (initial structure), making it difficult to sample structures far from the reference structure.
[0011] In one aspect, the present invention aims to provide a structural search program, a structural search method, and an information processing device that can expand the search range of a structural space. [Means for solving the problem]
[0012] In the first proposal, the computer performs the following steps: The computer replaces several atoms of the protein's first reference structure with coarse-grained particles and generates a first coarse-grained model by connecting each coarse-grained particle with a linear spring constant based on the first reference structure. The computer selects one of several second coarse-grained models with different structures, obtained by changing the positions of the coarse-grained particles in the first coarse-grained model in several patterns. The computer generates a new second reference structure corresponding to the selected second coarse-grained model by restoring the coarse-grained particles of the selected second coarse-grained model to their original atoms. The computer replaces several atoms of the second reference structure with coarse-grained particles and generates a third coarse-grained model by connecting each coarse-grained particle with a linear spring constant based on the second reference structure. The computer selects one of several fourth coarse-grained models with different structures, obtained by changing the positions of the coarse-grained particles in the third coarse-grained model in several patterns. [Effects of the Invention]
[0013] This allows us to expand the search range of the structural space. [Brief explanation of the drawing]
[0014] [Figure 1] Figure 1 is a diagram illustrating CG-ENM. [Figure 2]FIG. 2 is a diagram showing an example of sampling results by Adaptive CG-ENM. [Figure 3] FIG. 3 is a diagram showing the configuration of the information processing apparatus according to the present embodiment. [Figure 4] FIG. 4 is a flowchart showing the processing procedure of the information processing apparatus according to the present embodiment. [Figure 5] FIG. 5 is a diagram for explaining the effect of the information processing apparatus according to the present embodiment. [Figure 6] FIG. 6 is a diagram showing an example of the hardware configuration of a computer that realizes the same functions as the information processing apparatus according to the present embodiment. [Figure 7] FIG. 7 is a diagram showing the search results of the structural space according to the prior art. MODE FOR CARRYING OUT THE INVENTION
[0015] Hereinafter, embodiments of the structure search program, structure search method, and information processing apparatus disclosed in the present application will be described in detail based on the drawings. Note that the present invention is not limited by this embodiment. EXAMPLE
[0016] Before describing this embodiment, CG-ENM will be described more specifically. CG-ENM is a model in which the structure of a protein is represented by coarse-grained particles and the coarse-grained particles are connected by linear springs. FIG. 1 is a diagram for explaining CG-ENM. In the example shown in FIG. 1, the protein structure 10 has the structure of Crambin. The CG-ENM 11 is a CG-ENM generated based on the protein structure 10. In the CG-ENM 11, residues with strong correlation are connected by linear springs.
[0017] When generating CG-ENM11, the position of the Cα atom in each amino acid residue of protein structure 10 is used as the position of the coarse-grained particle, and one coarse-grained particle represents the amino acid residue. In conventional Tirion-type CG-ENM, a linear spring is set between pairs of amino acid residues that form a native contact in the reference structure (initial structure). A predetermined spring constant is used for this linear spring.
[0018] For example, a native contact is defined as follows: If the distance between one of the heavy atoms in the i-th residue and one of the atoms in the j-th residue is less than or equal to a certain threshold, then a native contact is considered to have formed between the heavy atom in the i-th residue and the atom in the j-th residue. However, this excludes hydrogen (H). The threshold is usually "6.5 Å".
[0019] Conventional Tirion-type CG-ENMs set up linear springs between residue pairs that form native contacts in the reference structure, making it difficult to sample structures that are far removed from the reference structure.
[0020] On the other hand, Adaptive CG-ENM determines the residue pairs and spring constants (linear spring constants) for setting linear springs using Bayesian optimization based on a dynamic cross-correlation map (DCCM) between residues. For example, Adaptive CG-ENM sets a stronger spring constant for residue pairs with a stronger correlation based on the DCCM, and does not set a linear spring for residue pairs whose correlation is below a threshold.
[0021] Adaptive CG-ENM performs MD (millennial MD) on all atoms from the reference structure and defines the spring constant based on DCCM between residues. However, even in this case, DCCM is strongly influenced by the reference structure, making it difficult to sample structures that are far removed from the reference structure.
[0022] In the following explanation, the CG-ENM (the model itself) generated from the reference structure will be referred to as "CG-ENM". Furthermore, the calculation that samples multiple structures (structures represented by coarse-grained particles) by applying various changes (such as applying heat) to the coarse-grained particles of the CG-ENM will be referred to as "CG-ENM calculation".
[0023] Figure 2 shows an example of sampling results obtained by Adaptive CG-ENM calculation. In the example shown in Figure 2, the Apo state of ADK (Adenylate kinase) is used as the reference structure, and the results of sampling by Adaptive CG-ENM calculation are obtained. The horizontal axis of graph G2 corresponds to the first principal component (PC1) in principal component analysis. The vertical axis of graph G2 corresponds to the second principal component (PC2) in principal component analysis. Furthermore, the color of each point corresponds to the RMSD (Root Mean Square Deviation), and the closer the color of a point is to the color with a higher numerical value set in bar 5, the greater the structural change from the reference structure.
[0024] Point p10 corresponds to the reference structure (initial structure, Apo structure). Point p11 corresponds to the Holo structure.
[0025] For example, a device that performs sampling using Adaptive CG-ENM calculations (hereinafter referred to as "device") performs sampling in the following steps (1) to (3).
[0026] (1) The apparatus uses the Apo state as its initial structure and acquires DCCM by whole-atom MD.
[0027] (2) Based on DCCM, the instrument determines the residue pairs and spring constants for setting the linear spring using Bayesian optimization. Based on the determined residue pairs and spring constants, the instrument performs an Adaptive CG-ENM calculation for a shorter period than usual and calculates a score. This score is a value calculated based on the Q-score and RMSD for the current CG-ENM structure. The instrument repeats (2) to identify the residue pairs and spring constants that yield the highest score as the optimal parameters.
[0028] (3) The apparatus performs Adaptive CG-ENM calculations over a normal period using optimal parameters and samples multiple structures.
[0029] As shown in Figure 2, even with Adaptive CG-ENM calculations, the sampling region is concentrated around point p10, making it difficult to sample structures far removed from the reference structure.
[0030] The above provides a more detailed explanation of CG-ENM.
[0031] Next, the information processing device according to this embodiment will be described. The information processing device according to this embodiment will be referred to as "information processing device 100". In CG-ENM calculations, the information processing device 100 expands the search range of the structure space by switching reference structures. For example, the information processing device 100 performs a CG-ENM calculation on a reference structure, returns the coarse-grained model obtained back to the structure before coarse-graining, and repeatedly performs a CG-ENM calculation using the returned structure as the new reference structure.
[0032] In a single CG-ENM calculation, the search range is limited to the vicinity of the reference structure. In contrast, the information processing device 100 can expand the search range of the structure space by taking a new coarse-grained structure sampled from the previous CG-ENM calculation, converting it back into the corresponding total atomic structure to create a new reference structure, and sequentially switching between these reference structures.
[0033] Here, an example of the configuration of the information processing device 100 in this embodiment will be described. Figure 3 is a diagram showing the configuration of the information processing device according to this embodiment. As shown in Figure 3, this information processing device 100 has a communication unit 110, an input unit 120, a display unit 130, a storage unit 140, and a control unit 150.
[0034] The communication unit 110 performs data communication with external devices, etc., via the network. The communication unit 110 is a NIC (Network Interface Card), etc. For example, the communication unit 110 may obtain reference structure data 141, etc., from external devices, etc.
[0035] The input unit 120 is an input device that inputs various types of information to the control unit 150 of the information processing device 100. For example, the input unit 120 can be a keyboard, mouse, touch panel, etc.
[0036] The display unit 130 is a display device that displays information output from the control unit 150.
[0037] The storage unit 140 has reference structure data 141 and sampling data 142. The storage unit 140 is a memory or the like.
[0038] Reference structure data 141 contains information about the reference structure (initial structure) of a real protein. For example, reference structure data 141 contains information about multiple amino acid residues, such as protein structure 10 shown in Figure 1.
[0039] Sampling data 142 is information about the structures of multiple proteins sampled by the control unit 150.
[0040] The control unit 150 includes a preprocessing unit 151, an extraction unit 152, a back mapping unit 153, a calculation unit 154, and a determination unit 155. The control unit 150 is a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), etc.
[0041] The preprocessing unit 151 acquires reference structure data 141 and generates a CG-ENM corresponding to the reference structure by setting linear springs between coarse-grained particles included in the protein structure. The preprocessing unit 151 may generate the CG-ENM in the same manner as the conventional Tirion-type CG-ENM described above, or it may generate the CG-ENM in the same manner as the Adaptive CG-ENM.
[0042] Furthermore, the preprocessing unit 151 performs CG-ENM calculations using the generated CG-ENM corresponding to the reference structure and samples the structures of multiple proteins (structures represented by coarse-grained particles). The preprocessing unit 151 outputs the sampling results to the extraction unit 152.
[0043] Each time the extraction unit 152 acquires a sampling result, it extracts one structure from the structures of multiple proteins included in the sampling result (structures represented by coarse-grained particles) as a new reference structure. For example, the extraction unit 152 extracts one structure as a new reference structure based on the first extraction criterion, the second extraction criterion, or the third extraction criterion.
[0044] The first extraction criterion will now be explained. The extraction unit 152 randomly selects one structure from among the multiple protein structures included in the sampling results to be used as a new reference structure.
[0045] The second extraction criterion will now be explained. The extraction unit 152 calculates the Q-score and RMSD for each of the multiple protein structures included in the sampling results. From the multiple protein structures whose Q-score is above the first threshold and whose RMSD is above the second threshold, the extraction unit 152 randomly selects one protein structure as the new reference structure. The extraction unit 152 may also select the new reference structure using either the Q-score or the RMSD.
[0046] The third extraction criterion will now be explained. When the extraction unit 152 repeats the process of extracting one structure as a new reference structure, it retains the principal component analysis results of structures extracted in the past. The extraction unit 152 compares the principal component analysis results (A) of multiple protein structures included in the sampling results with the principal component analysis results (B) of multiple structures extracted in the past. The extraction unit 152 extracts the protein structure with the principal component analysis result (A) that is furthest from the distribution of the principal component analysis result (B) as the new reference structure.
[0047] The extraction unit 152 outputs information about the newly extracted reference structure (structure represented by coarse-grained particles) to the back-mapping unit 153.
[0048] The back-mapping unit 153 back-maps the reference structure (structure represented by coarse-grained particles) extracted by the extraction unit 152 to the reference structure (whole atomic structure).
[0049] For example, the backmapping unit 153 may perform backmapping using a template-based method, a force field-based method, a machine learning-based method, or the like.
[0050] In the template-based method, the back-mapping unit 153 uses known atomic-level structures to position atoms in the coarse-grained particles.
[0051] In the force-field-based method, the back-mapping unit 153 uses an energy minimization algorithm to position atoms in physically reasonable locations.
[0052] In machine learning-based methods, the backmapping unit 153 uses a Neural Network (NN) or other machine learning algorithms to predict and position the atoms.
[0053] The backmapping unit 153 may perform backmapping using other existing methods. The backmapping unit 153 outputs the reference structure (total atomic structure) to the calculation unit 154.
[0054] The back mapping unit 153 repeatedly performs the above process each time it obtains reference structure information from the extraction unit 152.
[0055] The calculation unit 154 obtains a reference structure (total atomic structure) and generates a CG-ENM corresponding to the reference structure by setting linear springs between coarse-grained particles included in the protein structure. The calculation unit 154 may generate the CG-ENM in the same manner as the conventional Tirion-type CG-ENM described above, or it may generate the CG-ENM in the same manner as the Adaptive CG-ENM.
[0056] The calculation unit 154 performs CG-ENM calculations using the generated CG-ENM corresponding to the reference structure and samples the structures of multiple proteins (structures represented by coarse-grained particles). The calculation unit 154 registers the sampling results in the sampling data 142 of the storage unit 140.
[0057] The calculation unit 154 repeatedly performs the above process each time it obtains a reference structure (total atomic structure).
[0058] The determination unit 155 determines whether or not to terminate the structure search based on the sampling results registered in the sampling data 142.
[0059] For example, the determination unit 155 calculates the RMSD for the structures of multiple proteins registered in the sampling data 142, and determines to terminate the structure search if the RMSD value is equal to or greater than a preset threshold. The determination unit 155 may also use Holo structures or Apo structures as comparison targets for the structures of the sampling results when calculating the RMSD.
[0060] The determination unit 155 may determine whether or not to terminate the structure search based on predetermined structures (structures of proteins whose existence has been confirmed) recorded in the Protein Data Bank.
[0061] If the determination unit 155 determines, based on the sampling results registered in the sampling data 142, to continue the structure search, it outputs the most recent sampling result calculated by the calculation unit 154 to the extraction unit 152, and causes the processing units 152 to 154 to execute the above process again.
[0062] Next, an example of the processing procedure of the information processing device 100 according to this embodiment will be described. Figure 4 is a flowchart of the processing procedure of the information processing device according to this embodiment. As shown in Figure 4, the preprocessing unit 151 of the information processing device 100 acquires the reference structure data 141 (step S101).
[0063] The preprocessing unit 151 generates a CG-ENM corresponding to the reference structure by setting linear springs between the coarse-grained particles contained in the protein structure (step S102). The preprocessing unit 151 performs a CG-ENM calculation using the CG-ENM corresponding to the reference structure and samples the structures of multiple proteins (structures represented by coarse-grained particles) (step S103).
[0064] The extraction unit 152 extracts one structure from the structures of multiple proteins included in the sampling results (structures represented by coarse-grained particles) as a new reference structure (step S104).
[0065] The back-mapping unit 153 back-maps the new reference structure (structure represented by coarse-grained particles) to the reference structure (total atomic structure) (step S105).
[0066] The calculation unit 154 obtains a reference structure (total atomic structure) and generates a CG-ENM corresponding to the reference structure by setting linear springs between coarse-grained particles included in the protein structure (step S106). The calculation unit 154 uses the CG-ENM corresponding to the reference structure to perform a CG-ENM calculation and samples the structures of multiple proteins (structures represented by coarse-grained particles) (step S107).
[0067] If the determination unit 155 determines that the structure search is insufficient (step S108, No), it proceeds to step S104. On the other hand, if the determination unit 155 determines that the structure search is sufficient (step S108, Yes), it terminates the process.
[0068] Next, the effects of the information processing device 100 according to this embodiment will be described. The information processing device 100 repeatedly performs CG-ENM calculations on a reference structure, returns the coarse-grained model obtained to the original all-atom structure, and then uses the returned all-atom structure as a new reference structure to perform CG-ENM calculations. This expands the search range of the structural space.
[0069] Figure 5 is a diagram illustrating the effects of the information processing device according to this embodiment. Graph G3 in Figure 5 shows the sampling results using the prior art method. Graph G4 shows the sampling results using the information processing device 100. The horizontal axis of graphs G3 and G4 corresponds to the first principal component (PC1) in principal component analysis. The vertical axis of graphs G3 and G4 corresponds to the second principal component (PC2) in principal component analysis. The color of each point in graphs G3 and G4 corresponds to RMSD, and the closer the color of a point is to the color with a higher numerical value set in bar 5, the greater the structural change from the reference structure.
[0070] Point p10 corresponds to the reference structure (initial structure, Apo structure). Point p11 corresponds to the Holo structure.
[0071] In the conventional technique, the reference structure at point p10 is used as the initial structure, and the above steps (1) to (3) are performed to perform sampling. On the other hand, the information processing device 100 uses the reference structure at point p10 as the initial structure, performs a CG-ENM calculation on the reference structure, and then returns the resulting coarse-grained model to the original total atomic structure before coarse-graining. This process is repeated using the returned total atomic structure as the new reference structure to perform the CG-ENM calculation.
[0072] Comparing graphs G3 and G4, it can be seen that structures that could not be sampled with conventional techniques (those with a large RMSD for Holo structures) can now be sampled. In other words, the information processing device 100 can expand the search range of the structure space.
[0073] The information processing device 100 calculates the Q-score and RMSD for each of the multiple protein structures included in the sampling results, and extracts a new reference structure using the Q-score and / or RMSD. This efficiently expands the search range of the structure search.
[0074] Next, an example of a computer hardware configuration that realizes the same functions as the information processing device 100 described above will be explained. Figure 6 is a diagram showing an example of a computer hardware configuration that realizes the same functions as the information processing device according to this embodiment.
[0075] As shown in Figure 6, the computer 200 includes a CPU 201 that performs various calculations, an input device 202 that receives data input from the user, and a display 203. The computer 200 also includes a communication device 204 and an interface device 205 that exchange data with external devices via a wired or wireless network. Furthermore, the computer 200 includes a RAM 206 for temporarily storing various information and a hard disk drive 207. Each of these devices 201 to 207 is connected to a bus 208.
[0076] The hard disk drive 207 includes a preprocessing program 207a, an extraction program 207b, a back mapping program 207c, a calculation program 207d, and a determination program 207e. The CPU 201 reads each of the programs 207a to 207e and loads them into the RAM 206.
[0077] The preprocessing program 207a functions as the preprocessing process 206a. The extraction program 207b functions as the extraction process 206b. The back mapping program 207c functions as the back mapping process 206c. The calculation program 207d functions as the calculation process 206d. The determination program 207e functions as the determination process 206e.
[0078] The processing of the preprocessing process 206a corresponds to the processing of the preprocessing unit 151. The processing of the extraction process 206b corresponds to the processing of the extraction unit 152. The processing of the back mapping process 206c corresponds to the processing of the back mapping unit 153. The processing of the calculation process 206d corresponds to the processing of the calculation unit 154. The processing of the determination process 206e corresponds to the processing of the determination unit 155.
[0079] Furthermore, programs 207a to 207e do not necessarily have to be stored on the hard disk drive 207 from the beginning. For example, each program could be stored on a "portable physical medium" such as a flexible disk (FD), CD-ROM, DVD, magneto-optical disk, or IC card inserted into the computer 200. Then, the computer 200 could read and execute each program 207a to 207e.
[0080] With regard to embodiments including each of the above examples, the following additional information is disclosed.
[0081] (Note 1) By replacing multiple atoms of the first reference structure of the protein with coarse-grained particles, and by using a linear spring with a spring constant based on the first reference structure to connect each coarse-grained particle, a first coarse-grained model is generated. From among several second coarse-graining models with different structures obtained by changing the position of the coarse-grained particles in the first coarse-graining model in multiple patterns, one second coarse-graining model is selected. By returning the coarse-grained particles of the selected second coarse-graining model back to the original atoms, a new second reference structure corresponding to the selected second coarse-graining model is generated. By replacing several atoms in the second reference structure with coarse-grained particles, and by using a linear spring with a spring constant based on the second reference structure to connect each coarse-grained particle, a third coarse-grained model is generated. From among several fourth coarse-graining models with different structures, obtained by changing the position of the coarse-grained particles in the third coarse-graining model in multiple patterns, one fourth coarse-graining model is selected. A structure search program characterized by having a computer perform the processing.
[0082] (Note 2) The structure search program described in Note 1, characterized in that the process of selecting the second coarse-grained model is to perform a CG-ENM (Coarse-Grained Elastic Network Model) calculation to select one of several second coarse-grained models with different structures obtained by changing the positions of the coarse-grained particles of the first coarse-grained model in multiple patterns.
[0083] (Note 3) The structure search program described in Note 2, characterized in that the process of selecting the second coarse-graining model is to select one of the multiple second coarse-graining models based on their RMSD (Root Mean Square Deviation).
[0084] (Note 4) The structure search program according to Note 2, characterized in that the process of selecting the second coarse-graining model is to select one of the multiple second coarse-graining models based on their Q-scores.
[0085] (Note 5) By replacing multiple atoms of the first reference structure of the protein with coarse-grained particles, and by using a linear spring with a spring constant based on the first reference structure to connect each coarse-grained particle, a first coarse-grained model is generated. From among several second coarse-graining models with different structures obtained by changing the position of the coarse-grained particles in the first coarse-graining model in multiple patterns, one second coarse-graining model is selected. By returning the coarse-grained particles of the selected second coarse-graining model back to the original atoms, a new second reference structure corresponding to the selected second coarse-graining model is generated. By replacing several atoms in the second reference structure with coarse-grained particles, and by using a linear spring with a spring constant based on the second reference structure to connect each coarse-grained particle, a third coarse-grained model is generated. From among several fourth coarse-graining models with different structures, obtained by changing the position of the coarse-grained particles in the third coarse-graining model in multiple patterns, one fourth coarse-graining model is selected. A structure search method characterized by having a computer perform the processing.
[0086] (Note 6) The structure search method according to Note 5, characterized in that the process of selecting the second coarse-grained model is to perform a CG-ENM (Coarse-Grained Elastic Network Model) calculation to select one of several second coarse-grained models with different structures obtained by changing the positions of the coarse-grained particles of the first coarse-grained model in multiple patterns.
[0087] (Note 7) The structure search method according to Note 6, characterized in that the process of selecting the second coarse-graining model is to select one of the multiple second coarse-graining models based on their RMSD (Root Mean Square Deviation).
[0088] (Note 8) The structure search method according to Note 6, characterized in that the process of selecting the second coarse-graining model is to select one of the multiple second coarse-graining models based on their Q-scores.
[0089] (Note 9) By replacing several atoms of the first reference structure of the protein with coarse-grained particles, and by using a linear spring with a spring constant based on the first reference structure to connect each coarse-grained particle, a first coarse-grained model is generated. From among several second coarse-graining models with different structures obtained by changing the position of the coarse-grained particles in the first coarse-graining model in multiple patterns, one second coarse-graining model is selected. By returning the coarse-grained particles of the selected second coarse-graining model back to the original atoms, a new second reference structure corresponding to the selected second coarse-graining model is generated. By replacing several atoms in the second reference structure with coarse-grained particles, and by using a linear spring with a spring constant based on the second reference structure to connect each coarse-grained particle, a third coarse-grained model is generated. From among several fourth coarse-graining models with different structures, obtained by changing the position of the coarse-grained particles in the third coarse-graining model in multiple patterns, one fourth coarse-graining model is selected. An information processing device having a control unit that performs processing.
[0090] (Note 10) The information processing apparatus according to Note 9, wherein the process of selecting the second coarse-graining model is characterized by performing a CG-ENM (Coarse-Grained Elastic Network Model) calculation to select one of several second coarse-graining models with different structures obtained by changing the positions of the coarse-grained particles of the first coarse-graining model into multiple patterns.
[0091] (Note 11) The information processing apparatus according to Note 10, characterized in that the process of selecting the second coarse-graining model is to select one of the multiple second coarse-graining models based on their RMSD (Root Mean Square Deviation).
[0092] (Appendix 12) The information processing apparatus according to Appendix 10, characterized in that the process of selecting the second coarse-graining model is to select one of the multiple second coarse-graining models based on their Q-scores. [Explanation of Symbols]
[0093] 100 Information Processing Devices 110 Communications Department 120 Input section 130 Display section 140 Storage section 141 Reference Structure Data 142 sampled data 150 Control Unit 151 Preprocessing Unit 152 Extraction part 153 Back Mapping Section 154 Calculation Department 155 Judgment section
Claims
1. By replacing multiple atoms of the protein's first reference structure with coarse-grained particles, and by using a linear spring with a spring constant based on the first reference structure to connect the coarse-grained particles, a first coarse-grained model is generated. From among several second coarse-graining models with different structures obtained by changing the position of the coarse-grained particles in the first coarse-graining model in multiple patterns, one second coarse-graining model is selected. By returning the coarse-grained particles of the selected second coarse-graining model back to the original atoms, a new second reference structure corresponding to the selected second coarse-graining model is generated. By replacing each of the atoms in the second reference structure with coarse-grained particles, and by using a linear spring with a spring constant based on the second reference structure to connect each coarse-grained particle, a third coarse-grained model is generated. From among several fourth coarse-graining models with different structures obtained by changing the position of the coarse-grained particles in the third coarse-graining model in multiple patterns, one fourth coarse-graining model is selected. A structure search program characterized by having a computer perform the processing.
2. The structure search program according to claim 1, characterized in that the process of selecting the second coarse-graining model is to perform a CG-ENM (Coarse-Grained Elastic Network Model) calculation to select one of several second coarse-graining models with different structures obtained by changing the positions of the coarse-grained particles of the first coarse-graining model in multiple patterns.
3. The structure search program according to claim 2, characterized in that the process of selecting the second coarse-graining model selects one of the multiple second coarse-graining models based on their RMSD (Root Mean Square Deviation).
4. The structure search program according to claim 2, characterized in that the process of selecting the second coarse-graining model selects one of the multiple second coarse-graining models based on their Q-scores.
5. By replacing multiple atoms of the protein's first reference structure with coarse-grained particles, and by using a linear spring with a spring constant based on the first reference structure to connect the coarse-grained particles, a first coarse-grained model is generated. From among several second coarse-graining models with different structures obtained by changing the position of the coarse-grained particles in the first coarse-graining model in multiple patterns, one second coarse-graining model is selected. By returning the coarse-grained particles of the selected second coarse-graining model back to the original atoms, a new second reference structure corresponding to the selected second coarse-graining model is generated. By replacing each of the atoms in the second reference structure with coarse-grained particles, and by using a linear spring with a spring constant based on the second reference structure to connect each coarse-grained particle, a third coarse-grained model is generated. From among several fourth coarse-graining models with different structures obtained by changing the position of the coarse-grained particles in the third coarse-graining model in multiple patterns, one fourth coarse-graining model is selected. A structure search method characterized by having a computer perform the processing.
6. By replacing multiple atoms of the protein's first reference structure with coarse-grained particles, and by using a linear spring with a spring constant based on the first reference structure to connect the coarse-grained particles, a first coarse-grained model is generated. From among several second coarse-graining models with different structures obtained by changing the position of the coarse-grained particles in the first coarse-graining model in multiple patterns, one second coarse-graining model is selected. By returning the coarse-grained particles of the selected second coarse-graining model back to the original atoms, a new second reference structure corresponding to the selected second coarse-graining model is generated. By replacing each of the atoms in the second reference structure with coarse-grained particles, and by using a linear spring with a spring constant based on the second reference structure to connect each coarse-grained particle, a third coarse-grained model is generated. From among several fourth coarse-graining models with different structures obtained by changing the position of the coarse-grained particles in the third coarse-graining model in multiple patterns, one fourth coarse-graining model is selected. An information processing device having a control unit that performs processing.