Information processing program, information processing method, and information processing device
Patent Information
- Application Number
- JP2025032353
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2026-09-09
AI Technical Summary
【0009】 1態様によれば、3次元密度マップに対して適切な初期位置への分子の初期構造の配置が可能となる。
Smart Images

Figure 2026144824000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing program, an information processing method, and an information processing apparatus. [Background technology]
[0002] In drug discovery and development, analyzing the structure of macromolecules such as proteins is crucial. Cryo-electron microscopy (cryo-EM) can be used for molecular structure analysis. When a sample is analyzed with cryo-EM, voxel data called a three-dimensional density map is obtained. However, this three-dimensional density map only reveals the shape of the protein; the positions and connections of the individual atoms that make up the protein remain unknown. If the positions of the atoms are known, simulations can be used to predict the temporal changes in structure and the binding of ligands (candidate drug compounds).
[0003] Therefore, the three-dimensional structure of a molecule is reconstructed based on a three-dimensional density map. The three-dimensional structure shows the positions of the atoms that make up the molecule. For example, there is a technique (fitting) that obtains a three-dimensional structure corresponding to the three-dimensional density map by deforming the known structure of a molecule to match the three-dimensional density map.
[0004] Regarding fitting-based technologies, techniques for peptide / nucleic acid compositions for oral / mucosal bimodal activation of the immune defense system have been proposed. Methods for imaging protein structures using cryo-EM have also been proposed. As a cryo-EM technique, a method for modifying target nucleic acids using mutant CRISPR (Clustered Regularly Interspaced Short Palindromic Repeat)-Cas effector polypeptides has also been proposed. Furthermore, as a cryo-EM technique, a method for producing synthetic single-domain monoclonal antibody libraries using humanized ramana nanobody framework sequences has been proposed. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Special Publication No. 2012-519006 [Patent Document 2] U.S. Patent Application Publication No. 2023 / 0108717 [Patent Document 3] U.S. Patent Application Publication No. 2023 / 0407276 Specification [Patent Document 4] Special Publication No. 2024-535249 [Overview of the project] [Problems that the invention aims to solve]
[0006] The accuracy of the three-dimensional structure obtained through fitting depends on the initial position of the existing structure during fitting. For example, the existing structure of an object is moved to the initial fitting position by a rigid body transformation that matches the three-dimensional density map. However, conventional rigid body transformations may move the existing structure to an inappropriate initial position. If the initial position is inappropriate, it may not be possible to obtain an accurate three-dimensional structure of the target molecule even after fitting.
[0007] In one aspect, this project aims to enable the placement of the initial structure of molecules at appropriate initial positions relative to a three-dimensional density map. [Means for solving the problem]
[0008] One proposal provides an information processing program that instructs a computer to perform the following tasks: A computer acquires first structure information indicating a first structure of a molecule and second structure information indicating a second structure estimated based on a three-dimensional density map of the molecule. The computer calculates, for each of a plurality of second atoms in at least a part of the second structure, a first value indicating the reliability of a partial structure including the second atom based on a comparison result between the first structure and the second structure. Then, the computer weights a value for evaluating the position of each of a plurality of first atoms of the first structure with the first value of the second atom corresponding to the evaluated first atom, and determines the position and orientation of the first structure to be superimposed on the three-dimensional density map based on a second value obtained by calculation using the weighted value. [Effects of the Invention]
[0009] According to one aspect, an initial structure of a molecule can be arranged at an appropriate initial position with respect to a three-dimensional density map. [Brief Description of Drawings]
[0010] [Figure 1] It is a diagram illustrating an example of an information processing method according to a first embodiment. [Figure 2] It is a diagram illustrating an example of a system configuration according to a second embodiment. [Figure 3] It is a diagram illustrating an example of positions of information transferred between devices. [Figure 4] It is a diagram illustrating an example of hardware of a server. [Figure 5] It is a diagram illustrating an example of positions of structural changes in a protein. [Figure 6] It is a diagram illustrating an example of fitting. [Figure 7] It is a diagram illustrating an example of the relationship between an initial position and ease of fitting. [Figure 8] It is a diagram illustrating an example of functions of a server for protein structure analysis. [Figure 9] It is a diagram illustrating an example of structure information. [Figure 10] It is a diagram illustrating an example of three-dimensional density map information. [Figure 11] It is a flowchart showing an example of a procedure for analysis processing of three-dimensional structures. [Figure 12] It is a diagram showing an example of reference structure generation by a main chain tracing method. [Figure 13] It is a diagram showing an example of reliability calculation. [Figure 14] It is a diagram showing an example of a calculation method for RMSD. [Figure 15] It is a diagram showing an example of rotation of an initial structure using reliability. [Figure 16] It is a diagram showing an example of rigid body transformation. [Figure 17] It is a diagram showing an example of fitting processing. [Figure 18] It is a diagram showing an example of the relationship between an initial position and a fitting result. MODE FOR CARRYING OUT THE INVENTION
[0011] Hereinafter, the present embodiment will be described with reference to the drawings. Note that a plurality of embodiments can be implemented in combination within a consistent scope. [First Embodiment] The first embodiment is an information processing method that enables an initial structure of a molecule to be arranged at an appropriate initial position with respect to a three-dimensional density map.
[0012] FIG. 1 is a diagram showing an example of the information processing method according to the first embodiment. FIG. 1 shows an information processing apparatus 10 for implementing the information processing method according to the first embodiment. The information processing apparatus 10 can implement the information processing method according to the first embodiment, for example, by executing a predetermined information processing program.
[0013] The information processing device 10 includes a storage unit 11 and a processing unit 12. The storage unit 11 is, for example, a memory or storage device of the information processing device 10. The processing unit 12 is, for example, a processor of the information processing device 10. The information processing device 10 may have multiple processors. Some of the multiple processes performed by the information processing device 10 may be executed on different processors.
[0014] The memory unit 11 stores molecular information 1 and 3D density map information 2. Molecular information 1 is information that shows the atoms contained in the molecule to be analyzed and the connections between atoms. If the molecule to be analyzed is a protein, information showing the amino acid sequence, for example, is used as molecular information 1. 3D density map information 2 is information obtained by analyzing the molecule to be analyzed with cryo-EM. In 3D density map information 2, the density of the molecule to be analyzed present in each voxel obtained by dividing the 3D space is shown.
[0015] The processing unit 12 analyzes the structure of the molecule to be analyzed based on molecular information 1 and 3D density map information 2. In doing so, the processing unit 12 superimposes the molecular structure (first structure 3) estimated from molecular information 1 onto the 3D density map 4 shown in 3D density map information 2. Then, the processing unit 12 deforms the first structure 3 so that it overlaps with the 3D density map 4. This deformation process is also called fitting.
[0016] The accuracy of the molecular structure obtained by fitting depends on the position and orientation of the first structure 3 when it is moved and superimposed onto the 3D density map 4. Therefore, the processing unit 12 determines the position and orientation of the first structure 3 by following the procedure below so that the first structure 3 is placed in an appropriate initial position relative to the 3D density map 4.
[0017] The processing unit 12 obtains first structural information showing the first structure 3 of the molecule and second structural information showing the second structure 5 estimated based on the three-dimensional density map 4 of the molecule. For example, the processing unit 12 generates first structural information showing the first structure 3 based on molecular information 1. The processing unit 12 also generates second structural information showing the second structure 5 based on three-dimensional density map information 2. The processing unit 12 may also accept input of structural information showing a known structure of the molecule to be analyzed and use that structural information as the first structural information.
[0018] Based on the comparison results between the first structure 3 and the second structure 5, the processing unit 12 calculates a first value indicating the reliability of the substructure containing the second atom for each of the multiple second atoms 5a to 5i, which are at least a portion of the second structure 5. The substructure containing the second atom is, for example, a residue that makes up a protein.
[0019] For example, the processing unit 12 sets the first value to a large value for the second atoms 5d to 5f in the region where the distances between corresponding atoms in the first structure 3 and the second structure 5 are similar. Also, the processing unit 12 sets the first value to a small value for the second atoms 5a to 5c, 5g to 5i in the region where the distances between corresponding atoms in the first structure 3 and the second structure 5 are significantly different.
[0020] Specifically, the processing unit 12 identifies multiple second atoms 5a to 5i corresponding to each of the multiple first atoms 3a to 3i from the second structure 5. The processing unit 12 calculates a first value for one of the multiple second atoms 5a to 5i based on the positional relationship between that second atom and the surrounding atoms, and the positional relationship between the first atom corresponding to that second atom and the surrounding atoms. For example, the processing unit 12 assigns a larger first value to a second atom if its distance from the surrounding atoms is similar to that of the corresponding first atom.
[0021] The processing unit 12 determines the position and orientation of the first structure 3 for superimposing onto the 3D density map 4 based on a second value obtained by weighting the values that evaluate the positions of each of the multiple first atoms 3a to 3i of the first structure 3 by the first value. The values that evaluate the positions of each of the multiple first atoms 3a to 3i are, for example, the distance to the corresponding second atom in the second structure 5.
[0022] The second value is information indicating the similarity between, for example, the first structure 3 and the second structure 5. The second value indicating similarity can be calculated as follows. The processing unit 12 uses a value (e.g., the square of the distance) corresponding to the distance between the atomic pairs in the first structure 3 and the second structure 5 as a value to evaluate the position of the first atom constituting that atomic pair. The processing unit 12 calculates a value for each of the first atoms 3a to 3i by multiplying the value that evaluates the position of the first atom constituting the atomic pair by the weight of that first atom. For example, the processing unit 12 sets the second value to a smaller value the smaller the multiplication result for each of the first atoms 3a to 3i is. In this case, a smaller second value indicates a higher degree of similarity.
[0023] If the second value indicates similarity, the processing unit 12 determines the position and orientation of the first structure 3 that maximizes the similarity between the first structure 3 and the second structure 5, as represented by the second value. For example, if a smaller second value indicates higher similarity, the processing unit 12 determines the position and orientation of the first structure 3 that minimizes the second value.
[0024] Once the position and orientation of the first structure 3 are determined, the processing unit 12 places the first structure 3 in the determined orientation and at the determined position. The placement of the first structure 3 is performed by rigid body transformation, and the first structure 3 is not deformed. Then, the processing unit 12 deforms (fits) the placed first structure 3 according to the 3D density map 4.
[0025] In this way, the appropriate position and orientation of the first structure 3 is determined by determining its position and orientation. That is, the values used to evaluate the position of each of the multiple first atoms 3a to 3i of the first structure 3 are weighted by the first value. The first value indicates the confidence level of the substructure containing the second atom corresponding to the evaluated first atom, and the position of the first atom corresponding to the second atom of a substructure with high confidence level has a greater influence on the second value. Therefore, by the processing unit 12 determining the position and orientation of the first structure 3 based on the second value, the position and orientation of the first structure 3 is determined so that the substructure of the first structure 3 corresponding to the highly confident substructure in the second structure 5 correctly overlaps with the 3D density map 4. As a result, in fitting the first structure 3 placed at the determined position and orientation, less deformation is required in parts that already have high confidence level, and the accuracy of the final generated structure is improved.
[0026] The processing unit 12 may, for example, move the first structure 3, which is temporarily placed around the second structure 5, based on the second value, and determine the final position and orientation of the first structure 3. In this case, the first value of each of the second atoms 5a to 5i is used as a weight for the value that evaluates the position of the corresponding first atom, so that the processing unit 12 determines how to move the first structure 3 according to the first value. For example, the processing unit 12 moves the first structure 3 so that the more reliable the substructure, the more it overlaps with the 3D density map 4.
[0027] The processing unit 12 generates, for example, a rotation matrix to rotate the first structure 3 and a translation vector to translate the first structure 3, as a way to move the first structure 3. Once the way to move the first structure 3 is determined, the processing unit 12 rotates and translates the first structure 3 according to the determined way to move it to an initial position suitable for fitting.
[0028] The multiple first atoms 3a to 3i of the first structure 3 are, for example, multiple atoms that make up the main chain within the first structure 3. The processing unit 12 can calculate the appropriate movement with less computation by determining how to move the first structure 3 based on the atoms that make up the main chain.
[0029] [Second Embodiment] The second embodiment is a computer system for reconstructing the three-dimensional structure of a protein based on a three-dimensional density map (sometimes called a cryomap) obtained by analyzing the protein using cryo-EM.
[0030] Figure 2 shows an example of the system configuration of the second embodiment. Server 100 is connected to terminal device 30 via network 20. Terminal device 30 is a computer used by users who wish to perform protein structure analysis. cryo-EM31 is connected to terminal device 30. cryo-EM31 is a transmission electron microscope that analyzes samples at low temperatures. Server 100 is a computer that can estimate the three-dimensional structure of a protein based on its three-dimensional density map.
[0031] Figure 3 shows an example of the location of information exchanged between devices. When a protein is analyzed by cryo-EM31, a 3D density map 41 is generated. Then, 3D density map information 42, which represents the 3D density map 41, is transmitted from cryo-EM31 to the terminal device 30. The 3D density map information 42 is 3D voxel data; for example, the 3D density map information 42 consists of X × Y × Z (X, Y, Z are natural numbers) voxels, and each voxel has a density (scalar value). The density represents the density of atoms present within that voxel. In the 3D density map 41 in Figure 3, voxels that are further back when viewed from the front are represented by darker colors.
[0032] The user inputs instructions to the terminal device 30 for protein structure analysis based on a three-dimensional density map 41. The terminal device 30 then sends a request to the server 100 to generate a three-dimensional structure 45 based on the three-dimensional density map 41. The request to generate the three-dimensional structure 45 includes, for example, three-dimensional density map information 42 and amino acid sequence information 43 indicating the amino acid sequence of the protein to be analyzed.
[0033] Server 100 generates a three-dimensional structure 45 in accordance with the request to generate a three-dimensional structure 45. The structural information 44 representing the three-dimensional structure 45 is a set of three-dimensional coordinates (number of atoms × (x, y, z)). Server 100 transmits the generated structural information 44 representing the three-dimensional structure 45 to the terminal device 30.
[0034] The terminal device 30 visualizes the three-dimensional structure 45 of the protein, for example, based on structural information 44. The terminal device 30 can also perform simulations using the structural information 44 to predict the temporal changes in the protein's structure and its binding to ligands. The terminal device 30 can also have the simulations run on the server 100 or another computer (not shown).
[0035] Figure 4 shows an example of server hardware. Server 100 is controlled as a whole by a processor 101. The processor 101 is connected to memory 102 and several peripheral devices via a bus 109.
[0036] Server 100 may be a multiprocessor system having multiple processors. A collection of multiple processors in a multiprocessor system can be called a processor 101. A processor 101 may also be called a processor circuitry. Each of the multiple processors can execute some or all of the processes executed by Server 100. When there are multiple related processes, two or more of those processes may be executed by different processors.
[0037] The processor 101 is, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or a DSP (Digital Signal Processor). At least some of the functions that the processor 101 implements by executing a program may be implemented by electronic circuits such as an ASIC (Application Specific Integrated Circuit) or a PLD (Programmable Logic Device).
[0038] Memory 102 is used as the main memory of the server 100. Memory 102 temporarily stores at least a portion of the OS (Operating System) program and application programs that are to be executed by the processor 101. Memory 102 also stores various data used for processing by the processor 101. For memory 102, a volatile semiconductor memory device such as RAM (Random Access Memory) is used.
[0039] Peripheral devices connected to bus 109 include a storage device 103, a graphics controller 104, an input interface 105, an optical drive device 106, a device connection interface 107, and a network interface 108.
[0040] The storage device 103 electrically or magnetically writes and reads data from its built-in recording medium. The storage device 103 is used as an auxiliary storage device for the server 100. The storage device 103 stores the OS program, application programs, and various data. For example, the storage device 103 can be an HDD (Hard Disk Drive) or an SSD (Solid State Drive).
[0041] The graphics controller 104 is an arithmetic unit that performs image processing. The graphics controller 104 is, for example, a GPU (Graphics Processing Unit). A monitor 21 is connected to the graphics controller 104. The graphics controller 104 displays images on the screen of the monitor 21 according to instructions from the processor 101. The monitor 21 can be an OLED (Electroluminescence) display device or a liquid crystal display device. If a GPU is used as the graphics controller 104, the graphics controller 104 can also perform complex numerical calculations such as matrix calculations.
[0042] The input interface 105 is connected to a keyboard 22 and a mouse 23. The input interface 105 transmits signals from the keyboard 22 and mouse 23 to the processor 101. Note that the mouse 23 is just one example of a pointing device; other pointing devices can also be used. Other pointing devices include touch panels, tablets, touchpads, and trackballs.
[0043] The optical drive device 106 uses laser light or the like to read data recorded on the optical disc 24 or write data to the optical disc 24. The optical disc 24 is a portable recording medium on which data is recorded in a way that makes it readable by the reflection of light. Examples of optical discs 24 include DVD (Digital Versatile Disc), DVD-RAM, CD-ROM (Compact Disc Read Only Memory), and CD-R (Recordable) / RW (ReWritable).
[0044] The device connection interface 107 is a communication interface for connecting peripheral devices to the server 100. For example, a memory device 25 and a memory reader / writer 26 can be connected to the device connection interface 107. The memory device 25 is a recording medium equipped with a communication function with the device connection interface 107. The memory reader / writer 26 is a device that writes data to or reads data from the memory card 27. The memory card 27 is a card-type recording medium.
[0045] The network interface 108 is connected to the network 20. The network interface 108 transmits and receives data to and from other computers or communication devices via the network 20. The network interface 108 is a wired communication interface, for example, connected by cable to a wired communication device such as a switch or router. Alternatively, the network interface 108 may be a wireless communication interface, connected by radio waves to a wireless communication device such as a base station or access point.
[0046] The server 100 can implement the processing functions of the second embodiment using the hardware described above. The information processing device 10 shown in the first embodiment can also be implemented using hardware similar to that of the server 100 shown in Figure 4.
[0047] The server 100 implements the processing functions of the second embodiment by executing a program recorded on, for example, a computer-readable recording medium. The program describing the processing to be executed by the server 100 can be recorded on various recording media. For example, the program to be executed by the server 100 can be stored in the storage device 103. The processor 101 loads at least a portion of the program in the storage device 103 into the memory 102 and executes the program. Alternatively, the program to be executed by the server 100 can be recorded on a portable recording medium such as an optical disc 24, a memory device 25, or a memory card 27. The program stored on the portable recording medium becomes executable after being installed in the storage device 103, for example, under control from the processor 101. The processor 101 can also directly read and execute the program from the portable recording medium.
[0048] With this type of hardware server 100, it is possible to generate the three-dimensional structure of a protein from a three-dimensional density map obtained by analyzing the protein using cryo-EM.
[0049] The importance and difficulties of generating the three-dimensional structure of proteins will be explained below, with reference to Figures 5 to 7. Figure 5 shows examples of the locations of structural changes in a protein. The atoms that make up a protein are constantly fluctuating. Therefore, proteins can exist in various states with different structures. In the example in Figure 5, three states, state A, state B, and state C, are shown. Proteins change their structure and transition between states in response to environmental changes such as temperature, pressure, and time.
[0050] The state of a protein can be determined by analyzing its three-dimensional structure 51-53. Information about the state of a protein under various environmental conditions can be effectively used, for example, in drug discovery and development using that protein. When observing a protein in a specific environment, there is an approximately 80% probability of observing state A and an approximately 20% probability of observing state B (the probability of being in state C is negligible). If the probability of being in state A is high, searching for ligands (drug candidates) that readily bind to state A allows for efficient searching for ligands that readily bind to proteins in a specific environment.
[0051] Thus, it is important to know how the three-dimensional structure of the protein being analyzed transitions depending on the environment in which it is placed. The three-dimensional structure of a protein under specific environmental conditions can be estimated, for example, by fitting the protein's known three-dimensional structure to its three-dimensional density map.
[0052] Figure 6 shows an example of fitting. The three-dimensional structures 46 and 47 shown in Figure 6 were visualized using protein data and viewers disclosed in the following paper.
[0053] Reference 1 (3D structure viewer Mol*): David Sehnal, Sebastian Bittrich, Mandar Deshpande, Radka Svobodova, Karel Berka, Vaclav Bazgier, Sameer Velankar, Stephen K Burley, Jaroslav Koca, Alexander S Rose, “Mol* Viewer: modern web app for 3D visualization and analysis of large biomolecular structures”, Nucleic Acids Research, Volume 49, Issue W1, 2 July 2021, Pages W431-W437 (Svobodova's 'a' and the first 'a' in Vaclav have acute accents, Koca's 'c' has a Harchek accent) Reference 2 (Viewer provided by RCSB PDB): Helen M. Berman, John Westbrook, Zukang Feng, Gary Gilliland, TN Bhat, Helge Weissig, Ilya N. Shindyalov, Philip E. Bourne, “The Protein Data Bank”, Nucleic Acids Research, Volume 28, Issue 1, 1 January 2000, Pages 235-242 Reference 3 (Protein 4ake): CW Muller, GJ Schlauderer, J Reinstein, GE Schulz, “Adenylate kinase motions during catalysis: an energetic counterweight balancing substrate binding”, Structure, Volume 4, Issue 2, February 1996, Pages 147-156 (Muller's 'u' has an umlaut). For example, the server 100 generates three-dimensional structures 47 of the protein under analysis in various states by deforming the initial three-dimensional structure 46. Furthermore, the server 100 converts the deformed three-dimensional structures 47 into a three-dimensional density map 48.
[0054] Server 100 targets a three-dimensional density map 41 obtained by analyzing the target protein using cryo-EM. Server 100 repeatedly deforms the three-dimensional structure 46 so that the generated three-dimensional density map 48 approaches the target. Server 100 then estimates that the three-dimensional structure 47 when the difference between the generated three-dimensional density map 48 and the target three-dimensional density map 41 falls below a predetermined value represents the structure of the target protein.
[0055] In this fitting process, the server 100 first moves the initial structure, a 3D solid structure 46, to an initial position that overlaps with the target's 3D density map 41. This movement involves both translation and rotation. At this point, the accuracy of the fitted 3D solid structure 47 depends on the initial position of the initial structure.
[0056] Figure 7 shows an example of the relationship between the initial position and the ease of fitting. Server 100 performs a rigid body transformation (rotation and translation) of the initial 3D stereostructure 55 so that it overlaps with the 3D density map 54 of the target. In the example in Figure 7, the residues at the upper end 54a of the 3D density map 54 correspond to the residues at the upper end 55a of the 3D stereostructure 55 before the rigid body transformation. Similarly, the residues at the lower end 54b of the 3D density map 54 correspond to the residues at the lower end 55b of the 3D stereostructure 55 before the rigid body transformation.
[0057] In the first initial position example, the upper end 55a of the 3D solid structure 55 before rigid body transformation is superimposed on the upper end 54a of the 3D density map 54, and the lower end 55b of the 3D solid structure 55 before rigid body transformation is superimposed on the lower end 54b of the 3D density map 54. When the 3D solid structure 55 is placed in such an initial position, fitting becomes easier. That is, it becomes possible to obtain a highly accurate 3D solid structure through fitting.
[0058] On the other hand, in the second initial position example, the upper end 55a of the 3D solid structure 55 before rigid body transformation is superimposed on the lower end 54b of the 3D density map 54, and the lower end 55b of the 3D solid structure 55 before rigid body transformation is superimposed on the upper end 54a of the 3D density map 54. When the 3D solid structure 55 is placed in such an initial position, fitting becomes difficult. That is, even if the deformation of the 3D solid structure 55 is repeated, it becomes difficult to achieve high-precision fitting to the 3D density map 54.
[0059] As shown in the example in Figure 7, if the orientation of the initial position differs significantly from that of the target 3D density map 54, fitting may not yield a 3D solid structure corresponding to the 3D density map 54. For example, if a rigid body transformation is performed to match the external shape, and there are common externally distinctive parts in the 3D density map 54 and the 3D solid structure 55, these parts can be used as markers to align the orientations of the 3D density map 54 and the 3D solid structure 55. On the other hand, there are cases where the external shapes of the 3D density map 54 and the 3D solid structure 55 differ significantly, or where, as shown in Figure 7, there are multiple rigid body transformation patterns that allow for the overlapping of the external shapes. In these cases, it is difficult to align the orientations of the 3D density map 54 and the 3D solid structure 55 using only external features in a rigid body transformation.
[0060] The server 100 then uses some information about the three-dimensional structure corresponding to the target's three-dimensional density map 54 to rotate the initial structure. By effectively utilizing the information obtained from the target's three-dimensional density map 54, it becomes possible to rotate the three-dimensional structure 55 appropriately.
[0061] Figure 8 shows an example of the functions a server has for protein structure analysis. Server 100 has a storage unit 110, a request retrieval unit 120, a three-dimensional structure generation unit 130, a reliability calculation unit 140, a rigid body conversion unit 150, and a fitting unit 160.
[0062] The memory unit 110 stores various data used for analyzing the three-dimensional structure of proteins. For example, the memory unit 110 stores amino acid sequence information 111, structural information 112, and three-dimensional density map information 113. The amino acid sequence information 111 is information that shows the sequence of amino acids that make up the protein being analyzed. The amino acid sequence information 111 is, for example, a sequence of symbols that represent amino acids. The structural information 112 is information that shows the three-dimensional structure of the protein being analyzed. The three-dimensional density map information 113 is voxel data obtained by analyzing the protein being analyzed with cryo-EM.
[0063] The request acquisition unit 120 receives an analysis request from the terminal device 30 instructing the terminal device 30 to perform a three-dimensional structural analysis of a protein. The analysis request may include, for example, information identifying the protein to be analyzed, and three-dimensional density map information 113. The analysis request may also include amino acid sequence information 111, structural information 112, etc. The request acquisition unit 120 stores the information included in the analysis request in the storage unit 110 and instructs the three-dimensional structure generation unit 130 to generate a reference structure. The request acquisition unit 120 also acquires structural information showing the generated three-dimensional structure from the fitting unit 160 and transmits that structural information to the terminal device 30 as an analysis result.
[0064] The three-dimensional structure generation unit 130 infers the protein structure based on the amino acid sequence information 111. However, if information indicating the structure of a known protein is input, the three-dimensional structure generation unit 130 does not need to infer the protein structure.
[0065] Furthermore, the three-dimensional structure generation unit 130 generates a substructure of the three-dimensional structure of the protein to be analyzed based on the three-dimensional density map information 113, which shows the three-dimensional density map of the target. Hereinafter, the substructure generated by the three-dimensional structure generation unit 130 will be referred to as the reference structure. The reference structure is information that shows, for example, the position in three-dimensional space of atoms (Cα atoms) that make up the main chain of a protein. The Cα atom is the carbon atom closest to the carboxyl group of an amino acid.
[0066] The three-dimensional structure generation unit 130 generates a reference structure, for example, by the main chain tracing method. The main chain tracing method is a method for generating a partial structure by tracing the main chain of a protein. The main chain is the longest sequence of covalently bonded atoms. The reference structure generated by the main chain tracing method may be an incomplete structure with missing atoms or incorrect connections. The three-dimensional structure generation unit 130 may also accept manual input of atom arrangement from the user and generate a reference structure according to the input. The three-dimensional structure generation unit 130 transmits the generated reference structure to the reliability calculation unit 140.
[0067] The confidence calculation unit 140 calculates the confidence value of each residue included in the initial structure of the generated three-dimensional stereostructure. The confidence value is, for example, the pLDDT (predicted local distance difference test) value. The pLDDT value represents the confidence of the position of each residue, with a value from 0 to 100 for each residue. The pLDDT value is used as an indicator of how well the reference structure reproduces the initial structure. The higher the pLDDT value of a residue, the more stable the structure is, and the more likely it is that the residue constitutes a major part of the protein. The lower the pLDDT value of a residue, the more likely it is that it constitutes a flexible part. The confidence calculation unit 140 transmits the confidence value of each residue to the rigid body conversion unit 150.
[0068] The rigid body transformation unit 150 performs a rigid body transformation of the initial three-dimensional structure based on the reliability of each residue. For example, the rigid body transformation unit 150 weights the pLDDT value and determines the rotation matrix and translation vector such that the value of the Root Mean Square Deviation (RMSD) is small.
[0069] Specifically, the rigid body transformation unit 150 extracts Cα atoms of the same amino acid residue from both the initial structure and the reference structure. The rigid body transformation unit 150 superimposes the Cα atoms of the initial structure onto the Cα atoms of the reference structure in such a way that the RMSD value is smallest. The rigid body transformation unit 150 then calculates a rotation matrix and translation vector to move the initial structure to the position where the RMSD is minimized. In this process, the rigid body transformation unit 150 calculates an RMSD weighted by the pLDDT value. This provides a rotation matrix and translation vector that allows for highly accurate superimposition of stable regions. The rigid body transformation unit 150 then rotates the initial structure of the protein according to the obtained rotation matrix and translates it according to the obtained translation vector.
[0070] The fitting unit 160 deforms the initial structure, which is initially placed in its initial position, by rigid body transformation and fits it to the 3D density map obtained by cryo-EM. The fitting unit 160 transmits structural information showing the 3D density structure after fitting to the request acquisition unit 120.
[0071] With a server 100 having such functions, it is possible to generate highly accurate three-dimensional structures based on a three-dimensional density map of proteins. The functions of each element shown in Figure 8 can be realized, for example, by having the processor 101 execute the program module corresponding to that element.
[0072] Figure 9 shows an example of structural information. Structural information 112 has, for example, a record for each atom that makes up a protein. Each record contains data such as the sequence number of the atom, the type of element, the name of the atom, the type of residue, the x coordinate, y coordinate, z coordinate, and the sequence number of the residue.
[0073] The atomic serial number is a number that uniquely identifies the corresponding atom. The element type is the element symbol of the corresponding atom. The atomic name is the name of the corresponding atom in the protein structure. The atom listed as "CA" is a Cα atom. The residue type is the type of residue to which the corresponding atom belongs.
[0074] The x, y, and z coordinates are the three-dimensional coordinates of the corresponding atom. The residue number is a unique identifier for the residue to which the corresponding atom belongs. Figure 10 shows an example of 3D density map information. The 3D density map information 113 includes, for example, voxel number information 113a and density information 113b. Voxel number information 113a indicates the number of voxels in each axis direction of the 3D density map. The result of multiplying the x-axis voxel number, y-axis voxel number, and z-axis voxel number is the total number of voxels in the 3D density map.
[0075] Density information 113b is information indicating the density of each voxel. In density information 113b, each voxel is uniquely identified by a numerical sequence (x_id, y_id, z_id) that indicates which voxel it is in each axis direction.
[0076] Next, we will explain the procedure for analyzing the three-dimensional structure of the protein being analyzed based on its three-dimensional density map. Figure 11 is a flowchart showing an example of the procedure for analyzing a three-dimensional structure. The process shown in Figure 11 will be explained below according to the step numbers.
[0077] [Step S101] When a request for analysis of a three-dimensional structure is input to the request acquisition unit 120, it acquires amino acid sequence information 111 and three-dimensional density map information 113 from the analysis request. The request acquisition unit 120 stores the acquired amino acid sequence information 111 and three-dimensional density map information 113 in the storage unit 110. The three-dimensional density map shown in the acquired three-dimensional density map information 113 is the target three-dimensional density map.
[0078] [Step S102] The three-dimensional structure generation unit 130 infers the protein structure (structure A) based on the amino acid sequence information 111. The protein structure inference can be performed, for example, using a machine learning model to predict the protein's folded structure according to the width of atoms. The three-dimensional structure generation unit 130 uses the three-dimensional structure output by the inference as the initial structure for fitting. Alternatively, a known three-dimensional structure may be obtained without generating a three-dimensional structure.
[0079] [Step S103] The 3D structure generation unit 130 generates a reference structure (structure B) corresponding to the target 3D density map based on the 3D density map information 113. For example, the 3D structure generation unit 130 generates the reference structure using a main chain tracing method.
[0080] [Step S104] The confidence calculation unit 140 performs correspondence between pairs of amino acid residues between structure A and structure B. For example, the confidence calculation unit 140 performs the correspondence by amino acid sequence alignment. Alternatively, the confidence calculation unit 140 may perform correspondence between pairs of atoms between structure A and structure B. For example, the pairs of atoms are pairs of Cα atoms.
[0081] Structure A is generated based solely on amino acid sequence information 111 and has the correct amino acid sequence (e.g., "MRIILLGAPGA..."). On the other hand, although structure B utilizes amino acid sequence information 111, it may have an incorrect amino acid sequence (e.g., "MAIILRGAPGA...") in order to match the 3D density map information 113. In amino acid sequence alignment, for example, amino acid residues are compared sequentially from the beginning of the amino acid sequences of structure A and structure B, and identical amino acid residues are matched together.
[0082] [Step S105] The confidence calculation unit 140 extracts the Cα atoms from each pair of amino acid residues from structure A and structure B. The confidence calculation unit 140 generates structure A', which consists only of the Cα atoms extracted from structure A. Furthermore, the confidence calculation unit 140 generates structure B', which consists only of the Cα atoms extracted from structure B. Since the Cα atom is part of the main chain, structure A' and structure B' represent the structure of the main chain, respectively.
[0083] [Step S106] The reliability calculation unit 140 calculates the reliability of each Cα atom in structure B' based on structure A'. The reliability is, for example, the value of pLDDT. [Step S107] The rigid body transformation unit 150 performs a rigid body transformation and superimposes structure A' onto structure B'. For example, the rigid body transformation unit 150 finds a rotation matrix and a translation vector. In this case, the rigid body transformation unit 150 performs the superposition in such a way that the RMSD value, weighted by the pLDDT value, becomes small. The problem of finding the rotation matrix and translation vector for such a superposition is a nonlinear optimization problem in which one finds continuous variables (rotation matrix and translation vector) that satisfy the conditions of a nonlinear function (RMSD). For example, the rotation matrix and translation vector for rigid body transformation can be found using methods such as the Kabsch method or the ICP (Iterative Closest Point) method.
[0084] By calculating the rotation matrix and translation vector using only Cα atoms in this way, we can reliably superimpose the main chains while also reducing the computational complexity. It is also possible to calculate the rotation matrix and translation vector using all the atoms in each pair. Using all the atoms in each pair allows for superimposition that also considers the side chains.
[0085] [Step S108] Once the rotation matrix and translation vector are determined, the rigid body transformation unit 150 rotates structure A by the rotation matrix and then translates it by the translation vector. As a result, structure A is superimposed on structure B. The position of structure A after the translation is the initial position.
[0086] [Step S109] The fitting unit 160 fits the structure A in its initial position to the three-dimensional density map shown in the three-dimensional density map information 113. That is, the fitting unit 160 deforms the structure A so that it is as similar as possible to the three-dimensional density map.
[0087] In this way, the initial structure (structure A) can be fitted from an appropriate initial position, making it possible to obtain a highly accurate three-dimensional structure of the protein to be analyzed. In the example above, the main chain tracing method is used to generate the reference structure. However, the main chain tracing method does not guarantee that all atoms constituting each amino acid shown in the amino acid sequence information are included in the reference structure. For example, some atoms constituting an amino acid may be missing. Furthermore, the amino acid itself may be missing.
[0088] Figure 12 shows an example of reference structure generation using the main chain tracing method. For example, the three-dimensional structure generation unit 130 recognizes the types of atoms constituting the protein based on the amino acid sequence shown in the amino acid sequence information 111. Furthermore, the three-dimensional structure generation unit 130 determines that atoms exist in voxels with high density in the three-dimensional density map 41 shown in the three-dimensional density map information 113. Based on this information, the three-dimensional structure generation unit 130 identifies the positions of atoms that can be located in the three-dimensional density map 41 among the atoms constituting the protein. The information indicating the positions of the identified atoms will be missing information for atoms whose positions cannot be located in the three-dimensional density map 41.
[0089] The three-dimensional structure generation unit 130 connects atoms whose positions have been identified based on the trajectory of the main chain. However, errors in connection may occur during this process. The structure formed by connecting the identified atoms becomes the reference structure 61.
[0090] Thus, although the reference structure 61 contains atomic deficiencies and connection errors, the main chain is reproduced with high accuracy. Therefore, the reliability calculation unit 140 compares the Cα atoms constituting the main chain of the reference structure 61 with the corresponding Cα atoms in the initial structure, thereby obtaining a confidence level regarding the accuracy of the positions of the Cα atoms in the initial structure.
[0091] Figure 13 shows an example of confidence calculation. The pLDDT value can be used as the confidence level. The pLDDT value is a score calculated using only Cα in the lDDT (The Local Distance Difference Test). lDDT is a score that indicates how well the predicted structure reproduces the correct structure when there is a correct structure and a predicted structure. The confidence calculation unit 140 calculates the pLDDT value, which indicates the confidence level, using the initial structure 62 as the correct structure and the reference structure 61 as the predicted structure. The calculation result shows how well the reference structure 61 reproduces the initial structure 62.
[0092] Initial structure 62 and reference structure 61 each contain three Cα atoms: "A", "B", and "C". Cα atom "A" is contained in residue "X". Cα atom "B" is contained in residue "Y". Cα atom "C" is contained in residue "Z".
[0093] The confidence calculation unit 140 creates pairs of Cα atoms that are close together (within the threshold R (where R is a positive real number)) within the initial structure 62. In the example in Figure 13, R = 15 Å, and three pairs {(A, B), (A, C), (B, C)} are created. The confidence calculation unit 140 also creates the same pairs of atoms in the reference structure 61.
[0094] The reliability calculation unit 140 calculates the difference between the distance in the reference structure 61 and the distance in the initial structure 62 for each pair of atoms. The difference in distance d for pair (A,B) AB The difference in distance between pairs (A,C) is d. AC The difference in distance between the pair (B,C) is d. BC The equation is "2.2 - 4.5 = -2.3 Å".
[0095] The confidence calculation unit 140 determines whether the absolute value of the difference for each pair of atoms is within the threshold D (where D is a positive real number). If the absolute value of the difference is within the threshold D, it is determined that the distance between the pairs of atoms is preserved in the initial structure 62. For example, the threshold D is calculated for each of the four patterns "0.5 Å, 1 Å, 2 Å, 4 Å". In this case, when the threshold D is "0.5 Å, 1 Å, 2 Å", the distance (d difference d) between the pairs (B, C) is determined. BC The distance (=-2.3 Å) is not conserved. Otherwise, the distance between each pair is conserved.
[0096] The reliability calculation unit 140 determines that the distance between atomic pairs is not preserved if one or both of the pairs of Cα atoms within a distance of threshold R in the initial structure 62 are not present in the reference structure 61.
[0097] The reliability calculation unit 140 calculates the conservation rate of the distance between each atom in a pair of atoms. For Cα atom "A", the conservation rate is 100% for all thresholds D. For Cα atom "B", the conservation rate is 50% when the threshold D is "0.5 Å, 1 Å, 2 Å", and 100% when the threshold D is "4.0 Å". For Cα atom "C", the conservation rate is 50% when the threshold D is "0.5 Å, 1 Å, 2 Å", and 100% when the threshold D is "4.0 Å".
[0098] The confidence calculation unit 140 uses the average of the preservation rates for each threshold D value of the Cα atom as the final pLDDT value of that Cα atom. In the example in Figure 13, the pLDDT value of Cα atom "A" contained in residue "X" is "100%". The pLDDT value of Cα atom "B" contained in residue "Y" is "62.5%". The pLDDT value of Cα atom "C" contained in residue "Z" is "62.5%".
[0099] In this way, the pLDDT value of each Cα atom is calculated. The pLDDT value of the Cα atom serves as the reliability. Since each amino acid residue has one Cα atom, it can also be said that the reliability is calculated for each residue. The reliability for each Cα atom is used as a weight for RMSD used when generating a rotation matrix and a translation vector.
[0100] FIG. 14 is a diagram showing an example of an RMSD calculation method. The distance between a pair of Cα atoms "i" (i is a symbol representing an atom) of an initial structure 62 and a reference structure 61 is defined as "x i ". The distance between the Cα atom "A" of the initial structure 62 and the reference structure 61 is "x A ". The distance between the Cα atom "B" of the initial structure 62 and the reference structure 61 is "x B ". The distance between the Cα atom "C" of the initial structure 62 and the reference structure 61 is "x C ". In the example of FIG. 14, the distance "x" of the Cα atom "B" B is shorter than the distances of other Cα atoms (x B ≫x A ≒x C ).
[0101] The weight of the Cα atom "i" is defined as "w i ". The weight of the Cα atom "A" of the initial structure 62 is "w A ". The weight of the Cα atom "B" of the initial structure 62 is "w B ". The weight of the Cα atom "C" of the initial structure 62 is "w C ".
[0102] RMSD is obtained by arithmetically averaging the squares of the distances of respective Cα atoms between the reference structure 61 and the initial structure 62, then taking the square root of the result. Unweighted RMSD is calculated by the following formula.
[0103]
Chem.
[0104] N is the number of Cα atoms included in the initial structure 62. When formula (1) is used, to minimize RMSD, for example, the Cα atom "wA Even if the reliability of "w" is extremely low compared to other Cα atoms, A This reduces the distance between the atoms. In other words, in equation (1), the structure is affected to the same extent as the highly reliable atoms.
[0105] The weighted RMSD is expressed by the following formula:
[0106]
number
[0107] In equation (2), the weight of Cα atoms with low confidence is reduced. As a result, the influence of Cα atoms with low confidence is suppressed when minimizing the RMSD. In other words, the rotation matrix and translation vector are determined so that the distance between pairs of Cα atoms with high confidence is reduced.
[0108] Figure 15 shows an example of initial structure rotation using confidence levels. This figure assumes fitting the initial structure 62 to a 3D density map 63. Nine Cα atoms 61a to 61i are extracted from the reference structure 61 corresponding to the 3D density map 63. Similarly, nine Cα atoms 62a to 62i are extracted from the initial structure 62 by extracting Cα atoms from amino acid residues.
[0109] The Cα atoms 61a to 61c, indicated by dashed circles in reference structure 61, are paired with the Cα atoms 62a to 62c, each indicated by a dashed circle in initial structure 62. The Cα atoms 61d to 61f, indicated by solid circles in reference structure 61, are paired with the Cα atoms 62d to 62f, each indicated by a solid circle in initial structure 62. The Cα atoms 61g to 61i, indicated by double-line circles in reference structure 61, are paired with the Cα atoms 62g to 62i, each indicated by a double-line circle in initial structure 62.
[0110] Here, the confidence levels of each of the Cα atoms 61a to 61i in reference structure 61 are calculated. In this case, the confidence levels of Cα atoms 61d to 61f are assumed to be larger (higher confidence) than those of the other Cα atoms 61a to 61c and 61g to 61i. That is, in the region containing Cα atoms 61d to 61f, the distance between Cα atoms is maintained to be about the same as in the initial structure 62.
[0111] The lower left of Figure 15 shows an example where the initial structure 62 is superimposed on the reference structure 61 in such a way that the value of the unweighted RMSD (see equation (1)) is small. In this case, the initial structure 62 is oriented to be pulled towards the region containing Cα atoms 62a~62c and the region containing Cα atoms 62g~62i. When using the unweighted RMSD, the central helix region (the region containing Cα atoms 62d~62f) is significantly outside the 3D density map 63, making subsequent fitting difficult.
[0112] The lower right of Figure 15 shows an example of the case where the initial structure 62 is superimposed on the reference structure 61 in such a way that the weighted RMSD (see equation (2)) value is small. In this case, the area containing the less reliable Cα atoms 62a~62c, 62g~62i is easily moved by the rotation of the initial structure 62. That is, even if the rotation of the initial structure 62 moves the Cα atoms 62a~62c, 62g~62i of the initial structure away from the Cα atoms 61a~61c, 61g~61i of the reference structure, the effect on the RMSD is small. On the other hand, the RMSD can be reduced by bringing the area containing the more reliable Cα atoms 62d~62f closer to the Cα atoms 61d~61f of the reference structure.
[0113] Therefore, superposition focusing particularly on highly reliable Cα atoms is performed, and important secondary structural parts of the protein in the initial structure 62 are superimposed on the 3D density map 63. In this case, in subsequent fitting, less reliable atoms, which tend to move due to changes over time, can be preferentially moved. By lowering the priority of moving highly reliable atoms, the disruption of the distance relationships between atoms caused by forcibly moving highly reliable atoms during fitting is suppressed. As a result, the loss of natural protein characteristics can be suppressed. Moreover, since fitting can be performed while maintaining the shape of the highly reliable parts of the initial structure, the difficulty of fitting is low.
[0114] Once the rotation matrix and translation vector that minimize the weighted RMSD are determined, a rigid body transformation is performed based on that rotation matrix and translation vector. Figure 16 shows an example of a rigid body transformation. The generated rotation matrix is a 3x3 matrix. The generated translation vector is a 1x3 vector. The rigid body transformation unit 150 rotates the initial structure 62 in three-dimensional space using the rotation matrix. Then, the rigid body transformation unit 150 translates the rotated initial structure 62 using the translation vector. As a result, the initial structure 62 moves to an initial position that is easy to fit.
[0115] Based on the initial structure moved to its initial position, the fitting unit 160 performs the fitting. Figure 17 shows an example of the fitting process. Fitting is performed using, for example, FFM (Fitting Fold to Map). FFM is a technique that uses an AI (Artificial Intelligence) model to fit the initial 3D structure to the target 3D density map by moving the 3D structure of the initial structure to match the target 3D density map. The AI model for FFM is divided into multiple submodels, for example, 72, 74.
[0116] The fitting unit 160 generates feature vectors 71 based on amino acid sequence information 111. The fitting unit 160 takes the generated feature vectors 71 as input to the first submodel 72 and performs calculations according to the submodel 72, obtaining intermediate feature vectors 73 as output.
[0117] The fitting unit 160 also generates a 3D density map 70 at step 0 based on the structural information 112a of the initial structure, in which the positions of each atom have been transformed to their initial coordinates by rigid body transformation. The fitting unit 160 calculates the error between the target 3D density map shown in the 3D density map information 113 and the generated 3D density map 70. Then, as a backpropagation process of the error, the fitting unit 160 updates the intermediate features 73. For example, the fitting unit 160 calculates how the loss function changes when the intermediate features 73 are changed by a small amount, and updates the intermediate features 73 to minimize the error based on that change.
[0118] The fitting unit 160 takes the updated intermediate features 73 from backpropagation as input to the second submodel 74 and performs calculations according to the submodel 72, obtaining the first step's three-dimensional structure 75a as output.
[0119] The fitting unit 160 generates a 3D density map 76a based on the 3D solid structure 75a and calculates the difference between the target 3D density map shown in the 3D density map information 113 and the generated 3D density map 70. Thereafter, the fitting unit 160 repeatedly updates the intermediate features 73 by backpropagation of errors, generates a 3D solid structure using a partial model 74, generates a 3D density map, and calculates the error between the target 3D density map and the generated 3D density map 70. This generates the 3D solid structures 75b,...,75n and 3D density maps 76b,...,76n from the second step onward.
[0120] This iterative process continues until a termination condition is met. The termination condition is, for example, that the error becomes less than or equal to a predetermined value. When the 3D density map 76n at step n (where n is a non-negative integer) satisfies the termination condition, the fitting unit 160 outputs structural information 64 representing the 3D three-dimensional structure 75n at step n as the fitting result.
[0121] In this way, a 3D solid structure that fits the target's 3D density map is obtained. By updating the intermediate features 73 through backpropagation, the knowledge reflected in the vast number of parameters of the FFM AI model can be utilized to the fullest extent. In other words, the destruction of existing parameters (forgetting of knowledge) caused by retraining the AI model is avoided, and it is possible to generate a highly reliable 3D solid structure.
[0122] Figure 18 shows an example of the relationship between the initial position and the fitting result. Figure 18 shows the results of estimating the three-dimensional structure from the three-dimensional density model of a protein for which the correct structure is known.
[0123] Graph 81 shows the evaluation results of the 3D solid structure generated by fitting when the initial structure 62 is moved to its initial position without using a reference structure. The horizontal axis of Graph 81 represents the number of times the intermediate features are updated during fitting. The vertical axis of Graph 81 represents the RMSD of the 3D solid structure generated using the updated intermediate features. The thick line 81a shows the change in RMSD when the generated 3D solid structure is compared with a normalized structure. The thin line 81b shows the change in RMSD when compared with an existing structure similar to the initial structure 62.
[0124] Graph 82 shows the evaluation results of the 3D solid structure generated by fitting when the initial position of the initial structure 62 is set to the position that minimizes the RMSD (including confidence-based weighting) with respect to the reference structure. The horizontal axis of Graph 82 represents the number of intermediate feature updates during fitting. The vertical axis of Graph 82 represents the RMSD of the 3D solid structure generated using the updated intermediate features. The thick line 82a shows the change in RMSD when the generated 3D solid structure is compared with a normalized structure. The thin line 82b shows the change in RMSD when compared with an existing structure close to the initial structure 62.
[0125] In the example in Graph 81, even with repeated updates of intermediate features during fitting, the generated 3D solid structure does not approach the correct structure. This is because the initial position of the initial structure 62 is inappropriate. In contrast, in the example in Graph 82, the generated 3D solid structure rapidly approaches the correct structure by repeatedly updating the intermediate features during fitting. This is because the initial structure 62 is positioned in an appropriate initial location.
[0126] By moving the initial structure 62 to a position and orientation that minimizes the weighted RMSD with respect to the reference structure and then performing a fitting, a highly accurate three-dimensional structure can be generated.
[0127] [Other embodiments] In the second embodiment, fitting is performed using FFM, but fitting can also be performed using technologies other than FFM. For example, fitting using MDFF (Molecular Dynamics Flexible Fitting) is also possible. MDFF is a method that uses molecular dynamics simulations to bring the initial structure (a structure in a different state from the target structure of the same protein as the target structure) closer to the 3D density map of the target. The target structure is the structure corresponding to the 3D density map. In MDFF, when calculating the forces acting on the particle, the difference between the latest structure and the 3D density map of the target structure is calculated, and forces are added based on the difference to bring it closer to the density map.
[0128] Although embodiments have been illustrated above, the configurations of each part shown in the embodiments can be replaced with others having similar functions. Furthermore, other arbitrary components or processes may be added. Moreover, any two or more configurations (features) from the embodiments described above may be combined. [Explanation of symbols]
[0129] 1 Molecular information 2. 3D density map information 3. The first structure 3a~3i First atom 4. 3D density map 5. The second structure 5a~5i Second atom 10 Information Processing Devices 11 Storage section 12 Processing Units
Claims
1. First structural information showing the first structure of the molecule and second structural information showing the second structure estimated based on the three-dimensional density map of the molecule are obtained. Based on the comparison results between the first structure and the second structure, a first value indicating the reliability of the substructure containing the second atom is calculated for each of at least some of the second atoms in the second structure. The position of each of the multiple first atoms in the first structure is evaluated, and the first value of the second atom corresponding to the evaluated first atom is weighted by this value. Based on the second value obtained by calculation using the weighted value, the position and orientation of the first structure to be superimposed on the three-dimensional density map are determined. An information processing program that instructs a computer to perform a task.
2. The structure of prepared 1 is placed in the determined orientation and at the determined position. The information processing program according to claim 1, which causes the computer to perform further processing.
3. The first structure that is positioned is deformed according to the three-dimensional density map. The information processing program according to claim 2, which causes the computer to perform further processing.
4. In the process of calculating the first value, the first value of the first second atom is calculated based on the positional relationship of one of the plurality of first atoms with respect to surrounding atoms and the positional relationship of the one first atom with respect to the corresponding second atom in the second structure. The information processing program according to claim 1.
5. In the process of calculating the first value, the plurality of atoms constituting the main chain in the second structure are referred to as the plurality of second atoms. The information processing program according to claim 1.
6. In the process of determining the position and orientation of the first structure, a rotation matrix for rotating the first structure and a translation vector for translating the first structure are generated based on the second value. The information processing program according to claim 1.
7. In the process of determining the position and orientation of the first structure, the second value indicating the similarity between the first structure and the second structure is calculated. The information processing program according to claim 1.
8. In the process of determining the position and orientation of the first structure, the position and orientation of the first structure that maximizes the similarity shown in the second value are determined. The information processing program according to claim 7.
9. First structural information showing the first structure of the molecule and second structural information showing the second structure estimated based on the three-dimensional density map of the molecule are obtained. Based on the comparison results between the first structure and the second structure, a first value indicating the reliability of the substructure containing the second atom is calculated for each of at least some of the second atoms in the second structure. The position of each of the multiple first atoms in the first structure is evaluated, and the first value of the second atom corresponding to the evaluated first atom is weighted by this value. Based on the second value obtained by calculation using the weighted value, the position and orientation of the first structure to be superimposed on the three-dimensional density map are determined. An information processing method in which a computer performs the processing.
10. A processing unit obtains first structural information showing a first structure of a molecule and second structural information showing a second structure estimated based on a three-dimensional density map of the molecule, calculates a first value indicating the reliability of a substructure containing at least some of a plurality of second atoms in the second structure based on the comparison result between the first structure and the second structure, weights the value evaluating the position of each of the plurality of first atoms in the first structure by the first value of the second atom corresponding to the evaluated first atom, and determines the position and orientation of the first structure to be superimposed on the three-dimensional density map based on a second value obtained by calculation using the weighted value. An information processing device having
Citation Information
Patent Citations
Multi-antigen delivery system using hepatitis E virus-like particles
JP2012519006A
A synthetic humanized llama nanobody library and its use to identify SARS-COV-2 neutralizing antibodies
JP2024535249A
Novel methods of creating a protein map and using said map to identify therapeutic targets
US20230108717A1
Crispr-cas effector polypeptides and methods of use thereof
US20230407276A1