Three-dimensional structure identification system, method, and program
The three-dimensional structure identification system addresses the limitations of conventional molecular representation methods by converting interatomic distance matrices into proximity score matrices and images, enhancing the accuracy and efficiency of molecular structure analysis and comparison.
Patent Information
- Application Number
- JP2025019327
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2045-02-07
AI Technical Summary
Conventional methods for representing molecular structures, such as SMILES notation and PDB format, struggle to accurately reflect the spatial arrangement and stereochemical features of molecules, leading to inefficiencies in molecular structure analysis and comparison, particularly in the pharmaceutical field.
A three-dimensional structure identification system that acquires spatial coordinate data, calculates an interatomic distance matrix, constructs a proximity score matrix by assigning higher scores to closer distances, and converts it into pixel values to form images for efficient comparison of molecular structures.
Enables accurate and efficient identification and comparison of three-dimensional molecular structures by emphasizing atomic interactions and reducing the influence of positional variations, facilitating molecular similarity searches and drug candidate screening.
Smart Images

Figure 0007715333000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a three-dimensional structure identification system, a three-dimensional structure identification method, and a program for identifying two or more solids composed of predetermined atoms. [Background technology]
[0002] For example, in the process of molecular design and molecular synthesis planning, the structures of new molecules and existing molecules are represented and compared for structural similarity. Particularly in the process of drug discovery in the pharmaceutical field, the structure of a drug candidate molecule is compared with the structure of an existing molecule to identify structural similarity.
[0003] Conventional methods for representing molecular structures include the Simplified Molecular Input Line Entry System (SMILES) notation, which converts the chemical structure of a molecule into a string of alphanumeric characters in ASCII code and represents it as a two-dimensional diagram or a three-dimensional model (e.g., Non-Patent Documents 1 and 2).
[0004] There is also a notation method for describing the coordinates of atoms constituting a molecule in Euclidean space. The PDB format is a widely used file format. Files in this PDB format can be obtained from the Protein Data Bank (PDB). The coordinates can be read using software such as RasMol (e.g., Non-Patent Document 3), PyMOL (e.g., Non-Patent Document 4), or VMD (e.g., Non-Patent Document 5), and the molecule can be displayed as a three-dimensional model. In addition to the PDB format, other formats such as the XYZ format and the mmCIF (Macromolecular Crystallographic Information File) format are also commonly used. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] 3. SMILES - A Simplified Chemical Language, [online], [searched on November 28, Reiwa 6], Internet, <https: / / daylight.com / dayhtml / doc / theory / theory.smiles.html>
Non - Patent Document 2
Non - Patent Document 3
Non - Patent Document 4
Non - Patent Document 5
Non - Patent Document 6
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, the SMILES notation is specialized in two-dimensionally describing the connection information of atoms and bonds within a molecule, and it is difficult to completely reproduce the spatial arrangement and stereochemical features of the molecule (e.g., chirality and stereoisomer information), and there is a limitation that it cannot directly represent the three-dimensional structure and conformation of the molecule.
[0007] Also, although the PDB (Protein Data Bank) notation can describe the three-dimensional structure of a molecule in detail as coordinate data, when the orientation of the molecular coordinate system and the setting of the origin are different, even the same molecule may have different notations. Therefore, it has been pointed out that when comparing molecular structures using the PDB notation, the difference in the coordinate system must be considered, which reduces the computational efficiency.
[0008] Since these limitations often become major obstacles when performing molecular structure analysis and comparison, in a wide range of fields such as molecular design and molecular synthesis, especially in the pharmaceutical field, there is a need to develop a method for more accurately reflecting the spatial features and stereochemical information of molecules and identifying the structural consistency or similarity between different molecules.
[0009] By the way, the minimum information required to define a molecular structure is the type of atoms that make up the molecule and their coordinates. Generally, the information recorded in the PDB format is also based on the type of atoms and their coordinates. What can be obtained with high precision by X-ray crystallography is mainly the coordinate information of heavy atoms (e.g., carbon, nitrogen, oxygen, etc.), and these information become important elements in molecular structure analysis.
[0010] Therefore, in order to efficiently compare molecular structure information as image data, the interatomic distance matrix is converted into a "Proximity Score Matrix (hereinafter also referred to as 'PSM')". In order to prevent information loss in this conversion and accurately retain the three-dimensional structure information of the molecule, in addition to covering the information of the interatomic distances, it is necessary to appropriately reflect the information regarding the type of atoms.
[0011] In view of these problems, an object of the present invention is to provide a three-dimensional structure identification system, a three-dimensional structure identification method, and a program for identifying two or more three-dimensional objects composed of predetermined atoms by extracting feature amounts that are not affected by the position or rotation of the three-dimensional objects such as molecules.
Means for Solving the Problems
[0012] The present invention provides the following solution means.
[0013] According to the invention according to the first feature, a three-dimensional structure identification system for identifying two or more three-dimensional objects composed of predetermined atoms, an acquisition unit that acquires spatial coordinate data of all atoms constituting each of the three-dimensional objects, a calculation unit that calculates, as an interatomic distance matrix, a matrix having distances between all atoms of each of the three-dimensional objects as elements based on the acquired spatial coordinate data, a construction unit that constructs a proximity score matrix by converting the calculated interatomic distance matrix so as to assign a higher score as the distance is closer, an image formation unit that converts the proximity score matrix into pixel values within a specific range and forms an image based on the pixel values, a comparison unit that compares the structures of all or part of the two or more three-dimensional objects by analyzing the image, and provides a three-dimensional structure identification system including the same.
[0014] According to the invention according to the first feature, in order to extract feature amounts that are not affected by the position or rotation of three-dimensional objects such as molecules, accurate and efficient identification of three-dimensional structures is always possible. In addition, since the three-dimensional structure of a molecule can be accurately represented as a matrix, comparison and generation of molecular structures can be efficiently performed by image recognition.
[0015] The invention according to the second feature is the invention according to the first feature, wherein the construction unit constructs a proximity score matrix by converting each element of the interatomic distance matrix into the reciprocal of the power of the element. A three-dimensional structure identification system is provided.
[0016] According to the invention related to the second feature, by converting each element of the interatomic distance matrix into the reciprocal of a power, the interaction between adjacent atoms is emphasized, the influence between atoms far apart is reduced, and the physical interactions within the molecular structure can be expressed more accurately.
[0017] The invention related to the third feature is the invention related to the first feature, The image forming unit normalizes the proximity score matrix within a specific range and converts it into pixel values within the range of 0 to 255. A three-dimensional structure identification system is provided.
[0018] According to the invention related to the third feature, by normalizing the proximity score matrix within a specific range (0 to 255), the scale of the data is unified, making it easier to compare different molecules and structures. And by further quantization, it is possible to perform efficient calculation processing while compressing the data volume.
[0019] The invention related to the fourth feature is the invention related to the second feature, The construction unit further constructs, as a reference center proximity score matrix, N sorted matrices in which each row element of the proximity score matrix is sorted in descending order with respect to each of the atoms (1, 2, 3,..., N) in order. The image forming unit converts the reference center proximity score matrix into pixel values within a specific range and forms an image based on the pixel values. A three-dimensional structure identification system is provided.
[0020] ]]According to the invention related to the fourth feature, since a set of matrices in which the distance scores between other atoms when each atom is the reference point are arranged in descending order can be obtained, by arranging in descending order of the distance scores, it becomes easier to extract characteristic patterns of the molecular structure (for example, the proximity and connectivity between specific atoms), and it is possible to efficiently search for data that meets specific conditions (for example, atom pairs within a certain distance range).
[0021] The invention according to the fifth feature is the invention according to the fourth feature, wherein the construction unit rearranges the elements of each row in descending order by swapping rows and columns so that the atom of interest becomes an element of one row and one column for each atom. A three-dimensional structure identification system is provided.
[0022] According to the invention according to the fifth feature, by rearranging the matrix based on the atom of interest, the characteristics of the molecular structure are emphasized, and it becomes possible to improve the consistency of data and the efficiency of analysis. Further, by rearranging in descending order, important elements with large distances and scores are aggregated at the beginning of the row, and unnecessary elements with small noise and influence are pushed to the rear, so that the influence of unnecessary data can be minimized.
[0023] The invention according to the fifth feature is the invention according to the third feature, wherein the image forming unit normalizes the reference center proximity score matrix within a specific range and converts it into pixel values within the range of 0 to 255. A three-dimensional structure identification system is provided.
[0024] According to the invention according to the fifth feature, by normalizing the reference center proximity score matrix within a specific range (0 to 255), the scale of the data is unified, so that comparison between different molecules and structures becomes easy. And by further quantizing, it becomes possible to perform efficient calculation processing while compressing the data volume.
[0025] The invention according to the sixth feature is the invention according to the fourth feature, wherein the image forming unit normalizes the reference center proximity score matrix within a specific range and converts it into pixel values within the range of 0 to 255. A three-dimensional structure identification system is provided.
[0026] According to the invention related to the sixth feature, by normalizing the reference center proximity score matrix within a specific range (0 to 255), the scale of the data is unified, making it easier to compare between different molecules and structures. And by further quantization, it becomes possible to perform efficient computational processing while compressing the data volume.
[0027] The invention related to the seventh feature is the invention related to the first feature, The construction unit associates, for each element of the proximity score matrix, atomic identification information based on the atomic number and molecular property information including valence, polarity, bond orientation, and interatomic angle, or one or more combinations thereof. The image forming unit assigns the pixel value to the first color channel, assigns the atomic identification information to the second color channel, assigns the molecular property information to the third color channel, and forms the image based on the first color channel, the second color channel, and the third color channel. Provide a three-dimensional structure identification system.
[0028] According to the invention related to the seventh feature, by providing information that constitutes an element of color for each element of the proximity score matrix, not only the distance between atoms but also the type of atom and molecular properties can be represented in the image, making it easier to visually and more accurately compare the structural similarities between each molecule, and enabling similarity search applying image processing technology. Also, in order to simultaneously hold the interatomic distance, atomic identification information, and molecular property information, it becomes possible to more accurately preserve the three-dimensional structure of the molecules in the PDB file data. Furthermore, since molecular similarity search using existing image comparison algorithms can be performed, it becomes possible to improve the efficiency of searching for new compounds and screening drug candidates.
[0029] Although the present invention belongs to the category of computer systems, it also exhibits similar actions and effects according to the category in other categories such as methods and programs.
Advantages of the Invention
[0030] According to the present invention, it is possible to provide a three-dimensional structure identification system, a three-dimensional structure identification method, and a program for identifying two or more three-dimensional objects composed of predetermined atoms by extracting feature amounts that are not affected by the position or rotation of the three-dimensional objects such as molecules.
Brief Description of the Drawings
[0031]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
[0032] Hereinafter, the best mode for carrying out the present invention will be described with reference to the drawings. Note that this is merely an example, and the technical scope of the present invention is not limited thereto.
[0033] [First Embodiment] [Outline of Three-Dimensional Structure Identification System 1] The outline of the three-dimensional structure identification system 1 according to the first embodiment of the present invention will be described with reference to FIG. 1. FIG. 1 is a diagram for explaining the outline of the three-dimensional structure identification system 1 according to the first embodiment of the present invention. The three-dimensional structure identification system 1 is a computer system configured from a computer 2 for identifying two or more three-dimensional objects.
[0034] The computer 2 of the three-dimensional structure identification system 1 is, for example, a computer such as a desktop personal computer, a notebook personal computer, or a server, a mobile terminal such as a smartphone or a tablet terminal, a wearable terminal such as a head-mounted display such as smart glasses or a smartwatch, and the like.
[0035] Also, the computer 2 of the three-dimensional structure identification system 1 may be realized by, for example, a single terminal device, a plurality of terminal devices, or a virtual device such as a cloud computer.
[0036] Also, the three-dimensional structure identification system 1 may be configured from the above-described terminal devices instead of the computer 2.
[0037] The computer 2 of the three-dimensional structure identification system 1 is connected to be capable of data communication with the above-described terminal devices, other terminals and devices, etc. via a public communication network or the like, and performs transmission and reception of necessary data and information.
[0038] Next, the outline of the processing executed by the three-dimensional structure identification system 1 will be described. First, the computer 2 of the three-dimensional structure identification system 1 acquires the spatial coordinate data of all atoms constituting each three-dimensional object (step S1). Specifically, the computer 2 extracts and acquires the spatial coordinate data 100 of the heavy atoms constituting the molecule from the PDB file data of two or more molecules to be identified. In this specification, heavy atoms refer to atoms other than hydrogen. Regardless of the method of obtaining the PDB file data, for example, if it is an actual molecular structure, it may be downloaded from an external device such as the RCSB PDB website (https: / / www.rcsb.org / ), and if it is a new molecular structure, it may be the RDB data obtained after the new molecule is generated. Also, the PDB file data may be stored in advance in the computer 2.
[0039] Next, the computer 2 calculates a matrix having the distances between all atoms of each three-dimensional object as elements as the interatomic distance matrix 200 based on the acquired spatial coordinate data 100 (step S2). Specifically, the computer 2 calculates, as the interatomic distance matrix 200, a matrix represented in matrix form with the Euclidean distance between the heavy atoms of each molecule as an element based on the spatial coordinate data 100 represented by the coordinates in the three-dimensional space acquired in step S1 above.
[0040] Next, the computer 2 constructs a proximity score matrix (PSM) 300 by converting the calculated interatomic distance matrix 200 so as to assign a higher score to a closer distance (step S3). Specifically, the computer 2 calculates a score for the distance by converting each element of the interatomic distance matrix 200 calculated in step S2 above into the reciprocal of the power, and constructs the proximity score matrix 300. Therefore, each element of the proximity score matrix 300 represents the proximity score between the corresponding atoms, and the closer the interatomic distance, the larger the element of the proximity score matrix 300 becomes.
[0041] In the above step S3, the computer 2 may associate, for each element of the proximity score matrix 300, atomic identification information based on the atomic number with molecular property information including valence, polarity, bond orientation, and interatomic angle, or one or more combinations thereof. Specifically, the computer 2 may associate, for each element of the proximity score matrix 300, a score for identifying a combination of atoms based on the atomic number with molecular property information including valence, polarity, bond orientation, and interatomic angle, or any one or more combinations thereof.
[0042] Next, the computer 2 converts the proximity score matrix 300 into pixel values within a specific range and forms an image based on the pixel values (step S4). Specifically, the computer 2 converts the value of each element of the proximity score matrix 300 constructed in step S3 above into an appropriate range as the value of a pixel of the image, and forms a proximity score image 500 based on the pixel values after the proximity scores are converted. Therefore, the shading of the formed image will be determined based on the elements of the corresponding matrix.
[0043] In the above step S4, the computer 2 may assign the converted pixel values to the first color channel, assign the atomic identification information associated with each element of the proximity score matrix 300 in step S3 to the second color channel, assign the molecular property information associated with each element of the proximity score matrix 300 in step S3 to the third color channel, and form an image based on the first color channel, the second color channel, and the third color channel. Specifically, the computer 2 may, for example, assign pixel values to the R channel based on the RGB model, assign atomic identification information to the G channel, assign molecular property information to the B channel, and generate a color image based on these channels. The color model may be a model other than the RGB model, such as the CMYK model or the HSV model, and is not particularly limited.
[0044] Next, the computer 2 compares the overall or partial structures of two or more solids by analyzing the images (step S5). Specifically, the computer 2 searches for the whole or part that matches or is similar in two or more molecular structures represented by the proximity score image 500 by comparing the whole or part of the proximity score image 500 obtained in step S4 above on a pixel-by-pixel basis.
[0045] The above is an overview of the processing executed by the three-dimensional structure identification system 1.
[0046] [System Configuration of the Three-Dimensional Structure Identification System 1] Based on FIG. 2, the system configuration of the three-dimensional structure identification system 1 of the present embodiment will be described. The three-dimensional structure identification system 1 is composed of a computer 2 and is a computer system for identifying two or more solids.
[0047] Note that the three-dimensional structure identification system 1 may include other terminals and devices. For example, a different computer 2 may be used for each user. In this case, the three-dimensional structure identification system 1 will execute each process described below by any one or a plurality of combinations of the computer 2 and other included terminals and devices.
[0048] The computer 2 may be realized by, for example, a single terminal device, a plurality of terminal devices, or a virtual device such as a cloud computer.
[0049] The computer 2 is, for example, a computer such as a desktop personal computer, a notebook personal computer, or a server, a mobile terminal such as a smartphone or a tablet terminal, or a wearable terminal such as a head-mounted display like smart glasses or a smartwatch.
[0050] The computer 2 includes, as a control unit, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a RAM (Random Access Memory), a ROM (Read Only Memory), etc.
[0051] The computer 2 includes, as a storage unit, data storage by means of a hard disk, a semiconductor memory, a recording medium, a memory card, etc. The data storage may be internal data storage and / or external data storage. The storage destination of the data may be a cloud service, a database, etc.
[0052] The computer 2 includes, as a communication unit, a device for enabling communication with other terminals, devices, etc. The communication method may be wireless or wired.
[0053] The computer 2 is assumed to include functions necessary for operating the computer 2 as an input unit. As an example for realizing input, it is possible to include a liquid crystal display for realizing a touch panel function, a keyboard, a mouse, a tablet, a hardware button on the device, a microphone for performing voice recognition, etc. The present invention is not particularly limited by the input method.
[0054] The computer 2 is assumed to include functions necessary for the user of the three-dimensional structure identification system 1 to operate the computer 2 as an output unit. As an example for realizing output, forms such as display on a liquid crystal display, a PC display, projection onto a projector, and voice output are conceivable. The present invention is not particularly limited by the output method.
[0055] The control unit, in cooperation with the processing unit, realizes a calculation unit 22 and a construction unit 23. Also, the control unit, in cooperation with the processing unit, the communication unit, and the storage unit, realizes an acquisition unit 21. Also, the control unit, in cooperation with the processing unit and the output unit, realizes an image formation unit 24 and a comparison unit 25.
[0056] The above is the system configuration of the three-dimensional structure identification system 1.
[0057] [Three-dimensional structure identification process] Based on FIG. 3, the three-dimensional structure identification process executed by the computer 2 will be described. FIG. 3 is a diagram showing a flowchart of the three-dimensional structure identification process executed by the computer 2. As shown in FIG. 3, the three-dimensional structure identification process is composed of steps S11 to S15 and is the process executed in the above-described steps S1 to S5.
[0058] First, the acquisition unit 21 of the computer 2 of the three-dimensional structure identification system 1 acquires the spatial coordinate data of all the atoms constituting each three-dimensional structure (step S11). Specifically, the acquisition unit 21 extracts and acquires the spatial coordinate data 100 of the heavy atoms constituting the molecule from the PDB file data of two or more molecules to be identified. The PDB file data is uniquely identified by a PDB ID indicating specific molecular structure data such as proteins, DNA, RNA, ligands, etc. Regardless of the method of obtaining the PDB file data, for example, if it is an actual molecular structure, it may be downloaded from an external device such as the RCSB PDB website (https: / / www.rcsb.org / ) to the computer 2 via the communication unit, and if it is a new molecular structure, it may be the PDB file data obtained after the new molecule is generated. Also, the PDB file data may be stored in advance in the storage unit of the computer 2. A detailed description of the data items indicated by the PDB file data will be omitted.
[0059] FIG. 4 is a diagram showing an example of PDB file data, and FIG. 5 is a diagram showing the HETATM record of the heavy atoms extracted from the PDB file data shown in FIG. 4. The acquisition unit 21 extracts, for example, the HETATM record of the heavy atoms from the PDB file data shown in FIG. 4 as shown in FIG. 5, and acquires the X, Y, and Z coordinate data in the three-dimensional space of the atoms included in the HETATM record as the spatial coordinate data 100.
[0060] Next, the calculation unit 22 of the computer 2 calculates a matrix with the distances between all atoms of each solid as elements as the interatomic distance matrix 200 based on the acquired spatial coordinate data 100 (step S12). Specifically, the calculation unit 22 calculates, as the interatomic distance matrix 200, a matrix represented in matrix form with the Euclidean distance between the heavy atoms of each molecule as an element based on the spatial coordinate data 100 represented by the coordinates in the three-dimensional space acquired in step S11 above.
[0061] FIG. 6 is a diagram for explaining the interatomic distance matrix 200. For example, for heavy atom i located at coordinates (x i , y i , z i ) and heavy atom j located at coordinates (x j , y j , z j ), the Euclidean distance r ij between heavy atom i and heavy atom j is calculated by the following formula.
Equation
[0062] Therefore, as shown in FIG. 6, for each heavy atom a, b, c, d,... constituting one molecule, the interatomic distances r ab , r ac , r ad ,... obtained by the formula shown in Equation 1 above are used as elements to calculate the interatomic distance matrix 200.
[0063] Next, the construction unit 23 of the computer 2 constructs a proximity score matrix 300 by converting the calculated interatomic distance matrix 200 so as to assign a higher score to closer distances (step S13). Specifically, the construction unit 23 calculates a score for the distance by converting each element of the interatomic distance matrix 200 calculated in step S12 above into its reciprocal, and constructs the proximity score matrix 300. Therefore, each element of the proximity score matrix 300 represents the proximity score between the corresponding atoms, and the closer the interatomic distance, the larger the element of the proximity score matrix 300.
[0064] FIG. 7 is a diagram for explaining the proximity score matrix 300. For example, for each element of the interatomic distance matrix 200 calculated in step S12 described above, the interatomic distance r ab , r ac , r ad , ··· is raised to the power of the scaling exponent k (k < 0), and the interatomic distance r ab , r ac , r ad , ··· is converted to its reciprocal, and scores shown by the following formulas 2 to 4 etc. are calculated.
Equation
Equation
Equation
[0065] Therefore, for the interatomic distance matrix 200 as shown in FIG. 6, as shown in FIG. 7, a matrix with scores shown by formulas 2 to 4 etc. as elements is calculated as the proximity score matrix 300.
[0066] The scaling exponent k may take any value as long as k < 0, but considering the simplicity of calculation, when rounding the value of the matrix element to one decimal place, it is preferably -2.15 ≦ k ≦ -1.95.
[0067] FIG. 8 is a diagram for explaining the molecular structure information represented by the proximity score matrix 300 when the scaling exponent k = -2.15. For example, for the interatomic distance matrix 200 of the molecular structure A, in the molecular structure A shown by the proximity score matrix 300 when K = -2.15, it represents the correlation up to 2 bonds, but does not represent the long-range correlation, and as shown in FIG. 8, it does not contain steric information.
[0068] FIG. 9 is a diagram for explaining the molecular structure information represented by the proximity score matrix 300 when the scaling index k = -2.0. According to the structure indicated by the proximity score matrix 300 when K = -2.0 for the interatomic distance matrix 200 of the molecular structure A, compared with the case of K = -2.15 shown in FIG. 8, since the proximity scores shown in bold are calculated, long-range correlations are displayed and short-range steric information is included. However, since there is a threshold near the distance between the substituents at positions 1 and 4 in the molecule, it is difficult to distinguish the relative steric configurations of these substituents.
[0069] FIG. 10 is a diagram for explaining the molecular structure information represented by the proximity score matrix 300 when the scaling index k = -1.95. According to the structure indicated by the proximity score matrix 300 when K = -1.95 for the interatomic distance matrix 200 of the molecular structure A, compared with the case of K = -2.0 shown in FIG. 9, since the proximity scores shown underlined are calculated, interpretable short-range steric information is included, so it becomes possible to distinguish whether the steric configurations of the substituents at positions 1 and 4 in the molecule are in the 1,4-trans conformation or the 1,4-gauche conformation.
[0070] In the above step S13, the construction unit 23 of the computer 2 may associate, with respect to each element of the proximity score matrix 300, atom identification information based on the atomic number and molecular property information including the valence number, polarity, bond orientation, and interatomic angle, or one or more combinations thereof. Specifically, the construction unit 23 may associate, with respect to each element of the proximity score matrix 300, a score for identifying a combination of atoms based on the atomic number and molecular property information including the valence number, polarity, bond orientation, and interatomic angle, or any one or any one or more combinations thereof.
[0071] The atomic identification information may be, for example, for each combination of atoms constituting each element of the proximity score matrix 300, an atomic number is assigned in three digits, and the two atomic numbers are combined to obtain an identification number as the score. For example, the score between carbon-carbon is "006006", the score between carbon-nitrogen is "006007", the score between nitrogen-carbon is "007006", and the score between carbon-oxygen is "006008".
[0072] The valence number, polarity, bond orientation, and interatomic angle that may be included in the molecular property information may be indirectly estimated or calculated from the PDB file data by using a specific application. In particular, for the valence number, it may be derived from the interatomic distance matrix 200 calculated in step S12 by the Bond Valency Method (bond atomic valence method) based on Non-Patent Document 6. Note that the parameters that may be included in the molecular property information are not limited to the valence number, polarity, bond orientation, and interatomic angle. Also, the estimation method and calculation method for these parameters are not particularly limited.
[0073] Next, the image forming unit 24 of the computer 2 converts the proximity score matrix 300 into pixel values within a specific range, and forms an image based on the pixel values (step S14). Specifically, the image forming unit 24 converts the value of each element of the proximity score matrix 300 constructed in step S13 above into an appropriate range as the pixel value of the image, and forms a proximity score image 500 based on the pixel values after the proximity score is converted. Therefore, the shading of the formed image is determined based on the elements of the corresponding matrix.
[0074] The appropriate range as the above pixel value may be a numerical range represented by 8×n (n: natural number) bits. However, considering the memory efficiency of the computer 2, the range of 0 - 255 is preferable. FIG. 11 is a diagram showing a proximity score image 500 corresponding to the pixel values obtained by converting a certain specific proximity score matrix 300. In FIG. 11, a proximity score image 500 corresponding to the pixel values obtained by converting the proximity score matrix 300 of the molecular structure B is shown. In the proximity score matrix 300, elements with a small score are represented by "dark colors (e.g., black)", and elements with a large score are represented by "bright colors (e.g., white)".
[0075] In the above step S14, the image forming unit 24 of the computer 2 may assign the converted pixel values to the first color channel, assign the atom identification information associated with each element of the proximity score matrix 300 in the above step S13 to the second color channel, assign the molecular property information associated with each element of the proximity score matrix 300 in the above step S13 to the third color channel, and form the proximity score image 500 based on the first color channel, the second color channel, and the third color channel. Specifically, the computer 2 may, for example, assign pixel values to the R channel based on the RGB model, assign atom identification information to the G channel, assign molecular property information to the B channel, and form the proximity score image 500 of the color image based on these channels. Which information is associated with which color channel may be arbitrary. The color model for the color channel may be a model other than the RGB model, such as the CMYK model or the HSV model, and is not particularly limited.
[0076] Next, the comparison unit 25 of the computer 2 compares the overall or partial structures of two or more three-dimensional objects by analyzing the image (step S15). Specifically, the comparison unit 25 searches for the overall or partial parts that match or are similar in two or more molecular structures represented by the proximity score image 500 by comparing the whole or a part of the proximity score image 500 obtained in the above step S14 in pixel units using a difference image.
[0077] FIG. 12 is a diagram showing an example of a proximity score image 500 of molecular structures to be compared. For example, as shown in FIG. 12, by comparing the proximity score images 500 of molecular structure C and molecular structure D, in the proximity score image 500 of molecular structure D, the portion within the enclosed frame is searched as the portion that coincides with the pixels of the entire proximity score image 500 of molecular structure C.
[0078] The proximity score matrix 300 constructed in step S13 described above may be used in cooperation with a large language model (LLM) to construct an AI system. Specifically, the constructed proximity score matrix 300 may be vectorized and learned by the LLM. Also, using principal component analysis (PCA) or a convolutional neural network (CNN), important features may be extracted from the proximity score matrix 300 and learned by the LLM. Note that the learning process may be executed at any timing as long as it is after step S13 described above.
[0079] In step S14 described above, when the proximity score image 500 is created based on a color channel, in addition to the image comparison algorithm based on pixels, the proximity score image 500 may be compared by existing image comparison algorithms such as comparison considering the similarity of structures such as SSIM (Structural Similarity Index) and feature quantity-based comparison such as histogram comparison, and the image comparison algorithm to be used is not particularly limited.
[0080] The above is the three-dimensional structure identification process.
[0081] Therefore, according to the three-dimensional structure identification system 1, in order to extract feature quantities that are not affected by the position or rotation of a three-dimensional object such as a molecule, it is always possible to accurately and efficiently identify the three-dimensional structure. Also, since the three-dimensional structure of a molecule can be accurately represented as a matrix, it is possible to efficiently compare and generate molecular structures by image recognition.
[0082] According to the three-dimensional structure identification system 1, further, by converting each element of the interatomic distance matrix into the reciprocal of a power, the interaction between adjacent atoms is emphasized, the influence between atoms far apart is reduced, and the physical interactions within the molecular structure can be more accurately represented.
[0083] According to the three-dimensional structure identification system 1, further, by normalizing the proximity score matrix within a specific range (0 to 255), the scale of the data is unified, making it easier to compare different molecules and structures. Then, by further quantization, it is possible to perform efficient computational processing while compressing the data volume.
[0084] [Second Embodiment] [Overview of the Three-Dimensional Structure Identification System 1] The overview of the three-dimensional structure identification system 1 according to the second embodiment of the present invention will be described with reference to FIG. 13. FIG. 13 is a diagram for explaining the overview of the three-dimensional structure identification system 1 according to the second embodiment. The three-dimensional structure identification system 1 is composed of a computer 2 and is a computer system for identifying two or more three-dimensional objects. Note that the same reference numerals are given to the same functions and configurations as those in the above-described first embodiment, and the description thereof is omitted. The difference from the above-described first embodiment is that after step S3, a reference center proximity score matrix 310 is constructed (step S31), an image is formed based on the reference center proximity score matrix 310 (step S41), and the structure of the whole or a part of two or more three-dimensional objects is compared in step S51 by analyzing the image (step S51).
[0085] Since the computer 2 of the three-dimensional structure identification system 1 is the same as that in the above-described first embodiment, the description thereof is omitted.
[0086] Next, an overview of the processing executed by the three-dimensional structure identification system 1 will be described. First, the computer 2 of the three-dimensional structure identification system 1 acquires the spatial coordinate data of all atoms constituting each three-dimensional object (step S1). Since this step is the same as that in the above-described first embodiment, the description thereof is omitted.
[0087] Next, based on the acquired spatial coordinate data 100, computer 2 calculates an interatomic distance matrix 200 having, as elements, the distances between all atoms of each solid (step S2). Since this step is the same as that of the first embodiment described above, its description is omitted.
[0088] Next, computer 2 constructs a proximity score matrix 300 by converting the calculated interatomic distance matrix 200 so as to assign a higher score to a closer distance (step S3). Since this step is the same as that of the first embodiment described above, its description is omitted.
[0089] Next, for the proximity score matrix 300, computer 2 further constructs, as a reference center proximity score matrix 310, N sorted matrices in which each atom (1, 2, 3, …, N) is sequentially targeted and the elements of each row are arranged in descending order (step S31). Specifically, computer 2 further constructs, as a reference center proximity score matrix 310, N sorted matrices in which each atom (1, 2, 3, …, N) is sequentially targeted and the elements of each row are arranged in descending order for the proximity score matrix 300 constructed in step S3 above.
[0090] Next, computer 2 converts the reference center proximity score matrix 310 into pixel values within a specific range, and forms an image based on the pixel values (step S41). Specifically, computer 2 converts the value of each element of the reference center proximity score matrix 310 constructed in step S31 above into an appropriate range as the value of a pixel of the image, and forms a reference center proximity score image 510 with the converted reference center proximity score as the pixel value. Therefore, the shading of the formed image is determined based on the elements of the corresponding matrix.
[0091] Next, the computer 2 compares the overall or partial structure of two or more solids by analyzing the images (step S51). Specifically, the computer 2 searches for the overall or partial structure that matches or is similar in two or more molecular structures represented by the reference center proximity score image 510 by comparing all or part of the reference center proximity score image 510 obtained in step S41 above on a pixel-by-pixel basis.
[0092] In addition, in this embodiment, the processes of steps S4 - S5 of the first embodiment described above may be executed, and the timing for executing the processes of these steps may be any timing as long as it is after step S3.
[0093] The above is an overview of the processes executed by the three-dimensional structure identification system 1.
[0094] [System Configuration of the Three-Dimensional Structure Identification System 1] Since the system configuration of the three-dimensional structure identification system 1 according to the second embodiment of the present invention is the same as that of the first embodiment described above, the description thereof is omitted.
[0095] [Three-Dimensional Structure Identification Process] Based on FIG. 14, the three-dimensional structure identification process executed by the computer 2 will be described. FIG. 3 is a diagram showing a flowchart of the three-dimensional structure identification process executed by the computer 2. As shown in FIG. 14, the three-dimensional structure identification process is composed of steps S11 to S151 and is the process executed in steps S1 to S51 described above.
[0096] First, the acquisition unit 21 of the computer 2 acquires the spatial coordinate data of all the atoms constituting each solid (step S11). This step is the same as that of the first embodiment described above, so the description thereof is omitted.
[0097] Next, the calculation unit 22 of the computer 2 calculates an interatomic distance matrix 200 having, as elements, the distances between all atoms of each solid based on the acquired spatial coordinate data 100 (step S12). Since this step is the same as that of the first embodiment described above, the description thereof is omitted.
[0098] Next, the construction unit 23 of the computer 2 constructs a proximity score matrix 300 by converting the calculated interatomic distance matrix 200 so as to assign a higher score to closer distances (step S13). Since this step is the same as that of the first embodiment described above, the description thereof is omitted.
[0099] Next, the construction unit 23 of the computer 2 further constructs, as a reference center proximity score matrix 310, N sorted matrices obtained by sorting the elements of each row in descending order for each atom (1, 2, 3, …, N) with respect to the proximity score matrix 300 (step S131). Specifically, the construction unit 23, for example, moves the i-th row of the proximity score matrix 300 composed of N rows and N columns (N: natural number) to the first row and moves the i-th column to the first column. Next, the columns with a value of 0 in the first row are deleted, and the columns after the second column are sorted so that the values in the first row are in descending order. Then, the columns with a value of 0 in the first row are deleted, and the columns after the second column are sorted so that the values in the first row are in descending order, thereby constructing the reference center proximity score matrix 310. By repeating this for i = 1 to N, N reference center proximity score matrices 310 are constructed.
[0100] FIGS. 15 to 18 are diagrams for explaining the construction process of the reference center proximity score matrix 310 composed of 13 rows and 13 columns. As shown in bold in FIG. 15, the 13th row is moved to the 1st row, and as shown in bold in FIG. 16, the 13th column is moved to the 1st column. Next, as shown in bold in FIG. 17, the construction unit 23 deletes the columns where the value in the first row is 0, sorts the columns after the second column so that the values in the first row are in descending order, and further, as shown in bold in FIG. 18, deletes the columns where the value in the first row is 0, and sorts the columns after the second column so that the values in the first row are in descending order. As a result, the reference center proximity score matrix 310 in the case of i = 13 is constructed. For the sake of convenience in explanation, the case of i = 13 is described as an example. Actually, before the process of i = 13, the processes of i = 1 to 12 are sequentially performed, and the reference center proximity score matrices 310 in the cases of i = 1 to 12 are respectively constructed.
[0101] Next, the image forming unit 24 of the computer 2 converts the reference center proximity score matrix 310 into pixel values within a specific range, and forms an image based on the pixel values (step S141). Specifically, the image forming unit 24 converts the values of each element of the N reference center proximity score matrices 310 constructed in step S131 described above into an appropriate range as the values of the pixels of the image, and forms a reference center proximity score image 510 using the converted reference center proximity scores as pixel values. Therefore, the shading of the formed image is determined based on the elements of the corresponding matrix.
[0102] The appropriate range as the above-described pixel values may be a numerical range expressed in 8×n (n: natural number) bits. Considering the memory efficiency of the computer 2, the range of 0 to 255 is preferable. FIG. 19 is a diagram showing the reference center proximity score image 510 corresponding to the pixel values obtained by converting the reference center proximity score matrix 310. In FIG. 19, a proximity score image 500 corresponding to the pixel values obtained by converting the N (N = 13) reference center proximity score matrices 310 of the molecular structure B is shown. In the reference center proximity score matrix 310, elements with a small score are represented by "dark colors (e.g., black)", and elements with a large score are represented by "bright colors (e.g., white)".
[0103] In step S141 described above, the image forming unit 24 of the computer 2 assigns the converted pixel values to the first color channel, assigns the atom identification information associated with each element of the reference center proximity score matrix 310 to the second color channel, assigns the molecular property information associated with each element of the reference center proximity score matrix 310 to the third color channel, and may form a reference center proximity score image 510 based on the first color channel, the second color channel, and the third color channel. Specifically, for example, the computer 2 may assign pixel values to the R channel based on the RGB model, assign atom identification information to the G channel, assign molecular property information to the B channel, and form a reference center proximity score image 510 of a color image based on these channels. Which information is associated with which color channel may be arbitrary. The color model for the color channel may be a model other than the RGB model, such as the CMYK model or the HSV model, and is not particularly limited.
[0104] Next, the comparison unit 25 of the computer 2 compares the overall or partial structures of two or more three-dimensional objects by analyzing the images (step S151). Specifically, the comparison unit 25 compares all or part of the N (N: natural number) reference center proximity score images 510 obtained in step S141 described above pixel by pixel using a difference image, and searches for all or part that match or are similar in two or more molecular structures represented by the reference center proximity score image 510.
[0105] In addition, in this embodiment, the processes of steps S14 - S15 of the first embodiment described above may be executed, and the timing of executing the processes of these steps may be any timing as long as it is after step S13.
[0106] The proximity score matrix 300 constructed in the above step S131 may be used to build an AI system in cooperation with a large language model (LLM). Specifically, the constructed proximity score matrix 300 may be vectorized and learned by the LLM. Also, important features may be extracted from the proximity score matrix 300 using principal component analysis (PCA) or convolutional neural network (CNN) and then learned by the LLM. Note that this learning process may be executed at any timing as long as it is after the above step S13.
[0107] In the above step S141, when the reference center proximity score image 510 is created based on the color channel, in addition to the image comparison algorithm based on pixels, the reference center proximity score image 510 may be compared by existing image comparison algorithms such as comparison considering the structural similarity like SSIM (Structural Similarity Index) or feature-based comparison like histogram comparison, and the image comparison algorithm to be used is not particularly limited.
[0108] Regarding the reference center proximity score matrix 310 constructed in the above step S131, for the reference center proximity score matrix 310 related to a molecular structure with a clear three-dimensional structure, a database of the atomic configuration of real molecules (ACRM) may be constructed and stored in the storage unit of the computer 2. This storage process of the reference center proximity score matrix 310 may be executed at any timing as long as it is after the above step S131. Therefore, for a new molecular structure, it is possible to optimize it as a feasible molecular structure by searching the ACRM database.
[0109] The above is the three-dimensional structure identification process.
[0110] Therefore, according to the three-dimensional structure identification system 1, since a set of matrices arranged in descending order of the distance score between other atoms when each atom is used as a reference point can be obtained, by arranging in descending order of the distance score, it becomes easier to extract characteristic patterns of the molecular structure (for example, proximity and connectivity between specific atoms), and it becomes possible to efficiently search for data that meets specific conditions (for example, atom pairs within a certain distance range).
[0111] According to the three-dimensional structure identification system 1, furthermore, by rearranging the matrices based on the atoms of interest, the characteristics of the molecular structure are emphasized, and it becomes possible to improve the consistency of the data and the efficiency of analysis. Also, by arranging in descending order, important elements with large distances and scores are aggregated at the beginning of the row, and unnecessary elements with little noise and influence are pushed to the back, so it becomes possible to minimize the influence of unnecessary data.
[0112] According to the three-dimensional structure identification system 1, furthermore, by normalizing the reference center proximity score matrix within a specific range (0 to 255), the scale of the data is unified, so it becomes easier to compare between different molecules and structures. And by further quantization, it becomes possible to perform efficient computational processing while compressing the data volume.
[0113] According to the three-dimensional structure identification system 1, furthermore, by giving information that constitutes the color to each element of the proximity score matrix, not only the distance between atoms but also the type of atom and molecular characteristics can be represented in an image. Therefore, it becomes easier to visually and more accurately compare the structural similarities between each molecule, and similarity search applying image processing technology becomes possible. Also, in order to simultaneously hold the interatomic distance, atom identification information, and molecular characteristic information, it becomes possible to more accurately preserve the three-dimensional structure of the molecules in the PDB file data. Furthermore, since molecular similarity search using existing image comparison algorithms can be performed, it becomes possible to improve the efficiency of searching for new compounds and screening drug candidates.
[0114] According to the three-dimensional structure identification system 1, further, by normalizing the reference center proximity score matrix 310 within a specific range (0 to 255), the scale of the data is unified, making it easier to compare different molecules and structures. Then, by further quantizing, it is possible to perform efficient computational processing while compressing the data volume.
[0115] The above-described means and functions are realized by a computer (including a CPU, an information processing device, and various terminals) reading and executing a predetermined program. The program is provided, for example, in a form provided from one or more computers via a network (cloud service, SaaS: Software as a Service). Also, the program is provided, for example, in a form recorded on a computer-readable recording medium. In this case, the computer reads the program from the recording medium, transfers it to an internal recording device or an external recording device, records it, and then executes it. Further, the program may be pre-recorded in a recording device (recording medium) such as a magnetic disk, an optical disk, or a magneto-optical disk, and provided to the computer via a communication line from the recording device.
[0116] As described above, the embodiments of the present invention have been explained, but the present invention is not limited to these embodiments described above. Also, the effects described in the embodiments of the present invention are merely an enumeration of the most suitable effects resulting from the present invention, and the effects according to the present invention are not limited to those described in the embodiments of the present invention.
Explanation of Reference Numerals
[0117] 1 Three-dimensional structure identification system, 2 Computer, 21 Acquisition unit, 22 Calculation unit, 23 Construction unit, 24 Image formation unit, 25 Comparison unit, 100 Spatial coordinate data, 200 Interatomic distance matrix, 300 Proximity score matrix, 310 Reference center proximity score matrix, 500 Proximity score image, 510 Reference center proximity score image
Claims
1. A three-dimensional structure identification system for identifying the consistency or similarity of structures in two or more three-dimensions composed of specified atoms, comprising: an acquisition unit that acquires spatial coordinate data of all atoms constituting each of the three-dimensions; a calculation unit that calculates, based on the acquired spatial coordinate data, a matrix having distances between all atoms of each of the three-dimensions as elements as an inter-atomic distance matrix; a construction unit that constructs a proximity score matrix by converting the calculated inter-atomic distance matrix so as to assign a higher score as the distance is closer; an image formation unit that converts the proximity score matrix into pixel values within a specific range and forms an image based on the pixel values; a comparison unit that compares the structures of all or part of the two or more three-dimensions by analyzing the image; and the construction unit constructs a proximity score matrix by converting each element of the inter-atomic distance matrix into the reciprocal of the power of the element; the construction unit further constructs, for the proximity score matrix, N sorted matrices obtained by sequentially targeting each of the atoms (1, 2, 3,..., N) and sorting the elements of each row in descending order as a reference center proximity score matrix; the image formation unit converts the reference center proximity score matrix into pixel values within a specific range and forms an image based on the pixel values, a three-dimensional structure identification system.
2. The image formation unit normalizes the proximity score matrix within a specific range and converts it into pixel values within the range of 0 to 255, The three-dimensional structure identification system according to Claim 1.
3. The construction unit sorts the elements of each row in descending order by swapping rows and columns so that the atom of interest becomes the element of the first row and the first column for each of the atoms, The three-dimensional structure identification system according to Claim 1.
4. The image formation unit normalizes the reference center proximity score matrix within a specific range and converts it into pixel values within the range of 0 to 255, The three-dimensional structure identification system according to Claim 1.
5. The construction unit associates, for each of the elements of the proximity score matrix, atomic identification information based on atomic numbers with molecular property information including valence, polarity, bond orientation, and inter-atomic angle, or one or more combinations thereof, The image forming unit assigns the pixel values to a first color channel, assigns the atom identification information to a second color channel, assigns the molecular property information to a third color channel, and forms the image based on the first color channel, the second color channel, and the third color channel. The three-dimensional structure identification system according to claim 1.
6. A three-dimensional structure identification method for identifying the consistency or similarity of structures in two or more three-dimensional objects composed of predetermined atoms executed by a computer, comprising: acquiring spatial coordinate data of all atoms constituting each of the three-dimensional objects; calculating, based on the acquired spatial coordinate data, a matrix having distances between all atoms of each of the three-dimensional objects as elements as an inter-atomic distance matrix; constructing a proximity score matrix by converting the calculated inter-atomic distance matrix so as to assign a higher score as the distance is closer; converting the proximity score matrix into pixel values within a specific range, and forming an image based on the pixel values; comparing the structures of all or part of the two or more three-dimensional objects by analyzing the image; and comprising: The constructing step constructs a proximity score matrix by converting each element of the inter-atomic distance matrix into the reciprocal of the power of the element. The constructing step further constructs, for the proximity score matrix, N sorted matrices obtained by sequentially targeting each of the atoms (1, 2, 3,..., N) and arranging the elements of each row in descending order as a reference center proximity score matrix. The step of forming the image is a three-dimensional structure identification method that converts the reference center proximity score matrix into pixel values within a specific range and forms an image based on the pixel values.
7. On a computer, acquiring spatial coordinate data of all atoms constituting each of the three-dimensional objects; calculating, based on the acquired spatial coordinate data, a matrix having distances between all atoms of each of the three-dimensional objects as elements as an inter-atomic distance matrix; constructing a proximity score matrix by converting the calculated inter-atomic distance matrix so as to assign a higher score as the distance is closer; converting the proximity score matrix into pixel values within a specific range, and forming an image based on the pixel values; comparing the structures of all or part of two or more three-dimensional objects by analyzing the image; A computer-readable program for causing execution, wherein the constructing step constructs a proximity score matrix by converting each element of the interatomic distance matrix into the reciprocal of a power thereof, the constructing step further constructs, as a reference center proximity score matrix, N sorted matrices obtained by sorting the elements of each row in descending order for each of the atoms (1, 2, 3, ..., N) in order with respect to the proximity score matrix, the step of forming the image is a computer-readable program for forming an image based on pixel values obtained by converting the reference center proximity score matrix within a specific range.
Citation Information
Patent Citations
Drug-protein interaction prediction model based on convolutional neural network
CN113593633A
Method an device for identifying molecular attributes, and method and device for training identification model
CN113823360A
Method and system for converting a protein data bank file into a two-dimensional numerical matrix
US20240221869A1