Characteristic quantity calculation method, screening method and compound creation method

By calculating the cross-sectional area characteristic amount between the object structure and the probe, combined with the machine learning generator, the problem of inaccurate chemical properties of compounds in the prior art is solved, and efficient screening of medical candidate compounds and three-dimensional structure creation is achieved.

CN116157680BActive Publication Date: 2025-08-26FUJIFILM CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202180063534.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-30
Filing Date
2021-09-28
Publication Date
2025-08-26
Estimated Expiration
2041-09-28

AI Technical Summary

Technical Problem

In the prior art, the characteristic amount fails to accurately represent the chemical properties of the compound, resulting in low efficiency in screening and stereoscopic structure creation of medical candidate compounds, especially when facing complex target proteins, the number of compounds is huge and it is difficult to investigate bonding forces.

Method used

By calculating the cross-sectional area characteristic quantity between the object structure and the probe, combined with the machine learning generator, the three-dimensional structure of the compound is constructed to achieve accurate representation and efficient screening of the characteristic quantity.

Benefits of technology

Accurate representation of the chemical properties of the compound is achieved, the screening efficiency and stereoscopic structure creation efficiency of medical candidate compounds are improved, and compounds that are bound to the target protein can be effectively discovered.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116157680B_ABST
    Figure CN116157680B_ABST
Patent Text Reader

Abstract

One embodiment of the present invention provides a feature quantity calculation method capable of calculating a feature quantity accurately representing the chemical properties of a target structure, a screening method capable of efficiently screening candidate pharmaceutical compounds using the feature quantity, and a compound creation method capable of efficiently creating the stereostructure of a candidate pharmaceutical compound using the feature quantity. In one embodiment of the present invention, the feature quantity calculation method comprises: a target structure specifying step for specifying a target structure composed of multiple unit structures having chemical properties; a stereostructure acquisition step for acquiring the stereostructure of the target structure based on the multiple unit structures; and a probe feature quantity calculation step for calculating a feature quantity representing the cross-sectional area of ​​one or more probes relative to the target structure, the probes being structures having real charges and having multiple points separated by van der Waals forces.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a drug discovery support technology, and in particular to the calculation of characteristic quantities, the screening of drug candidate compounds, and the creation of the three-dimensional structure of drug candidate compounds. Background Art

[0002] In the past, computer-based drug discovery and development has involved preparing a library of tens to hundreds of thousands of existing compounds and providing the structural formulas of the compounds to investigate their binding forces to a target protein, thereby searching for candidate pharmaceutical compounds (hereinafter referred to as "hits"). For example, Patent Document 1 below describes providing the structural formula of a compound to predict its binding force. Furthermore, Patent Document 2 describes repeating the generation of structural formulas and prediction of binding forces to gradually search for compounds with the desired binding force (repeated experiments).

[0003] Furthermore, Patent Document 3 describes searching using descriptors called "compound fingerprints." "Descriptors" refer to information derived from a compound's structural formula, and "compound fingerprints" indicate information such as the presence or absence of various functional groups. A characteristic of these descriptors is that similar descriptors indicate similar compound structures.

[0004] Furthermore, searches for structures of compounds having desired physical property values ​​have traditionally been conducted primarily by solving "forward problems" (providing a molecular structure as a cause of the problem to determine the resulting physical property values). However, with the recent development of informatics, research into solutions to "inverse problems" (providing physical property values ​​to determine the molecular structure having those physical property values) has been rapidly advancing. For example, non-patent document 1 is known regarding structure searches based on solving inverse problems.

[0005] Prior art literature

[0006] Patent Literature

[0007] Patent Document 1: U.S. Patent No. 9,373,059

[0008] Patent Document 2: Japanese Patent No. 5946045

[0009] Patent Document 3: Japanese Patent No. 4564097

[0010] Non-patent literature

[0011] Non-Patent Literature 1: “Bayesian molecular design with a chemical language”, Hisaki Ikebata et al. [retrieved July 17, 2020], Internet (https: / / www.ncbi.nlm.nih.gov / pubmed / 28281211) Summary of the Invention

[0012] Technical issues to be solved by the invention

[0013] In recent years, target proteins with high demand have become more complex and difficult to identify, making it difficult to find a hit simply by screening a library. On the other hand, the theoretical number of compounds, even if limited to low molecules with a molecular weight of less than 500, is (10 to the power of 60). When expanded to medium molecules with a molecular weight of around 1,000, the number increases further. If we consider that the number of compounds synthesized throughout history is around (10 to the power of 9), it is still possible to find a hit. However, investigating the bonding strength of such an astronomical number of compounds as a whole is almost impossible, not only in experiments but also in simulations. Even when investigating the bonding strength of a portion of the compounds, the efficiency is low when repeated experiments are repeated as described in Patent Documents 1 and 2. Moreover, in the case of existing descriptors (feature quantities) such as fingerprints described in Patent Document 3, even for compounds that show the same pharmacological effect, their feature quantities are not necessarily similar, and the feature quantities do not accurately represent the chemical properties of the target structure, so searches using feature quantities are inefficient.

[0014] Furthermore, the inverse quantitative structure-property relationship (iqspr) method described in the aforementioned non-patent document 1 suffers from a problem of immediate loss of search efficiency due to its structure update algorithm (a particle filter based on Bayesian estimation). Specifically, while the Bayesian-based structure update approaches the target value for physical properties, the diversity of search structures decreases, causing the search to become trapped in a local minimum. Even with repeated trials, the local minimum cannot be easily escaped (thus preventing the final structure from being reached).

[0015] As described above, in the conventional technology, the feature values ​​do not accurately represent the chemical properties of the target structure, and thus the efficiency of screening and three-dimensional structure creation using the feature values ​​is low.

[0016] The present invention has been developed in light of this situation, and its purpose is to provide a method for calculating feature quantities that can accurately represent the chemical properties of a target structure. Furthermore, the present invention aims to provide a screening method that can effectively screen candidate pharmaceutical compounds using the feature quantities. Furthermore, the present invention aims to provide a compound creation method that can effectively create the three-dimensional structure of a candidate pharmaceutical compound using the feature quantities.

[0017] Means for solving technical problems

[0018] In order to achieve the above-mentioned purpose, the feature quantity calculation method involved in the first embodiment of the present invention has: an object structure specifying step, specifying an object structure composed of multiple unit structures with chemical properties; a three-dimensional structure acquisition step, acquiring a three-dimensional structure based on multiple unit structures for the object structure; and a probe feature quantity calculation step, calculating a feature quantity representing the cross-sectional area of ​​one or more probes relative to the object structure, where the probe is a structure having a real charge and having multiple points separated by the van der Waals force.

[0019] The chemical properties of the target structure are represented by the interaction between the target structure and one or more surrounding probes. Therefore, similarity in the characteristic quantities representing the cross-sectional area between target structures indicates similar chemical properties of these target structures. In other words, target structures with similar characteristic quantities calculated using the first method exhibit similar chemical properties. Therefore, using the first method, it is possible to calculate characteristic quantities that accurately represent the chemical properties of the target structure. In addition, in the first method and the following methods, "cross-sectional area" includes scattering cross-sectional area (differential scattering cross-sectional area), reaction cross-sectional area, and absorption cross-sectional area.

[0020] The feature quantity calculation method according to the second aspect is that, in the probe feature quantity calculation step in the first aspect, the cross-sectional area, the closest distance, and the scattering angle are calculated as the feature quantity. The second aspect specifically defines the "feature quantity indicating the cross-sectional area" in the first aspect.

[0021] The feature quantity calculation method according to the third aspect is characterized in that, in the probe feature quantity calculation step, a feature quantity depending on the type, number, combination, collision diameter, and incident energy of the probes is calculated as the feature quantity.

[0022] The feature quantity calculation method according to a fourth aspect is any one of the first to third aspects, wherein the three-dimensional structure acquisition step acquires the three-dimensional structure of the designated target structure by generating the three-dimensional structure.

[0023] The feature quantity calculation method according to the fifth embodiment, in any of the first to fourth embodiments, comprises the following steps: in the target structure specifying step, a compound is specified as the target structure; in the stereostructure acquisition step, the stereostructure of the compound based on multiple atoms as multiple unit structures is acquired; and in the probe feature quantity calculation step, the first feature quantity is calculated using amino acids as probes for the compound acquired in the stereostructure acquisition step. In the fifth embodiment, the "probe" in the first embodiment is an amino acid, the "target structure" in the first embodiment is a compound, and the "multiple unit structures" in the first embodiment are multiple atoms. Furthermore, the amino acid for quantifying the degree of aggregation is not limited to a single type; a peptide containing two or more amino acids bonded together may be used.

[0024] The feature value calculation method according to the sixth aspect, in the fifth aspect, further includes an invariance step of calculating a first invariant feature value by making the first feature value rotationally invariant with respect to the compound. In the sixth aspect, making the first feature value rotationally invariant with respect to the compound facilitates feature processing and reduces data capacity. The invariance of the first feature value can be performed by angular averaging of the potential.

[0025] The feature quantity calculation method according to the seventh aspect is configured such that, in the sixth aspect, the first feature quantity is calculated for two different amino acids in the probe feature quantity calculation step, and the first feature quantity for the two different amino acids is used in the invariance step to calculate the first invariant feature quantity. According to the seventh aspect, in the calculation of the first invariant feature quantity, the first feature quantity for the two different amino acids is used to perform invariance while maintaining information on the interaction between the amino acids. This allows accurate comparison (drug efficacy determination) of compounds based on the feature quantity (the first invariant feature quantity).

[0026] The feature quantity calculation method according to the eighth embodiment, in any of the first to fourth embodiments, comprises: in the target structure designation step, designating a pocket structure that bonds to the pocket serving as the active site of the target protein as the target structure; in the three-dimensional structure acquisition step, acquiring the three-dimensional structure of the pocket structure based on a plurality of virtual spheres; and in the probe feature quantity calculation step, calculating a second feature quantity for the pocket structure acquired in the three-dimensional structure acquisition step using amino acids as probes. In the eighth embodiment, the "probe" in the first embodiment is an amino acid, the "target structure" in the first embodiment is a pocket structure, and the "unit structure" in the first embodiment is a plurality of virtual spheres. The "active site" of the target protein refers to a site where the activity of the target protein is promoted or inhibited by bonding to the pocket structure, and the "virtual sphere" can be considered to have chemical properties such as a van der Waals radius and charge.

[0027] In the fifth embodiment, the degree of amino acid aggregation relative to a provided compound is calculated. In contrast, in the eighth embodiment, a characteristic quantity (second characteristic quantity) representing the cross-sectional area of ​​amino acids relative to a pocket structure bound to a pocket of a provided target protein is calculated. Pocket structures with similar characteristic quantities according to the eighth embodiment exhibit similar chemical properties, and therefore, the eighth embodiment allows calculation of a characteristic quantity that accurately represents the chemical properties of the pocket structure. Furthermore, the pocket structure corresponds to a compound bound to a pocket of a target protein. Furthermore, in the eighth embodiment, simulations based on actual measurement results of the target protein's three-dimensional structure, pocket positional information, and the like can be used to calculate the second characteristic quantity. Furthermore, the three-dimensional structure of the target protein can be determined by any measurement technique (such as X-ray crystallography, NMR (NMR: Nuclear Magnetic Resonance), or cryo-TEM (TEM: Transmission Electron Microscopy)) as long as the resolution of each amino acid residue can be recognized.

[0028] The feature quantity calculation method according to the ninth aspect, in the eighth aspect, further includes an invariant step for calculating a second invariant feature quantity by making the second feature quantity invariant with respect to rotation of the pocket structure. According to the ninth aspect, similar to the sixth aspect, feature quantity processing can be facilitated and data capacity can be reduced. Similar to the sixth aspect, invariant quantization of the second feature quantity can be performed by angular averaging of the potential.

[0029] The feature quantity calculation method according to the tenth embodiment is configured such that, in the ninth embodiment, the second feature quantity is calculated for two different amino acids in the probe feature quantity calculation step, and the second feature quantity for the two different amino acids is used in the invariance step to calculate the second invariant feature quantity. According to the tenth embodiment, in the calculation of the second invariant feature quantity, invariance can be performed while maintaining information on the interaction between the amino acids by using the second feature quantity for the two different amino acids. This enables accurate comparison (drug efficacy determination) of compounds based on the feature quantity (the second invariant feature quantity).

[0030] The characteristic quantity calculation method involved in the 11th embodiment is any one of the first to fourth embodiments, wherein in the target structure specifying step, a compound is specified as the target structure, in the three-dimensional structure acquisition step, a three-dimensional structure of the compound based on multiple atoms is generated, and in the probe characteristic quantity calculation step, for the three-dimensional structure of the compound acquired in the three-dimensional structure acquisition step, one or more nucleic acid bases, one or more lipid molecules, one or more monosaccharide molecules, water, and one or more ions are used as probes to calculate a third characteristic quantity. In the 11th embodiment, the "probe" in the first embodiment is set to one or more nucleic acid bases, etc. (which can be any type, number, or combination), the "target structure" in the first embodiment is set to the compound, and the "multiple unit structures" in the first embodiment are set to multiple atoms. In addition, the ion can be a monatomic ion or an ion composed of multiple atoms.

[0031] The feature quantity calculation method according to the twelfth aspect, in the eleventh aspect, further includes an invariance step, wherein the third feature quantity is invariant to rotation of the compound to calculate a third invariant feature quantity. According to the twelfth aspect, similar to the sixth and ninth aspects, feature quantity processing can be facilitated and data capacity can be reduced. Regarding the invariance of the third feature quantity, similar to the sixth and ninth aspects, the invariance can be performed by angular averaging of the potential.

[0032] To achieve the above-mentioned object, a screening method according to a thirteenth embodiment of the present invention extracts a first target compound that binds to a target protein and / or a second target compound that does not bind to the target protein from a plurality of compounds. The screening method comprises: a storage step of associating, for each of the plurality of compounds, a first feature value for the compound's structure, calculated using the feature value calculation method according to the fifth embodiment, with a first feature value for the compound's structure, calculated using the feature value calculation method according to the fifth embodiment; a screening feature value calculation step of calculating a first feature value for a ligand, which is a compound confirmed to be bound to the target protein; a similarity calculation step of calculating the similarity between the first feature values ​​for the plurality of compounds and the first feature value for the ligand; and a compound extraction step of extracting the first target compound and / or the second target compound from the plurality of compounds based on the similarity. As described above in the fifth embodiment, if the first feature values ​​of the ligand and the target compound are similar, then the pharmacological effects of the two are similar. Therefore, according to the thirteenth embodiment, target compounds (the first target compound and / or the second target compound) having pharmacological effects similar to those of the ligand are extracted based on the first feature value, thereby enabling efficient screening of candidate pharmaceutical compounds.

[0033] In order to achieve the above-mentioned purpose, the screening method involved in the 14th embodiment of the present invention extracts the first target compound bonded to the target protein and / or the second target compound not bonded to the target protein from multiple compounds, and the screening method comprises: a storage process, for each of the multiple compounds, establishing an association between the stereoscopic structure of the compound based on multiple atoms and the first invariant feature quantity for the stereoscopic structure of the compound calculated using the feature quantity calculation method involved in the 6th embodiment and storing the association; a screening feature quantity calculation process, calculating the first invariant feature quantity for the ligand of the compound confirmed to be bonded to the target protein; a similarity calculation process, calculating the similarity between the first invariant feature quantity for the multiple compounds and the first invariant feature quantity for the ligand; and a compound extraction process, extracting the first target compound and / or the second target compound from the multiple compounds based on the similarity. In terms of calculating the characteristic quantity for the ligand, the 14th method is similar to the 13th method. In the 14th method, target compounds (the first target compound and / or the second target compound) with similar efficacy to the ligand are extracted based on the similarity of the first invariant characteristic quantity, which can effectively screen candidate pharmaceutical compounds.

[0034] In order to achieve the above-mentioned purpose, the screening method involved in the 15th embodiment of the present invention extracts the first target compound bonded to the target protein and / or the second target compound not bonded to the target protein from multiple compounds, and the screening method comprises: a storage process, for each of the multiple compounds, establishing an association between the three-dimensional structure of the compound based on multiple atoms and the first feature quantity calculated using the feature quantity calculation method involved in the 5th embodiment and storing the association; a screening feature quantity calculation process, using the feature quantity calculation method involved in the 8th embodiment to calculate the second feature quantity for the pocket structure of the target protein; a similarity calculation process, calculating the similarity between the first feature quantity for multiple compounds and the second feature quantity for the pocket structure; and a compound extraction process, extracting the first target compound and / or the second target compound from the multiple compounds based on the similarity.

[0035] As described above in the eighth embodiment, if the second characteristic value of the pocket structure and the target compound are similar, then the chemical properties of the two are similar. Therefore, according to the fifteenth embodiment, target compounds (first target compound and / or second target compound) with similar chemical properties to the pocket structure are extracted, and the screening of drug candidate compounds can be effectively performed. In addition, the pocket structure corresponds to the compound bound to the target protein, so the characteristic value for the pocket structure (second characteristic value) and the characteristic value for the compound (first characteristic value) can be compared and the similarity can be calculated.

[0036] To achieve the above-mentioned object, a screening method according to a sixteenth aspect of the present invention extracts a first target compound that binds to a target protein and / or a second target compound that does not bind to the target protein from a plurality of compounds. The screening method comprises: a storage step of associating and storing, for each of the plurality of compounds, a three-dimensional structure of the compound based on multiple atoms with a first invariant feature calculated using the feature calculation method according to the sixth aspect; a screening feature calculation step of calculating a second invariant feature for a pocket structure of the target protein using the feature calculation method according to the ninth aspect; a similarity calculation step of calculating the similarity between the first invariant feature for the plurality of compounds and the second invariant feature for the pocket structure; and a compound extraction step of extracting the first target compound and / or the second target compound from the plurality of compounds based on the similarity. In the sixteenth aspect, target compounds (the first target compound and / or the second target compound) having chemical properties similar to those of the pocket structure are extracted using the first and second invariant feature values, thereby enabling efficient screening of candidate pharmaceutical compounds. Furthermore, similarly to the above-described content of the fifteenth aspect, the feature value for the pocket structure (the second invariant feature value) and the feature value for the compound (the first invariant feature value) can be compared to calculate the similarity.

[0037] To achieve the above-mentioned object, a screening method according to a seventeenth aspect of the present invention extracts target compounds that are bound to a target biopolymer other than a protein from a plurality of compounds. The screening method comprises: a storage step of associating, for each of the plurality of compounds, a three-dimensional structure of the compound based on multiple atoms with a third feature quantity for the compound's three-dimensional structure calculated using the feature quantity calculation method according to the eleventh aspect, and storing the associated structure; a feature quantity calculation step of calculating the third feature quantity for a bonded compound confirmed to be bound to a target biopolymer other than a protein; a similarity calculation step of calculating the similarity between the third feature quantity for the plurality of compounds and the third feature quantity for the bonded compound; and a compound extraction step of extracting target compounds from the plurality of compounds based on the similarity. As described above in the eleventh aspect, the present invention can use DNA, etc., as a target biopolymer other than a protein. If the third feature quantity of the bonded compound bound to the target biopolymer and the target compound are similar, then the drug efficacy of the two compounds is similar. Therefore, according to the seventeenth aspect, target compounds having drug efficacy similar to that of the bonded compound are extracted based on the third feature quantity, enabling efficient screening of candidate pharmaceutical compounds.

[0038] Furthermore, in the 13th through 17th embodiments, the compound extraction step can extract compounds with a similarity above a threshold. This threshold can be set based on the screening objective, accuracy, and other conditions, or it can be set to a user-specified value. Furthermore, in the 13th through 17th embodiments, the compound extraction step can extract compounds in descending order of similarity. This extraction method allows for efficient screening of candidate pharmaceutical compounds.

[0039] In order to achieve the above-mentioned purpose, the compound creation method involved in the 18th embodiment of the present invention creates a three-dimensional structure of a target compound bonded to a target protein from multiple compounds, and the compound creation method comprises: a storage process, for each of the multiple compounds, establishing an association between the three-dimensional structure of the compound based on multiple atoms and the first feature quantity calculated using the feature quantity calculation method involved in the 5th embodiment and storing the association; a feature quantity calculation process, calculating the first feature quantity for the ligand of the compound confirmed to be bonded to the target protein; a generator construction process, constructing a generator by machine learning with the three-dimensional structures of multiple compounds set as training data and the first feature quantity set as an explanatory variable; and a compound three-dimensional structure generation process, using the generator to generate the three-dimensional structure of the target compound from the first feature quantity of the ligand.

[0040] In the screening methods involved in the above-mentioned 13th to 17th methods, compounds that match the ligand or target protein are found from multiple compounds whose structural formulas have been determined (written down). Therefore, after calculating the characteristic quantity of the compound, a method of extracting the compound based on the similarity with the characteristic quantity of the pocket structure of the ligand or target protein calculated separately is adopted, that is, a retrieval method. Therefore, as long as the correspondence between the structural formula of the compound and the characteristic quantity is recorded in advance, a structural formula with a high similarity (or above a threshold) can be found. In contrast, in the 18th method, a structural formula of a compound having a characteristic quantity similar to the characteristic quantity (first characteristic quantity) of the ligand (therefore, similar in efficacy) is generated without searching.

[0041] Regarding the generation of a structural formula when a feature value is provided, a generator constructed by machine learning can be used. Specifically, in the 18th embodiment, a generator is constructed by machine learning (the learning method is not particularly limited) using the three-dimensional structure of the compound as training data and the first feature value as an explanatory variable, and the three-dimensional structure of the target compound is generated from the first feature value of the ligand using this generator. In the 18th embodiment, since no search is performed, the three-dimensional structure of the compound can be generated even in the case where the result of the screening-based search is "no solution", thereby effectively creating the three-dimensional structure of the pharmaceutical candidate compound.

[0042] Furthermore, the stereostructures generated in the eighteenth embodiment are influenced by the characteristics of the compounds provided as training data. Therefore, by selecting the characteristics of the compounds provided as training data, compounds with different stereostructures can be generated. For example, by providing easily synthesized compounds as training data, compounds with easily synthesized stereostructures can be generated.

[0043] To achieve the above-mentioned object, a compound creation method according to a 19th aspect of the present invention creates a stereostructure of a target compound bound to a target protein from a plurality of compounds. The compound creation method comprises: a storage step of associating and storing, for each of the plurality of compounds, the stereostructure of the compound based on multiple atoms with a first invariant feature calculated using the feature calculation method according to the 6th aspect; a creation feature calculation step of calculating the first invariant feature for a ligand, which is a compound confirmed to be bound to the target protein; a generator construction step of constructing a generator using machine learning using the stereostructures of the plurality of compounds as training data and the first invariant feature as an explanatory variable; and a compound stereostructure generation step of using the generator to generate the stereostructure of the target compound from the first invariant feature of the ligand. In the 19th aspect, similar to the 18th aspect, a structural formula of a compound having a feature similar to the feature (first invariant feature) of the ligand (thus, having similar pharmacological efficacy) is generated without performing a search, thereby enabling efficient creation of the stereostructure of a candidate pharmaceutical compound. Furthermore, similarly to the eighteenth aspect, by selecting the features of the compound provided as training data, compounds having three-dimensional structures with different features can be generated.

[0044] To achieve the above-mentioned object, a compound creation method according to a 20th aspect of the present invention creates a three-dimensional structure of a target compound bound to a target protein from a plurality of compounds. The compound creation method comprises: a storage step of associating and storing, for each of the plurality of compounds, the three-dimensional structure of the compound based on multiple atoms with a first feature value calculated using the feature value calculation method according to the fifth aspect; a feature value creation step of calculating a second feature value for a pocket structure of the target protein using the feature value calculation method according to the eighth aspect; a generator construction step of constructing a generator using machine learning using the three-dimensional structures of the plurality of compounds as training data and the first feature value as an explanatory variable; and a compound three-dimensional structure generation step of using the generator to generate the three-dimensional structure of the target compound from the second feature value of the pocket structure. According to the 20th aspect, similar to the 18th and 19th aspects, a structural formula of a compound having a feature value similar to the feature value (second feature value) of the pocket structure (thus, having similar pharmacological efficacy) is generated without performing a search, thereby enabling efficient creation of the three-dimensional structure of a candidate pharmaceutical compound. Furthermore, similarly to the eighteenth and nineteenth aspects, by selecting the features of the compound provided as training data, compounds having three-dimensional structures with different features can be generated.

[0045] To achieve the above-mentioned object, a compound creation method according to a 21st aspect of the present invention creates a three-dimensional structure of a target compound bound to a target protein from a plurality of compounds. The compound creation method comprises: a storage step of associating and storing, for each of the plurality of compounds, the three-dimensional structure of the compound based on multiple atoms with a first invariant feature calculated using the feature calculation method according to the 6th aspect; a feature calculation step of calculating a second invariant feature for a pocket structure of the target protein using the feature calculation method according to the 9th aspect; a generator construction step of constructing a generator using machine learning using the three-dimensional structures of the plurality of compounds as training data and the first invariant feature as an explanatory variable; and a compound three-dimensional structure generation step of using the generator to generate the three-dimensional structure of the target compound from the second invariant feature of the pocket structure. According to the 21st aspect, similar to the 18th to 20th aspects, a structural formula of a compound having a feature similar to the feature (second invariant feature) of the pocket structure (thus, having similar pharmacological efficacy) is generated without performing a search, thereby enabling efficient creation of the three-dimensional structure of a candidate pharmaceutical compound. Furthermore, similarly to the eighteenth to twentieth aspects, by selecting the features of the compound provided as training data, compounds having three-dimensional structures with different features can be generated.

[0046] To achieve the above-mentioned object, a compound creation method according to a 22nd aspect of the present invention creates a three-dimensional structure of a target compound bonded to a target biopolymer other than a protein from a plurality of compounds. The compound creation method comprises: a storage step of associating and storing, for each of the plurality of compounds, the three-dimensional structure of the compound based on multiple atoms with a third feature value calculated using the feature value calculation method according to the 11th aspect; a creation feature value calculation step of calculating the third feature value for a bonded compound that is a compound confirmed to be bonded to a target biopolymer other than a protein; a generator construction step of constructing a generator using machine learning using the three-dimensional structures of the plurality of compounds as training data and the third feature value as an explanatory variable; and a compound three-dimensional structure generation step of using the generator to generate the three-dimensional structure of the target compound from the third feature value of the bonded compound. According to the 22nd aspect, similar to the 18th to 21st aspects, a structural formula of a compound having a feature value similar to the feature value (third feature value) of the bonded compound (thus, having similar pharmacological efficacy) is generated without performing a search, thereby enabling efficient creation of the three-dimensional structure of a candidate pharmaceutical compound. Furthermore, similarly to the eighteenth to twenty-first aspects, by selecting the features of the compound provided as training data, compounds having three-dimensional structures with different features can be generated.

[0047] In order to achieve the above-mentioned purpose, the compound creation method involved in the 23rd embodiment of the present invention creates a three-dimensional structure of a target compound bonded to a target protein, and the compound creation method comprises: an input process for inputting the chemical structure of one or more compounds, a first feature quantity for the chemical structure calculated using the feature quantity calculation method involved in the 5th embodiment, and a first feature quantity for a ligand of the compound confirmed to be bonded to the target compound as a target value of the first feature quantity; a candidate structure acquisition process for changing the chemical structure and acquiring a candidate structure; a feature quantity calculation process for creating a candidate structure, for calculating the first feature quantity using the feature quantity calculation method involved in the 5th embodiment; a candidate structure adoption process for adopting or abandoning the candidate structure, and the like. In the process, a first adoption process is performed to determine whether to adopt the candidate structure based on whether the first feature value of the candidate structure is close to the target value due to the change in the chemical structure. If the candidate structure is not adopted through the first adoption process, a second adoption process is performed to determine whether to adopt the candidate structure based on whether the structural diversity of the structure group composed of the chemical structure and the candidate structure is increased due to the change in the chemical structure. If the candidate structure is not adopted through the first adoption process and the second adoption process, an abandonment process is performed to abandon the change in the chemical structure and return to the chemical structure before the change; and a control process is performed until the end condition is met, repeating the processes in the input process, the candidate structure acquisition process, the feature value calculation process and the candidate structure adoption process.

[0048] In the compound creation method according to the 23rd aspect, by promoting escape from local minima based on structural diversity, it is possible to efficiently search for structures of compounds having desired physical property values ​​(target values ​​of the first feature quantity). Furthermore, in the 23rd aspect, since no search is performed, as in the 18th aspect, even when the result of the screening-based search is "unsolvable," the three-dimensional structure of a compound (a compound having a feature quantity similar to that of the ligand (the first feature quantity), and therefore having similar pharmacological efficacy) can be generated, thereby enabling efficient creation of the three-dimensional structure of a candidate pharmaceutical compound.

[0049] In order to achieve the above-mentioned purpose, the compound creation method involved in the 24th embodiment of the present invention creates a three-dimensional structure of a target compound bonded to a target protein, and the compound creation method comprises: an input process for inputting the chemical structure of one or more compounds, a first invariant feature quantity for the chemical structure calculated using the feature quantity calculation method involved in the 6th embodiment, and a first invariant feature quantity for a ligand of the compound confirmed to be bonded to the target compound as a target value of the first invariant feature quantity; a candidate structure acquisition process for changing the chemical structure and acquiring a candidate structure; a feature quantity calculation process for creating a candidate structure, for calculating the first invariant feature quantity using the feature quantity calculation method involved in the 6th embodiment; a candidate structure adoption process for adopting or abandoning the candidate structure. A candidate structure, wherein a first adoption process is performed to determine whether to adopt the candidate structure based on whether a first invariant feature value of the candidate structure approaches a target value due to a change in the chemical structure; if the candidate structure is not adopted by the first adoption process, a second adoption process is performed to determine whether to adopt the candidate structure based on whether the structural diversity of a structure group consisting of the chemical structure and the candidate structure increases due to the change in the chemical structure; if the candidate structure is not adopted by the first adoption process and the second adoption process, a abandonment process is performed to abandon the change in the chemical structure and return to the chemical structure before the change; and a control step is performed to repeat the processes of the input step, the candidate structure acquisition step, the creation feature value calculation step, and the candidate structure adoption step until an end condition is satisfied. In the 24th embodiment, as in the 23rd embodiment, the structure of a compound having a desired physical property value (the target value of the first invariant feature value) can be efficiently searched, and the three-dimensional structure of the pharmaceutical candidate compound can be efficiently established.

[0050] In order to achieve the above-mentioned purpose, the compound creation method involved in the 25th embodiment of the present invention creates a three-dimensional structure of a target compound bonded to a target protein, and the compound creation method comprises: an input process, in which the chemical structure of one or more compounds is input, a second feature quantity for the chemical structure calculated using the feature quantity calculation method involved in the 8th embodiment, and a second feature quantity for the pocket structure confirmed to be bonded to the pocket of the active site of the target protein as the target value of the second feature quantity; a candidate structure acquisition process, in which the chemical structure is changed and a candidate structure is acquired; a feature quantity calculation process is created, in which the second feature quantity is calculated for the candidate structure using the feature quantity calculation method involved in the 8th embodiment; a candidate structure adoption process, in which the candidate is adopted or abandoned. A structure wherein a first adoption process is performed to determine whether to adopt the candidate structure based on whether the second feature value of the candidate structure approaches a target value due to a change in the chemical structure; if the candidate structure is not adopted by the first adoption process, a second adoption process is performed to determine whether to adopt the candidate structure based on whether the structural diversity of the structure group consisting of the chemical structure and the candidate structure increases due to the change in the chemical structure; if the candidate structure is not adopted by the first adoption process and the second adoption process, a abandonment process is performed to abandon the change in the chemical structure and return to the chemical structure before the change; and a control step is performed to repeat the processes of the input step, the candidate structure acquisition step, the creation feature value calculation step, and the candidate structure adoption step until an end condition is satisfied. In the 25th embodiment, similar to the 23rd and 24th embodiments, the structure of a compound having a desired physical property value (target value of the second feature value) can be efficiently searched, and the three-dimensional structure of the pharmaceutical candidate compound can be efficiently established.

[0051] In order to achieve the above-mentioned purpose, the compound creation method involved in the 26th embodiment of the present invention creates a three-dimensional structure of a target compound bonded to a target protein, and the compound creation method comprises: an input process, in which the chemical structure of one or more compounds is input, a second invariant feature quantity for the chemical structure calculated using the feature quantity calculation method involved in the 9th embodiment, and a second invariant feature quantity for a pocket structure confirmed to be bonded to a pocket serving as an active site of the target protein as a target value of the second invariant feature quantity; a candidate structure acquisition process, in which the chemical structure is changed and a candidate structure is acquired; a feature quantity calculation process is created, in which the second invariant feature quantity is calculated for the candidate structure using the feature quantity calculation method involved in the 9th embodiment; and a candidate structure adoption process is adopted. A method for selecting or rejecting a candidate structure comprises: performing a first adoption process to determine whether to adopt the candidate structure based on whether the second invariant feature value of the candidate structure approaches a target value due to a change in the chemical structure; performing a second adoption process to determine whether to adopt the candidate structure based on whether the structural diversity of the structure group consisting of the chemical structure and the candidate structure increases due to the change in the chemical structure; and performing a rejection process to abandon the change in the chemical structure and return to the chemical structure before the change if the candidate structure is not adopted by the first and second adoption processes; and repeating the input process, the candidate structure acquisition process, the creation feature value calculation process, and the candidate structure adoption process until an end condition is satisfied. In the 26th embodiment, as in the 23rd to 25th embodiments, it is possible to efficiently search for the structure of a compound having a desired physical property value (target value of the second invariant feature value) and to efficiently establish the three-dimensional structure of a pharmaceutical candidate compound.

[0052] In order to achieve the above-mentioned purpose, the compound creation method involved in the 27th embodiment of the present invention creates a three-dimensional structure of a target compound bonded to a target biopolymer other than a protein, and the compound creation method comprises: an input process for inputting the chemical structure of one or more compounds, a third feature quantity for the chemical structure calculated using the feature quantity calculation method involved in the 11th embodiment, and a third feature quantity for the bonded compound as a compound confirmed to be bonded to a target biopolymer other than a protein as a target value of the third feature quantity; a candidate structure acquisition process for changing the chemical structure and acquiring a candidate structure; a feature quantity calculation process for creating a candidate structure, for calculating the third feature quantity using the feature quantity calculation method involved in the 11th embodiment; a candidate structure adoption process , adopting or rejecting a candidate structure, wherein a first adoption process is performed to determine whether to adopt the candidate structure based on whether the third feature value of the candidate structure approaches a target value due to a change in the chemical structure; if the candidate structure is not adopted by the first adoption process, a second adoption process is performed to determine whether to adopt the candidate structure based on whether the structural diversity of the structure group consisting of the chemical structure and the candidate structure increases due to the change in the chemical structure; if the candidate structure is not adopted by the first adoption process and the second adoption process, a rejection process is performed to abandon the change in the chemical structure and return to the chemical structure before the change; and a control process is repeated until an end condition is satisfied, and the processes in the input process, the candidate structure acquisition process, the creation feature value calculation process, and the candidate structure adoption process are repeated. In the 27th embodiment, as in the 23rd to 26th embodiments, the structure of a compound having a desired physical property value (target value of the third feature value) can be effectively searched, and the three-dimensional structure of the pharmaceutical candidate compound can be effectively established.

[0053] In the 23rd to 27th aspects, the "chemical structure" includes not only the structure in the initial state (initial structure) but also a structure in which the initial structure changes due to repetition of treatment.

[0054] Effects of the Invention

[0055] As described above, the feature quantity calculation method of the present invention can calculate feature quantities that accurately represent the chemical properties of a target structure. Furthermore, the screening method of the present invention can effectively screen for candidate pharmaceutical compounds. Furthermore, the compound creation method of the present invention can effectively create the three-dimensional structure of a candidate pharmaceutical compound. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 This is a block diagram showing the configuration of the screening device according to the first embodiment.

[0057] Figure 2 It is a block diagram showing the structure of the processing unit.

[0058] Figure 3 It is a diagram showing information stored in the storage unit.

[0059] Figure 4 This is a diagram showing a state in which structural information of a compound and feature quantities are associated and stored.

[0060] Figure 5 This is a diagram showing the state of obtaining the differential scattering cross-sectional area.

[0061] Figure 6 is a flow chart showing the steps for calculating AAS descriptors for a compound.

[0062] Figure 7 This is a diagram showing an example of three-dimensionalization of a structural formula.

[0063] Figure 8 It is a diagram showing an example of a differential scattering cross-sectional area.

[0064] Figure 9 This is another diagram showing an example of a differential scattering cross-sectional area.

[0065] Figure 10 This is a flowchart showing the steps for calculating the AAS descriptor for a pocket structure.

[0066] Figure 11 This is a conceptual diagram showing the relationship between the target protein and the pocket structure.

[0067] Figure 12 This is a diagram showing an example of the ease of finding hits for various invariant quantized AAS descriptors.

[0068] Figure 13 This is a graph showing the ease of finding hits for compounds that bind to the target protein and compounds that do not bind to the target protein.

[0069] Figure 14 is a flow chart showing the steps of compound extraction based on ligand AAS descriptors.

[0070] Figure 15 This is a table showing examples of screening results of ligand input.

[0071] Figure 16 This is a flowchart showing the steps of filtering using AAS descriptors for pocket structures.

[0072] Figure 17 This is a table showing examples of screening results of target protein input.

[0073] Figure 18 This is a block diagram showing the structure of the compound creation device according to the second embodiment.

[0074] Figure 19 It is a diagram showing the structure of the processing unit.

[0075] Figure 20 It is a diagram showing information stored in the storage unit.

[0076] Figure 21 This is a flowchart showing the steps of generating a three-dimensional structure when a ligand is input.

[0077] Figure 22 This is a conceptual diagram showing the state of a generator built using machine learning.

[0078] Figure 23 This is a diagram showing an example of generating a three-dimensional structure using a generator.

[0079] Figure 24 This is a flowchart showing the steps of generating a three-dimensional structure when a target protein is input.

[0080] Figure 25 This is a diagram showing the structure of a compound creation device when creating the three-dimensional structure of a compound based on structural diversity.

[0081] Figure 26 This is a flowchart showing the steps of a three-dimensional structure generation process based on structural diversity.

[0082] Figure 27 This is a block diagram showing the structure of a pharmaceutical candidate compound search device according to a third embodiment.

[0083] Figure 28 It is a diagram showing the structure of the processing unit.

[0084] Figure 29 It is a diagram showing information stored in the storage unit. DETAILED DESCRIPTION

[0085] Hereinafter, embodiments of the feature quantity calculation method, screening method, and compound creation method of the present invention will be described in detail with reference to the accompanying drawings.

[0086] <First embodiment>

[0087] Figure 1 This is a block diagram showing the structure of the screening device 10 (feature quantity calculation device, screening device) involved in the first embodiment. The screening device 10 is a device that calculates the feature quantity of the compound (target structure) and / or pocket structure (target structure) and extracts (screens) the target compound, and can be implemented using a computer. Figure 1As shown, the screening device 10 includes a processing unit 100 (processor), a storage unit 200, a display unit 300 and an operating unit 400, and is interconnected to send and receive required information. Various configuration methods can be adopted for these components. Each component can be set in one place (inside a frame, in a room, etc.), or it can be set in a separate location and connected via a network. In addition, the screening device 10 is connected to an external server 500 and an external database 510 such as PDB (Protein Data Bank: large database) via a network NW such as the Internet, and can obtain information such as the structural formula of the compound and the crystal structure of the protein as needed.

[0088] <Structure of the processing unit>

[0089] Figure 2 This is a block diagram showing the configuration of a processing unit 100 (processor). The processing unit 100 includes an information input unit 110, a feature value calculation unit 120, a similarity calculation unit 130, a compound extraction unit 140, a display control unit 150, a CPU 160 (CPU: Central Processing Unit), a ROM 170 (ROM: Read Only Memory), and a RAM 180 (RAM: Random Access Memory).

[0090] The information input unit 110 inputs information such as the structural formula of the compound, the X crystal structure of the target protein, and the position of the pocket via a recording medium such as a magneto-optical disk or semiconductor memory (not shown) and / or a network NW. The feature quantity calculation unit 120 (target structure designation unit, three-dimensional structure generation unit, feature quantity calculation unit) calculates the feature quantities involved in the present invention (first feature quantity, first invariant feature quantity, second feature quantity, second invariant feature quantity, third feature quantity, and third invariant feature quantity). The similarity calculation unit 130 (similarity calculation unit) calculates the similarity between the calculated feature quantities. The compound extraction unit 140 (compound extraction unit) extracts the target compound from multiple compounds based on the similarity. The display control unit 150 controls the display of the input information and processing results on the monitor 310. The details of the feature quantity calculation and target compound screening process using these functions of the processing unit 100 will be described later. In addition, the processing based on these functions is performed under the control of the CPU 160.

[0091] The functions of each part of the processing unit 100 can be implemented using various processors. Among the various processors, for example, there is a CPU, which is a general-purpose processor that executes software (programs) to implement various functions. In addition, among the various processors mentioned above, there are also programmable logic devices (PLDs) such as FPGAs (Field Programmable Gate Arrays) that are processors whose circuit structure can be changed after manufacturing. In addition, dedicated circuits such as ASICs (Application Specific Integrated Circuits) that are processors with circuit structures specifically designed to perform specific processing are also included in the various processors mentioned above.

[0092] The functions of each part can be realized by one processor or by combining multiple processors. In addition, multiple functions can be realized by one processor. As an example of how multiple functions are constituted by one processor, there is the following method: as represented by computers such as clients and servers, one processor is constituted by a combination of one or more CPUs and software, and multiple functions are realized by this processor. The second method is as represented by a system on chip (SoC) or the like, where a processor that realizes the functions of the entire system is used by one IC (Integrated Circuit) chip. In this way, various functions are constituted by using one or more of the above-mentioned various processors as a hardware structure. Moreover, more specifically, the hardware structure of these various processors is a circuit (circuitry) that combines circuit elements such as semiconductor elements.

[0093] When the processor or circuit executes the software (program), the processor (computer) readable code of the software to be executed is stored in the ROM 170 (see Figure 2) or other non-temporary recording media, so that the processor refers to the software. The software pre-stored in the non-temporary recording medium includes a program (feature quantity calculation program, screening program, and compound creation program) for executing the feature quantity calculation method, screening method, and compound creation method involved in the present invention. The code can also be recorded in non-temporary recording media such as various optical magnetic recording devices and semiconductor memories instead of in ROM170. When performing processing using software, for example, RAM180 is used as a temporary storage area, and it is also possible to refer to data stored in, for example, an EEPROM (Electronically Erasable and Programmable Read Only Memory) not shown in the figure.

[0094] <Storage Unit Structure>

[0095] The storage unit 200 is composed of non-temporary recording media such as DVD (Digital Versatile Disk), hard disk, various semiconductor memories and their control units, and stores Figure 3 The image and information shown. The structural information 210 includes the structural formula of the compound, the three-dimensional structure of the target protein and the pocket position. The three-dimensional structure information 220 is information on the three-dimensional structure of the compound and / or pocket structure, which can be generated by the structural information 210, or the three-dimensional information can be input. The AAS descriptor 230 (the first feature quantity, the second feature quantity, the third feature quantity) is a feature quantity representing the cross-sectional area of ​​one or more probes relative to the target structure such as the compound and the pocket structure, and is calculated by the feature quantity calculation method described later. The invariant AAS descriptor 240 (the first invariant feature quantity, the second invariant feature quantity, the third invariant feature quantity) is a feature quantity that makes the AAS descriptor 230 rotationally invariant for the compound or pocket structure. The similarity information 250 is information representing the similarity between the feature quantities, and the compound extraction result 260 is information representing the target compound extracted based on the similarity.

[0096] Figure 4 2 is a diagram showing a state in which structural information 210, stereostructure information 220, AAS descriptors 230, and invariant AAS descriptors 240 for N compounds (N is an integer greater than or equal to 2) are associated and stored in the storage unit 200. Figure 4 In the example, the structural formula can be set as the structural information 210, and the three-dimensional structural formula (described later) can be set as the three-dimensional structural information 220. Figure 4 For each compound, the AAS descriptor 230 (recorded as "V a(r)”; a is a subscript indicating the type of amino acid) and the invariant quantized AAS descriptor 240 (recorded as “V a (r)"; a is a subscript indicating the type of amino acid, r = |r|; r is the absolute value of the vector r) and are associated and stored. In addition, as described later, the AAS descriptor 230 and the invariant AAS descriptor 240 can be expressed in the form of the nearest distance, scattering angle, differential scattering cross-section, etc., but in Figure 4 For the sake of convenience, these expressions are collectively referred to as “V a (r)” and “V a (r)". Furthermore, the AAS descriptors 230 and the invariant AAS descriptors 240 may be stored for a portion of amino acids according to the number of descriptors used for screening, rather than for all 20 amino acids.

[0097] In the storage unit 200, a plurality of Figure 4 In addition, Figure 4 , the information stored for the target protein can also be stored in the same structure. The calculation method of the AAS descriptor and / or the invariant AAS descriptor using this structural information and stereostructure information will be described later.

[0098] <Display and Operation Section Structure>

[0099] The display unit 300 includes a monitor 310 (display device) and can display input images, images and information stored in the storage unit 200, and results of processing by the processing unit 100. The operation unit 400 includes a keyboard 410 and a mouse 420 as input devices and / or pointing devices. The user can use these devices and the screen of the monitor 310 to perform operations required for executing the feature quantity calculation method involved in the present invention and extracting target compounds (described later). Examples of user-operable operations include specifying the processing mode, the type of descriptor to be calculated, the descriptor used for screening, and the threshold value for similarity.

[0100] <Processing in the screening device>

[0101] In the screening device 10 having the above configuration, calculation of feature values ​​(descriptors) and / or extraction of target compounds can be performed in accordance with instructions from the user via the operation unit 400. The details of each process will be described below.

[0102] Calculation of characteristic quantities

[0103] The screening device 10 can calculate the AAS descriptor and / or the invariant quantized AAS descriptor in accordance with an instruction from the user via the operation unit 400 .

[0104] Calculation of AAS descriptors for compounds

[0105] The AAS descriptor is the differential scattering cross-section (cross-sectional area, scattering cross-sectional area) when a probe such as an amino acid (20 types such as alanine and valine) collides with and scatters a compound (target structure). This differential scattering cross-section can be calculated by performing a simulation (executing the characteristic quantity calculation method of the present invention) using the screening device 10. In the simulation, Figure 5 As shown in FIG. 1 , it is assumed that a probe 902 (probe) such as an amino acid collides with and scatters a compound 900 disposed at the origin of a coordinate system.

[0106] The data (scalar) obtained from the simulation, i.e., the data obtained by the following equation (1), is the target descriptor. As mentioned above, this descriptor is called the "amino acid scattering descriptor (AAS descriptor)." Furthermore, since the probes, such as amino acids, are in a scattering state, the total energy (interaction energy + kinetic energy) is positive.

[0107] [Formula 1]dσ / dΩ(E,b,a)(1)

[0108] Among them, E is the independent variable used to determine the incident energy of the probe, b is the independent variable used to determine the collision diameter of the probe, and a is the independent variable used to determine the type of probe. Figure 5 The case of scattering of one amino acid and a compound was described in the above simulation, but in the above simulation, a peptide connected with two or more amino acids can also be used as a probe. In this case, "a" in formula (1) refers to an independent variable for determining the type of peptide.

[0109] For ease of processing, the above dσ / dΩ(E, b, a) can be constructed to be invariant to compound rotations. This descriptor is referred to as an "invariant amino acid scattering descriptor (invariant AAS descriptor)." For example, the value obtained by averaging the compound at all angles and calculating dσ / dΩ(E, b, a) relative to the averaged compound is invariant to rotations and is a representation of an "invariant amino acid scattering descriptor (invariant AAS descriptor)." The invariance of AAS descriptors will be discussed later.

[0110] Regarding AAS descriptors and invariant AAS descriptors, amino acids are one example of probes, but other substances can also be probes, as described later. Probes are structures with multiple points separated by a real charge and generating van der Waals forces. In the present invention, the feature quantity when the scope of probes is expanded to include substances other than amino acids is also referred to as an "AAS descriptor" or "invariant AAS descriptor."

[0111] Figure 6 This flowchart illustrates the steps for calculating an AAS descriptor for a compound (target structure). In step S100, the information input unit 110 (processor) inputs the structural formula of a compound in accordance with user instructions via the operation unit 400. The compound represented by the input chemical formula is then designated as the target structure (target structure designation step).

[0112] The feature value calculation unit 120 (processor) converts the input structural formula into three dimensions, thereby generating a three-dimensional structure of the compound based on multiple atoms (multiple unit structures having chemical properties) (step S102: three-dimensional structure acquisition step). Various methods are known for converting structural formulas into three dimensions, and the method used in step S102 is not particularly limited. Figure 7 This is a diagram showing an example of a three-dimensional structural formula. Figure 7 Part (a) represents the input structural formula, Figure 7 Part (b) represents the three-dimensionalized structural formula. Alternatively, instead of inputting the structural formula as in steps S100 and S102 to convert it into three dimensions, a three-dimensional structure that has already been converted into three dimensions can be obtained (input) (three-dimensional structure acquisition step). The feature quantity calculation unit 120 places the center of gravity of the three-dimensionalized structural formula at r(x, y, z) = 0 (the origin of the coordinate system) (step S102: three-dimensional structure acquisition step).

[0113] The feature value calculation unit 120 calculates the interaction energy V felt by each atom “μ” of the amino acid “a” (a is a number indicating the type of amino acid; 1 to 20) aμ (r) (Step S104; probe feature value calculation process). aμ In (r), “r” is a vector. As V aμ The calculation method of (r) can be molecular dynamics (MD) method, but is not limited to this. aμ The amino acid (r) may be a predetermined type, or may be determined according to a user's instruction (it may be one or more types, or may be multiple types).

[0114] The feature quantity calculation unit 120 calculates the feature quantity based on V aμ(r) Calculate the interaction energy V felt by the center of gravity of the amino acid a (r) (Step S106: Probe feature value calculation process), and relative to V a (r) Calculate V by taking the average angle around r = 0 ((x, y, z) = (0, 0, 0)) a (r) (Step S108: Probe feature value calculation step) r = |r| (r is the absolute value of vector r), that is, V a r in (r) is V aμ The absolute value of the vector r in (r).

[0115] Then, the feature quantity calculation unit 120 calculates the V a (r) Calculate the closest distance r min,a and scattering angle θ a As a function of the incident energy E and the collision diameter b (step S110: probe feature value calculation process). As will be described later, the closest distance r min,a and scattering angle θ a is an expression of the AAS descriptor. The feature quantity calculation unit 120 calculates the feature quantity according to the nearest distance r min,a and scattering angle θ a The differential scattering cross section dσ / dΩ(E, b, a) is calculated as a function of the incident energy E and the collision diameter b (step S112: probe characteristic quantity calculation step). The differential scattering cross section dσ / dΩ(E, b, a) is also an expression of the AAS descriptor.

[0116] Figure 8 It means targeting Figure 7 Plot of the differential scattering cross-sections for the indicated compounds. Figure 8 Part (a) is a graph of the differential scattering cross-section (AAS descriptor, first characteristic quantity) for alanine (amino acid). Figure 8 Part (b) is a graph of the differential scattering cross-section (first characteristic quantity) for phenylalanine (amino acid). Figure 9 In other forms, it is expressed Figure 7 The characteristic amount of the compound shown in FIG. (the probe is alanine). Specifically, Figure 9 Part (a) represents the closest distance r min,a (first feature quantity) graph, Figure 9 Part (b) shows the scattering angle θ a chart.

[0117] Calculation of AAS Descriptors for Pocket Structures

[0118] In the screening device 10, a pocket structure that binds to a target protein can also be designated as a target structure, and a feature value (AAS descriptor; second feature value) for the pocket structure can be calculated. The pocket structure is the target structure that binds to the pocket that serves as the active site of the target protein, and the "active site" refers to the site where the activity of the target protein is promoted or inhibited by the binding of the pocket structure. Figure 10 is a flowchart showing the calculation steps of the AAS descriptor for the pocket structure. Figure 11 This is a conceptual diagram showing the relationship between the target protein and the pocket structure.

[0119] exist Figure 10 In the flowchart of FIG. 1 , the information input unit 110 inputs the actual measurement of the three-dimensional structure of the target protein and the position information of the pocket (step S200 : target structure specifying step). Figure 11 Part (a) shows the pocket PO in the target protein TP. Through the process of step S200, the pocket structure is designated as the target structure.

[0120] The feature value calculation unit 120 inserts multiple virtual spheres (multiple unit structures with chemical properties) into the pocket of the target protein (step S202: target structure designation step, three-dimensional structure acquisition step). The "virtual spheres" can be assumed to have chemical properties such as van der Waals radius and charge, and the "insertion of virtual spheres" can be performed using simulation (e.g., molecular dynamics methods). Through step S202, the collection of inserted virtual spheres (three-dimensional structure) can be obtained as the three-dimensional structure of the pocket structure (target structure) (step S204: three-dimensional structure generation step). Figure 11 Part (b) shows an example of the pocket structure PS for the target protein TP.

[0121] The feature quantity calculation unit 120 uses the acquired three-dimensional structure to calculate the cross-sectional area (second feature quantity; one form of AAS descriptor) of one or more amino acids relative to the pocket structure (step S206: probe feature quantity calculation process). In fact, it is possible to calculate how the amino acid is scattered by the pocket structure. In addition, the amino acids for which the second feature quantity is calculated only need to be one or more (it can also be multiple). Moreover, the calculation of the second feature quantity can be performed for a predetermined type of amino acid or for an amino acid set according to a user's operation. The feature quantity calculation unit 120 associates the calculated AAS descriptor as an AAS descriptor 230 with the structural information (structural information 210) and the three-dimensional structure information (three-dimensional structure information 220) of the compound and stores it in the storage unit 200 (reference Figure 3 , Figure 4; Storage step). When the invariant quantization AAS descriptor described later has been calculated, the feature amount calculation unit 120 associates the AAS descriptor with the invariant quantization AAS descriptor.

[0122] <Calculation of AAS Descriptors Using Nucleobases, Etc. as Probes>

[0123] In the present invention, biopolymers (compounds) other than proteins, such as DNA (deoxyribonucleic acid), RNA (ribonucleic acid), cell membranes, and polysaccharides, can be used as pharmaceutical targets. When calculating characteristic quantities (third characteristic quantities; one form of AAS descriptors) for these target compounds, the probe is set to another substance (a component of each target) instead of an amino acid. Specifically, when the targets are DNA, RNA, cell membranes, and polysaccharides, the probes are set to one or more nucleic acid bases, one or more lipid molecules, and one or more monosaccharide molecules, respectively. Furthermore, when calculating characteristic quantities for these probes, water and one or more ions can be considered. From a local perspective, the pharmacological efficacy of a compound (the binding strength with a target such as DNA) is expressed as the result of the interaction between the compound and the nucleic acid base, etc. (probe). Therefore, as long as the characteristic quantities representing the cross-sectional area of ​​the nucleic acid base, etc., are similar between compounds, it means that the binding strength of these compounds with the target is similar. In other words, compounds with similar third characteristic quantities exhibit similar pharmacological efficacy. Therefore, the chemical properties of the compound can be accurately determined by the third characteristic quantity. In addition, the third characteristic quantity can be calculated in the same manner as the first and second characteristic quantities (refer to Figure 5 、 Figure 6 and their descriptions, etc.).

[0124] Invariant Quantization of AAS Descriptors

[0125] The above-mentioned AAS descriptors represent the cross-sectional area of ​​amino acids, etc., but even if the compound is the same, the value will change if rotation occurs. Therefore, in the screening device 10 involved in the first embodiment, the feature calculation unit 120 (processor) can calculate "invariant AAS descriptors that make the AAS descriptors invariant with respect to the rotation of the compound" (first invariant feature, second invariant feature, third invariant feature) in addition to or instead of the AAS descriptors. In addition, regardless of whether it is the case of a compound or a pocket structure, invariance can be performed using the same steps. When the AAS descriptors (first feature, third feature) for the compound are used, the invariant AAS descriptors (first invariant feature, third invariant feature) for the compound can be obtained, and when the AAS descriptors (second feature) for the pocket structure are used, the invariant AAS descriptors (second invariant feature) for the pocket structure can be obtained.

[0126] According to the above interaction energy V a (r) (refer to step S106) The calculated closest distance r min,a and scattering angle θ a , and the differential scattering cross section dσ / dΩ(E,b,a) is an example of an AAS descriptor before invariant quantization. The feature quantity calculation unit 120 calculates the value of V a (r) V obtained by angular averaging a (r) (r = |r|, r is the absolute value of vector r; refer to step S108), the invariant quantized AAS descriptor (the closest distance r) can be calculated. min,a and scattering angle θ a , and the differential scattering cross-section dσ / dΩ(E,b,a); the first to third invariants. Furthermore, the AAS descriptor is invariant to translation from the outset, and the only invariant is rotation.

[0127] The feature calculation unit 120 associates the calculated invariant quantized AAS descriptor as the invariant quantized AAS descriptor 240 with the structural information (structural information 210), the stereoscopic structure information (stereoscopic structure information 220) of the compound, and the original AAS descriptor 230, and stores the associated information in the storage unit 200 (refer to FIG. Figure 3 , Figure 4 ; storing step). In addition, when the invariant AAS descriptor is calculated using the AAS descriptors for two different amino acids, it is also possible to have a plurality of AAS descriptors associated with the invariant AAS descriptor.

[0128] According to the above-mentioned invariant AAS descriptors, compounds with similar descriptors show similar pharmacodynamics (e.g., bonding to the target protein), thereby accurately representing the chemical properties of the target structure (compound, pocket structure, biopolymer). In addition, the invariant AAS descriptors obtained by invariantly quantizing the AAS descriptors, for example, by using AAS descriptors for two different amino acids for invariant quantization, can accurately perform comparisons of descriptor-based compounds (pharmaceutical efficacy determination) while easily processing feature quantities and reducing data capacity. Moreover, according to the invariant AAS descriptors, hits can be easily found.

[0129] Ease of finding hits based on invariant quantized AAS descriptors

[0130] The ease of finding hits based on the invariant quantized AAS descriptors was evaluated through the following steps 1 to 5.

[0131] (Step 1) X hit compounds and Y non-hit compounds are mixed with respect to a certain target (target protein, etc.).

[0132] (Step 2) Calculate the invariant quantitative AAS descriptor for the (X+Y) compounds as a whole.

[0133] (Step 3) Calculate the similarity of each descriptor.

[0134] (Step 4) Group the (X+Y) compounds according to the similarity of their invariant quantized AAS descriptors.

[0135] (Step 5) Check whether the group in which the hits are clustered is generated mechanically.

[0136] For the groups prepared for protein ABL1 (kinase), the ease of finding hits (= expected value; number of hits × hit content) for each group is shown in the table (comparison results with the case of random grouping). Figure 12 In addition, Figure 12 Part (a) shows the results of the following experiments: (1) invariant AAS descriptor (probes are amino acids; first invariant feature quantity), (2) invariant polyion (probes are all monatomic ions of Na + and Cl - ; 3rd invariant feature quantity), (3) invariant AAS descriptor and ion (probes are amino acids and Na + and Cl -; 4th invariant feature), (4) invariant dipole (probe is dipole; 5th invariant feature), (5) invariant AAS descriptor and dipole (probe is amino acid and dipole; 6th invariant feature), (6) invariant multi-ion and dipole (probe is Na + and Cl - and dipoles; the 7th invariant feature), (7) invariant AAS descriptors, multiple ions and dipoles (probes are amino acids, Na + and Cl - and dipole; the expected value of the hit number of the 8th invariant quantized feature quantity). And, Figure 12 Part (b) shows the results of the analysis of (1) invariant AAS descriptors (probes are amino acids; the first invariant feature quantity and Figure 12 (a) Same as (1)), (8) Invariant monatomic ions (probe is Na + ; 3rd invariant feature), (9) invariant AAS descriptors and monatomic ions (probes are amino acids and Na + ; 4th invariant characteristic quantity), (10) invariant monatomic ions and dipoles (probe is Na + and dipoles; the 7th invariant feature), (11) invariant AAS descriptors, monatomic ions and dipoles (probes are amino acids, Na + and dipole; the expected value of the number of hits of the 8th invariant quantized feature quantity).

[0137] according to Figure 12 As a result, it can be seen that when the invariant quantization AAS descriptor is used, more hit groups are generated than random grouping. Figure 12 In the example, the group numbers vary depending on the grouping method (random, invariant quantization AAS descriptor), so the quality of the grouping is judged by "whether it contains groups with high expected values ​​(containing more hits)" rather than by comparing the expected values ​​in the same group number.

[0138] Invariant AAS descriptors for compounds bound to and not bound to target proteins

[0139] According to the feature quantity used in the present invention (AAS descriptor, invariant quantized AAS descriptor, amino acid scatter descriptor), for example Figure 12 As described above, the target compound that is bound to the target protein can be extracted and created. However, in addition to this, for example, the target compound that is not bound to the target protein can also be extracted and created. Figure 13 Part (a) shows the correlation between the target protein (and Figure 12Example of the expected value of the number of hits (comparison with the case of random grouping) of the compound (first target compound) bound to the same protein ABL1) (the probe is an amino acid), Figure 13 Part (b) shows an example of the expected number of hits of a compound that does not bind to the target protein (second target compound) calculated similarly based on the invariant AAS descriptor. Figure 13 It can be seen that by using the characteristic quantities of the present invention, it is possible to easily identify hits not only for compounds that bind to the target protein (first target compound) but also for compounds that do not bind to the target protein (second target compound). Binding strength can be measured, for example, using the IC50 (halfmaximal (50%) inhibitory concentration; 50% inhibition concentration), where the "binding / non-binding" threshold can be around 100 to 1000 μM. However, different indicators and values ​​can also be used depending on the subject (which property is being evaluated).

[0140] Furthermore, “not binding to a specific protein” means “effective for describing compounds with no toxicity (low toxicity)”, and thus, using the similarity between the AAS descriptor and the invariant AAS descriptor, compounds with no toxicity (low toxicity) can be searched or created.

[0141] <Effects of feature calculation methods>

[0142] As described above, in the screening device 10 (feature quantity calculation device, screening device) involved in the first embodiment, the feature quantity calculation method involved in the present invention and the program for executing the same (feature quantity calculation program) can be used to calculate the feature quantity (AAS descriptor, invariant AAS descriptor) that accurately represents the chemical properties of the object structure.

[0143] <Extraction (screening) of target compounds>

[0144] The extraction of target compounds (candidate pharmaceutical compounds) from a plurality of compounds using the aforementioned AAS descriptors and invariant AAS descriptors is described. Target compound extraction includes a mode based on ligand descriptors (AAS descriptors, invariant AAS descriptors) (mode 1), a mode based on target protein pocket structure descriptors (AAS descriptors, invariant AAS descriptors) (mode 2), and a mode based on bound compounds (compounds confirmed to be bound to target biopolymers other than proteins) descriptors (AAS descriptors, invariant AAS descriptors) (mode 3). The user can select which mode to use for extraction based on an operation performed via the operating unit 400.

[0145] <Screening of ligand input>

[0146] Figure 14 This is a flowchart showing the steps of screening (first mode) using the AAS descriptor of the ligand. When the process starts, the feature value calculation unit 120 calculates the AAS descriptor of the ligand (step S300: screening feature value calculation process). In addition, since the ligand is a compound that has been confirmed to be bonded to the target protein, the calculation of the AAS descriptor in step S300 can be performed by Figure 6 The steps are as shown in the flowchart.

[0147] like Figure 4 Based on the above, the screening device 10 associates the compound's three-dimensional structure based on multiple atoms with the AAS descriptor (first feature value) corresponding to the three-dimensional structure for each of the multiple compounds and stores the associated information in the storage unit 200. The similarity calculation unit 130 calculates the similarity between the compound's AAS descriptor and the ligand's AAS descriptor calculated in step S400 (step S302: similarity calculation step). After calculating the similarity, the compound extraction unit 140 extracts target compounds based on the similarity (step S304: compound extraction step). As described above, similarity in the AAS descriptors indicates similar pharmacological effects (binding to the target protein). Therefore, using the similarity of the AAS descriptors allows the extraction of compounds with similar pharmacological effects as the ligand (i.e., target compounds as drug candidates). Specifically, the extraction of target compounds based on similarity (step S3404) can be performed by, for example, extracting compounds with a similarity above a threshold or extracting compounds in descending order of similarity.

[0148] exist Figure 14 In the above description, the filtering process using the AAS descriptor is described. However, the filtering process using the invariant quantization AAS descriptor can also be performed in the same manner. Specifically, the feature amount calculation unit 120 performs Figure 6 The invariant AAS descriptor (first invariant feature value) of the ligand is calculated using the steps of (2) and (3) above, and the similarity calculation unit 130 calculates the similarity with the invariant AAS descriptor of the compound stored in the storage unit 200. After the similarity is calculated, the compound extraction unit 140 extracts the target compound based on the similarity. The specific method of extracting the target compound based on the similarity can be performed in the same manner as the AAS descriptor.

[0149] Figure 15 This is a table showing examples of screening results of ligand input. Figure 15 Part (a) shows the results of the case where the AAS descriptor is used to "extract compounds with a similarity greater than or equal to a threshold value." Figure 15Part (b) shows the result of the case where the invariant quantized AAS descriptor is used to extract compounds in descending order of similarity. Figure 15 In part (a), the AAS descriptor for amino acid 1 (the same as that for Figure 4 The various expressions are collectively referred to as "V a (r)”), compounds can be extracted based on AAS descriptors for other amino acids (amino acids 2 to 20) (for example, V2(r)). Furthermore, the similarities of multiple AAS descriptors (for example, V1(r) and V2(r)) for different amino acids (the similarities between V1(r)s and the similarities between V2(r)s) can be calculated separately, and compounds can be extracted based on these. There may be one AAS descriptor used for compound extraction, but by using multiple AAS descriptors, compounds can be correctly extracted based on similarity. In addition, when multiple AAS descriptors are used, the combination of amino acids between these descriptors is not particularly limited (for example, it can be V1(r) and V2(r), or it can be V3(r) and V4(r)).

[0150] Similarly, in Figure 15 In part (b), compounds are extracted based on the invariant AAS descriptors (V1(r), V2(r)) for amino acids 1 and 2, but the amino acids for which the invariant AAS descriptors are calculated can be other combinations (for example, V3(r), V4(r) based on amino acids 3 and 4). Furthermore, compound extraction can be performed based on multiple invariant AAS descriptors with different amino acid combinations (for example, V1(r) and V2(r) and V3(r) and V4(r)) (for example, using the similarity of V1(r), V2(r) and the similarity of V3(r), V4(r)). There can be only one invariant AAS descriptor used for compound extraction, but by using multiple invariant AAS descriptors, it is possible to correctly extract compounds based on similarity. Furthermore, when using multiple invariant AAS descriptors, the combination of amino acids between these descriptors is not particularly limited (for example, it can be V1(r) and V2(r) and V3(r) and V4(r), or it can be V1(r) and V2(r) and V1(r) and V3(r)). The processing unit 100 (feature value calculation unit 120, similarity calculation unit 130, compound extraction unit 140) may determine which amino acid descriptor and similarity are calculated based on an instruction from the user via the operation unit 400, or it may determine the calculation based on no instruction from the user.

[0151] In addition, Figure 15In part (a), the threshold value of similarity is set to 80%, and in part (b), the number of extractions is set to 100. However, these values ​​are for example only, and the threshold value and the number of extractions can be set according to conditions such as the accuracy of the screening. The setting can be made according to the input made by the user through the operation unit 400. Figure 15 On the contrary, when the AAS descriptor is used, it can be set to "extract compounds in descending order of similarity", and when the invariant AAS descriptor is used, it can be set to "extract compounds with similarity greater than a threshold value". Figure 15 The extraction results shown are stored in the storage unit 200 as compound extraction results 260 (see Figure 3 ).

[0152] <Screening of target protein input>

[0153] Figure 16 This is a flowchart showing the steps of screening (second mode) using AAS descriptors for the pocket structure of the target protein. When the process starts, the feature value calculation unit 120 calculates the AAS descriptors for the pocket structure of the target protein (step S400: screening feature value calculation process). The calculation of the AAS descriptors in step S400 can be performed by Figure 11 The similarity calculation unit 130 calculates the similarity between the AAS descriptor for the compound and the AAS descriptor for the pocket structure calculated in step S400 (step S402: similarity calculation process). After calculating the similarity, the compound extraction unit 140 extracts the target compound based on the similarity (step S404: compound extraction process). Similar to the case of inputting the ligand described above, the extraction of the target compound based on the similarity (step S404) can be specifically performed by "extracting compounds with a similarity above a threshold value", "extracting compounds in order of high to low similarity", etc.

[0154] In the case of using the invariant quantization AAS descriptor, it is also possible to Figure 16 The target compounds were extracted using the same steps as in the flowchart.

[0155] Figure 17 This is a table showing examples of screening results of target protein input. Figure 17 Part (a) shows the results of the case where the AAS descriptor is used to "extract compounds with a similarity greater than or equal to a threshold value." Figure 17 Part (b) shows the result of the case where the invariant quantized AAS descriptor is used to extract compounds in the order of high to low similarity. The threshold of similarity and the number of extractions can be set according to the conditions such as the accuracy of the screening. The setting can be made according to the input made by the user through the operation unit 400. Figure 17 Conversely, when the AAS descriptor is used, the method may be to "extract compounds in descending order of similarity," and when the invariant AAS descriptor is used, the method may be to "extract compounds having a similarity greater than or equal to a threshold value."

[0156] The situation of screening with target protein input is also similar to that of screening with ligand input (refer to Figure 14 、 Figure 15 and its description), the type of amino acid can be changed, and multiple descriptors (AAS descriptors, invariant AAS descriptors) for different amino acids can be used. The descriptor used for compound extraction can be one, but by using multiple descriptors, the extraction of compounds based on similarity can be performed correctly. In addition, when multiple descriptors are used, the combination of amino acids between these descriptors is not particularly limited. Regarding which amino acid the descriptor and similarity are calculated for, it can be determined by the processing unit 100 (feature quantity calculation unit 120, similarity calculation unit 130, compound extraction unit 140) according to the instructions of the user via the operation unit 400, or it can be determined by the processing unit 100 without following the instructions of the user.

[0157] The compound extraction unit 140 will Figure 17 The extraction results shown are stored in the storage unit 200 as compound extraction results 260 (see Figure 3 ).

[0158] <Screening of target biopolymers other than proteins>

[0159] In the screening device 10 according to the first embodiment, target compounds that are bound to target biopolymers other than proteins can also be extracted. Figure 14 、 Figure 16 The same steps as in the flowchart are performed, and the third feature value is used for screening (third mode).

[0160] <Effects of screening equipment>

[0161] As described above, in the screening device 10 involved in the first embodiment, the feature quantities (AAS descriptors, invariant AAS descriptors) calculated by the feature quantity calculation method involved in the present invention (a program that enables a computer to execute the feature quantity calculation method) can be used, and the screening method involved in the present invention (and a program that enables a computer to execute the screening method) can be used to effectively screen for pharmaceutical candidate compounds.

[0162] <Second embodiment>

[0163] A compound creating device according to a second embodiment of the present invention will be described. Figure 18 1 is a block diagram showing the structure of the compound creation device 20 (feature value calculation device, compound creation device). Components identical to those in the first embodiment are denoted by the same reference numerals, and detailed descriptions thereof are omitted.

[0164] The compound creation device 20 includes a processing unit 101. The processing unit 101 is configured as follows Figure 19 The present invention also includes an information input unit 110, a feature value calculation unit 120 (creating a feature value calculation unit), a generator construction unit 132 (creating a generator construction unit), a compound stereostructure generation unit 142 (creating a compound stereostructure generation unit), and a display control unit 150. The functions of the information input unit 110, feature value calculation unit 120, and display control unit 150 are respectively the same as those of the information input unit 110, feature value calculation unit 120, and display control unit 150 in the aforementioned screening device 10. As with the aforementioned content of the screening device 10, the functions of these units can be implemented using various processors.

[0165] Figure 20 2 is a diagram showing information stored in the storage unit 201. The storage unit 201 stores the three-dimensional structure generation result 270 in the screening device 10 instead of the compound extraction result 260. Figure 4 Similar to the above contents, the information stored in the storage unit 201 is stored in association with each other.

[0166] <Generation of the stereostructure of the target compound>

[0167] The generation of the three-dimensional structure of the target compound (candidate drug compound) using the above-mentioned AAS descriptors and invariant AAS descriptors is described. In the generation of the three-dimensional structure of the target compound based on the compound creation device 20, since no search is performed, the three-dimensional structure of the compound can be generated even in the case where "the result of the retrieval based on the screening is unsolvable", thereby effectively creating the three-dimensional structure of the candidate drug compound. Regarding the generation of the three-dimensional structure, there are a mode (the first mode) based on the descriptor of the ligand (AAS descriptor, invariant AAS descriptor), a mode (the second mode) based on the descriptor of the pocket structure of the target protein (AAS descriptor, invariant AAS descriptor), and a mode (the third mode) based on the descriptor of the bonded compound (AAS descriptor, invariant AAS descriptor). It is possible to select which mode to use for the generation of the three-dimensional structure according to the operation performed by the user via the operation unit 400.

[0168] <Generation of the stereostructure of the input ligand>

[0169] Figure 21This is a flowchart showing the steps of generating a three-dimensional structure when a ligand is input. When the processing starts, the feature quantity calculation unit 120 calculates the descriptor (AAS descriptor) of the ligand (step S500: object structure designation step, three-dimensional structure generation step, creation feature quantity calculation step). As in the first embodiment, the processing of step S500 can be performed using the feature quantity calculation method (and a program for causing a computer to execute the feature quantity calculation method) according to the present invention (see Figures 6 to 9 and descriptions of these figures).

[0170] In step S502, the generator construction unit 132 constructs a generator by machine learning (generator construction process). Figure 22 The processing of step S502 is also described.

[0171] (Step 1) Figure 22 As shown in part (a), the feature quantity calculation unit 120 calculates AAS descriptors (first feature quantities) using amino acids as probes for a plurality of compounds, and creates pairs of a structural formula 912 obtained by stereoregulating the structural formula of compound 910 and an AAS descriptor 914 .

[0172] (Step 2) Figure 22 As shown in part (b), the generator construction unit 132 constructs a generator 916 by using machine learning such as deep learning with the three-dimensional structure of the compound (structural formula 912) as training data and the AAS descriptor 914 as an explanatory variable. The machine learning method is not limited to a specific method. For example, it can be a simple all-bonded neural network, a convolutional neural network (CNN), and a generative adversarial network (GAN). The accuracy of the three-dimensional structure generation depends on the learning method used. Therefore, it is preferable to select a learning method based on conditions such as the generation conditions and the required accuracy of the three-dimensional structure.

[0173] If the processing of step 1 and step 2 is completed, return to Figure 21Flowchart of . The compound stereostructure generation unit 142 uses the constructed generator to generate the stereostructure (stereomorphic structural formula) of the target compound (hit) from the AAS descriptor of the ligand (step S504: compound stereostructure generation process). In this way, it is possible to obtain the stereostructure of a compound having a similar pharmacological effect (bonding to the target protein) as the ligand, that is, a pharmaceutical candidate compound. In addition, there can be multiple stereostructures providing the same AAS descriptor. The compound stereostructure generation unit 142 associates the generated stereostructure with the AAS descriptor (AAS descriptor 230) as a stereostructure generation result 270 and stores it in the storage unit 201. In accordance with the instructions given by the user via the operation unit 400, the display control unit 150 can display the generated stereostructure on the display 310.

[0174] In addition, in the above steps, the amino acid for calculating the AAS descriptor for constructing the generator can be one or more. Among them, by calculating the AAS descriptors for multiple amino acids and using them for learning (construction of the generator), the accuracy of the generated three-dimensional structure can be improved. In addition, when using multiple AAS descriptors of different types of amino acids, the combination of amino acids between these descriptors is not particularly limited. As to which amino acid the AAS descriptor is calculated for and used for learning, it can be determined by the processing unit 100 (feature quantity calculation unit 120, similarity calculation unit 130, compound extraction unit 140) according to the instructions of the user via the operating unit 400, or it can be determined by the processing unit 100 without following the instructions of the user.

[0175] <Example of generating a three-dimensional structure>

[0176] Figure 23 An example of a three-dimensional structure generated using a generator constructed through machine learning is described. Figure 23 Part (a) is the correct data of the three-dimensional structure. Figure 23 Part (b) is an example of a three-dimensional structure generated using the generator. Figure 23 The compound to be created is Figure 7 、 Figure 22 Compound 910 in .

[0177] Relationship between the characteristics of training data and the generated 3D structure

[0178] The stereostructures generated through the above steps are influenced by the characteristics of the compounds provided as training data. Therefore, by selecting the characteristics of the compounds provided as training data, compounds with different stereostructures can be generated. For example, by providing AAS descriptors for compounds with easily synthesized stereostructures as training data, compounds with similar pharmacodynamics to the ligand and easily synthesized stereostructures can be generated. The choice of which compound to provide AAS descriptors for as training data can be made based on the characteristics of the compound to be generated.

[0179] <Generation of 3D Structure Using Invariant Quantization AAS Descriptors>

[0180] exist Figures 22 and 23 , generation of a stereostructure using an AAS descriptor (first feature) is described. In contrast, when an invariant AAS descriptor (first invariant feature) is used, the stereostructure of a target compound can be generated by machine learning (deep learning) using the invariant AAS descriptor as training data and the stereostructure (stereomorphic structural formula) as an explanatory variable, similarly to the case of using the AAS descriptor.

[0181] <Generate the three-dimensional structure of the target protein>

[0182] In addition to generating a three-dimensional structure based on the aforementioned ligand input, the compound creation device 20 can also generate a three-dimensional structure of a target compound by inputting a target protein. In this case, similar to the case of inputting a ligand, three-dimensional structure generation using an AAS descriptor (second feature quantity) and using an invariant AAS descriptor (second invariant feature quantity) can be performed.

[0183] Figure 24 This is a flowchart showing the steps of generating a three-dimensional structure when a target protein is input (assuming that an AAS descriptor is used). When the process starts, the feature value calculation unit 120 calculates the AAS descriptor (second feature value) of the pocket structure of the target protein (step S600: target structure designation step, three-dimensional structure generation step, creation feature value calculation step). As in the first embodiment, the process of step S600 can be performed using the feature value calculation method of the present invention (see Figure 9 and descriptions of these figures).

[0184] In step S602, as in the case of inputting a ligand, the generator construction unit 132 constructs a generator (generator construction process) by machine learning (deep learning). Regarding the construction of the generator, it can be carried out in the same manner as the above-mentioned steps 1 and 2. Specifically, the feature quantity calculation unit 120 calculates the AAS descriptor (the second feature quantity) using amino acids as probes for the pocket structure, and makes a pairing of the three-dimensional structure of the pocket structure and the AAS descriptor. The generator construction unit 132 sets the AAS descriptor as an explanatory variable and the three-dimensional structure of the pocket structure as training data, and constructs a generator. The compound three-dimensional structure generation unit 142 uses the constructed generator to generate the three-dimensional structure (stereoscopic structural formula) of the target compound (hit) from the AAS descriptor of the pocket structure (step S604: compound three-dimensional structure generation process). Thus, it is possible to obtain a compound having a pharmacological effect (bonding to the target protein) similar to that of the pocket structure, that is, the three-dimensional structure of the pharmaceutical candidate compound. In addition, there can be multiple three-dimensional structures providing the same AAS descriptor. The compound stereo structure generating unit 142 associates the generated stereo structure with the AAS descriptor (AAS descriptor 230) as a stereo structure generating result 270 and stores the result in the storage unit 201 (refer to Figure 20 ). In accordance with the user's instruction via the operation unit 400 , the display control unit 150 may display the generated three-dimensional structure on the display 310 .

[0185] Furthermore, when the second invariant quantized feature value (invariant quantized AAS descriptor) is used, a three-dimensional structure can be generated in the same manner.

[0186] <Generation of three-dimensional structures when inputting target biopolymers other than proteins>

[0187] In addition to the above-described method, the compound creation device 20 can also generate the three-dimensional structure of a target compound by inputting a target biopolymer other than a protein. In this case, similarly to the above-described method, three-dimensional structure generation using AAS descriptors (third feature quantities) and three-dimensional structure generation using invariant AAS descriptors (third invariant feature quantities) can be performed.

[0188] Creation of compounds based on structural diversity

[0189] In the above-described method, the three-dimensional structure of a pharmaceutical candidate compound is generated using a generator formed by machine learning. However, as described below, the three-dimensional structure of a compound can also be generated based on structural diversity.

[0190] <Additional structure of compound creation method>

[0191] The compound creation method described below corresponds to the 23rd to 27th aspects of the present invention described above, but the following structure (hereinafter referred to as "additional structure") is appropriately added to the 23rd to 27th aspects (hereinafter referred to as "basic structure").

[0192] (Additional structure: 1)

[0193] In the basic structure, in the candidate structure adoption process, as the first adoption processing, when the absolute value of the difference between the physical property value of the candidate structure and the target value of the physical property value is less than the absolute value of the difference between the physical property value of the chemical structure and the target value of the physical property value, the candidate structure adoption processing is performed, and when the absolute value of the difference between the physical property value of the candidate structure and the target value of the physical property value is greater than the absolute value of the difference between the physical property value of the chemical structure and the target value of the physical property value, the following processing is performed, that is, based on the difference between the physical property value of the candidate structure and the target value of the physical property value, the first adoption probability is calculated by the first function, and the candidate structure is adopted with the first adoption probability.

[0194] (Additional structure: Part 2)

[0195] In the additional structure 1, the first function is a monotonically decreasing function of the difference between the absolute value of the difference between the physical property value of the candidate structure and the target physical property value and the absolute value of the difference between the physical property value of the chemical structure and the target physical property value.

[0196] (Additional structure: Part 3)

[0197] In any one of the basic structure, additional structure 1 and structure 2, in the candidate structure adoption process, the following processing is performed as the second adoption processing, that is, the increase or decrease in the structural diversity of the structure group is calculated, and based on the increase or decrease, the second adoption probability is calculated by the second function, and the candidate structure is adopted with the second adoption probability.

[0198] (Additional structure: 4)

[0199] In the additional structure 3, the second function is a monotonically increasing function with respect to the increase or decrease amount of the structural diversity.

[0200] (Additional structure: No. 5)

[0201] In any one of the basic structure and the additional structures 1 to 4, in the candidate structure acquisition step, atoms or atomic groups are added or deleted from the chemical structure to generate a target structure, and the target structure is used as a candidate structure.

[0202] (Additional structure: No. 6)

[0203] In any one of the basic structure and additional structures 1 to 5, in the control process, when the number of times the chemical structure is changed reaches the specified number and / or the physical property value of the candidate structure reaches the target value, it is determined that the termination condition is met, and the processing of the input process, the candidate structure acquisition process, the physical property value calculation process and the candidate structure adoption process is terminated.

[0204] <Structure of compound creation device>

[0205] Figure 25 This is a diagram showing the structure of a compound creation device when creating a three-dimensional structure of a compound based on structural diversity. In this embodiment, the compound creation device 20 has a processing unit 103 (processor) instead of Figure 18 、 Figure 19 The processing unit 101 shown in FIG. The processing unit 103 includes an input unit 105, a candidate structure acquisition unit 107, a physical property value calculation unit 109, a candidate structure adoption unit 111, a control unit 113, a display control unit 115, a CPU 121, a ROM 123, and a RAM 125. Other structures are similar to Figure 18 In addition, by using a processing unit having the configuration of the processing unit 103 in addition to the configuration of the processing unit 101 , it is possible to execute the creation of compounds based on the generator and the creation of compounds based on structural diversity.

[0206] <Steps of compound creation method>

[0207] Figure 26 is a flowchart representing the steps of a method for creating compounds based on structural diversity.

[0208] <Data input>

[0209] The input unit 105 inputs the chemical structure (initial structure) of one or more compounds, one or more physical property values ​​in the chemical structure (initial structure), and target values ​​of the physical property values ​​(step S1010: input process). Regarding this data, data stored in the storage unit 201 can be used, or it can be obtained from the external server 500 and the external database 510 via the network NW. The input data can be determined according to the instruction input by the user through the operation unit 400. There can be one or more initial structures. In addition, there can be one or more physical property values.

[0210] <Physical properties and target values>

[0211] When creating compounds based on structural diversity, the values ​​of AAS descriptors (first to third feature quantities) and invariant AAS descriptors (first to third invariant feature quantities) can be set as "physical property values." These physical property values ​​can be calculated using the feature quantity calculation method of the present invention. Furthermore, the values ​​of AAS descriptors and invariant AAS descriptors for ligands, pocket structures, bonded compounds, and the like can be set as "target values" for the physical property values. Specific examples will be described later.

[0212] <Acquisition of candidate structures>

[0213] The candidate structure acquisition unit 107 randomly changes the chemical structure to acquire candidate structures (step S1020: candidate structure acquisition process). Any method can be used as long as it changes the chemical structure. For example, a method can be used in which atoms or atomic groups are added or deleted from a chemical structure to generate a target structure, and the target structure is used as a candidate structure. Specifically, the method is a method for generating a compound structure, comprising: (A) preparing a compound database and a compound structure (chemical structure) as a benchmark for evaluating synthetic suitability; (B) selecting either to add an atom or an atomic group to the compound structure or to delete an atom from the compound structure; (C) if adding an atom to the compound structure is selected, bonding the new atom to an atom selected from atoms included in the compound structure, or if deleting an atom from the compound structure is selected, deleting the atom selected from atoms included in the compound structure, and obtaining a modified compound structure; (D) determining the synthetic suitability of the modified compound structure based on information in the compound database; (E) probabilistically allowing the modification if the modified compound structure has synthetic suitability, and probabilistically abandoning the modification if the modified compound structure does not have synthetic suitability; and (F) repeating steps (B) to (E) until the compound structure that has undergone step (E) satisfies an end condition. Furthermore, the generated candidate structure can be displayed on the display 310 (display device) by the display control unit 115. Furthermore, when returning to step S1020 from the later-described step S1090, one or more structures whose physical property values ​​are close to the target values ​​among the structures generated last time are added to the compound database used to evaluate the synthetic suitability. In step S1020, structures having physical property values ​​close to the target values ​​can also be gradually generated easily.

[0214] <Evaluation of physical properties>

[0215] The physical property value calculation unit 109 calculates the physical property values ​​of the candidate structure (the structure that changed in step S1020) (step S1030: physical property value calculation step, creation feature value calculation step). The calculation of the physical property values ​​is preferably performed using the same method as that used to estimate the physical property values ​​of the initial structure.

[0216] <First Adoption Process>

[0217] The candidate structure adoption unit 111 determines whether the physical property value is close to the target value (step S1040: candidate structure adoption process). Specifically, when the physical property value before the structural change is set to f0, the physical property value after the structural change is set to f1, and the target physical property value is set to F, if |F-f1|≤|F-f0| holds (the absolute value of the difference (first difference) between the physical property value of the candidate structure and the target physical property value is less than the absolute value of the difference (second difference) between the physical property value of the chemical structure and the target physical property value), the physical property value is close to (not far from) the target value, so the process proceeds to step S1070 and the structural change is adopted (first adoption process). On the other hand, if |F-f1|>|F-f0| (the absolute value of the difference (first difference) between the physical property value of the candidate structure and the target physical property value is greater than the absolute value of the difference (second difference) between the physical property value of the chemical structure and the target physical property value), the process proceeds to step S1050.

[0218] In step S1050 (candidate structure adoption process), the candidate structure adoption unit 111 calculates a first adoption probability using a first function based on the difference between the physical property value of the candidate structure and the target physical property value (first adoption process). Specifically, the candidate structure adoption unit 111 provides a monotonically decreasing function P1(d) where d = |F-f1| - |F-f0| and estimates the probability p1 = P1(d). The monotonically decreasing function P1(d) corresponds to the "first function" in the present invention (a monotonically decreasing function of the absolute value of the difference between the physical property value of the candidate structure and the target physical property value, and the absolute value of the difference between the physical property value of the chemical structure and the target physical property value), and the probability p1 corresponds to the "first adoption probability" in the present invention.

[0219] Various functions can be used as the monotonically decreasing function P1(d), but for example, the function represented by the following equation (2) can be used. σ is a hyperparameter, and the degree of monotonically decreasing function can be adjusted by changing the value of σ. The value of the parameter can be changed by a user input via the operating unit 400.

[0220] [Formula 2]

[0221] P1(d)=exp[-d / σ] (2)

[0222] In the case of n targets (n physical property values ​​are input in step S1010 ), an index representing each target can be set to i, and functions represented by the following equations (3) and (4) can be used, for example.

[0223] [Formula 3]

[0224] w i =exp[-d i / σ i ] (3)

[0225] [Formula 4]

[0226]

[0227] The functions expressed in equations (3) and (4) serve as a benchmark for "using a structural change if there is a single physical property value close to the target," but various other functions can also be used. Furthermore, to put it more simply, a method can be considered in which the physical property values ​​of n targets are considered as n-dimensional vectors ff and FF, and d = |FF-ff1| - |FF-ff0| is estimated from the Euclid distance |FF-ff| to solve the problem as a single target (ff, ff0, ff1, FF are vectors). When adopting this approach, it is preferable to calculate the mean and variance of each physical property value based on existing data in advance, perform normalization, and then calculate the distance.

[0228] After calculating the probability p1, the candidate structure adoption unit 111 uses an appropriately generated random number to enter step S1070 with probability p1 and adopt the structural change, and enters step S1055 with probability (1-p1). That is, the candidate structure adoption unit 111 adopts the candidate structure with the first adoption probability (first adoption processing). This probability processing (even when the physical characteristic value is far from the target value, the structural change is adopted with probability p1) is to prevent falling into a local minimum. The local minimum is "a state in which the physical characteristic value is far from the target value no matter how the structure is changed." In order to escape from the local minimum and reach the global minimum, a structural change in which the physical characteristic value is far from the target value must be passed. This path can be ensured by the above-mentioned probability processing.

[0229] <Second Adoption Process>

[0230] If the candidate structure resulting from the first adoption process is not adopted in step S1050 (probability (1-p1)), the candidate structure adoption unit 111 performs a second adoption process to determine whether to adopt the candidate structure based on whether the structural diversity of the structure group consisting of the chemical structure and the candidate structure has increased due to the change in chemical structure (steps S1055, S1060, and S1070). The second adoption process is described below. Furthermore, the index representing the structure is set to j, and the structure group is represented as S = {sj}. The function that provides the structural diversity of the structure group S is expressed as V(S). The greater the structural diversity, the larger the value of V(S) is adopted.

[0231] <When multiple initial structures are provided>

[0232] Given N (>1) initial structures, the structural change of the kth chemical structure among the N chemical structures is considered for adoption or rejection. In the mth trial, the structural group Sk = {s(m-1)0, s(m-1)1, ..., smk, ..., s(m-1)N} of the kth chemical structure after structural change is defined based on the structural group Sm-1 = {s(m-1)j} before the structural change (m-1 rounds) and the structural group Sm = {smj} after the structural change (m-1 rounds), and dv = V(Sk) - V(Sm-1) is estimated. In other words, dv represents the increase or decrease in structural diversity due to the structural change. If dv ≥ 0 (diversity has increased due to the kth structural change; "Yes" in step S1055), a monotonically increasing function P2(dv) relative to dv (the increase or decrease in structural diversity) is provided, and the probability p2 = P2(dv) is calculated (step S1060: second adoption process). Then, using a suitably generated random number, the process proceeds to step S1070 with probability p2 (adopting the structural change; second adoption process), and proceeds to step S1080 with probability (1-p2) (discarding the structural change and returning to the original structure; discarding the process). The monotonically increasing function P2(dv) corresponds to the "second function" in the present invention, and the probability p2 corresponds to the "second adoption probability" in the present invention.

[0233] When structural diversity increases, the above-described probabilistic processing (calculating candidate structures using probability p2 calculated using the monotonically increasing function P2(dv)) is performed because, when the rule "structural changes must be adopted when structural diversity increases," even if the physical property values ​​are far from the target values, the frequency of adopting structural changes becomes excessively high, resulting in a situation where convergence toward the target physical property values ​​is slowed. This probabilistic processing accelerates convergence of physical property values ​​and allows for efficient searches for compound structures.

[0234] If dv<0 calculated in step S1060 (diversity decreases; "No" in step S1055), the process proceeds to step S1080 (abandoning the structural change and returning to the original structure; abandoning the process).

[0235] <When providing one initial structure>

[0236] In addition, when there is only one initial structure, the index representing the experiment is set to t, and the structure group Sprev = {st-1, st-2, ..., st-m} obtained through the past m experiments and the structure group Scurr = {st, st-1, ..., st-(m-1)} obtained by additionally considering the adopted or abandoned structure st are considered, and dv = V(Scurr) - V(Sprev) is calculated. In the same way as in the case where there are multiple initial structures, the probability p2 is calculated by the monotonically increasing function P2(dv) (step S1060: second adoption processing).

[0237] Function that provides structural diversity of a structural group

[0238] As the above-mentioned "function that provides the structural diversity of a structure group," for example, the following definition based on the Tanimoto coefficient (one of the indicators indicating the similarity of compounds) can be considered (various other definitions are also possible). Specifically, if the structure s is represented by a fingerprint (converted into a fixed-length vector according to a certain rule for the compound, and various generation methods are known) of a bit string (a sequence of 0s or 1s) is denoted as Fs, the definition of the Tanimoto coefficient is expressed by the following formula (5).

[0239] [Formula 5]

[0240]

[0241] Here, |Fs| is the number of bits that are 1 in Fs, and |Fs∩Fs'| is the number of bits that are 1 in both Fs and Fs'. When Fs and Fs' are completely identical, Ts,s' is 1, and when Fs and Fs' are completely inconsistent, Ts,s' is 0. Therefore, Ts,s' is an indicator of the similarity between structures s and s'. Since the desired result is dissimilarity, the dissimilarity between structures s and s', vs,s', is defined by the following equation (6).

[0242] [Formula 6]

[0243] v s,s′ =1-T s,s′ (6)

[0244] The dissimilarity vs,s′ can be used to define the dissimilarity of the structure group S (ie, the structural diversity of the structure group) by the following equation (7).

[0245] [Formula 7]

[0246]

[0247] V(S) takes values ​​ranging from 0 to 1, and a larger value indicates a higher structural diversity of the structural group.

[0248] Furthermore, as a monotonically increasing function P2(dv) relative to the increase or decrease in structural diversity dv, for example, the function represented by the following equation (8) can be used. σv and Cv are hyperparameters, and the degree of monotonous increase can be adjusted by changing their values. The values ​​of these parameters can be changed by user input via operating unit 400.

[0249] [Formula 8]

[0250] P2(d v )=C v (1-exp[-d v / σ v ]) (8)

[0251] As is clear from the functional form, P2 becomes Cv in the limit of dv→∞. Therefore, Cv is the probability of adopting a structural change when diversity is sufficiently increased.

[0252] <Duplication of processing>

[0253] The first adoption process, the second adoption process, and the abandonment process are performed for each of the provided initial structures. When the above processes are completed for all chemical structures, one experiment is completed.

[0254] After a candidate structure is adopted or rejected as a result of the first adoption process, the second adoption process, and the rejection process, the control unit 113 determines whether the termination condition is satisfied (step S1090: control process). For example, the termination condition can be determined to be satisfied when the number of chemical structure variations (number of trials) reaches a specified number and / or when the physical property values ​​of the candidate structure reach target values. If multiple chemical structures and / or physical property values ​​are calculated, the calculation can be terminated as soon as one chemical structure and / or physical property value reaches the target value, or the experiment can be repeated until all structures and / or physical property values ​​reach the target value. The control unit 113 repeats the processes from steps S1020 to S1080 (the input process, the candidate structure acquisition process, the physical property value calculation process, and the candidate structure adoption process) until the termination condition is satisfied ("No" in step S1090). When the termination condition is satisfied ("Yes" in step S1090), the compound creation method ends (step S1100).

[0255] <The effect of creating a three-dimensional structure based on structural diversity>

[0256] As described above, the compound creation method of creating a three-dimensional structure based on structural diversity can promote escape from local minimum values ​​and accelerate the convergence of physical property values, thereby effectively searching for the structure of a compound having desired physical property values.

[0257] <Specific physical property values, characteristic quantities, etc.>

[0258] Specific physical property values ​​and characteristic quantities in the above-described method of creating a three-dimensional structure (creating a compound) based on structural diversity will be described.

[0259] When creating a three-dimensional structure of a target compound bonded to a target protein (method 23), the feature quantity (first feature quantity, AAS descriptor) calculated using the feature quantity calculation method involved in method 5 is a "physical property value", and the first feature quantity for the ligand is a "target value of the physical property value".

[0260] When creating a three-dimensional structure of a target compound bonded to a target protein (method 24), the feature quantity (first invariant feature quantity, invariant AAS descriptor) calculated using the feature quantity calculation method involved in method 6 is a "physical property value", and the first invariant feature quantity for the ligand is a "target value of the physical property value".

[0261] In the case of creating a three-dimensional structure of a target compound bonded to a target protein (method 25), the feature quantity (second feature quantity, AAS descriptor) calculated using the feature quantity calculation method involved in method 8 is a "physical property value", and the second feature quantity of the pocket structure confirmed to be bonded to the pocket serving as the active site of the target protein is a "target value of the physical property value".

[0262] When creating a three-dimensional structure of a target compound bonded to a target protein (method 26), the feature quantity (second invariant feature quantity, invariant AAS descriptor) calculated using the feature quantity calculation method involved in method 9 is a "physical property value", and the second invariant feature quantity of the pocket structure confirmed to be bonded to the pocket serving as the active site of the target protein is a "target value of the physical property value".

[0263] In the case of creating a three-dimensional structure of a target compound bonded to a target biopolymer other than a protein (method 27), the feature quantity (third feature quantity, AAS descriptor) calculated using the feature quantity calculation method involved in method 11 is a "physical property value", and the third feature quantity of the bonded compound that is a compound confirmed to be bonded to a target biopolymer other than a protein is a "target value of the physical property value".

[0264] Effects of the Compound Creation Device

[0265] As described above, in the compound creation device 20 involved in the second embodiment, it is possible to use the feature quantities (AAS descriptors, invariant AAS descriptors) calculated by the feature quantity calculation method involved in the present invention, and to effectively create the three-dimensional structure of the pharmaceutical candidate compound through the compound creation method involved in the present invention (and the compound creation program that enables a computer to execute the method).

[0266] <Third embodiment>

[0267] The first embodiment described above is a method for calculating characteristic quantities and screening based on the characteristic quantities, and the second embodiment is a method for calculating characteristic quantities and creating a three-dimensional structure of a target compound based on the characteristic quantities. However, in addition to calculating characteristic quantities, both screening and creating a three-dimensional structure of a target compound can also be performed. Therefore, in the pharmaceutical candidate compound search device 30 involved in the third embodiment (characteristic quantity calculation device, screening device, compound creation device; reference Figure 27 ), with Figure 27 The processing unit 102 shown is replaced by Figure 1 The processing unit 100 of the screening device 10 shown in FIG. Figure 18 The processing unit 101 of the compound creation apparatus 20 shown in FIG. Figure 25 The processing unit 103 shown. Figure 28 As shown, the processing unit 102 includes a communication control unit 110A (communication control unit), a feature value calculation unit 120 (feature value calculation unit), a similarity calculation unit 130 (similarity calculation unit), a generator construction unit 132 (generator construction unit), a compound extraction unit 140 (compound extraction unit), a compound stereo structure generation unit 142 (compound stereo structure generation unit), a display control unit 150 (display control unit), a CPU 160, a ROM 170, and a RAM 180, and is capable of calculating feature values, screening, and creating stereo structures of compounds. Furthermore, the drug candidate compound search device 30 stores the information required for these processes and the results of the processes in the storage unit 202. Specifically, as Figure 29 As shown, the information stored in the storage unit 200 and the storage unit 201 (reference Figure 3 、 Figure 20) are stored in the storage unit 202 accordingly.

[0268] Other requirements and Figure 1 The screening device 10 shown, Figure 18 The compound creation device 20 shown in FIG. 1 is the same as that shown in FIG. 1 , and therefore the same reference numerals are given and detailed description thereof is omitted. In addition, when creating a compound based on structural diversity, the processing unit 102 has a corresponding processing unit 103 (refer to FIG. 1 ). Figure 25 ) and stores information corresponding to the creation of compounds based on structural diversity (physical property values, target values, created three-dimensional structures, etc.) in the storage unit 202.

[0269] Through the above-mentioned structure, in the drug candidate compound search device 30 involved in the third embodiment, similar to the screening device 10 and the compound creation device 20, it is possible to calculate the characteristic quantities that accurately represent the chemical properties of the object structure, effectively screen drug candidate compounds, and effectively create the three-dimensional structure of drug candidate compounds.

[0270] As mentioned above, although embodiment of this invention was described, this invention is not limited to the said form, As shown below, various deformation|transformation is possible.

[0271] <Target of Treatable Medicine>

[0272] In the present invention, as the target of medicine, in addition to protein, DNA (Deoxyribonucleic Acid), RNA (Ribonucleic Acid), cell membrane, polysaccharide can also be used. Among them, the probe (amino acid) when it is necessary to change the protein is changed to other substances. Specifically, in the case of DNA, amino acid is changed to nucleic acid base, in the case of RNA, amino acid is changed to nucleic acid base, in the case of cell membrane, amino acid is changed to lipid molecule, in the case of polysaccharide, amino acid is changed to monosaccharide molecule. Below, the reason why DNA, RNA, cell membrane, polysaccharide can also be processed by this change in the present invention is explained.

[0273] Protein, DNA, RNA, cell membrane, polysaccharide are collectively referred to as biopolymers, and are composed of inherent components. Specifically, the components of protein are amino acids, the components of DNA are nucleic acid bases, the components of RNA are nucleic acid bases in the same way, the components of cell membrane are lipid molecules, and the components of polysaccharides are monosaccharide molecules. Similar to proteins, there are pockets as active sites in DNA, RNA, cell membranes, and polysaccharides. Therefore, when the target of medicine is DNA, RNA, cell membrane, or polysaccharide, the present invention can also be dealt with by changing the amino acids in the embodiment shown in the case of protein to components of the target. In addition, when quantifying the degree of aggregation of amino acids, nucleic acid bases, lipid molecules, and monosaccharide molecules around a compound or pocket structure, water can also be considered.

[0274] <Processable Activities>

[0275] In the present invention, in addition to common activities such as "activity of the compound on a target biomolecule alone", "activity of the compound on cells of a complex composed of other biomolecules in addition to the target biomolecule" can also be used.

[0276] Explanation of symbols

[0277] 10 Screening device

[0278] 20 Compound Creation Device

[0279] 30 Pharmaceutical candidate compound search device

[0280] 100 Processing Department

[0281] 101 Processing Department

[0282] 102 Processing Department

[0283] 103 Processing Department

[0284] 105 Input unit

[0285] 107 Candidate Structure Acquisition Department

[0286] 109 Physical property value calculation section

[0287] 110 Information Input Unit

[0288] 110A Communication Control Unit

[0289] 111 Candidate Structure Adoption Department

[0290] 113 Control Department

[0291] 115 Display control unit

[0292] 120 Feature Calculation Unit

[0293] 121 CPU

[0294] 123 ROM

[0295] 125 RAM

[0296] 130 Similarity Calculation Unit

[0297] 132 Generator Construction Department

[0298] 140 Compound Extraction Department

[0299] 142 Compound Stereostructure Generation Section

[0300] 150 Display control unit

[0301] 160 CPU

[0302] 170 ROM

[0303] 180 RAM

[0304] 200 Storage Department

[0305] 201 Storage Department

[0306] 202 Storage Department

[0307] 210 Structural Information

[0308] 220 three-dimensional structure information

[0309] 230 AAS descriptor

[0310] 240 Invariant Quantization AAS Descriptor

[0311] 250 similarity information

[0312] 260 compound extraction results

[0313] 270 Stereostructure Generation Results

[0314] 300 Display Unit

[0315] 310 Monitor

[0316] 400 Operation Department

[0317] 410 Keyboard

[0318] 420 Mouse

[0319] 500 External Server

[0320] 510 External Database

[0321] 900 compounds

[0322] 902 Probe

[0323] 910 compound

[0324] 912 structural formula

[0325] 914 AAS Descriptor

[0326] 916 Generator

[0327] NW Network

[0328] PO pocket

[0329] PS three-dimensional structure

[0330] S100~S112 Steps of the feature quantity calculation method

[0331] S200~S206 Steps of feature quantity calculation method

[0332] S300~S304 Steps of compound extraction method

[0333] S400~S404 Steps of compound extraction method

[0334] S500~S504 Steps of compound creation method

[0335] S600~S604 Steps of compound creation method

[0336] S1010~S1100 Steps of compound creation method

[0337] TP target protein

[0338] B Collision diameter

[0339] Rmin closest distance

[0340] θ a Scattering angle

Claims

1. A method for calculating a feature quantity by a computer, the computer executing the following steps: An object structure specifying step of specifying an object structure composed of a plurality of unit structures having chemical properties; a three-dimensional structure acquisition step of acquiring a three-dimensional structure based on the plurality of unit structures for the target structure; and a probe feature value calculation step of calculating a feature value representing a cross-sectional area of ​​one or more probes relative to the target structure; The probe is a structure having real charges and having a plurality of points generating van der Waals forces that are arranged separately.

2. The feature quantity calculation method according to claim 1, wherein: In the probe characteristic value calculation step, a differential scattering cross-sectional area is calculated as the characteristic value, or the differential scattering cross-sectional area is calculated from the closest distance and the scattering angle.

3. The feature quantity calculation method according to claim 1 or 2, wherein: In the probe feature value calculation step, data obtained by the following formula (1) is calculated as the feature value, dσ / dΩ(E,b,a)(1) Here, E is an independent variable for determining the incident energy of the probe, b is an independent variable for determining the collision diameter of the probe, and a is an independent variable for determining the type of the probe.

4. The feature quantity calculation method according to claim 1 or 2, wherein: In the three-dimensional structure acquisition step, the acquisition is performed by generating a three-dimensional structure of the designated target structure.

5. The feature quantity calculation method according to claim 1 or 2, wherein: In the target structure specifying step, a compound is specified as the target structure, In the stereostructure acquisition step, the stereostructure of the compound based on a plurality of atoms as the plurality of unit structures is acquired. In the probe feature value calculation step, a first feature value is calculated using amino acids as the probe for the compound acquired in the three-dimensional structure acquisition step. 6 . The feature value calculation method according to claim 5 , further comprising an invariant quantization step of calculating a first invariant feature value by quantizing the first feature value with respect to rotation invariance of the compound.

7. The feature quantity calculation method according to claim 6, wherein: In the probe characteristic value calculation step, the first characteristic value is calculated for two different amino acids. In the invariant step, the first invariant feature value is calculated using the first feature values ​​for the two different amino acids.

8. The feature quantity calculation method according to claim 1 or 2, wherein: In the target structure specifying step, a pocket structure that bonds to a pocket that is an active site of a target protein is specified as the target structure. In the three-dimensional structure acquisition step, the three-dimensional structure of the pocket structure is acquired based on a plurality of virtual spheres. In the probe feature value calculation step, a second feature value is calculated using amino acids as the probe for the pocket structure acquired in the three-dimensional structure acquisition step.

9. The feature quantity calculation method according to claim 8, further comprising an invariant quantization step of calculating a second invariant feature quantity by quantizing the second feature quantity to be rotationally invariant with respect to the pocket structure.

10. The feature quantity calculation method according to claim 9, wherein: In the probe characteristic value calculation step, the second characteristic value is calculated for two different amino acids. In the invariant step, the second invariant feature value is calculated using the second feature values ​​for the two different amino acids.

11. The feature quantity calculation method according to claim 1 or 2, wherein: In the target structure specifying step, a compound is specified as the target structure, In the stereostructure acquisition step, a stereostructure of the compound based on a plurality of atoms is generated. In the probe characteristic quantity calculation process, for the stereostructure of the compound obtained in the stereostructure acquisition process, one or more nucleic acid bases, one or more lipid molecules, one or more monosaccharide molecules, water, and one or more ions are used as the probe to calculate the third characteristic quantity. 12 . The feature value calculation method according to claim 11 , further comprising a step of quantizing the third feature value so as to be rotationally invariant to the compound to calculate a third quantized feature value.

13. A computer-based screening method for extracting a first target compound that binds to a target protein and / or a second target compound that does not bind to the target protein from a plurality of compounds, wherein the computer performs the following steps: a storing step of associating, for each of the plurality of compounds, a three-dimensional structure of the compound based on a plurality of atoms with the first feature value of the three-dimensional structure of the compound calculated using the feature value calculation method according to claim 5 and storing the associated feature value; a screening feature value calculation step of calculating the first feature value for the ligand as a compound confirmed to be bound to the target protein; a similarity calculation step of calculating similarities between the first feature quantity for the plurality of compounds and the first feature quantity for the ligand; and The compound extraction step extracts the first target compound and / or the second target compound from the plurality of compounds based on the similarity.

14. A computer-based screening method for extracting a first target compound that binds to a target protein and / or a second target compound that does not bind to the target protein from a plurality of compounds, wherein the computer performs the following steps: a storing step of associating, for each of the plurality of compounds, a three-dimensional structure of the compound based on a plurality of atoms with the first invariant quantized feature value for the three-dimensional structure of the compound calculated using the feature value calculation method according to claim 6 and storing the associated feature value; a screening feature value calculation step of calculating the first invariant feature value for the ligand as a compound confirmed to be bound to the target protein; a similarity calculation step of calculating similarities between the first invariant quantized feature quantity for the plurality of compounds and the first invariant quantized feature quantity for the ligand; and The compound extraction step extracts the first target compound and / or the second target compound from the plurality of compounds based on the similarity.

15. A computer-based screening method for extracting a first target compound that binds to a target protein and / or a second target compound that does not bind to the target protein from a plurality of compounds, wherein the computer performs the following steps: a storing step of associating, for each of the plurality of compounds, a three-dimensional structure of the compound based on a plurality of atoms with the first feature quantity calculated using the feature quantity calculation method according to claim 5 and storing the associated features; a screening feature quantity calculation step of calculating the second feature quantity for the pocket structure of the target protein using the feature quantity calculation method according to claim 8; a similarity calculation step of calculating similarities between the first feature quantity for the plurality of compounds and the second feature quantity for the pocket structure; and The compound extraction step extracts the first target compound and / or the second target compound from the plurality of compounds based on the similarity.

16. A computer-based screening method for extracting a first target compound that binds to a target protein and / or a second target compound that does not bind to the target protein from a plurality of compounds, wherein the computer performs the following steps: a storing step of associating, for each of the plurality of compounds, a three-dimensional structure of the compound based on a plurality of atoms with the first invariant feature value calculated using the feature value calculation method according to claim 6 and storing the associated features; a screening feature quantity calculation step of calculating the second invariant feature quantity for the pocket structure of the target protein using the feature quantity calculation method according to claim 9; a similarity calculation step of calculating the similarity between the first invariant quantized feature value for the plurality of compounds and the second invariant quantized feature value for the pocket structure; and The compound extraction step extracts the first target compound and / or the second target compound from the plurality of compounds based on the similarity.

17. A computer-based screening method for extracting a target compound that binds to a target biopolymer other than a protein from a plurality of compounds, wherein the computer performs the following steps: a storing step of associating, for each of the plurality of compounds, a three-dimensional structure of the compound based on a plurality of atoms with the third feature value of the three-dimensional structure of the compound calculated using the feature value calculation method according to claim 11 and storing the associated feature value; a feature quantity calculation step of calculating the third feature quantity for a bonded compound that is a compound confirmed to be bonded to the target biopolymer other than the protein; a similarity calculation step of calculating similarities between the third feature quantity for the plurality of compounds and the third feature quantity for the bonded compound; and The compound extraction step extracts the first target compound and / or the second target compound from the plurality of compounds based on the similarity.

18. A computer-generated compound creation method, wherein the computer creates a three-dimensional structure of a target compound bound to a target protein from a plurality of compounds, wherein the computer performs the following steps: a storing step of associating, for each of a plurality of compounds, a three-dimensional structure of the compound based on a plurality of atoms with the first feature quantity calculated using the feature quantity calculation method according to claim 5 and storing the associated features; creating a feature quantity calculation step of calculating the first feature quantity for a ligand that is a compound confirmed to be bound to the target protein; a generator construction step of constructing a generator by machine learning using the three-dimensional structures of the plurality of compounds as training data and the first feature value as an explanatory variable; and The compound stereostructure generating step generates the stereostructure of the target compound from the first characteristic value of the ligand using the generator.

19. A computer-generated compound creation method, wherein the computer creates a three-dimensional structure of a target compound bound to a target protein from a plurality of compounds, wherein the computer performs the following steps: a storing step of associating, for each of a plurality of compounds, a three-dimensional structure of the compound based on a plurality of atoms with the first invariant feature value calculated using the feature value calculation method according to claim 6 and storing the associated features; creating a feature quantity calculation step of calculating the first invariant feature quantity for a ligand that is a compound confirmed to be bound to the target protein; a generator construction step of constructing a generator by machine learning using the three-dimensional structures of the plurality of compounds as training data and the first invariant quantitative feature as an explanatory variable; and The compound stereostructure generating step generates the stereostructure of the target compound from the first invariant feature value of the ligand using the generator.

20. A computer-based method for creating a compound, wherein the computer creates a three-dimensional structure of a target compound bound to a target protein from a plurality of compounds, wherein the computer performs the following steps: a storing step of associating, for each of the plurality of compounds, a three-dimensional structure of the compound based on a plurality of atoms with the first feature quantity calculated using the feature quantity calculation method according to claim 5 and storing the associated features; creating a feature quantity calculation step for calculating the second feature quantity using the feature quantity calculation method according to claim 8 for the pocket structure of the target protein; a generator construction step of constructing a generator by machine learning using the three-dimensional structures of the plurality of compounds as training data and the first feature value as an explanatory variable; and The compound three-dimensional structure generating step generates the three-dimensional structure of the target compound from the second characteristic value of the pocket structure using the generator.

21. A computer-generated compound creation method, wherein the computer creates a three-dimensional structure of a target compound bound to a target protein from a plurality of compounds, wherein the computer performs the following steps: a storing step of associating, for each of the plurality of compounds, a three-dimensional structure of the compound based on a plurality of atoms with the first invariant feature value calculated using the feature value calculation method according to claim 6 and storing the associated features; creating a feature quantity calculation step for calculating the second invariant feature quantity using the feature quantity calculation method according to claim 9 for the pocket structure of the target protein; a generator construction step of constructing a generator by machine learning using the stereostructures of the plurality of compounds as training data and the first invariant quantitative feature as an explanatory variable; and The compound three-dimensional structure generating step generates the three-dimensional structure of the target compound from the second invariant feature value of the pocket structure using the generator.

22. A computer-generated compound creation method for creating a three-dimensional structure of a target compound bound to a target biopolymer other than a protein from a plurality of compounds, the computer performing the following steps: a storing step of associating, for each of a plurality of compounds, a three-dimensional structure of the compound based on a plurality of atoms with the third feature quantity calculated using the feature quantity calculation method according to claim 11 and storing the associated features; creating a feature quantity calculation step of calculating the third feature quantity for a bonded compound that is a compound confirmed to be bonded to the target biopolymer other than the protein; a generator construction step of constructing a generator by machine learning using the three-dimensional structures of the plurality of compounds as training data and the third feature value as an explanatory variable; and The compound stereostructure generating step generates the stereostructure of the target compound from the third characteristic value of the bonded compound using the generator.

23. A computer-generated method for creating a compound, wherein the computer generates a three-dimensional structure of a target compound bound to a target protein, wherein the computer performs the following steps: an input step of inputting a chemical structure of one or more compounds, the first characteristic quantity for the chemical structure calculated using the characteristic quantity calculation method according to claim 5, and the first characteristic quantity for a ligand as a compound confirmed to be bonded to the target compound, which is a target value for the first characteristic quantity; a candidate structure acquisition step, which changes the chemical structure and acquires a candidate structure; Creating a feature quantity calculation step for calculating the first feature quantity using the feature quantity calculation method according to claim 5 for the candidate structure; a candidate structure adoption step for adopting or rejecting the candidate structure, wherein a first adoption process is performed to determine whether to adopt the candidate structure based on whether the first feature value of the candidate structure approaches the target value due to the change in the chemical structure; if the candidate structure is not adopted by the first adoption process, a second adoption process is performed to determine whether to adopt the candidate structure based on whether the structural diversity of the structure group consisting of the chemical structure and the candidate structure increases due to the change in the chemical structure; and if the candidate structure is not adopted by the first and second adoption processes, a rejection process is performed to reject the change in the chemical structure and return to the chemical structure before the change; and The control step repeats the processing in the input step, the candidate structure acquisition step, the creation feature value calculation step, and the candidate structure adoption step until an end condition is satisfied.

24. A computer-generated method for creating a compound, wherein the method creates a three-dimensional structure of a target compound bound to a target protein, the computer performing the following steps: an input step of inputting a chemical structure of one or more compounds, the first invariant feature quantity for the chemical structure calculated using the feature quantity calculation method according to claim 6, and the first invariant feature quantity for a ligand of a compound confirmed to be bonded to the target compound, which serves as a target value for the first invariant feature quantity; a candidate structure acquisition step, which changes the chemical structure and acquires a candidate structure; Creating a feature quantity calculation step for calculating the first invariant feature quantity using the feature quantity calculation method according to claim 6 for the candidate structure; a candidate structure adoption step for adopting or rejecting the candidate structure, wherein a first adoption process is performed to determine whether to adopt the candidate structure based on whether the first invariant quantitative feature value of the candidate structure approaches the target value due to the change in the chemical structure; if the candidate structure is not adopted by the first adoption process, a second adoption process is performed to determine whether to adopt the candidate structure based on whether the structural diversity of the structure group consisting of the chemical structure and the candidate structure increases due to the change in the chemical structure; and if the candidate structure is not adopted by the first and second adoption processes, a rejection process is performed to abandon the change in the chemical structure and return to the chemical structure before the change; and The control step repeats the processing in the input step, the candidate structure acquisition step, the creation feature value calculation step, and the candidate structure adoption step until an end condition is satisfied.

25. A computer-generated method for creating a compound, wherein the method creates a three-dimensional structure of a target compound bound to a target protein, the computer performing the following steps: an input step of inputting a chemical structure of one or more compounds, the second characteristic quantity for the chemical structure calculated using the characteristic quantity calculation method according to claim 8, and the second characteristic quantity for a pocket structure confirmed to be bonded to a pocket serving as an active site of the target protein as a target value for the second characteristic quantity; a candidate structure acquisition step, which changes the chemical structure and acquires a candidate structure; Creating a feature quantity calculation step for calculating the second feature quantity using the feature quantity calculation method according to claim 8 for the candidate structure; a candidate structure adoption step for adopting or rejecting the candidate structure, wherein a first adoption process is performed to determine whether to adopt the candidate structure based on whether the second feature value of the candidate structure approaches the target value due to the change in the chemical structure; if the candidate structure is not adopted by the first adoption process, a second adoption process is performed to determine whether to adopt the candidate structure based on whether the structural diversity of the structure group consisting of the chemical structure and the candidate structure increases due to the change in the chemical structure; and if the candidate structure is not adopted by the first and second adoption processes, a rejection process is performed to reject the change in the chemical structure and return to the chemical structure before the change; and The control step repeats the processing in the input step, the candidate structure acquisition step, the creation feature value calculation step, and the candidate structure adoption step until an end condition is satisfied.

26. A computer-generated method for creating a compound, wherein the method creates a three-dimensional structure of a target compound bound to a target protein, the computer performing the following steps: an input step of inputting a chemical structure of one or more compounds, the second invariant feature quantity for the chemical structure calculated using the feature quantity calculation method according to claim 9, and the second invariant feature quantity for a pocket structure confirmed to be bonded to a pocket serving as an active site of the target protein as a target value for the second invariant feature quantity; a candidate structure acquisition step, which changes the chemical structure and acquires a candidate structure; creating a feature quantity calculation step for calculating the second invariant feature quantity using the feature quantity calculation method according to claim 9 for the candidate structure; a candidate structure adoption step for adopting or rejecting the candidate structure, wherein a first adoption process is performed to determine whether to adopt the candidate structure based on whether the second invariant quantitative feature value of the candidate structure approaches the target value due to the change in the chemical structure; if the candidate structure is not adopted by the first adoption process, a second adoption process is performed to determine whether to adopt the candidate structure based on whether the structural diversity of the structure group consisting of the chemical structure and the candidate structure increases due to the change in the chemical structure; and if the candidate structure is not adopted by the first and second adoption processes, a rejection process is performed to abandon the change in the chemical structure and return to the chemical structure before the change; and The control step repeats the processing in the input step, the candidate structure acquisition step, the creation feature value calculation step, and the candidate structure adoption step until an end condition is satisfied.

27. A computer-generated method for creating a compound, the method comprising: generating a three-dimensional structure of a target compound bonded to a target biopolymer other than a protein, the computer performing the following steps: an input step of inputting a chemical structure of one or more compounds, the third feature quantity for the chemical structure calculated using the feature quantity calculation method according to claim 11, and the third feature quantity for a bonded compound confirmed to be bonded to the target biopolymer other than the protein, as a target value for the third feature quantity; a candidate structure acquisition step, which changes the chemical structure and acquires a candidate structure; Creating a feature quantity calculation step for calculating the third feature quantity using the feature quantity calculation method according to claim 11 for the candidate structure; a candidate structure adoption step for adopting or rejecting the candidate structure, wherein a first adoption process is performed to determine whether to adopt the candidate structure based on whether the third feature value of the candidate structure approaches the target value due to the change in the chemical structure; if the candidate structure is not adopted by the first adoption process, a second adoption process is performed to determine whether to adopt the candidate structure based on whether the structural diversity of the structure group consisting of the chemical structure and the candidate structure increases due to the change in the chemical structure; and if the candidate structure is not adopted by the first and second adoption processes, a rejection process is performed to reject the change in the chemical structure and return to the chemical structure before the change; and The control step repeats the processing in the input step, the candidate structure acquisition step, the creation feature value calculation step, and the candidate structure adoption step until an end condition is satisfied.

Citation Information

Patent Citations

  • Semiconductor integrated circuit device

    JP1984046045A

  • Systems and methods for applying a convolutional network to spatial data

    US9373059B1

  • Feature quantity calculating method, feature quantity calculating program and feature quantity calculating device, screening method, screening program and screening device, compound creating method, compound creating program and compound creating device

    CN111279419A

  • Probe, probe design device, and probe design program

    JP2011004621A