Molecule analysis system, program, and molecule analysis method
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2024-03-22
- Publication Date
- 2025-09-25
AI Technical Summary
Existing molecular analysis methods provide only chemical structural information and fail to predict in vivo behavior of molecules.
A molecular analysis system and method that utilizes quantum chemical calculations to acquire three-dimensional structural information, sets information acquisition points within or around the molecule, and acquires data on pre-set parameters at these points using machine learning models.
Enables accurate prediction of molecular behavior in vivo by capturing spatial information, supporting applications in drug discovery and materials science with reduced computational costs.
Abstract
Description
Molecular analysis system, program, and molecular analysis method
[0001] The present invention relates to a molecular analysis system, a program, and a molecular analysis method.
[0002] Conventionally, methods have been known for analyzing molecules by acquiring one-dimensional structural information such as molecular formula, molecular weight, number of atoms, number of functional groups, number of hydrogen bonds, etc., acquiring two-dimensional structural information such as chemical bond information of the molecule and adjacent functional groups, and acquiring three-dimensional structural information such as three-dimensional information of the molecule, molecular surface area information, electron density, molecular orbital, etc. Furthermore, Patent Document 1 discloses a method for generating a generative model using a machine learning technique with input information being the chemical characteristics of a compound and a training set including biological or chemical information related to the compound.
[0003] Special table 2019-502988 publication
[0004] The above-mentioned molecular analysis methods that obtain one-dimensional, two-dimensional, and three-dimensional structural information have the problem that they are meaningful only from a chemical perspective and do not lead to predictions of behavior in vivo, for example.
[0005] The present disclosure has been made in consideration of these points, and aims to provide a new molecular analysis system, program, and information processing method for extracting spatial information of molecules.
[0006] The molecular analysis system of the present disclosure is characterized by comprising: a first acquisition means for acquiring information on the three-dimensional structure of a molecule through quantum chemical calculations; a setting means for setting a plurality of information acquisition points inside or around the three-dimensional structure of the molecule; and a second acquisition means for acquiring data on one or more pre-set parameters at each of the information acquisition points set by the setting means.
[0007] In the molecular analysis system of the present disclosure, the setting means may divide the three-dimensional structure of the molecule acquired by the first acquisition means into predetermined cubes, and set an information acquisition point for each cube.
[0008] Furthermore, in the molecular analysis system of the present disclosure, the setting means may set an axis connecting the start point and end point of the three-dimensional structure of the molecule acquired by the first acquisition means, and when arranging information acquisition points around the axis, may set a virtual three-dimensional curve from the set start point to the set end point, and set multiple information acquisition points at intervals on the set virtual three-dimensional curve.
[0009] Furthermore, in the molecular analysis system of the present disclosure, the setting means may set a longitudinal direction in the three-dimensional structure of the molecule acquired by the first acquisition means, and set the start point and the end point near the ends along the longitudinal direction.
[0010] In the molecular analysis system of the present disclosure, the setting means may set the start point and the end point so that the linear distance between them is greatest.
[0011] In the molecular analysis system of the present disclosure, the setting means may set the virtual three-dimensional curve so as to cover a part or all of the three-dimensional structure of the molecule acquired by the first acquisition means.
[0012] In the molecular analysis system of the present disclosure, the setting means may set a plurality of the information acquisition points at equal intervals on the virtual three-dimensional curve.
[0013] The molecular analysis system of the present disclosure may further include a receiving means for receiving information relating to a chemical formula of a molecule, and a generating means for generating a three-dimensional structure of the molecule by quantum chemical calculation from the chemical formula of the molecule received by the receiving means.
[0014] In addition, the molecular analysis system of the present disclosure may further include a determination means for determining the position of the information acquisition point from the three-dimensional structure of the molecule acquired by the first acquisition means using a first model generated by machine learning using training data including the three-dimensional structure of the molecule, the position of the information acquisition point, and a score related to the molecular analysis result, and the setting means may set the position of the information acquisition point determined by the determination means.
[0015] In addition, the molecular analysis system of the present disclosure may further include a parameter type setting means for setting the type of a parameter from the three-dimensional structure of the molecule acquired by the first acquisition means or the position of each information acquisition point set by the setting means, using a second model generated by machine learning using training data including the three-dimensional structure of the molecule, the position of the information acquisition point, the type of parameter, and a score related to the analysis result of the molecule, and the second acquisition means may acquire data at each information acquisition point for one or more parameters set by the parameter type setting means.
[0016] The program of the present disclosure is a program that functions as a first acquisition means, a setting means, and a second acquisition means when executed by a computer, wherein the first acquisition means acquires information on the three-dimensional structure of a molecule through quantum chemical calculations, the setting means sets a plurality of information acquisition points inside or around the three-dimensional structure of the molecule, and the second acquisition means acquires data on one or a plurality of pre-set parameters at each of the information acquisition points set by the setting means.
[0017] The molecular analysis method of the present disclosure is a molecular analysis method executed by a computer, and is characterized by comprising the steps of: acquiring information on the three-dimensional structure of a molecule through quantum chemical calculations; setting a plurality of information acquisition points inside or around the three-dimensional structure of the molecule; and acquiring data on one or more pre-set parameters at each of the set information acquisition points.
[0018] According to the present disclosure, it is possible to provide a new molecular analysis system, program, and information processing method for extracting spatial information of molecules.
[0019] FIG. 1 is a diagram schematically illustrating a configuration of a molecular analysis system according to an embodiment of the present disclosure. FIG. 2 is a diagram illustrating an example of the flow of information processing in a molecular analysis system and an information processing method according to an embodiment of the present disclosure, and is a diagram illustrating an example of a case where information acquisition points and parameters are acquired by first and second model generation means of the molecular analysis system. FIG. 3 is a diagram for explaining a method of setting information acquisition points in a molecular analysis system and an information processing method according to an embodiment of the present disclosure. FIG. 4 is a diagram for explaining another method of setting information acquisition points in a molecular analysis system and an information processing method according to another embodiment of the present disclosure. FIG. 5 is a diagram illustrating an exemplary flow of information processing in an information processing method according to an embodiment of the present disclosure. FIG. 6 is a diagram illustrating another exemplary flow of information processing in an information processing method according to an embodiment of the present disclosure.
[0020] 1 to 6 are diagrams illustrating a molecular analysis system 1 and an information processing method according to the present disclosure.
[0021] [Molecular Analysis System 1] The molecular analysis system 1 according to this embodiment shown in FIG. 1 extracts spatial information of molecules using a novel method not previously available. The molecular analysis system 1 is a computer that can be connected to a communication network (not shown) such as the Internet, and can be configured from an industrial computer, a personal computer, a tablet terminal, or the like. Alternatively, the molecular analysis system 1 may be a virtual server within a physical server accessible via the Internet. Furthermore, these hardware and virtual servers may be multiple, or may be combined in any desired manner. As shown in FIG. 1, the molecular analysis system 1 includes a control unit 10, a memory unit 28, a communication unit 30, a display unit 32, and an operation unit 34.
[0022] (Control Unit 10) The control unit 10 is composed of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), etc., and controls each process and operation within the molecular analysis system 1. Specifically, the control unit 10 executes programs stored in a storage unit 28 (described later) to function as a reception unit 12, a generation unit 14, a first acquisition unit 16, a setting unit 18, a second acquisition unit 20, a determination unit 22, a parameter type setting unit 24, a score generation unit 26, a first model generation unit 200, a second model generation unit 300, etc. Note that these functions may be achieved by executing one or more independent programs or applications. Furthermore, these programs and applications may be provided on a single terminal (the molecular analysis system 1) or may be distributed across multiple terminals (including the molecular analysis system 1), and in the latter case, may be connected to each other via a wired cable or a communication network. The above-mentioned units will be described later.
[0023] (Storage Unit 28) The storage unit 28 is configured, for example, with a hard disk drive (HDD), random access memory (RAM), read only memory (ROM), solid state drive (SSD), etc. Furthermore, the storage unit 28 is not limited to being built into the molecular analysis system 1, but may be a storage medium (for example, a USB memory) that can be detachably attached to the molecular analysis system 1. Furthermore, instead of providing the storage unit 28, various information may be stored in other storage means (such as a cloud server).
[0024] The memory unit 28 and the memory unit 40 are configured to store information such as the program executed by the control unit 10, the first trained model 220, the second trained model 320, the three-dimensional structure of each molecule, information acquisition points, parameters, etc.
[0025] (Communication Unit 30) The communication unit 30 connects the control unit 10 to an external device (e.g., a calculation unit not shown) wirelessly or via a wire so that the control unit 10 can communicate with the external device. The control unit 10 transmits and receives information to and from the external device via the communication unit 30.
[0026] (Display Unit 32, Operation Unit 34) The display unit 32 is, for example, a monitor or the like, and displays (outputs) various screens upon receiving a display command signal from the control unit 10. The operation unit 34 is, for example, a keyboard or the like, and can give (input) various commands to the control unit 10. Note that in some embodiments, a display operation unit such as a touch panel in which the display unit 32 and operation unit 34 are integrated may be provided.
[0027] (Details of the control unit 10) (Receiving means 12) The receiving means 12 receives information related to the chemical formula of a molecule, for example, via the communication unit 30 or the operation unit 34. In this specification, the term "molecule" is used in its ordinary sense and may include, for example, a drug, a material itself, a structural isomer, a stereoisomer, an optical isomer (enantiomer), etc., without being particularly limited. In the analysis according to the present disclosure, the stability and flexibility of the molecule during the analysis are not important.
[0028] (Generation means 14) The generation means 14 generates a three-dimensional structure of a molecule (100 in FIG. 2) by quantum chemical calculation from the chemical formula of the molecule accepted by the acceptance means 12 (see FIGS. 3 and 4). As the "quantum chemical calculation" for obtaining the three-dimensional structure of the molecule, HF (Hartree-Fock), DFT (Density Functional Theory), CI method (Configuration Interaction Method), CC (Coupled Cluster), MBPT (Many Body Perturbation Theory), etc. can be used.
[0029] (First Acquisition Unit 16) The first acquisition unit 16 acquires information on the three-dimensional structure of a molecule through quantum chemical calculations. The first acquisition unit 16 can also acquire information on the three-dimensional structure of a molecule from an external source, separately from the above-mentioned acceptance unit 12 and generation unit 14.
[0030] (Setting Means 18) The setting means 18 sets multiple information acquisition points within or around the three-dimensional structure of the molecule (110 in Figure 2). As used herein, "information acquisition points" refer to spatial coordinates arbitrarily set within or around the three-dimensional structure. As described below, information such as electron density (hereinafter also referred to as "parameters (values)") is acquired and set for each of these information acquisition points (coordinates). In order to find the optimal information acquisition points for each molecule, the positions and number of information acquisition points can be appropriately set and changed. Once the optimal information acquisition points for each molecule are found for its intended use, the same information acquisition points can be applied to other molecules with similar structures. Here, "optimal" can vary depending on the use and can be determined arbitrarily. If the intended effect is obtained through actual application of the molecule (for example, if the molecular compound binds to a specific receptor), it may be evaluated as "optimal."
[0031] The method for setting the information acquisition points is not particularly limited, but the setting means 18 can set an axis connecting the start point and end point of the three-dimensional structure of the molecule acquired by the first acquisition means 16, and when arranging information acquisition points around the axis (i.e., around the three-dimensional structure), set a virtual three-dimensional curve from the set start point to the end point, and set multiple information acquisition points at intervals on the set virtual three-dimensional curve (FIG. 3). These start points and end points may also be used as information acquisition points.
[0032] In this specification, the "axis connecting the starting point and the end point" refers to a line passing through the target molecule, and the "starting point" and "end point" are its two ends. For example, the "axis connecting the starting point and the end point" is depicted by a two-dot chain line in the upper left of Figure 3. Note that, like the information acquisition points, the "starting point" and "end point" do not necessarily have to be set on the three-dimensional structure, but may be set inside or around the three-dimensional structure. The setting means 18 may set the starting point and the end point (and the axis) so that the linear distance between them is maximized, or may set a longitudinal direction in the three-dimensional structure of the molecule acquired by the first acquisition means 16 and set the starting point and the end point near the end along that longitudinal direction. In cases where multiple longitudinal directions can be set, such as in the case of water molecules, any of the longitudinal directions may be selected.
[0033] The setting means 18 may also set a virtual three-dimensional curve to cover part or all of the three-dimensional structure of the molecule acquired by the first acquisition means 16. A "virtual three-dimensional curve" is a mathematical curve on which information acquisition points are arranged. Zero to multiple information acquisition points can be set as intermediate points excluding the "start point" and "end point." Note that the parameter values at the information acquisition points are used to predict the binding ability, interactions, side effects, and other behaviors of the molecule or functional group. Therefore, the more information available, the better. However, depending on the molecule, it is not necessary to have enough information acquisition points (parameter values) to cover the entire three-dimensional structure. For example, it is possible to estimate pharmacological effects and side effects by focusing on highly reactive parts such as functional groups. More specifically, it is possible to predict cases where the pharmacological effects and side effects of molecules (compounds) that share a common functional group differ due to differences other than the functional group. In other words, it is also possible to focus on the partial structure.
[0034] Furthermore, the "shape" of the virtual three-dimensional curve may be, but is not limited to, a spiral shape (for example, the upper right side of FIG. 3 ), a parabolic shape, etc. If multiple information acquisition points are set as described above, a straight line may be included in part of the virtual three-dimensional curve.
[0035] The "interval" at which the multiple information acquisition points are set on the virtual three-dimensional curve may be set at equal intervals, and the distance of each interval may be set arbitrarily. Furthermore, the information acquisition points do not necessarily have to be arranged on the virtual three-dimensional curve, and may be set along or near the virtual three-dimensional curve.
[0036] As another method for setting information acquisition points, the setting means 18 can divide the three-dimensional structure of the molecule acquired by the first acquisition means 16 into predetermined cubes and set information acquisition points for each cube. The "predetermined cubes" correspond to voxels. Each voxel can be treated as a single information acquisition point (one voxel = one information acquisition point). In this case, the information acquisition point may be the center point of the voxel, or multiple coordinates within a voxel may be integrated using summary statistics such as the mean, median, or mode and treated as a single information acquisition point. The size of the voxel is not particularly limited and can be set arbitrarily. In this voxel method, voxels of the same size are arranged to form a collection in a coordinate system, and the three-dimensional structure of the molecule is represented by this collection of unit voxels. To set information acquisition points in the voxel method, voxels are arranged inside or around the three-dimensional structure of the molecule, and then information acquisition points (coordinates) are set for the arranged voxels, similar to when information acquisition points are arranged on a virtual three-dimensional curve. Alternatively, among the collection of voxels corresponding to the three-dimensional structure of the molecule, voxels immediately outside the outermost voxel may be appropriately selected as information acquisition points, or voxels may be selected as information acquisition points by focusing on a part of the three-dimensional structure of the molecule (e.g., a highly reactive part such as a functional group). Alternatively, the three-dimensional structure of the molecule may be divided into a lattice (the three-dimensional structure may be divided at equal intervals in the XY, YZ, and ZX planes; i.e., voxels are defined by lattices and intervals), and the points where the three-dimensional structure intersect with the lattice may be set as information acquisition points. These setting methods eliminate the need to set start and end points, making it easy to set information acquisition points.
[0037] (Second Acquisition Means 20) The second acquisition means 20 acquires data on one or more pre-set parameters at each information acquisition point set by the setting means 18 (120 in Figure 2). Here, "parameters" can include multidimensional information such as steric, electronic, and chemical information related to the binding ability of a molecule. Steric information provides orientation to the arrangement of the molecule's three-dimensional structure, enabling the differentiation of stereoisomers, which have been largely ignored until now, and enabling analysis of even flexible molecules. Furthermore, electronic and steric information enables the calculation of intermolecular interactions. Examples of parameters include charge state, electron density, density of states (DOS), molecular orbitals, electrostatic potential, and space-filling factor. Furthermore, "pre-set" means that parameter data can be acquired at the positions of information acquisition points set within or around the three-dimensional structure of an existing molecule, and therefore the parameter data can be selected. Specifically, parameters such as electron density can be acquired by ab initio calculating the molecular electronic wave function using a quantum chemistry calculation program. The quantum chemical calculation here may be independent of the quantum chemical calculation for obtaining the three-dimensional structure of the molecule. Examples of programs or software that perform quantum chemical calculations to obtain parameter values include Gaussian, ADF (Amsterdam Density Functional), and Psi4.
[0038] The storage unit 28 can store the three-dimensional structure, information acquisition points, and parameters of each molecule as described above in association with each other as a "data set." Axes, virtual three-dimensional curves, voxels, etc. may also be stored in association with each other. The heat map-like tables in FIGS. 3 and 4 are also a type of data set.
[0039] Furthermore, the information acquisition points and parameters can be expressed, for example, by a heat map-like table such as those shown in Figures 3 and 4 (as an example, the horizontal axis of the table represents the position of the information acquisition points, and the vertical axis represents the contents of the corresponding quantified parameters). For example, electron density as a parameter can be expressed by a color that indicates a specific numerical value. Alternatively, the information acquisition points and parameters can be expressed by a single matrix.
[0040] (Determination means 22) The determination means 22 determines the positions of the information acquisition points from the three-dimensional structure of the molecule acquired by the first acquisition means 16 using a first model (first trained model 220 in FIG. 2 ) generated by machine learning using teacher data 210 including the three-dimensional structure of the molecule acquired as described above, the positions of the information acquisition points, and scores related to the molecular analysis results described below (which may further include various parameters). Specifically, the determination means 22 "inputs" the three-dimensional structure of the molecule acquired by the first acquisition means 16 into the first trained model 220, and can cause the first trained model 220 to "output" the positions of the information acquisition points with high scores. As a result, the setting means 18 can set the positions of the information acquisition points determined by the determination means 22.
[0041] (Parameter Type Setting Means 24) The parameter type setting means 24 sets the type of parameters from the three-dimensional structure of the molecule acquired by the first acquisition means 16 or the position of each information acquisition point set by the setting means 18 (or determined by the determination means 22) using a second model (second trained model 320 in FIG. 2 ) generated by machine learning using teacher data 310 including the three-dimensional structure of the molecule acquired as described above, the positions of the information acquisition points, the types of parameters, and scores related to the molecular analysis results described below. Specifically, the parameter type setting means 24 "inputs" the three-dimensional structure of the molecule acquired by the first acquisition means 16 or the positions of each information acquisition point set by the setting means 18 into the second trained model 320, and causes the second trained model 320 to "output" the type of parameter with a high score. As a result, the second acquisition means 20 can acquire data at each information acquisition point for one or more parameters set by the parameter type setting means 24.
[0042] (Score Generation Means 26) For the above learning, the score generation means 26 generates scores as analysis results for different molecules using a dataset including the three-dimensional structure, information acquisition points, and parameters of a certain molecule. For example, assume that the three-dimensional structures of molecules A and B are similar. In this case, it is highly likely that the information acquisition points and parameters of one molecule (molecule A) can be used as the information acquisition points and parameters of the other molecule (molecule B) in verifying a specific application. Specifically, the information acquisition points and parameter values of molecule A are compared with the corresponding information acquisition points and parameter values in the three-dimensional structure of molecule B. When comparing, it is preferable that the positions of the information acquisition points are the same. If the comparison results in a small difference in the parameter values, a high score is generated for both the information acquisition points and the parameter values. Conversely, if the difference in the parameter values is large, a low score is generated for both the information acquisition points and the parameter values. If the score is low, the positions of the information acquisition points are changed (by shifting the virtual three-dimensional curve, changing the shape, the spiral winding pattern, spacing, number, voxel size, etc.), and the parameter values are acquired again. The high or low score can be arbitrarily determined based on the difference between the compared parameter values, whether the parameter is suitable for the application being verified, etc. Information about these scores may also be stored in association with the aforementioned data set.
[0043] (First model generation means 200, second model generation means 300) The first model generation means 200 and the second model generation means 300 generate a first trained model 220 and a second trained model 320 through machine learning using teacher data 210 and 310, respectively. More specifically, the first model generation means 200 trains the first trained model 220 so that when the three-dimensional structure of the molecule acquired by the first acquisition means 16 is "input," the first trained model 220 "outputs" the positions of information acquisition points with high scores. The second model generation means 300 trains the second trained model 320 so that when the three-dimensional structure of the molecule acquired by the first acquisition means 16 or the positions of each information acquisition point set by the setting means 18 is "input," the second trained model 320 "outputs" the type of parameter with high scores. The "machine learning" is not particularly limited as long as it can perform the above-mentioned input and output, and various techniques such as deep learning, random forests, and support vector machines can be used.
[0044] The first trained model 220 and / or the second trained model 320 may function by the control unit 10 executing a program stored in the memory unit 28, or may function by a device other than the molecular analysis system 1 executing a predetermined program. In the latter case, the first trained model 220 and / or the second trained model 320 generated by the other device are transmitted to the computer 3 and stored in the memory unit 40 of the computer 3.
[0045] [Information Processing Method 1] Next, information processing method 1 in the molecular analysis system 1 described above will be described with reference to Figure 5. Here, it is assumed that information acquisition points are set using a virtual three-dimensional curve (curve method) and the three-dimensional structure of the molecule is acquired within the molecular analysis system 1. In the following description, components with the same reference numerals are the same as the components described above, and duplicated descriptions will be omitted as appropriate. Note that the processing described below is performed by executing a program stored in the storage unit 28, but the information processing method according to the present disclosure is not limited to this.
[0046] First, the control unit 10 (receiving means 12) receives information relating to the chemical formula of a molecule (step S1).
[0047] Next, the control unit 10 (generation means 14) generates a three-dimensional structure of the molecule from the received chemical formula of the molecule by quantum chemical calculation (step S2).
[0048] Next, the control unit 10 (setting means 18) sets an axis connecting the start point and the end point of the generated and acquired three-dimensional structure of the molecule (step S3). Here, as described above, the setting means 18 may set the longitudinal direction in the generated three-dimensional structure of the molecule, and set the start point and the end point near the end along the longitudinal direction.
[0049] Next, when arranging information acquisition points around the axis, the control unit 10 (setting means 18) sets a virtual three-dimensional curve from the set start point to the end point, and sets multiple information acquisition points at intervals on the set virtual three-dimensional curve (step S4). Note that multiple information acquisition points may be set inside or around the three-dimensional structure of the molecule, or the start point and end point may be set so that the linear distance between them is the greatest. In one embodiment, in this step, a virtual three-dimensional curve may first be set (FIG. 3).
[0050] Next, the control unit 10 (second acquisition means 20) acquires data on one or more parameters that are set in advance at each of the set information acquisition points (step S5).
[0051] [Information Processing Method 2] Next, information processing method 2 in the molecular analysis system 1 will be described with reference to Fig. 6. Note that this assumes that information acquisition points are set using voxels (voxel method) and information on the three-dimensional structure of molecules is acquired from outside.
[0052] First, the control unit 10 (first acquisition means 16) acquires the three-dimensional structure of the molecule from the chemical formula of the molecule accepted by the acceptance means 12 through quantum chemical calculations (step S11) (corresponding to steps S1 and S2).
[0053] Next, the control unit 10 (setting means 18) divides the three-dimensional structure of the molecule acquired by the first acquisition means 16 into predetermined cubes (step S12).
[0054] Next, the control unit 10 (setting means 18) sets information acquisition points at the intersections of each cube (step S13).
[0055] Next, the control unit 10 (second acquisition means 20) acquires data on one or more parameters that are set in advance at each of the set information acquisition points (step S14).
[0056] [Uses and Application Examples] The three-dimensional structure, information acquisition points, and parameters acquired by information processing methods 1 and 2 as described above can be used to analyze the use of a new molecule. For example, they can be used as training data 210 to generate a first trained model 220 by machine learning. Such a first trained model 220 can determine the positions of suitable information acquisition points from the three-dimensional structure of the new molecule. Furthermore, the three-dimensional structure, information acquisition points, and parameters acquired by information processing methods 1 and 2 can be used as training data 310 to generate a second trained model 320 by machine learning. Such a second trained model 320 can set suitable parameter types from the three-dimensional structure of the new molecule or the first trained model 220 or the positions of each arbitrarily determined information acquisition point.
[0057] Furthermore, conventional techniques such as those described in Patent Document 1 involve simulations and preprocessing to evaluate the physical properties of compounds themselves, but because all unnecessary data is extracted, the computational costs are enormous (and this method may be used only for academic compound research and is not practical). Furthermore, a stable structure is determined from the compound, and data is then acquired only when the structure is stable, failing to address structural flexibility. Furthermore, intermolecular interactions are not assumed. Furthermore, the one-dimensional, two-dimensional, and three-dimensional structural information provided by conventional methods is meaningful only from a chemical perspective, but does not lead to predictions of in vivo behavior. What is important is the reaction (binding) with proteins, etc., and spatial information about the surroundings of the compound is crucial.
[0058] To address the above issues, the invention disclosed herein plots the space surrounding a compound using information acquisition points. This method has not yet been used in any field. Conventional methods involve placing a molecule in a space where its structure is likely to be most stable and extracting information from it. However, the method disclosed herein is based on the completely opposite concept of defining a molecule using the space surrounding it (the information acquisition points set therein). This allows for the method to be unconstrained by the stable state of the structure. Furthermore, molecular characteristics can be accurately captured at low computational cost. Furthermore, utilizing the characteristics of molecular structure is expected to be useful in a wide range of industrial fields beyond drug discovery.
[0059] An example application of the invention disclosed herein is the design of new drugs. By calculating the similarity of extracted chemical information, it is possible to create drugs with properties similar to a desired drug (ligand-based drug discovery). Furthermore, by constructing a machine learning model using chemical information extracted from a compound as input data and the compound's interactions with biological substances such as proteins, pharmacokinetic indicators such as the compound's metabolism and absorption in the body, and the compound's pharmacological effects and side effects as output data, the invention can be used to predict various in vivo behaviors of unknown compounds. Another example application is predicting the solubility of compounds. Specifically, solubility can be predicted based on the ratio of hydrophobic and hydrophilic regions around the molecule. Furthermore, applications in materials science are also possible. Specifically, this disclosure is expected to lead to the development of new materials such as plastics, paints, rubber, and fibers. It is also possible to investigate quantitative structure-activity relationships based on the chemical information extracted as parameters and design new materials with desired properties.
[0060] In the above application examples, machine learning can be used to identify and predict the uses of molecules. That is, experimental and predicted values for a wide variety of molecules can be obtained, and the suitability of the uses can be evaluated by scoring. If data and scores, such as information acquisition points and parameters, can be used as training data, scores can be output from the machine learning model to which the data is input, and the suitability of the molecule for uses such as drug discovery and materials, as well as the suitability of the method for setting the information acquisition points, can be determined. When calculating scores for a new dataset, it is preferable to standardize and unify the method for setting the information acquisition points (e.g., helix or voxel spacing) regardless of the molecule. If the score is poor, attempt to rearrange the information acquisition points. On the other hand, if the score (accuracy) is good, the judgment based on the information acquisition points disclosed herein is good, and the molecule can be presented, for example, as a candidate lead compound for a specific pathology.
[0061] Furthermore, according to the present disclosure, since information acquisition points and parameters can be associated with molecular compound structures, it is expected that information acquisition points and parameters that have achieved a high score for a certain application can be used to predict new drugs or new materials with three-dimensional structures that may achieve a high score for a similar application.
[0062] The molecular analysis system 1 (computer), program, and information processing method according to the present disclosure, configured as described above, are provided with a first acquisition means 16 that acquires information on the three-dimensional structure of a molecule through quantum chemical calculations, a setting means 18 that sets a plurality of information acquisition points inside or around the three-dimensional structure of the molecule, and a second acquisition means 20 that acquires data on one or a plurality of pre-set parameters at each of the information acquisition points set by the setting means 18.
[0063] The presently disclosed invention extracts spatial information about molecules using a novel, previously unseen method. Because the binding ability of a molecule depends on the presence of electrons surrounding it, the spatial information geometrically acquired by this disclosure (particularly the information acquisition points and parameters) can predict the binding ability of the molecule, thereby supporting research and development related to molecular applications such as drug discovery and new materials. Furthermore, because data is acquired around the molecule and the molecule is identified, a wide range of information can be expressed.
[0064] Furthermore, in the molecular analysis system 1, the program, and the information processing method according to the present disclosure, the setting means 18 may divide the three-dimensional structure of the molecule acquired by the first acquisition means 16 into predetermined cubes and set information acquisition points for each cube. As described above, this method makes it possible to easily set information acquisition points.
[0065] Furthermore, in the molecular analysis system 1, program, and information processing method according to the present disclosure, the setting means 18 may set an axis connecting the start point and end point of the three-dimensional structure of the molecule acquired by the first acquisition means 16, and when arranging information acquisition points around the axis, may set a virtual three-dimensional curve from the set start point to the end point, and set multiple information acquisition points at intervals on the set virtual three-dimensional curve. In this way, the minimum number of information acquisition points required can be set according to the three-dimensional structure of the molecule.
[0066] Furthermore, in the molecular analysis system 1, program, and information processing method according to the present disclosure, the setting means 18 may set a longitudinal direction in the three-dimensional structure of the molecule acquired by the first acquisition means 16, and set the start point and end point near the end along the longitudinal direction. The setting means 18 may also set the start point and end point so that the linear distance between them is the greatest. In this way, a suitable method for setting the start point and end point can be selected depending on the three-dimensional structure of the molecule.
[0067] Furthermore, in the molecular analysis system 1, the program, and the information processing method according to the present disclosure, the setting means 18 may set a virtual three-dimensional curve so as to cover part or all of the three-dimensional structure of the molecule acquired by the first acquisition means 16. In this way, the amount of information of the parameter values can be increased or decreased depending on the structure and size of the molecule to be analyzed.
[0068] In the molecular analysis system 1, the program, and the information processing method according to the present disclosure, the virtual three-dimensional curve set by the setting means 18 may have a helical shape. In this way, necessary information acquisition points can be set according to the three-dimensional structure of the molecule. Furthermore, a helical shape can reflect the difference between enantiomers.
[0069] In the molecular analysis system 1, program, and information processing method according to the present disclosure, the setting means 18 can also set multiple information acquisition points at equal intervals on a virtual three-dimensional curve. By narrowing the interval, information can be collected evenly. By widening the interval, only the minimum necessary information can be collected.
[0070] Furthermore, the molecular analysis system 1, the program, and the information processing method according to the present disclosure may further include a receiving means 12 that receives information related to the chemical formula of the molecule, and a generating means 14 that generates a three-dimensional structure of the molecule by quantum chemical calculation from the chemical formula of the molecule received by the receiving means 12. As described above, the three-dimensional structure of the molecule to be analyzed may be generated by the molecular analysis system 1, as well as being obtained from an external source.
[0071] Furthermore, the molecular analysis system 1, program, and information processing method according to the present disclosure may further include a determination means 22 that determines the position of the information acquisition point from the three-dimensional structure of the molecule acquired by the first acquisition means 16, using a first model (first trained model 220) generated by machine learning using training data 210 including the three-dimensional structure of the molecule, the position of the information acquisition point, and a score related to the molecular analysis result. In this way, the setting means 18 may set the position of the information acquisition point determined by the determination means 22.
[0072] Furthermore, the molecular analysis system 1, program, and information processing method according to the present disclosure may further include parameter type setting means 24 that sets the type of a parameter from the three-dimensional structure of the molecule acquired by the first acquisition means 16 or the position of each information acquisition point set by the setting means 18 (or the determining means 22) using a second model (second trained model 320) generated by machine learning using training data 310 including the three-dimensional structure of the molecule, the position of the information acquisition point, the type of parameter, and a score related to the analysis result of the molecule. In this way, the second acquisition means 20 may acquire data at each information acquisition point for one or more parameters set by the parameter type setting means 24.
[0073] The molecular analysis system 1, the program, and the information processing method according to the present disclosure are not limited to the above-described aspects and combinations, and various modifications can be made.
[0074] For example, it has been explained that the second acquisition means 20 acquires data on one or more parameters that are set in advance at each information acquisition point set by the setting means 18, but it is not necessary to set a parameter at a certain information acquisition point in the curve method, voxel method, etc. In other words, if more information acquisition points than necessary are set, it may not be necessary to obtain parameter values at all information acquisition points.
Claims
1. A molecular analysis system comprising: a first acquisition means for acquiring information on the three-dimensional structure of a molecule through quantum chemical calculations; a setting means for setting a plurality of information acquisition points inside or around the three-dimensional structure of the molecule; and a second acquisition means for acquiring data on one or more pre-set parameters at each of the information acquisition points set by the setting means.
2. A molecular analysis system according to claim 1, wherein said setting means divides the three-dimensional structure of the molecule acquired by said first acquisition means into predetermined cubes and sets information acquisition points for each of said cubes.
3. The molecular analysis system according to claim 1, wherein the setting means sets an axis connecting the start point and end point of the three-dimensional structure of the molecule acquired by the first acquisition means, and when arranging information acquisition points around the axis, sets a virtual three-dimensional curve from the set start point to the set end point, and sets multiple information acquisition points at intervals on the set virtual three-dimensional curve.
4. A molecular analysis system as described in claim 3, wherein the setting means sets a longitudinal direction in the three-dimensional structure of the molecule acquired by the first acquisition means, and sets the start point and the end point near the ends along that longitudinal direction.
5. A molecular analysis system according to claim 3, wherein said setting means sets said start point and said end point so that the linear distance between them is greatest.
6. A molecular analysis system according to claim 3, wherein said setting means sets said virtual three-dimensional curve so as to cover a part or all of the three-dimensional structure of the molecule acquired by said first acquisition means.
7. A molecular analysis system according to claim 3, wherein said setting means sets a plurality of said information acquisition points at equal intervals on said virtual three-dimensional curve.
8. The molecular analysis system according to claim 1, further comprising: a receiving means for receiving information relating to the chemical formula of a molecule; and a generating means for generating a three-dimensional structure of the molecule by quantum chemical calculation from the chemical formula of the molecule received by said receiving means.
9. A molecular analysis system as described in claim 1, further comprising a determination means for determining the position of the information acquisition point from the three-dimensional structure of the molecule acquired by the first acquisition means using a first model generated by machine learning using training data including the three-dimensional structure of the molecule, the position of the information acquisition point, and a score related to the molecular analysis result, and wherein the setting means sets the position of the information acquisition point determined by the determination means.
10. A molecular analysis system as described in claim 1, further comprising a parameter type setting means for setting the type of parameter from the three-dimensional structure of the molecule acquired by the first acquisition means or the position of each information acquisition point set by the setting means, using a second model generated by machine learning using training data including the three-dimensional structure of the molecule, the position of the information acquisition point, the type of parameter, and a score related to the molecular analysis result, wherein the second acquisition means acquires data at each information acquisition point for one or more parameters set by the parameter type setting means.
11. A program that functions as a first acquisition means, a setting means, and a second acquisition means when executed by a computer, wherein the first acquisition means acquires information on the three-dimensional structure of a molecule through quantum chemical calculations, the setting means sets a plurality of information acquisition points inside or around the three-dimensional structure of the molecule, and the second acquisition means acquires data on one or a plurality of pre-set parameters at each of the information acquisition points set by the setting means.
12. A molecular analysis method executed by a computer, comprising: a step of acquiring information on the three-dimensional structure of a molecule through quantum chemical calculation; a step of setting a plurality of information acquisition points inside or around the three-dimensional structure of the molecule; and a step of acquiring data on one or more predetermined parameters at each of the set information acquisition points.