Method for predicting physicochemical properties of indole derivative
Patent Information
- Application Number
- CN202511092276.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-21
AI Technical Summary
传统吲哚衍生物物理性质预测方法依赖大量实验,成本高、周期长,难以满足快速发展的研究需求。
采用基于吲哚衍生物分子结构图的拓扑指数和回归模型,通过神经网络训练预测物理化学性质,减少实验资源消耗。
实现了快速、低成本的吲哚衍生物物理化学性质预测,提高研究效率,适用于大规模分子库筛选。
Smart Images

Figure CN120998340A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mathematical chemistry, specifically to a method for predicting the physicochemical properties of indole derivatives. Background Technology
[0002] Indole is a heterocyclic compound composed of a benzene ring fused to a pyrrole ring. It is a fundamental building block of many natural products, including the essential amino acid tryptophan. Released into the atmosphere through natural processes (e.g., organic degradation) and anthropogenic activities (e.g., industrial combustion), indole participates in complex atmospheric photo-oxidation pathways involving hydroxyl radicals, ozone, and other reactive oxidants. These processes determine its environmental persistence, reactivity, and potential health effects, thus requiring robust methods to predict its behavior.
[0003] Indole derivatives are diverse, with over 40,000 known natural products and hundreds of thousands of indole and its derivatives recorded in chemical databases. These derivatives can be classified in various ways, including by their functional groups, substitution positions, and biological activities, encompassing a wide range of structural types from simple substituted indoles to complex alkaloids. In applications, indole derivatives not only serve as the core structure of many important drugs in the pharmaceutical field but are also used in materials science to manufacture organic semiconductors and optoelectronic devices, and participate in atmospheric photo-oxidation processes in environmental chemistry. With ongoing research, new indole derivatives are constantly being discovered and synthesized, further expanding their application potential across various fields.
[0004] Indole derivatives, due to their unique aromatic heterocyclic structure, have shown broad application potential in drug development, materials science, environmental chemistry and other fields. Predicting their physical properties can significantly accelerate materials design, optimize device performance and promote technological breakthroughs. However, traditional methods for predicting the physical properties of indole derivatives mainly rely on a large number of experiments, which are costly, time-consuming and have great limitations. In view of this, this invention proposes a method for predicting the physicochemical properties of indole derivatives. Summary of the Invention
[0005] The purpose of this invention is to propose a method for predicting the physicochemical properties of indole derivatives, which can predict the physicochemical properties of molecules without conducting a large number of experiments, significantly reducing the consumption of experimental resources and improving research efficiency.
[0006] To achieve the above objectives, the technical solution of the present invention is: a method for predicting the physicochemical properties of indole derivatives, specifically including the following steps:
[0007] S1. Determine the physicochemical properties to be predicted;
[0008] S2. Obtain the molecular structures of known indole derivatives from public datasets, and obtain the corresponding properties of known indole derivatives based on the physicochemical properties to be predicted;
[0009] S3. Based on the known molecular structure of indole derivatives, construct a molecular structure graph G = (V(G), E(G)) with atoms as vertices and chemical bonds as connecting edges, where V(G) is the set of vertices and E(G) is the set of edges.
[0010] The degree of the vertex is calculated based on the molecular structure diagram of the known indole derivatives, and multiple different correlation indices are generated for each known indole derivative using the Norm method.
[0011] S4. Based on the corresponding properties of the known indole derivatives obtained in S2 and the relevant indices of the known indole derivatives generated in S3, establish a dataset and use the dataset to train the neural network.
[0012] S5. Generate multiple different correlation indices for the indole derivative to be predicted using the same method as in S3, and input them into the trained neural network to predict the correlation properties.
[0013] Preferably, the physicochemical properties to be predicted include density, boiling point, melting point, water solubility, molecular weight, and vapor pressure.
[0014] Preferably, the Norm method specifically involves traversing all edges of the indole derivative molecular structure diagram, calculating a relevant index based on the degree of each endpoint of an edge, and excluding any endpoint of an edge with a degree of 1 from the calculation.
[0015]
[0016] Where G is the molecular structure diagram of the indole derivative; u and v are vertices in the molecular structure diagram; d u and d v Let u and v be the degrees of vertices u and v, respectively; φ is the adjustment coefficient, φ∈{0,1}; n and m are both natural numbers; k is a non-zero real number, and G, φ, n, m, and k are all parameters to be received. By changing the values of parameters φ, n, m, and k, we can obtain the correlation indices of several different indole derivatives.
[0017] Preferably, the neural network includes an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is equal to the number of correlation indices for each indole derivative. The hidden layer adopts a three-layer fully connected structure. The number of neurons in the output layer is consistent with the number of physicochemical properties to be predicted.
[0018] Preferably, the hidden layer adopts a three-layer fully connected structure with 64, 128, and 64 neurons respectively, and a ReLU activation function is set after each hidden layer.
[0019] Preferably, the neural network is trained using the mean squared error loss function.
[0020] Compared with the prior art, the present invention has the following beneficial effects:
[0021] This invention proposes a method for predicting the physicochemical properties of indole derivatives. By calculating topological indices and regression models, the physicochemical properties of molecules can be predicted without conducting a large number of experiments, significantly reducing the consumption of experimental resources and improving research efficiency. Furthermore, it provides a rapid and low-cost material screening method suitable for large-scale molecular libraries. Attached Figure Description
[0022] Figure 1 This is a diagram of a neural network structure in one embodiment of the present invention;
[0023] Figure 2 This is a structural diagram of an indole molecule in one embodiment of the present invention. Detailed Implementation
[0024] The following is in conjunction with the appendix Figure 1-2 The technical solution of the present invention will be described in detail below.
[0025] This invention proposes a method for predicting the physicochemical properties of indole derivatives, specifically including the following steps:
[0026] S1. Determine the physicochemical properties to be predicted;
[0027] S2. Obtain the molecular structures of known indole derivatives from public datasets, and obtain the corresponding properties of known indole derivatives based on the physicochemical properties to be predicted;
[0028] S3. Based on the known molecular structure of indole derivatives, construct a molecular structure graph G = (V(g), E(G)) with atoms as vertices and chemical bonds as connecting edges, where V(G) is the set of vertices and E(G) is the set of edges.
[0029] The degree of the vertex is calculated based on the molecular structure diagram of the known indole derivatives, and multiple different correlation indices are generated for each known indole derivative using the Norm method.
[0030] S4. Based on the corresponding properties of the known indole derivatives obtained in S2 and the relevant indices of the known indole derivatives generated in S3, establish a dataset and use the dataset to train the neural network.
[0031] S5. Generate multiple different correlation indices for the indole derivative to be predicted using the same method as in S3, and input them into the trained neural network to predict the correlation properties.
[0032] The physicochemical properties and molecular structure diagrams of relevant indole derivatives can be obtained from publicly available datasets. Based on the above method, the datasets shown in Table 1 can be generated:
[0033] Table 1
[0034] substance Feature 1 Feature 2 Feature 3 ...... Index 1 Index 2 Index 3 ...... Substance 1 Substance 2 Substance 3 ......
[0035] The "substances" in this table refer to indole derivatives, and the "properties" include, but are not limited to, density, boiling point, melting point, water solubility, molecular weight, and vapor pressure. The "indices" are the various results generated using Norm transform parameters. For this dataset, a fully connected neural network was used for analysis; the neural network structure is referenced [reference needed]. Figure 1 .
[0036] In this embodiment, the Norm method specifically involves traversing all edges of the indole derivative molecular structure diagram, calculating a relevant index based on the degree of each endpoint of an edge, and excluding any endpoint of an edge with a degree of 1 from the calculation.
[0037]
[0038] Where G is the molecular structure diagram of the indole derivative; u and v are vertices in the molecular structure diagram; d u and d v Let u and v be the degrees of vertices u and v, respectively, where the degree of a vertex is the number of edges connected to it; φ is an adjustment coefficient, φ∈{0,1}; n and m are both natural numbers; k is a non-zero real number, and G, φ, n, m, and k are all parameters to be received. By changing the values of parameters φ, n, m, and k, we can obtain the correlation indices of several different indole derivatives.
[0039] Current research uses graph theory to define the following topological indices:
[0040] For example, the first Zagreb index emphasizes the interaction between atomic connectivity and chemical behavior:
[0041]
[0042] Second Zagreb Index:
[0043]
[0044] The modified Zagreb index is used to quantify the complexity or degree of molecular branching of chemical compounds.
[0045]
[0046] Atomic bond connectivity (ABC) index:
[0047]
[0048] Sombor Index:
[0049]
[0050] Simplified Sombor Index
[0051]
[0052] The correspondence between the Norm method and existing indices is shown in Table 2:
[0053] Table 2
[0054] n m k φ type 1 1 1 0 First Zagreb Index 1 1 -1 0 Second Zagreb Index 1 -1 1 0 Improved Zagreb index 1 1 1 / 2 1 Atomic Bond Connectivity (ABC) Index 1 0 1 / 2 0 Sombor Index 1 0 1 / 2 1 Simplified Sombor Index
[0055] Six existing exponents are listed here. In practical applications, the Norm method can theoretically include an infinite number of exponents not defined in the prior art by adjusting the range of parameters φ, n, m, and k.
[0056] The Norm method process is as follows:
[0057]
[0058] In the code above, f and the adjustment coefficient φ refer to the same thing, and different f, n, m, and k form different exponents.
[0059] These indices are calculated using the following method.
[0060]
[0061]
[0062] The following calculations are illustrated using the indole molecular diagram as an example. (Indole molecular structure diagram referenced below.) Figure 2 .
[0063] First, calculate the degree of each edge, and then remove edges containing vertices with a degree of 1, resulting in Table 3:
[0064] Table 3
[0065] (du,dv) frequency (2,2) 5 (2,3) 4 (3,3) 1
[0066] When n, m, k, and φ take the values 1, 1, 1, and 0 respectively:
[0067]
[0068] This index is the first Zagreb index mentioned earlier. Changing different n, m, k, and φ will yield different indices. For ease of description, the form "n_m_k_φ" is used to name different indices. For example, the first Zagreb index can be named "1_1_1_0" in this study.
[0069] In this embodiment, the neural network includes an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is the number of correlation indices for each indole derivative. The hidden layer adopts a three-layer fully connected structure. The number of neurons in the output layer is consistent with the number of physicochemical properties to be predicted.
[0070] In this embodiment, the hidden layer adopts a three-layer fully connected structure with 64, 128 and 64 neurons respectively, and a ReLU activation function is set after each hidden layer.
[0071] In this embodiment, the neural network is trained using the mean squared error loss function.
[0072] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for predicting the physicochemical properties of indole derivatives, characterized in that, Specifically, the following steps are included: S1. Determine the physicochemical properties to be predicted; S2. Obtain the molecular structures of known indole derivatives from public datasets, and obtain the corresponding properties of known indole derivatives based on the physicochemical properties to be predicted; S3. Based on the known molecular structure of indole derivatives, construct a molecular structure graph G = (V(G), E(G)) with atoms as vertices and chemical bonds as connecting edges, where V(G) is the set of vertices and E(G) is the set of edges. The degree of the vertex is calculated based on the molecular structure diagram of the known indole derivatives, and multiple different correlation indices are generated for each known indole derivative using the Norm method. S4. Based on the corresponding properties of the known indole derivatives obtained in S2 and the relevant indices of the known indole derivatives generated in S3, establish a dataset and use the dataset to train the neural network. S5. Generate multiple different correlation indices for the indole derivative to be predicted using the same method as in S3, and input them into the trained neural network to predict the correlation properties.
2. The method for predicting the physicochemical properties of indole derivatives according to claim 1, characterized in that, The physicochemical properties to be predicted include density, boiling point, melting point, water solubility, molecular weight, and vapor pressure.
3. The method for predicting the physicochemical properties of indole derivatives according to claim 1, characterized in that, The Norm method specifically involves traversing all edges of the indole derivative molecular structure diagram and calculating a relevant index based on the degree of each endpoint of an edge. If the degree of an endpoint of an edge is 1, it is not included in the calculation. Where G is the molecular structure diagram of the indole derivative; u and v are vertices in the molecular structure diagram; d u and d v Let u and v be the degrees of vertices u and v, respectively; φ is the adjustment coefficient, φ∈{0,1}; n and m are both natural numbers; k is a non-zero real number, and G, φ, n, m, and k are all parameters to be received. By changing the values of parameters φ, n, m, and k, we can obtain the correlation indices of several different indole derivatives.
4. The method for predicting the physicochemical properties of indole derivatives according to claim 1, characterized in that, The neural network includes an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is equal to the number of correlation indices for each indole derivative. The hidden layer adopts a three-layer fully connected structure. The number of neurons in the output layer is the same as the number of physicochemical properties to be predicted.
5. The method for predicting the physicochemical properties of indole derivatives according to claim 4, characterized in that, The hidden layers employ a three-layer fully connected structure with 64, 128, and 64 neurons respectively, and a ReLU activation function is set after each hidden layer.
6. The method for predicting the physicochemical properties of indole derivatives according to claim 1, characterized in that, The neural network is trained using the mean squared error loss function.