Construction method of large-scale reaction molecule simulation visual full-molecular structure library
By constructing a large-scale reaction molecular simulation visualization library of complete molecular structures, the problem of molecular structure generation errors in existing technologies has been solved, enabling relatively accurate prediction of molecular structures and construction of multi-dimensional feature systems, supporting basic research in the fields of chemistry, materials, and biology.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-03
AI Technical Summary
Existing reaction molecular dynamics simulation methods, when dealing with intramolecular excessive bonding, charge or free radical anomalies, result in the generation of incorrect SMILES structures or non-standard structural information, making it difficult to interface with molecular information tools such as RDKit and unable to provide a standardized molecular information database.
Through steps such as overbonding correction, preliminary molecular structure identification, detection of special groups, and molecular structure output, a large-scale reaction molecular simulation visualization library of complete molecular structures is constructed. This includes correcting the number of atomic bonds, identifying special groups, setting bond order and charge, and free radicals. Molecular structures are output in SMILES and SDF formats, and then visualized and their properties predicted.
It achieves accurate identification and labeling of functional groups such as nitro and carboxyl groups, as well as free radicals, ensuring the correct expression of molecular charge states and functional group structures. It can interface with tools such as RDKit to construct multi-dimensional molecular feature systems and support the construction of visualized full molecular libraries.
Smart Images

Figure CN121789844A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of reaction molecular dynamics simulation technology, and specifically to a method for constructing a large-scale reaction molecular simulation visualization library of complete molecular structures. Background Technology
[0002] Reactive molecular dynamics (RMD) simulations can examine changes in molecular structure during complex reactions at the microscopic scale. The accuracy of the predicted molecular structure depends primarily on the reactive force field used. Although the accuracy is lower than that of quantum chemical calculations, it can examine complex chemical reactions and molecular structure evolution on relatively large spatial and temporal scales and under complex loading conditions. RMD simulations rely on reaction analysis tools to identify molecular structures, such as generating SMILES information. However, existing methods mainly predict atomic structures based on atomic positions or chemical bond order. Some methods consider the structural reduction of molecules divided by periodic boundaries, but they lack research on intramolecular overbonding, charge, or free radical anomalies. This often leads to the generation of incorrect SMILES structures or non-standard structural information, making it difficult to interface with molecular information tools such as RDKit and failing to provide a standardized molecular information database for machine learning and AI technologies. Therefore, developing a method for constructing a visualized full molecular structure library suitable for large-scale reactive molecular dynamics simulations to achieve relatively accurate prediction of molecular structures is of great significance for promoting basic science and materials applications in materials science, chemistry, biology, and related fields. Summary of the Invention
[0003] To address the aforementioned shortcomings of existing technologies, the purpose of this invention is to provide a method for constructing a large-scale reaction molecular simulation visualization library of complete molecular structures, thereby achieving relatively accurate prediction of molecular structures.
[0004] The technical solution of this invention to solve the above-mentioned technical problems is as follows: A method for constructing a large-scale reaction molecule simulation visualization library of complete molecular structures is provided, comprising the following steps: (1) Overbonding correction: Correct the number of bonds of all atoms in each trajectory. Remove overbonding atoms based on bond order value and the number of bonds of each element. The correction method is as follows:
[0005] in, BO a and BO a ’ It is an atom a Connect the sets of bond order values before and after chemical bond correction, where BO a ={bo1,bo2,...bo n}; na and n a, cutoff Atoms a The number of bonds and the bond number threshold; \ indicates the operation of removing chemical bonds; (2) Preliminary identification of molecular structure ① Based on bonding relationships, obtain the set of atom IDs of all molecules contained in each trajectory, and fuse them with the atomic trajectory information at the same moment to obtain the set of coordinate information of all atoms within the molecule. Position mol as follows:
[0006] in, ID 1, ID 2, ID 3… represents the numbering of all atoms in the molecule; element1, element2, element3… represent the element symbols of each atom in the molecule; x1, y1, z1 represent… ID The spatial coordinates of one atom, and so on; ②Based on the coordinate information of all atoms within the molecule, and fused with the atomic trajectory information at the same moment, a set of bond-level information of all atoms within the molecule is obtained. bond mol as follows:
[0007] in, ID 1, ID 2, ID 3…… represents the numbering of all atoms in the molecule; order1, order2…… represent the bond order information of each atom in the molecule, and so on; ③ The bond order values of each chemical bond are based on the ReaxFF force field. Using species and reaction space information analysis methods based on reaction molecular dynamics simulations, specific numerical values are assigned to chemical bonds based on rounding rules. For the bond order values of N≡N and C=O bonds, a bond order determination compensation algorithm is introduced: the N≡N bond order value is multiplied by a compensation coefficient of 1.05, and the C=O bond order value is multiplied by a compensation coefficient of 1.2. The results are then compared with the criteria for N≡N or C=O bonds to confirm whether an N≡N or C=O bond has formed. For other bond order values, the bond order determination compensation algorithm is not required; the obtained values are directly compared with the determination criteria, which are as follows: ; (3) Molecular structure recognition: ① The following judgment method is used for the detection of special groups: I. Identifying Special Groups: Special groups can be initially identified by the types of bonding atoms. II. Locate the central atom of the special group from step I; III. Verify the chemical bonds of the central atom in step II: Verify the special functional groups by confirming that the central atom is bonded to other atoms; IV. Set bond order, charge, and free radicals based on the chemical properties of the functional groups; ② General bond order and charge setting method: Set single bond, double bond, triple bond or fractional bond order according to the bonding situation of atoms, and set atomic charge according to the electronegativity of groups and resonance structure; ③ The general free radical detection and identification strategy is as follows: First, exclude hydrogen atoms, metal elements and atoms related to special groups; then optimize and adjust the atomic valence state. When the atomic valence bond is insufficient, try to borrow excess bond order from adjacent atoms; when the number of atomic bonds is less than the bonding threshold and no charge is marked, mark the atom as a free radical atom. (4) Molecular structure output: The molecular structure is output using two common formats, SMILES and SDF; (5) Molecular visualization and property prediction: Molecular visualization and molecular property prediction are carried out on the SMILES and SDF format data generated in step (4), and comprehensive feature extraction and quantitative analysis of molecular structure are performed to construct a multi-dimensional molecular feature system; (6) Construction of a visualized molecular library: Based on the molecular feature system of step (5), a visualized molecular structure library is constructed; The specific groups are nitro, cyano, sulfonic acid, carboxyl, phosphate, amino, hydroxy, amide, nitroso, carbonyl, and mercapto.
[0008] Furthermore, the threshold values for the number of bonds formed by each element in step (1) are as follows: .
[0009] Furthermore, in step (1) when n a > n a, cutoff When, remove atoms a Minimum bond level, i.e., min(BO) a The chemical bond of ) when n a = n a, cutoff Chemical bond correction is completed in time.
[0010] Furthermore, the detection of special groups in step (3) adopts the following judgment method, as detailed below: .
[0011] Furthermore, in step (4), the molecular structure file of SMILES contains information on chemical bonds, charges, and free radicals; implicit hydrogens are disabled, and only explicit hydrogens are retained; chiral and aromatic representations are supported.
[0012] Furthermore, the molecular structure file of SDF in step (4) includes the simulated trajectory number, atom ID, element symbol, three-dimensional coordinates, atomic charge and free radical, bonded atom pair ID, bond type, bond order value and special group mark.
[0013] Furthermore, special group labels include nitro (N), cyano (N), sulfonic acid (S), carboxyl (C), phosphate (P), amino (N), hydroxyl (O), amide (C), nitroso (N), carbonyl (C), and mercapto (S).
[0014] The present invention has the following beneficial effects: The visualized whole molecular structure library constructed by the present invention has the following advantages and positive effects: (1) Through the bond level compensation mechanism, it shows accuracy in processing chemical bonds such as metal coordination bonds and N≡N bonds, and realizes automatic identification and labeling of functional groups such as nitro groups and carboxyl groups and free radicals, ensuring the correct expression of molecular charge state and functional group structure; (2) Based on the custom atomic number threshold, it realizes the segmented processing and data acquisition of macromolecules, and realizes the acquisition of molecular information such as atom ID, element symbol, three-dimensional coordinates, two-dimensional structure (2D) and three-dimensional structure (3D); (3) It can interface SMILES and SDF formats with molecular information tools such as RDKit to construct a multi-dimensional molecular feature system and realize the construction of a visualized whole molecular library. This method can construct a visualized whole molecular structure library for large-scale reaction molecular dynamics simulation, realize the relatively accurate prediction of molecular structure, and can be widely used in basic research in chemistry, materials, biology and related fields. Attached Figure Description
[0015] Figure 1 This is a V3000 format RDX molecular structure information diagram; Figure 2 These are two-dimensional and three-dimensional structure diagrams of RDX generated based on SDF. Detailed Implementation
[0016] The examples given below are for illustrative purposes only and are not intended to limit the scope of the invention. Unless otherwise specified, conditions in the examples are performed under standard conditions or as recommended by the manufacturer. Reagents or instruments whose manufacturers are not specified are all commercially available products.
[0017] Example: A method for constructing a large-scale reaction molecular simulation visualization library of complete molecular structures. (1) Overbonding correction Reaction molecular dynamics simulations, which rely on bond order thresholds to determine chemical bonds, can lead to excessive bonding of atoms, causing errors in subsequent molecular structure normalization. To avoid this problem, the number of bonds for all atoms in each trajectory is corrected. Based on the bond order value and the bond quantity threshold for each element (see Table 1), excessive bonding of atoms is removed. The correction method is as follows:
[0018] Table 1. Thresholds for the number of bonds for each element
[0019] Note: The number of element types can be expanded according to the model. In the formula BO a and BO a ’ It is an atom a Connect the sets of bond order values before and after chemical bond correction, where BO a ={bo1,bo2,...bo n}; n a and n a, cutoff Atoms a The number of bonds and the bond number threshold; \ indicates the operation of removing chemical bonds. When n a > n a, cutoff When, remove atoms a Minimum bond level, i.e., min(BO) a The chemical bond of ) when n a = n a, cutoff Chemical bond correction is completed in time.
[0020] (2) Preliminary identification of molecular structure Based on bonding relationships, the set of atom IDs contained in all molecules along each trajectory is obtained, and fused with the atomic trajectory information at the same moment to obtain the set of coordinate information of all atoms within the molecule. Position mol as follows:
[0021] in, ID 1, ID 2, ID 3, ... represent the numbers of all atoms in the molecule; element1, element2, element3, ... represent the element symbols of each atom in the molecule; x1, y1, z1 represent the spatial coordinates of atom ID1, and so on.
[0022] Based on this, the set of bond-level information for all atoms within the molecule is obtained by fusing the coordinate information of all atoms at the same time with the atomic trajectory information. bond mol as follows:
[0023] in, ID 1, ID 2, ID 3, ... represent the numbering of all atoms in the molecule; order1, order2, ... represent the bond order information of each atom in the molecule, and so on.
[0024] The bond order values of each chemical bond are based on the ReaxFF force field. Using species and reaction space information analysis methods based on reaction molecular dynamics simulations, specific values are assigned to the bond order based on rounding rules. For the bond order values of N≡N bonds and C=O bonds, a bond order determination compensation algorithm is introduced: the N≡N bond order value is multiplied by a compensation coefficient of 1.05, and the C=O bond order value is multiplied by a compensation coefficient of 1.2. The bond order is then compared with the judgment criteria for N≡N bonds or C=O bonds to confirm whether N≡N bonds or C=O bonds are formed. The determination of other bond order values does not require the introduction of a bond order determination compensation algorithm. The obtained values are directly compared with the judgment criteria, which are shown in Table 2.
[0025] Table 2 Personalized Key-Level Thresholds
[0026] (3) Molecular structure recognition ① The following judgment method is used for the detection of special groups: I. Identifying Special Groups: Special groups can be initially identified by the types of bonding atoms. II. Locate the central atom of the special group in step I (atoms that connect to other molecular structures in the special group, such as the N atom in -NO3); III. Verification of the chemical bond of the central atom in step II: Verify the special functional groups by confirming that the central atom is bonded to other atoms; IV. Bond order, charge, and free radicals are set based on the chemical characteristics of the functional groups. Detailed methods for identifying special functional groups are shown in Table 3. Atoms of special functional groups will be individually marked in the SDF file. ② General bond order and charge setting method: Set single bond, double bond, triple bond or fractional bond order according to the bonding situation of atoms, and set atomic charge according to the electronegativity of groups and resonance structure; ③ The general free radical detection and identification strategy is as follows: First, exclude hydrogen atoms, metal elements, and atoms related to special groups; then optimize and adjust the atomic valence state. When the atomic valence bond is insufficient, try to borrow excess bond order from adjacent atoms; when the number of atomic bonds is less than the bonding threshold and no charge is marked, mark the atom as a free radical atom; among them, special groups are nitro, cyano, sulfonic acid, carboxyl, phosphate, amino, hydroxy, amide, nitroso, carbonyl, and mercapto.
[0027] If structural optimization is required, an appropriate force field (MMFF94, MMFF94s, UFF, etc.) can be selected based on the characteristics of the simulated molecule to perform energy minimization calculations. After optimization, the atomic coordinates are written into the molecular conformation object, and each atom is precisely located to its three-dimensional spatial coordinates according to the established index mapping table.
[0028] Table 3. Methods for identifying special functional groups
[0029] Note: Other atomic groups can be derived using the method described above. (4) Molecular structure output Molecular structures are output using two common formats: SMILES and SDF. SMILES structures include information on chemical bonds, charges, and free radicals; implicit hydrogens are disabled, retaining only explicit hydrogens; chirality and aromaticity representation are supported; to support rapid output of SMILES data for very large molecules or clusters, a customizable atom number threshold can be defined. Large molecules exceeding this threshold are segmented using asterisks and then connected as SMILES data. SDF structures include simulated trajectory numbers, atom IDs (consistent with trajectory information), element symbols, three-dimensional coordinates (consistent with trajectory information), atomic charges and free radicals, bonded atom pair IDs, bond types, bond order values (consistent with bond order information), and special group labels (nitro N, cyano N, sulfonic acid S, carboxyl C, phosphate P, amino N, hydroxyl O, amide C, nitroso N, carbonyl C, mercapto S, etc.), generating V3000 and V2000 formats.
[0030] (5) Molecular visualization and property prediction The generated SMILES and SDF format data are integrated with molecular information tools such as RDKit for molecular visualization and property prediction. SDF data can plot precise 2D or 3D structures, map atomic coordinates to trajectory files, include atom ID information, and detect functional groups and free radical states such as nitro (N), cyano (N), sulfonic acid (S), carboxyl (C), phosphate (P), amino (N), hydroxyl (O), amide (C), nitroso (N), carbonyl (C), and mercapto (S). RDKit and similar tools are used for comprehensive feature extraction and calculation of molecular structures to obtain basic molecular properties, topological descriptions, physicochemical properties, charge characteristics, topological parameters, and stereochemical features, constructing a multi-dimensional molecular feature system. Furthermore, it can also be integrated with other molecular structure or molecular dynamics analysis programs.
[0031] (6) Construction of a visual molecular library Based on the molecular feature system of step (5), a visual molecular structure library is constructed.
[0032] The practical application of the above method is as follows: (1) Based on the RDX thermal decomposition example (see / Examples / reaxff / RDX in the LAMMPS program) provided by the open-source simulation platform LAMMPS, add commands to output trajectory and bond level information. After running the example, obtain the calculation results, obtain the initial bond level (bonds.reax.bof) and trajectory (dump.trj) data, and perform data cleaning to obtain one-dimensional target data of trajectory and bond level. According to the bonding number threshold and correction method of each element in Table 1 of the embodiment step (1), perform atomic bonding number correction on the bond level data to obtain the corrected bond level data as shown in Table 4.
[0033] Table 4 Bond level data after bond number correction
[0034] (2) The corrected atom IDs and intramolecular bonding atom pair information are shown in Table 5. The corrected molecular information and SMILES data are shown in Table 6.
[0035] Table 5 Corrected atom IDs and intramolecular bonding atom pairs
[0036] Table 6 Corrected molecular information and SMILES data
[0037] (3) SDF information of V3000 format RDX molecules as follows Figure 1 As shown. The two-dimensional and three-dimensional structure diagrams of the RDX containing atomic IDs are as follows. Figure 2 As shown (numbers are atom IDs). The predicted molecular properties from the visualized molecular library are shown in Table 7.
[0038] Table 7 Predictions of Molecular Properties
[0039] In summary, the construction method of this invention can realize interactive functions such as efficient storage and retrieval of molecular data, filtering of time frame range, filtering of molecular weight range, selection of atom ID range, and intuitive display of molecular structure diagrams; at the same time, it can also integrate visualization functions of basic molecular properties, topological description, physicochemical properties, charge characteristics, and topological parameters, and support the analysis of two-dimensional (2D), three-dimensional (3D), and molecular structures at different times.
[0040] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for constructing a large-scale reaction molecular simulation visualization library of complete molecular structures, characterized in that, Includes the following steps: (1) Overbonding correction: Correct the number of bonds of all atoms in each trajectory. Remove overbonding atoms based on bond order value and the number of bonds of each element. The correction method is as follows: in, BO a and BO a ’ It is an atom a Connect the sets of bond order values before and after chemical bond correction, where BO a ={bo1,bo2,...bo n }; n a and n a, cutoff Atoms a The number of bonds and the bond number threshold; \ indicates the operation of removing chemical bonds; (2) Preliminary identification of molecular structure ① Based on bonding relationships, obtain the set of atom IDs of all molecules contained in each trajectory, and fuse them with the atomic trajectory information at the same moment to obtain the set of coordinate information of all atoms within the molecule. Position mol as follows: in, ID 1, ID 2, ID 3… represents the numbering of all atoms in the molecule; element1, element2, element3… represent the element symbols of each atom in the molecule; x1, y1, z1 represent… ID The spatial coordinates of one atom, and so on; ②Based on the coordinate information of all atoms within the molecule, and fused with the atomic trajectory information at the same moment, a set of bond-level information of all atoms within the molecule is obtained. bond mol as follows: in, ID 1, ID 2, ID 3…… represents the numbering of all atoms in the molecule; order1, order2…… represent the bond order information of each atom in the molecule, and so on; ③ The bond order values of each chemical bond are based on the ReaxFF force field. Using species and reaction space information analysis methods based on reaction molecular dynamics simulations, specific numerical values are assigned to chemical bonds based on rounding rules. For the bond order values of N≡N and C=O bonds, a bond order determination compensation algorithm is introduced: the N≡N bond order value is multiplied by a compensation coefficient of 1.05, and the C=O bond order value is multiplied by a compensation coefficient of 1.
2. The results are then compared with the criteria for N≡N or C=O bonds to confirm whether an N≡N or C=O bond has formed. For other bond order values, the bond order determination compensation algorithm is not required; the obtained values are directly compared with the determination criteria, which are as follows: ; (3) Molecular structure recognition: ① The following judgment method is used for the detection of special groups: I. Identifying Special Groups: Special groups can be initially identified by the types of bonding atoms. II. Locate the central atom of the special group from step I; III. Verify the chemical bonds of the central atom in step II: Verify the special functional groups by confirming that the central atom is bonded to other atoms; IV. Set bond order, charge, and free radicals based on the chemical properties of the functional groups; ② General bond order and charge setting method: Set single bond, double bond, triple bond or fractional bond order according to the atomic bonding situation, and set atomic charge according to the electronegativity of the group and resonance structure; ③ The general free radical detection and identification strategy is as follows: First, exclude hydrogen atoms, metal elements and atoms related to special groups; then optimize and adjust the atomic valence state. When the atomic valence bond is insufficient, try to borrow excess bond order from adjacent atoms; when the number of atomic bonds is less than the bonding threshold and no charge is marked, mark the atom as a free radical atom. (4) Molecular structure output: The molecular structure is output using two common formats, SMILES and SDF; (5) Molecular visualization and property prediction: Molecular visualization and molecular property prediction are carried out on the SMILES and SDF format data generated in step (4), and comprehensive feature extraction and quantitative analysis of molecular structure are performed to construct a multi-dimensional molecular feature system; (6) Construction of a visualized molecular library: Based on the molecular feature system of step (5), a visualized molecular structure library is constructed; The specific groups are nitro, cyano, sulfonic acid, carboxyl, phosphate, amino, hydroxy, amide, nitroso, carbonyl, and mercapto.
2. The construction method according to claim 1, characterized in that, The bonding thresholds for each element mentioned in step (1) are as follows: 。 3. The construction method according to claim 1, characterized in that, In step (1), when the aforementioned n a > n a, cutoff When, remove atoms a Minimum bond level, i.e., min(BO) a The chemical bond of ) when n a = n a, cutoff Chemical bond correction is completed in time.
4. The construction method according to claim 1, characterized in that, The detection of special groups in step (3) adopts the following judgment method, as detailed below: 。 5. The construction method according to claim 1, characterized in that, The molecular structure file of SMILES described in step (4) contains information on chemical bonds, charges, and free radicals; implicit hydrogens are disabled, and only explicit hydrogens are retained; chiral and aromatic representations are supported.
6. The construction method according to claim 1, characterized in that, The molecular structure file of the SDF mentioned in step (4) includes the simulated trajectory number, atom ID, element symbol, three-dimensional coordinates, atomic charge and free radical, bonded atom pair ID, bond type, bond order value and special group mark.
7. The construction method according to claim 5, characterized in that, The special group markings include nitro (N), cyano (N), sulfonic acid (S), carboxyl (C), phosphate (P), amino (N), hydroxyl (O), amide (C), nitroso (N), carbonyl (C), and mercapto (S).