Method for constructing virtual screening space of organic molecules based on chemical feasibility verification
The virtual screening space for organic molecules is constructed through chemical feasibility verification, which solves the problem of insufficient judgment of molecular infeasibility in the prior art, and achieves efficient virtual screening and material development efficiency improvement.
Patent Information
- Application Number
- CN202211441121.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-17
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-11-17
AI Technical Summary
The existing virtual screening space construction methods cannot effectively judge the chemical feasibility of molecules, resulting in a large number of unfeasible molecules being screened out, affecting the efficiency and hit rate of virtual screening.
Through the chemical feasibility verification method, a virtual screening space for organic molecules is constructed, including molecular cleavage, breadth traversal and pattern matching library establishment, to verify the rationality of molecules, and the RECAP method is used to obtain the molecular backbone library and pattern matching library, and the frequency of group-skeleton fragments is counted to form a reasonable virtual screening space.
The hit rate and efficiency of virtual screening are improved. The generated virtual screening space focuses on specific functions, can quickly generate chemically feasible molecules, and improve the efficiency of organic material development.
Smart Images

Figure CN116110514B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of drug molecular design and energetic material design, and particularly relates to a method for constructing a virtual screening space of organic molecules based on chemical feasibility verification. Background Art
[0002] The virtual screening technology based on high-throughput computing and machine learning has brought about a qualitative leap in the design of new materials. By using a computer to quickly screen out potential molecules from a large compound library, the number of compounds entering the chemical experiment stage can be significantly reduced, which can effectively improve the success rate and efficiency of drug R & D and new material discovery.
[0003] The primary prerequisite for virtual screening is to have a reasonable virtual screening space. The quality of the virtual screening space determines the quality of the screened molecules. There are many existing virtual screening spaces available for selection, and their performances vary, resulting in widely different success rates of virtual screening in actual use. Moreover, since the objects faced by virtual screening may be tens of millions of compounds, the efficiency of virtual screening is also a key factor affecting the application of virtual screening methods. Molecular fragment assembly is a commonly used method for constructing a virtual screening space of organic molecules. A large number of molecules are obtained by enumerating the attachment sites of molecular fragments. Generally, the chemical feasibility of the molecular structure is simply judged by the valence rule. Although the finally generated chemical molecules conform to the valence rule, the proportion of effective molecules in the large number of generated molecules is very small. A large number of compounds may have unreasonable structures, unstable molecules, or cannot be synthesized, etc. The defect of the valence rule is that it cannot judge the chemical feasibility of molecules from the thermodynamic and kinetic levels, resulting in a large number of chemically infeasible molecules in the virtual screening space, which will lead to the screening of unreasonable target molecules. Therefore, constructing a reasonable virtual screening space is the key to improving the hit rate and efficiency of virtual screening. Summary of the Invention
[0004] Aiming at the problems existing in the prior art, the present invention provides a method for constructing a virtual screening space of organic molecules based on chemical feasibility verification, which can verify the feasibility of molecules and construct a high-quality virtual screening space of organic molecules. At the same time, this method has the ability to be freely extended, and the compound libraries mentioned in the present invention can all be expanded and replaced.
[0005] The technical solution adopted by the present invention is as follows: A method for constructing a virtual screening space of organic molecules based on chemical feasibility verification is provided, including the following steps:
[0006] A method for constructing a virtual screening space of organic molecules based on chemical feasibility verification includes the following steps:
[0007] S1. Obtain the reported organic molecule dataset and get the SMILES sequences of all molecules;
[0008] S2. Use the RECAP method to cleave each molecule in the dataset to obtain a large number of molecular skeletons, and organize and record them to form a molecular skeleton library;
[0009] S3. For each molecule in the dataset, use the concerned group as the vertex, perform a breadth-first traversal on the molecule, cleave the molecule after reaching the specified number of layers to obtain group-skeleton fragments, and count their frequencies to form a pattern matching library;
[0010] S4. Select skeletons from the molecular skeleton library, and assemble these skeletons with groups to form complete molecules;
[0011] S5. Cut the generated molecules in the manner described in S3 to obtain the group-skeleton fragments of the molecules, search for the fragments in the pattern matching library, if the structure can be matched in the pattern matching library, the molecule is reasonable, otherwise the molecule is unreasonable.
[0012] A further technical solution is that step S2 specifically includes the following steps:
[0013] The first step: Molecular cleavage. Use the RECAP method to fragment the molecules. The specific method is as follows:
[0014] STEP1: Collect the structures active against specific targets and analyze to obtain 11 cleavable molecular bonds;
[0015] STEP2: Read the input molecule and break one of the cleavable molecular bonds mentioned in STEP1 to obtain multiple sub-nodes;
[0016] STEP3: Take the sub-nodes after cleavage in STEP2. If the sub-node has a molecular bond mentioned in STEP1, repeat the steps in STEP2, otherwise mark the sub-node as a leaf node; if no new sub-nodes are generated after cleavage of all sub-nodes, the molecular cleavage is completed;
[0017] STEP4: Only retain the final fragments during fragmentation without retaining the intermediate process. After the fragments are analyzed, they are merged into building blocks and put into the fragmentation results.
[0018] The second step: Select specific skeletons from all the molecular fragmentation results to generate a molecular skeleton library. The skeleton selection rules are as follows: If the fragment only contains small functional groups, the fragment is not retained; retain the ring structure; retain the fragments containing double bonds and triple bonds.
[0019] A further technical solution is that in step S3, a breadth-first traversal is performed on the molecule, and the molecule is cleaved after reaching the specified number of layers. The rules are as follows: Perform a breadth-first traversal on the molecule with the specified group as the vertex. When the specified number of layers is reached, obtain the molecular bonds between the molecules in the current layer and the next layer, and break the obtained bonds to obtain molecular group-skeleton fragments; during the traversal, if an atom is within a ring, stop the traversal, retain the ring structure, and break the other bonds on the ring that are not connected to the specified group to obtain group-ring fragments.
[0020] A further technical solution is that the specific process of step S3 is as follows:
[0021] STEP1: Search for the concerned group in the molecule, obtain the serial number of the group. Set A represents the set of atoms containing the concerned group, set B represents the set of edges to be cleaved, and set C represents the set of atoms in the current layer. Add the obtained serial number of the group to sets A and B;
[0022] STEP2: Traverse all atoms in set C, obtain the neighbor nodes of each node. If the neighbor node is on a ring, add all atoms on the ring to set A, and add all edges connected to the ring except the edge connected to this atom to set C. If the neighbor node is not in A, add the node to set A; clear set C, add the neighbor node to set C, and execute STEP2 again until the number of traversed layers reaches the set threshold. Obtain the neighbor nodes of each atom in set C and the corresponding edges, and add the edges to set B;
[0023] STEP3: Break the edges in set B, cleave the molecule into multiple fragments, and find the target fragments according to set A;
[0024] STEP4: Collect the target fragments, convert the target fragments into SMILES sequences, remove the irrelevant symbols in the sequences, convert the processed SMILES sequences into SMARTS codes for matching, count the frequencies of various codes, and finally form a pattern matching library.
[0025] A further technical solution is that the specific process of the skeleton and the group being assembled with each other to form a complete molecule in step S4 is as follows: Select n skeletons, and the set of the number of substitution sites of each skeleton is S = {S1, S2,..., S n}, select m groups. The n skeletons are connected and combined with each other at any substitution site to form a larger skeleton, and then the groups are connected to the remaining one or more substitution sites of the skeleton to form a complete molecule. The number of molecules that can be generated after substitution is:
[0026] A further technical solution is that in step S5, the molecule is verified for effectiveness, and the specific steps are as follows:
[0027] Select the molecule to be verified, cut the molecule in the manner described in S3 to obtain group-skeleton fragments, and use the RDKIT toolkit to match the fragments in the pattern matching library. If there are similar fragments in the pattern matching library, it proves that the molecule is valid; otherwise, the molecule is invalid.
[0028] This method divides the construction of the virtual screening space into two steps, adding a sub-chemical feasibility verification step on the basis of the commonly used molecular fragment assembly. The present invention uses the organic molecules generated by molecular fragment assembly as the verification object, and abandons the infeasible molecules through chemical feasibility verification to obtain the final virtual screening space. In the chemical reliability verification, for any specified parent body or substituent, by collecting the reported organic molecular structures as the basic data source, the occurrence frequency of the chemical environment permitted by the fragment is counted, and a pattern matching library for the fragment and the chemical environment is constructed accordingly. On this basis, the chemical feasibility of the molecule is verified by means of substructure matching. The organic molecular structure data source and the pattern matching library involved in the present invention both have the ability to be freely expanded, including supplementation, deletion, and replacement. The present invention realizes the discrimination of the chemical feasibility of organic molecules based on statistical methods, and can provide a chemically feasible virtual screening space for the computer-aided design of organic molecules such as drugs, energetic materials, and optoelectronic materials, which is of great significance for improving the research and development efficiency of organic materials.
[0029] Compared with the prior art, the present invention has the following beneficial effects: Through the scheme of traversing the group breadth and cleaving at a specific layer, the binding information between the group and the skeleton is effectively retained. By using this scheme to cleave a large number of existing molecular data sets, a large number of group-skeleton fragments are obtained, these fragments are sorted and recorded, and the occurrence frequency of the fragments is counted. The finally formed pattern matching library can effectively verify the rationality of the generated molecule, and the occurrence frequency of the fragments indicates the feasibility of the structure. The molecular skeleton library and the pattern matching library mentioned in the construction of the virtual screening space can both be custom-built. Selecting the concerned skeleton and the groups with specific properties can quickly generate a series of molecules with specific chemical properties or chemical structures, making the generated virtual screening space more focused on specific functions, improving the hit rate and screening efficiency of the virtual screening. Brief Description of the Drawings
[0030] Figure 1 It is the flowchart of molecule generation of the present invention.
[0031] Figure 2 It is the flowchart of group breadth traversal and cleavage of the present invention. Detailed Embodiments
[0032] The following further describes the present invention with reference to the drawings.
[0033] Example 1
[0034] As Figure 1 shown, the present invention provides a method for constructing an organic molecule virtual screening space based on chemical feasibility verification. By different molecular cleavage methods, a molecular fragment library and a pattern matching library are obtained. The skeletons and specific groups of the fragment library are used to generate new molecules through molecular fragment groups, and then the validity of the new molecules is verified according to the pattern matching library, and finally an organic molecule virtual screening space is formed. The specific steps are as follows:
[0035] S1. Obtain the reported organic molecule dataset and get the SMILES sequences of all molecules;
[0036] S2. Cleave each molecule in the dataset using the RECAP method to obtain a large number of molecular skeletons, and organize and record them to form a molecular skeleton library;
[0037] S3. For each molecule in the dataset, taking the concerned group as the vertex, perform a breadth-first traversal of the molecule. After reaching the specified layer, cleave the molecule to obtain group-skeleton fragments, and count their frequencies to form a pattern matching library;
[0038] S4. Select skeletons from the molecular skeleton library and assemble these skeletons with groups to form complete molecules;
[0039] S5. Cut the generated molecules in the manner described in S3 to obtain the group-skeleton fragments of the molecules. Search for the fragments in the pattern matching library. If the structure can be matched in the pattern matching library, the molecule is reasonable; otherwise, the molecule is unreasonable.
[0040] Specifically, for the molecular cleavage scheme mentioned in the process of forming the molecular skeleton library in S2, the RECAP or BRICS scheme can be adopted. In this paper, the RECAP method is used for molecular fragmentation, and its specific steps are as follows:
[0041] 1. Collect a series of structures active against specific targets and analyze to obtain 11 cleavable molecular bonds.
[0042] 2. Read the input molecule and break one of the cleavable molecular bonds mentioned in STEP1 to obtain multiple sub-nodes.
[0043] 3. Take the sub-nodes after cleavage in step 2. If the sub-node has a molecular bond mentioned in STEP1, repeat step 2; otherwise, mark the sub-node as a leaf node. If no new sub-nodes are generated after all sub-nodes are cleaved, the molecular cleavage is completed.
[0044] 4. Only retain the final fragments during fragmentation without retaining the intermediate process. After the fragments are analyzed, they are merged into building blocks and placed in the fragmentation results.
[0045] The molecules are fragmented into multiple fragments by the RECAP method. A specific skeleton is selected from all the molecular fragment result sets to generate a molecular skeleton library. The skeleton selection rules are as follows: if a fragment contains only small functional groups, the fragment is not retained; ring structures are retained; fragments containing double bonds or triple bonds are retained.
[0046] Specifically, for the construction of the pattern matching library mentioned in S3, the process of molecular cleavage based on breadth-first traversal of the molecules is as Figure 2 shown:
[0047] 1. Search for the concerned groups in the molecule to obtain the serial numbers of the groups. Set A represents the set of atoms containing the concerned groups, set B represents the set of edges to be cleaved, and set C represents the set of atoms at the current layer. Add the obtained serial numbers of the groups to sets A and B.
[0048] 2. Traverse all the atoms in set C to obtain the neighbor nodes of each node. If a neighbor node is on a ring, add all the atoms on the ring to set A, and add all the edges connected to the ring except the edge connected to this atom to set C. If a neighbor node is not in A, add the node to set A; empty set C and add the neighbor node to set C. Execute STEP2 again until the traversed layer reaches the set threshold, obtain the neighbor nodes of each atom in set C and the corresponding edges, and add the edges to set B.
[0049] 3. Break the edges in set B to fragment the molecule into multiple fragments, and find the target fragments according to set A.
[0050] 4. Collect the target fragments, convert the target fragments into SMILES sequences, and remove the irrelevant symbols in the sequences, such as the wildcard '*', the symbols ' / ' and '\' representing the spatial structure, etc. Convert the processed SMILES sequences into SMARTS codes for matching, count the frequencies of various codes, and finally form a pattern matching library.
[0051] The pattern matching library stores the SMARTS codes of the fragments and their occurrence frequencies. In the molecular validity verification, the credibility of the structure validity verification results can be quantified according to the fragment occurrence frequencies.
[0052] Specifically, for the assembly of the skeleton and groups mentioned in S4 to form a complete molecule, the specific process is as follows:
[0053] Select n skeletons. The set of the number of substitution sites for each skeleton is S = {S1, S2,..., S n}, and select m groups. The n skeletons are connected and combined with each other at any substitution site to form a larger skeleton, and then groups are connected to one or more remaining substitution sites of the skeleton to form a complete molecule. The number of molecules that can be generated after substitution is:
[0054]
[0055] Specifically, in step S5, the validity of the molecule is verified. The molecule to be verified is selected, and the molecule is cleaved in the manner described in S3 to obtain group-skeleton fragments. The RDKIT toolkit is used to match the fragments in the pattern matching library. If there are similar fragments in the pattern matching library, it proves that the molecule is valid; otherwise, the molecule is invalid.
[0056] Although the present invention has been described herein with reference to the illustrative embodiments of the present invention, the above embodiments are only preferred embodiments of the present invention. The embodiments of the present invention are not limited by the above embodiments. It should be understood that those skilled in the art can design many other modifications and embodiments, and these modifications and embodiments will fall within the scope of the principles and spirit disclosed in this application.
Claims
1. A method for constructing a virtual screening space of organic molecules based on chemical feasibility verification, characterized in that It includes the following steps: S1. Obtain the reported organic molecule dataset to get the SMILES sequences of all molecules; S2. Use the RECAP method to cleave each molecule in the dataset to obtain a large number of molecular skeletons, and organize and record them to form a molecular skeleton library; S3. For each molecule in the dataset, with the concerned group as the vertex, perform a breadth-first traversal of the molecule. After reaching the specified layer, cleave the molecule to obtain group-skeleton fragments, and count their frequencies to form a pattern matching library; S4. Select skeletons from the molecular skeleton library, and assemble these skeletons with groups to form complete molecules; S5. Cut the generated molecules in the way of S3 to obtain the group-skeleton fragments of the molecules. Search for the fragments in the pattern matching library. If the fragments can be matched in the pattern matching library, the molecule is reasonable; otherwise, the molecule is unreasonable; In step S3, when performing a breadth-first traversal of the molecule and cleaving the molecule after reaching the specified layer, the rule is: perform a breadth-first traversal of the molecule with the specified group as the vertex. When reaching the specified layer, obtain the molecular bonds between the molecules in the current layer and the next layer, and break the obtained bonds to get the molecule group-skeleton fragments; during the traversal, if the atom is in a ring, stop the traversal, retain the ring structure, and break the other bonds on the ring that are not connected to the specified group to get the group-ring fragments.
2. The method for constructing an organic molecule virtual screening space based on chemical feasibility verification according to claim 1, wherein Step S2 specifically includes the following steps: The first step: Molecular cleavage. Use the RECAP method to perform molecular fragmentation. The specific method is as follows: STEP1: Collect the structures active against specific targets and analyze to obtain 11 cleavable molecular bonds; STEP2: Read the input molecule and break one of the cleavable molecular bonds mentioned in STEP1 to obtain multiple sub-nodes; STEP3: Take the sub-nodes after cleavage in STEP2. If the sub-node has the molecular bond mentioned in STEP1, repeat STEP2; otherwise, mark the sub-node as a leaf node; if no new sub-nodes are generated after cleavage of all sub-nodes, the molecular cleavage is completed; STEP4: Only retain the final fragments without retaining the intermediate process during fragmentation. After the fragments are analyzed, merge them into building blocks and put them into the fragmentation results; The second step: Select specific skeletons from all the molecular fragmentation results to generate a molecular skeleton library. The skeleton selection rule is: if the fragment only contains small functional groups, the fragment is not retained; retain the ring structure; retain the fragments containing double bonds and triple bonds.
3. The method for constructing an organic molecule virtual screening space based on chemical feasibility verification according to claim 1, wherein The specific process of step S3 is as follows: STEP1: Search for the concerned group in the molecule, obtain the serial number of the group. Set A represents the set of atoms containing the concerned group, set B represents the set of edges to be cleaved, and set C represents the set of atoms in the current layer. Add the serial number of the searched group to sets A and B; STEP 2: Traverse all atoms in set C and obtain the neighbor nodes of each node. If the neighbor node is on a ring, add all atoms on the ring to set A and add all edges connected to the ring except those connected to the atom to set C. If the neighbor node is not in A, add the node to set A. Clear set C and add the neighbor node to set C. Repeat STEP 2 until the number of traversed layers reaches the set threshold. Obtain the neighbor nodes of each atom in set C and the corresponding edges, and add the edges to set B. STEP 3: Break the edges in set B, split the molecule into multiple fragments, and find the target fragment according to set A; STEP 4: Collect target fragments, convert them into SMILES sequences, remove irrelevant symbols in the sequences, convert the processed SMILES sequences into SMARTS codes for matching, count the frequencies of various codes, and finally form a pattern matching library.
4. The method for constructing a virtual screening space of organic molecules based on chemical feasibility verification according to claim 1, characterized in that: The specific process of assembling the skeleton and the group to form a complete molecule in step S4 is as follows: n skeletons are selected, and the number of substitution sites of each skeleton is set to , select m groups, n skeletons are connected to each other at any substitution site to form a larger skeleton, and then the remaining one or more substitution sites of the skeleton are connected with groups to form a complete molecule. The number of molecules that can be generated after substitution is: .
Citation Information
Patent Citations
Protein degradation drug molecular library and construction method thereof
CN109841263A
Method and device for partitioning a molecule
US20050228592A1