Molecular fragmentation method, fragment dataset, adaptation method, and medium
Through the slicing method based on wildgazine structure, the problem of inflexible molecular slicing in the prior art is solved, ensuring that the slicing results contain specific fragments, expanding the diversity of molecular fragments, and suitable for molecular design and drug design.
Patent Information
- Application Number
- PCT/CN2024/142470
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-08
- Filing Date
- 2024-12-25
- Publication Date
- 2025-07-17
AI Technical Summary
The existing molecular slicing technology based on rotatable bonds is not flexible enough to ensure that specific substructures can be obtained after slicing, especially in fragment-based molecular design, it is impossible to effectively collect specific fragment information.
Using a slicing method based on wildgazine structure, by establishing a wildgazine structure table, screening and merging fragment substructures, using SMARTs wildcards to define wildgazine structures, combining PH-PMI three-dimensional spatial mapping, ensuring that the slicing results contain the molecular fragments of interest and expand the diversity of molecular fragments.
The slicing results are implemented to include molecular fragments of interest, expand the diversity of molecular fragments, suitable for real chemical spaces, and support fragment-based molecular design and drug design tasks.
Smart Images

Figure CN2024142470_17072025_PF_FP_ABST
Abstract
Description
Molecular segmentation methods, fragment datasets, diversion methods and media Technical Field
[0001] The present application relates to the fields of biomedicine and computer processing, and in particular to a molecular segmentation method, a fragment data set, a conversion method, and a medium. Background Art
[0002] Most existing molecular segmentation technologies are based on rotatable bonds to decompose molecules, such as RECAP segmentation in rdkit and BRICS splitting. These segmentation methods use a binary tree bisection strategy, taking the entire molecule as the root node and gradually decomposing the molecule into substructures that cannot be further decomposed. These substructures are leaf nodes.
[0003] The inventors discovered that the above fragmentation strategies are inflexible and are primarily used for data generation and verification in retrosynthetic analysis. They cannot guarantee that the fragmented molecules will have the desired specific structure, such as the inclusion of specific substructures. However, in fragment-based molecular design, the acquisition of specific fragment information is crucial, and constructing training datasets for specific fragments requires fragment-based molecular decomposition. Summary of the Invention
[0004] To address the current inflexibility of molecular segmentation based solely on rotatable bonds and the inability to guarantee specific substructures after segmentation, this application proposes a molecular segmentation method, a fragment dataset, a conversion method, and a medium. The technical solutions are as follows:
[0005] In one aspect, a molecular segmentation method is provided, comprising the steps of:
[0006] S1: establishing a wildcard substructure table, which includes multiple wildcard substructures;
[0007] S2: Match the wildcard substructures in sequence and select one or more fragment substructures matching each wildcard substructure in the molecule;
[0008] S3: De-overlapping the fragment substructures and merging two fragment substructures whose overlap rate exceeds a set threshold;
[0009] S4: Classify the fragment substructures after de-overlapping to obtain larger main fragments and smaller edge fragments, and merge the edge fragments into the main fragments connected to them.
[0010] The wildcard substructure table is defined by using SMARTs wildcards, and the wildcard substructure table includes three-dimensional structures.
[0011] Among them, the wildcard substructure has priority, and the wildcard substructure at the front has higher matching priority.
[0012] The overlap ratio is calculated as follows:
[0013] The matching result of a molecule for one of the wildcard substructures is n fragment substructures, and the i-th fragment substructure is recorded as f i ,i∈[1,n];
[0014] Compare the fragment substructures f one by one i With other fragment substructures f j , j∈[1,n] and j≠i;
[0015] Calculate f i With f j The intersection of i∩j ;
[0016] Calculate f i∩j In f j The proportion of is recorded as the overlap rate.
[0017] The steps include:
[0018] Each molecule is mapped to a molecular coordinate point in the PH-PMI three-dimensional space with coordinates (PH, npr1, npr2). The calculation method of PH is as follows:
[0019] TPSA is the topological polar area of the molecule, ClogP is the oil-water partition coefficient, H A is the number of hydrogen bond acceptors of the molecule, H D is the number of hydrogen bond donors in the molecule, and a, b, c, α and β are configurable constants, where a∈[-10,10], b∈[-10,10], c∈[-10,10], α∈[-1,1], β∈[-1,1].
[0020] The steps include:
[0021] After the molecule is divided into n fragment substructures, each fragment substructure is mapped to a fragment coordinate point in the PH-PMI three-dimensional space, so that each molecule corresponds to n fragment coordinate points in the PH-PMI three-dimensional space;
[0022] In the coordinate system of the PH-PMI three-dimensional space, the fragment coordinate points corresponding to the fragment substructures are connected according to their connection mode in the molecule, so that each molecule can be represented by a broken line or a molecular coordinate point;
[0023] The differences and similarities between molecules can be viewed intuitively through the broken lines or molecular coordinate points in the PH-PMI three-dimensional space, as well as the local differences and similarities within the molecules.
[0024] On the other hand, a molecular fragment dataset obtained by one or more molecular segmentation methods is provided, wherein the molecules include a ligand molecule dataset in a protein.
[0025] In another aspect, a molecular fragment dataset obtained by one or more molecular segmentation methods is provided, wherein the molecules include a drug molecule dataset.
[0026] On the other hand, a method for converting more than one molecular segmentation method is provided, including performing amino acid-based segmentation on protein peptide chains; wherein the wildcard structure table is replaced by an amino acid structure table, and the molecule is replaced by a protein peptide chain.
[0027] In another aspect, a method for adapting one or more molecular segmentation methods is provided, comprising performing nucleic acid-based segmentation on deoxynucleotide chains; wherein the wildcard structure table is replaced by a nucleic acid structure table, and the molecule is replaced by a deoxynucleotide chain.
[0028] On the other hand, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements any of the above molecular segmentation methods, or implements a method for converting any of the above molecular segmentation methods.
[0029] The beneficial effects of this application are: proposing a molecular segmentation method that transforms segmentation based on rotatable bonds to segmentation based on wildcard structures, ensuring that the segmentation results all contain molecular fragments of interest and greatly expanding the diversity of molecular fragments. These molecular fragments are defined in the structure table using wildcards and can be modified and adjusted according to different needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] FIG1 is a schematic flow chart of an embodiment of a molecular segmentation method of the present application;
[0031] FIG2 is a schematic diagram of an embodiment of a wildgamete structure table of the present application;
[0032] FIG3 is a schematic diagram of an embodiment of a fragment substructure table of the present application;
[0033] FIG4 is a schematic diagram of a sub-process of an embodiment of the molecular segmentation method of the present application;
[0034] FIG5 is a schematic diagram of a detailed sub-process of FIG4 ;
[0035] FIG6 is a schematic diagram of segmentation statistics of an embodiment of a segmentation dataset of the present application;
[0036] FIG7 is a schematic diagram of segmentation statistics of an embodiment of a segmentation dataset of the present application;
[0037] FIG8 is a schematic diagram of an embodiment of an amino acid structure table of the present application;
[0038] FIG9 is a schematic diagram of an embodiment of the peptide chain cleavage effect of the present application;
[0039] FIG10 is a schematic diagram of an embodiment of a nucleic acid structure table of the present application;
[0040] FIG11 is a schematic diagram of an embodiment of the deoxynucleotide chain cleavage effect of the present application;
[0041] FIG12A is a schematic diagram of an embodiment of a molecular PMI of the present application;
[0042] FIG12B is a schematic diagram of an embodiment of molecular PH-PMI mapping of the present application;
[0043] FIG13 is a schematic diagram of an embodiment of PH-PMI mapping after molecular segmentation of the present application;
[0044] FIG14 is a schematic diagram of the molecular structure corresponding to FIG13;
[0045] FIG15 is a schematic diagram of an embodiment of PH-PMI mapping after molecular segmentation of the present application;
[0046] FIG16 is a schematic diagram of the molecular structure corresponding to FIG15;
[0047] FIG17 is a schematic diagram of an embodiment of PH-PMI mapping of DrugSpace dataset segments of the present application;
[0048] FIG18 is a schematic diagram of the structure of the hardware operating environment of an embodiment of the method of the present application;
[0049] FIG19 is a schematic diagram showing a comparison of the mapping positions of two molecules in PH-PMI and PMI according to an embodiment of the present application;
[0050] FIG20 shows the distribution of three molecules in PMI according to an embodiment of the present application;
[0051] FIG21 shows the distribution of three molecules in PH-PMI according to an embodiment of the present application. DETAILED DESCRIPTION
[0052] To facilitate understanding of the present application, the present application is described in more detail below with reference to the accompanying drawings and specific embodiments. The accompanying drawings provide preferred embodiments of the present application. However, the present application can be implemented in many different forms and is not limited to the embodiments described in this specification. Rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of the present application.
[0053] It should be noted that, unless otherwise defined, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit this application. For example, the term "plurality" includes two or more.
[0054] The following are some explanations of terms used in this application:
[0055] SMARTS: SMiles ARbitrary Target Specification, a molecular description language based on SMILES, which allows the use of wildcards to represent atoms and chemical bonds.
[0056] Rdkit: An open source Python-based chemical information toolkit that uses machine learning methods to generate compound descriptors, fingerprints, calculate compound structure similarity, and display molecules.
[0057] Molecular PMI: A molecular shape analysis and display method that uses principal moments of inertia (PMI) to classify molecular shapes and uses PMI diagrams to provide a two-dimensional visual description of the shapes.
[0058] PH-PMI: A molecular shape analysis and display method that adds a dimension to the molecular PMI for three-dimensional visual description.
[0059] On the one hand, referring to FIG1 , a molecular segmentation method is provided, comprising the steps of:
[0060] S1: establishing a wildcard substructure table, which includes multiple wildcard substructures;
[0061] S2: Match the wildcard substructures in sequence and select one or more fragment substructures matching each wildcard substructure in the molecule;
[0062] S3: De-overlapping the fragment substructures and merging two fragment substructures whose overlap rate exceeds a set threshold;
[0063] S4: Classify the fragment substructures after de-overlapping to obtain larger main fragments and smaller edge fragments, and merge the edge fragments into the main fragments connected to them.
[0064] Specifically, please refer to Figure 2, which is a schematic diagram of an embodiment of a wildcard substructure table. As shown in Figure 2, the wildcard substructure table includes 39 (numbered 0 to 38) wildcard substructures.
[0065] Please refer to Figure 3, which is a fragment substructure table obtained by matching the wildcard substructure numbered 18 in Figure 2, including 42 fragment substructures. The label of each fragment substructure is the number of the matched fragment substructures.
[0066] Please refer to Figures 4 and 5 for detailed sub-flowcharts of the molecular segmentation method of this embodiment. As shown in Figures 4 and 5, after structural matching, overlap removal, segment classification, and merging smaller edge segments into the larger main segment connected to them, the final segmentation result is obtained.
[0067] The molecular segmentation method of this embodiment changes from segmentation based on rotatable bonds to segmentation based on wildcard structures, ensuring that the segmentation results all contain molecular fragments of interest and greatly expanding the diversity of molecular fragments. These molecular fragments are defined in the structure table using wildcards and can be modified and adjusted according to different needs.
[0068] The wildcard substructure table is defined by using SMARTs wildcards, and the wildcard substructure table includes three-dimensional structures.
[0069] To enhance the diversity and flexibility of matching segments, we used SMARTs wildcards to define a wildcard substructure table. For example, all five-membered rings can be represented as "*1~*~*~*~*~1," and all aromatic six-membered rings can be represented as "a1aaaaa1." We also defined several three-dimensional structures, such as "*12~*3~*4~*~1~*5~*~4~*~3~*~5~2" and "*12~*~*3~*~*(~*~2)~*~*(~*~3)~*~1." All the structures in the wildcard substructure table are shown in Figure 2.
[0070] Among them, the wildcard substructure has a priority. The higher the priority, the closer its label is, and the wildcard substructure at the front has a higher priority in matching.
[0071] Furthermore, in the wildcard substructure table, the subsequent wildcard substructure cannot be included in the previous wildcard substructure, thereby ensuring that no repeated matching is performed and that larger fragment substructures are obtained through limited matching.
[0072] Specifically, the wildcard substructure table is defined as above, and the fragments in the table have a priority. The earlier the fragment, the higher the priority. The rdkit toolkit can be used to match the fragments of molecules and substructures. According to the priority defined above, the fragments are matched one by one from high to low. When a part of the molecule matches a fragment, it is marked. In future matches, this part will no longer be searched until the molecule is completely marked or all the fragments in the substructure table are matched. If all the fragments in the substructure table are matched, there may be some unmarked parts left in the molecule to be matched. These parts are marked as remaining atoms. When the fragments are merged, they will continue to be merged into the main fragments connected to them. In this way, a larger fragment substructure can be obtained.
[0073] The overlap ratio in step S3 is calculated as follows:
[0074] The matching result of a molecule for one of the wildcard substructures is n fragment substructures, and the i-th fragment substructure is recorded as f i ,i∈[1,n];
[0075] Compare the fragment substructures f one by one i With other fragment substructures f j , j∈[1,n] and j≠i;
[0076] Calculate f i With f j The intersection of i∩j ;
[0077] Calculate f i∩j In f j The proportion of is recorded as the overlap rate.
[0078] Specifically, in the above-mentioned fragment matching, it is very likely that there are many matching substructures for the same specific structure in a molecule. As shown in Figure 4, this application has developed a de-overlapping strategy to handle the overlapping matched fragments. Specifically, the de-overlapping process is as follows:
[0079] Assume that the matching result of a molecule to a certain shape in the substructure table is n fragments, and the i-th fragment is denoted as f i ,i∈[1,n]. We compare each fragment f one by one i With other fragments f j , j∈[1,n] and j≠i. Calculate f i With f j The intersection of i∩j , and calculate f i∩j In f j The proportion of is recorded as the overlap rate. i∩j With f jThe overlap rate of the fragment f exceeds a certain threshold, indicating that the fragment f j With fragment f i Most of them are overlapping, then f j Merge into f i Among them, that is, f i =f i ∪f j If f i∩j With f j If the overlap rate is lower than a certain threshold, it means that f j With f i The overlapping parts are small and do not constitute the merging conditions, so continue to check f j+1 The threshold for measuring the overlap rate will also change with different fragments. When f j When there are more than 5 atoms, the threshold is set to 0.5. j When the number of atoms is 5 or less, the threshold is set to 0. The threshold here can be set according to different needs.
[0080] Furthermore, the entire overlap removal algorithm is shown in Algorithm 1.
[0081] During the above fragmentation process, many fragments are obtained, and these fragments all conform to the shapes in the substructure table. However, in terms of molecular design and functional group expression, larger ring-forming fragments are more important, so here we merge the above fragments and the remaining atoms.
[0082] Among them, fragments with atomic numbers less than 4 are identified as edge fragments, and other larger fragments are identified as main fragments. The edge fragments are merged into the main fragments connected to them, so as to form greater fragment diversity.
[0083] Furthermore, the merging algorithm is shown in Algorithm 2.
[0084] Furthermore, the overall process of molecular segmentation is shown in Algorithm 3 below:
[0085] where removeOverlap is the de-overlapping function shown in Algorithm 2.
[0086] Applying the above-mentioned molecular segmentation method, the applicant also established a molecular fragment data set and made statistics on the segmentation results, as shown below.
[0087] On the other hand, a molecular fragment dataset obtained by one or more molecular segmentation methods is provided, wherein the molecules include a ligand molecule dataset in a protein.
[0088] In another aspect, a molecular fragment dataset obtained by one or more molecular segmentation methods is provided, wherein the molecules include a drug molecule dataset.
[0089] Specifically, protein PDB files derived from real experiments were downloaded in batches from the https: / / www.rcsb.org website, and their resolution accuracy was within In the above, we applied the molecular segmentation method to the ligand molecules in proteins, generating the fragment dataset LigandFragmentsDataset. We also screened a dataset of clinically active and marketed drug molecules from the CHEMBL official website https: / / www.ebi.ac.uk / chembl / . This dataset includes approximately 8,000 molecules, including small molecules, oligonucleotides, oligosaccharides, and substances in some locations, excluding proteins, genes, and cells. We also applied the molecular segmentation method to this dataset, generating the fragment dataset FragmentsDataset.
[0090] The LigandFragmentsDataset fragment dataset can be used for fragment-based protein pocket drug design tasks. Its detailed information can be seen in Figure 6; the FragmentsDataset fragment dataset is shown in Figure 7.
[0091] On the other hand, a method for converting more than one molecular segmentation method is provided, including performing amino acid-based segmentation on protein peptide chains; wherein the wildcard structure table is replaced by an amino acid structure table, and the molecule is replaced by a protein peptide chain.
[0092] Specifically, we define the amino acid table as shown in Figure 8. The amino acid table here can also be modified according to different task requirements. We selected 20 common amino acids and one other amino acid.
[0093] According to the mutual inclusion relationship in the structure, the matching priority of amino acids is shown in Figure 8.
[0094] By replacing the wildcard structure table in step 1 in the above specific implementation method with the amino acid structure table shown in FIG8 , the amino acid segmentation of the peptide chain can be achieved. The effect after segmentation is shown in FIG9 .
[0095] In another aspect, a method for adapting one or more molecular segmentation methods is provided, comprising performing nucleic acid-based segmentation on deoxynucleotide chains; wherein the wildcard structure table is replaced by a nucleic acid structure table, and the molecule is replaced by a deoxynucleotide chain.
[0096] Wherein, by replacing the wildcard structure table in step 1 in the above specific implementation method with the nucleic acid structure table shown in FIG10 , nucleic acid segmentation of the deoxynucleotide chain can be achieved. The effect after segmentation is shown in FIG11 .
[0097] Specifically, the nucleic acid table in FIG10 includes five nucleic acids, namely, adenine, guanine, cytosine, thymine, and uracil. At the same time, in order to cope with special situations, some different matching forms are further defined for these five nucleic acids.
[0098] It can be seen that according to the conversion method of the above-mentioned grouping and segmentation method, it is also possible to perform amino acid-based segmentation on protein peptide chains and nucleic acid-based segmentation on deoxynucleotide chains. The process only requires changing the above-mentioned Figure 2 to an amino acid matching table or a nucleic acid matching table.
[0099] In addition, based on this segmentation method, we can also achieve an extension of molecular PMI.
[0100] As shown in Figure 12A, traditional PMI maps molecules into a two-dimensional coordinate system with npr1 as the horizontal axis and npr2 as the vertical axis. All molecules fall within the triangular region defined by the points (0,1), (1,1), and (0.5,0.5). Molecules closer to (1,1) tend to be more spherical; those closer to (0.5,0.5) tend to be more flat; and those closer to (0,1) tend to be more linear.
[0101] Based on this, this application proposes a three-dimensional description method PH-PMI. On the basis of traditional PMI calculation, an additional dimension is added. The value of this dimension is calculated by the polar surface area of the molecule and the number of hydrogen bond donors and acceptors of the molecule, and is recorded as PH.
[0102] Wherein, the molecular segmentation method further comprises the steps of:
[0103] Each molecule is mapped to a molecular coordinate point in the PH-PMI three-dimensional space with coordinates (PH, npr1, npr2). The calculation method of PH is as follows:
[0104] TPSA is the topological polar area of the molecule, ClogP is the oil-water partition coefficient (which can be calculated by rdkit), H A is the number of hydrogen bond acceptors of the molecule, H Dis the number of hydrogen bond donors in the molecule, and a, b, c, α and β are configurable constants, where a∈[-10,10], b∈[-10,10], c∈[-10,10], α∈[-1,1], β∈[-1,1].
[0105] The steps include:
[0106] After the molecule is divided into n fragment substructures, each fragment substructure is mapped to a fragment coordinate point in the PH-PMI three-dimensional space, so that each molecule corresponds to n fragment coordinate points in the PH-PMI three-dimensional space;
[0107] In the coordinate system of the PH-PMI three-dimensional space, the fragment coordinate points corresponding to the fragment substructures are connected according to their connection mode in the molecule, so that each molecule can be represented by a broken line or a molecular coordinate point;
[0108] The differences and similarities between molecules can be viewed intuitively through the broken lines or molecular coordinate points in the PH-PMI three-dimensional space, as well as the local differences and similarities within the molecules.
[0109] Specifically, a, b, c, α, and β are configurable constants that can be adjusted according to specific circumstances. In this embodiment, a = 0, b = 0, α = 1, and β = 0.5. Therefore, in the pH-PMI mapping of this embodiment, each molecule can be mapped to a point in three-dimensional space with coordinates (PH, npr1, npr2). PH is calculated as described in the above formula. The calculation of npr1 and npr2 remains consistent with traditional PMI and can be solved using the rdkit library.
[0110] Specifically, as shown in FIG12B , the applicant randomly sampled 1,000 molecules in DrugSpaceX and compared the effects of PMI and PH-PMI. The left side of the figure shows the molecular mapping of traditional PMI, and the right side of the figure shows the molecular mapping of PH-PMI.
[0111] Furthermore, combined with the above-mentioned molecular segmentation method, the molecule can be divided into n substructure fragments, and the PH-PMI, npr1 and npr2 of each fragment can also be calculated. In this way, each fragment can also be mapped to the PH-PMI three-dimensional space, so each molecule can correspond to n points in the PH-PMI three-dimensional space.
[0112] Furthermore, in the PH-PMI coordinate system, the points corresponding to the fragments are connected according to their connection method in the molecule, so that each molecule can be represented by a broken line or a point (when the molecule is very small and cannot be divided, there is only one fragment, which is itself). In this three-dimensional coordinate system, we can intuitively see the differences and similarities of molecules through the molecular broken line, and we can also see the local differences and similarities of molecules.
[0113] Specifically, Figures 13-16 show the PH-PMI mapping of two groups of molecules. The larger points in the PH-PMI coordinate system represent the PH-PMI of the molecule itself, and the smaller points represent the PH-PMI of each fragment after the molecule is divided into fragments.
[0114] As shown in Figures 13 and 14, the three molecules m0, m1, and m2 have only one part with significant differences in structure. This is reflected in the PH-PMI coordinate system as the coordinates of one fragment point of the three molecules differ greatly, while the coordinates of the other fragment points are very close or even overlap.
[0115] It can be seen from this that by dividing molecules into fragments and displaying them in PH-PMI, we can intuitively compare their similarities or local differences, which is more specific and detailed than looking at the differences between single-point molecules.
[0116] Furthermore, pH is calculated from the polar surface area of the molecule, the number of hydrogen bond donors and acceptors, and the topological shape index of the molecule. The calculation method is as follows:
[0117] TPSA is the topological polar area of the molecule, ClogP is the oil-water partition coefficient calculated by rdkit, and H A is the number of hydrogen bond acceptors of the molecule, H D is the number of hydrogen bond donors of the molecule, kappa is the topological shape index of the molecule, which can also be calculated by rdkit, and a, b, c, d, α and β are configurable constants, where a∈[-10,10], b∈[-10,10], c∈[-10,10], d∈[-10,10], α∈[-1,1], β∈[-1,1].
[0118] The parameters can be adjusted according to specific circumstances. In one embodiment, they can be set to: a=0, b=0, c=1, d=0, α=1, β=0.5.
[0119] The following example demonstrates the role of kappa in PH-PMI. When calculating the kappa value, different kappa values can be selected, namely kappa1, kappa2, and kappa3. They focus on measuring different aspects. The calculation of kappa3 is used as an example for demonstration.
[0120] Please refer to Figure 19. When a=0, b=0, c=1, d=1, α=1, and β=0.5, the value of the ordinate is the result of a mixed calculation of TPSA, Kappa, and the number of hydrogen bond donors and hydrogen bond acceptors. Figure 19 shows the comparison between PH-PMI and PMI. It can be seen that some disadvantages of PMI are avoided.
[0121] Specifically, Figure 19 is a comparison of the mapping positions of the two molecules in PH-PMI and PMI. It can be seen from Figure 19 that the positions of the two molecules with little structural similarity in the PMI calculation (left coordinate system) are very close, without forming a big distinction. However, in the PH-PMI calculation (right coordinate system), they have a large difference in Z value and are far apart in space, forming a good distinction.
[0122] To further demonstrate the role of Kappa, assume a = 0, b = 0, c = 0, d = 1, α = 0, and β = 0. In this case, the ordinate of the molecule is calculated solely from the molecular topological shape index, kappa. Figures 20 and 21 show the same three molecules in PH-PMI and PMI, respectively. From left to right, the three test molecules are A, B, and C.
[0123] Figure 20 shows the distribution of the test molecules in PMI. In Figure 20, the mapping positions of the three molecules ABC in PMI are very close, while the topological structures of molecules A and BC are very different, and such differentiation is not captured in the PMI calculation.
[0124] Figure 21 shows the distribution of the test molecules in the PH-PMI. As can be seen, in the PH-PMI coordinate space, molecules B and C are very close to each other, while molecules A are far away. This mapping calculation captures the topological characteristics of the molecules, keeping structurally similar molecules mapped close together while also distinguishing molecules with significant structural differences that are indistinguishable in the PMI. Furthermore, when further differentiation between molecules B and C is needed, the TPSA weights can be added.
[0125] Furthermore, the present application applies the above-mentioned molecular segmentation method to fragment more than 2,000 molecules in the DrugSpace dataset, and maps the obtained molecular fragments into the PH-PMI space, as shown in FIG17 .
[0126] From this, we can see the tendency of the fragment shapes in spatial distribution. At the same time, for molecules with heteroatoms, their polar surface area will increase accordingly. Through calculation, their points in PH-PMI will be distributed in the direction of increasing vertical coordinates, which can be greatly distinguished from molecules without heteroatoms but with similar molecular geometric shapes.
[0127] On the other hand, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements any of the above molecular segmentation methods, or implements a method for converting any of the above molecular segmentation methods.
[0128] Specific reference is made to FIG18 . In practical applications, FIG18 is a schematic diagram of the structure of the hardware operating environment involved in the method of the present application.
[0129] As shown in Figure 18, the hardware operating environment may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0130] Those skilled in the art will understand that the hardware structure for running the method described in this application shown in Figure 18 does not constitute a limitation on the device for running the method, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0131] As shown in Figure 18, memory 1005, a readable storage medium, may include an operating system, a network communication module, a user interface module, and a computer program. The operating system is a management and control program that supports the operation of the network communication module, the user interface module, the computer program, and other programs or software. The network communication module is used to manage and control the network interface 1004; the user interface module is used to manage and control the user interface 1003.
[0132] In the hardware structure shown in Figure 18, the network interface 1004 is mainly used to connect to the background server and communicate data with the background server; the user interface 1003 is mainly used to connect to the client (user end) and communicate data with the client; the processor 1001 can call the computer program stored in the memory 1005 and execute the steps of the aforementioned method.
[0133] In summary, this application has the following beneficial effects:
[0134] The molecular segmentation method of the present invention changes from segmentation based on rotatable bonds to segmentation based on wildcard structures, ensuring that the segmentation results all contain molecular fragments of interest. Moreover, unlike the fixed-shape molecular structure vocabulary, the wildcard structure of the present invention can greatly expand the diversity of molecular fragments, making it more applicable in the real chemical space (approximately 10 60 order of magnitude).
[0135] Furthermore, these molecular fragments are defined using wildcards in a wildcard substructure table, which can be modified and adjusted according to different needs. When a new molecule needs to be segmented or a substructure of a new shape is required, the wildcard substructure table can be further modified and adjusted to meet the needs.
[0136] Furthermore, a high-quality molecular fragment dataset was constructed through this segmentation method, which can be further applied to fragmentation-based molecular drug design methods and models.
[0137] Furthermore, two specific applications are provided, namely, amino acid cleavage of peptide chains and nucleic acid cleavage of deoxynucleotide chains.
[0138] Furthermore, PH-PMI was proposed based on the traditional molecular PMI. Combined with the above-mentioned molecular segmentation method, the similarities and differences of molecules can be viewed and compared in more detail.
Claims
1. A molecular segmentation method, characterized in that, Including the steps: S1: Establish a general gamete structure table, where the general gamete structure table includes multiple general gamete structures; S2: Sequentially match according to the general gamete structures, and screen out one or more fragment sub-structures in the molecule that match each general gamete structure; S3: Perform an overlapping removal process on the fragment sub-structures, and merge two fragment sub-structures with an overlapping rate exceeding a set threshold; S4: Classify the fragment sub-structures after the overlapping removal process to obtain larger main fragments and smaller edge fragments, and merge the edge fragments onto the main fragments connected thereto.
2. The molecular segmentation method according to claim 1, wherein The general gamete structure table is defined using SMARTs wildcards, and the general gamete structure table includes three-dimensional structures.
3. The molecular segmentation method according to claim 1, wherein The general gamete structures have priorities, and the earlier the general gamete structure, the more preferentially it is matched.
4. The molecular segmentation method according to claim 1, characterized in that The calculation of the overlapping rate is as follows: The matching result of a molecule for one of the said general gamete structures is n fragment sub-structures, and the i-th fragment sub-structure is denoted as f i , where i ∈ [1, n]; Compare the fragment sub-structures f one by one i with other fragment sub-structures f j , where j ∈ [1, n] and j ≠ i; Calculate f i The intersection with f j is denoted as f i∩j ; Calculate f i∩j In f j The proportion is recorded as the overlap rate.
5. The molecular segmentation method according to any one of claims 1-4, characterized in that, Also including the steps: Map each of the said molecules to a molecular coordinate point in the PH-PMI three-dimensional space, whose coordinates are (PH, npr1, npr2), and the calculation method of PH is as follows: TPSA is the topological polar surface area of the molecule, ClogP is the octanol-water partition coefficient, H A is the number of hydrogen bond acceptors of the molecule, H D is the number of hydrogen bond donors of the molecule, a, b, c, α, and β are adjustable constants, where a ∈ [-10, 10], b ∈ [-10, 10], c ∈ [-10, 10], α ∈ [-1, 1], β ∈ [-1, 1].
6. The molecular segmentation method according to claim 5, characterized in that, The calculation method of pH is as follows: kappa is the topological shape index of the molecule, d is a settable constant, where d ∈ [-10, 10].
7. The molecular segmentation method according to claim 5, characterized in that Also including the steps: After splitting the molecule into n fragment sub-structures, map each fragment sub-structure to a fragment coordinate point in the PH-PMI three-dimensional space, so that each molecule corresponds to n fragment coordinate points in the PH-PMI three-dimensional space; In the coordinate system of the PH-PMI three-dimensional space, connect the fragment coordinate points corresponding to the fragment sub-structures according to their connection manners in the molecule, so that each molecule can be represented by a broken line or molecular coordinate points; Through the broken line or molecular coordinate points in the PH-PMI three-dimensional space, the differences and similarities between the molecules can be visually viewed, and at the same time, the local differences and similarities within the molecules can be viewed.
8. A molecular fragment data set obtained by the molecular segmentation method according to any one of claims 1-4, characterized in that, The molecule includes a ligand molecule data set in a protein.
9. A molecular fragment data set obtained by the molecular segmentation method according to any one of claims 1-4, characterized in that, The molecule includes a drug molecule data set.
10. A method of adapting the molecular segmentation method according to any one of claims 1-4, characterized in that, Including performing amino acid-based splitting on a protein peptide chain; wherein, replacing the general gamete structure table with an amino acid structure table and replacing the molecule with a protein peptide chain.
11. A method of adapting the molecular segmentation method according to any one of claims 1-4, characterized in that, Including performing nucleic acid-based splitting on a deoxyribonucleotide chain; wherein, replacing the general gamete structure table with a nucleic acid structure table and replacing the molecule with a deoxyribonucleotide chain.
12. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, and when the computer program is executed by a processor, it implements the molecule splitting method according to any one of claims 1-7, or implements a conversion method of the molecule splitting method according to claim 10 or 11.
Citation Information
Patent Citations
Three-dimensional substructure-based drug molecule comparison method
CN107657146A
Molecular generation method and device, molecular design method and device and electronic equipment
CN115691704A
Biosynthetic pathway prediction method and system based on deep learning
CN116825202A
Biological inverse synthesis prediction method and device based on deep search and electronic equipment
CN116935969A
Molecular segmentation method, fragment data set, transformation method and medium
CN117789844A