Natural-like product molecular structure generation method and system
Through molecular cleavage and constraint generation models, the intrinsic semantic information of natural product fragment sequences is captured, and the problem of natural product structure processing in the prior art is solved. The generated molecular structure is excellent in drug similarity and synthesis accessibility, and the biological correlation of natural products is retained.
Patent Information
- Application Number
- CN202510440257.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-24
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing generative models are difficult to deal with complex natural product structures and cannot effectively retain their biological correlations, resulting in difficulty in generating effective natural product heuristic molecular structures.
Using a fragment-enhanced sampling method, the intrinsic semantic information of natural product fragment sequences is captured through molecular cleavage and constraint generation models to generate natural product heuristic molecular structures that meet the expected standards.
The accuracy of molecular reconstruction was achieved to reach 99.9%, effectively retaining the biological correlation of natural products, and the generated molecular structure performed excellently in drug similarity and synthesis accessibility.
Smart Images

Figure CN120260701A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning and pharmacology, and more specifically, to a method and system for generating molecular structures of natural product-like compounds. Background Art
[0002] Natural products (NPs) have been evolutionarily selected over millions of years and can tightly bind to biological macromolecules, thus possessing important biological activities and medicinal values. With the rapid development of pharmacology and synthetic technologies, natural products have gradually become an important research source, providing bioactive compounds with novel molecular skeletons. Natural products, with their high structural diversity and complexity, such as a high proportion of sp3 carbon atoms, stereocenters, and diverse ring systems, represent an incompletely explored chemical space and have great potential for drug design.
[0003] In traditional drug research and development, researchers mainly design and optimize lead compounds through experiments and their own experience. This research and development method has the disadvantages of high cost, high risk, and long cycle. Moreover, screening drugs only through experimental methods is time-consuming and laborious. In recent years, with the booming development of artificial intelligence technology, drug research and development using deep learning processes large-scale data sets through deep learning models such as neural networks. Introducing molecular generation models into the drug discovery process has become a research hotspot in molecular design and synthesis, opening up new perspectives and approaches for drug design and development. Currently, commonly used generation models include generative adversarial networks (GANs), variational autoencoders (VAEs), and diffusion models, etc. Among them, constrained generation models aim to generate molecules with specific physicochemical properties or optimize the properties of molecules, expecting the newly generated molecules to have the same or similar activities.
[0004] However, current generation models are more applied in the synthesis of molecules, and there are fewer molecular generation methods focusing on natural products. Existing generation models face many challenges when dealing with natural products. First, they lack the ability to handle the complex natural product structures including stereochemical information. Second, natural products often have biologically relevant molecular scaffolds and pharmacophore patterns, and existing molecular generation models often ignore the transfer of such important structural features to compound libraries inspired by natural products, that is, they cannot inherit the biological relevance of natural products. Therefore, generating effective natural product-inspired molecular structures remains a challenging task. Summary of the Invention
[0005] To overcome the above-mentioned defects in the prior art of ignoring the biological relevance of natural products and being difficult to generate effective natural product-inspired molecules, the present invention provides a method and system for generating molecular structures of natural product-like compounds based on fragment-enhanced sampling.
[0006] To solve the above technical problems, the technical solution of the present invention is as follows:
[0007] A method for generating the molecular structure of a natural product-like substance, comprising:
[0008] Collecting natural product structure data and cleaning and preprocessing it to obtain a molecular data set;
[0009] Performing molecular cleavage on the molecular data set to obtain a natural product fragment sequence, and performing molecular reconstruction verification on the natural product fragment sequence;
[0010] Inputting the natural product fragment sequence that has passed the molecular reconstruction verification into a constraint generation model, where the constraint generation model is used to capture the intrinsic semantic information of the natural product fragment sequence through a self-attention mechanism and output a natural product fragment generation sequence;
[0011] Reconstructing the natural product fragment generation sequence based on molecular reconstruction rules to generate the molecular structure of a natural product-like substance.
[0012] Furthermore, the present invention also proposes a natural product-inspired molecular structure generation system based on fragment-enhanced sampling, comprising:
[0013] A data collection module for collecting natural product structure data and cleaning and preprocessing it to obtain a molecular data set;
[0014] A molecular cleavage module for performing molecular cleavage on the molecular data set to obtain a natural product fragment sequence, and performing molecular enhanced sampling and reconstruction verification on the natural product fragment sequence;
[0015] A sequence generation module equipped with a constraint generation model for inputting the normalized natural product fragment sequence that has passed the molecular reconstruction verification into the constraint generation model; the constraint generation model takes Transformer as the core and captures the intrinsic semantic information of the natural product fragment sequence through a self-attention mechanism and outputs a natural product fragment generation sequence;
[0016] A reconstruction module for reconstructing the natural product fragment generation sequence based on molecular reconstruction rules to generate the molecular structure of a natural product-like substance.
[0017] Furthermore, the present invention also proposes a device, comprising a memory and a processor, where computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, the processor is caused to execute all or part of the steps of the method for generating the molecular structure of a natural product-like substance proposed by the present invention.
[0018] Furthermore, the present invention also provides a storage medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, all or part of the steps of the method for generating a molecular structure of a natural product-like proposed by the present invention are implemented.
[0019] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0020] The present invention provides a natural product-inspired generation model based on fragment-enhanced sampling. Inspired by existing natural product library synthesis strategies, such as bio-guided synthesis and pseudo-natural product strategies, the model adopts a fragment-based paradigm for multi-objective generation and optimization to generate natural product-inspired molecular structures that meet the expected standards.
[0021] Through a special molecular cleavage method, the present invention first generates disordered initial fragments and then generates a canonical natural product fragment sequence through a precise reconstruction algorithm, ensuring that the accuracy of molecular reconstruction reaches 99.9%. By constraining the generation model to capture the intrinsic semantic information of the natural product fragment sequence, more specific substructures or general structural features in the natural product are retained, which helps to inherit the biological relevance of the natural product. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flowchart of a method for generating a molecular structure of a natural product-like according to an embodiment of the present invention.
[0023] Figure 2 It is a conceptual diagram of the generation of a molecular structure of a natural product-like according to an embodiment of the present invention.
[0024] Figure 3 It is a flowchart of the pretreatment of natural product fragments according to an embodiment of the present invention.
[0025] Figure 4 It is a flowchart of fragment sequence enhanced sampling and normalization of molecules according to an embodiment of the present invention.
[0026] Figure 5 It is a schematic diagram of an example of molecular reconstruction from a natural product fragment sequence according to an embodiment of the present invention.
[0027] Figure 6 It is a schematic diagram of the characteristic distribution of a natural product-inspired molecular set according to an embodiment of the present invention.
[0028] Figure 7 It is a schematic diagram of the molecular mass evaluation results of different generation models according to an embodiment of the present invention.
[0029] Figure 8 It is a schematic diagram of the structure-oriented molecular generation performance evaluation according to an embodiment of the present invention.
[0030] Figure 9 Schematic diagram of the overall performance of the model of the present invention and the benchmark model in the task of generating anti-malaria activity-guided molecules.
[0031] Figure 10 Architecture diagram of the natural product-like molecule structure generation system shown according to an embodiment of the present invention. Detailed implementation manners
[0032] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0033] The terms used in the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0034] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0035] The present invention will be described in detail below with reference to the drawings and specific implementation manners.
[0036] Embodiment 1
[0037] This embodiment proposes a method for generating natural product-like molecule structures, as Figure 1 shown, which is a flowchart of the method for generating natural product-like molecule structures in this embodiment.
[0038] In the method for generating natural product-like molecule structures proposed in this embodiment, the following steps are included:
[0039] S1. Collect natural product structure data and clean and preprocess it to obtain a molecular data set;
[0040] S2. Sequentially perform molecular cleavage on the molecular dataset to obtain a canonical natural product fragment sequence, and perform molecular reconstruction verification on the fragment sequence;
[0041] S3. Input the fragment sequence verified by molecular reconstruction into a constrained generation model, which is used to capture the intrinsic semantic information of the natural product fragment sequence through the self-attention mechanism and output a heuristic generation sequence of natural product fragments;
[0042] S4. Reconstruct the generated sequence of natural product fragments based on the molecular reconstruction rules to generate a molecular structure of a natural product-like compound.
[0043] As Figure 2 shown, this embodiment focuses on a molecular deep generation model of natural products combined with an efficient fragment extraction and reconstruction algorithm to generate pseudo-natural product virtual compounds that meet characteristic constraints, so as to explore the heuristic chemical space of natural products.
[0044] The model proposed in this embodiment generates disordered initial fragments through a special molecular cleavage method, and then generates a canonical natural product fragment sequence through an accurate reconstruction algorithm, ensuring that the accuracy of molecular reconstruction reaches 99.9%. This effectively avoids common connection problems in traditional fragment-based models, such as "where to connect new fragments" and "which chemical bond to choose". This embodiment further captures the intrinsic semantic information of the natural product fragment sequence through a constrained generation model, retaining more specific substructures (such as the core scaffold) or general structural features (such as functional groups and ring systems) in natural products, which helps to inherit the biological relevance of natural products. In addition, the uniqueness and accuracy of molecular preprocessing in this embodiment can effectively ensure the reliability of this generation paradigm.
[0045] In an alternative embodiment, in step S1, natural product structure data is collected from available datasets in the public domain, and the available datasets include COCONUT and TeroKIT, and the collected natural product structures are cleaned and preprocessed.
[0046] In an alternative embodiment, in step S1, cleaning and preprocessing the collected natural product structures includes the following steps:
[0047] Collect all natural products with stereochemical information in the form of canonical SMILES strings;
[0048] Eliminate the stereochemical information of natural products;
[0049] Neutralize the charge;
[0050] Remove glycosylation modifications, that is, polysaccharide side chains (>2 polysaccharides);
[0051] Filter the data according to molecular normalization and remove duplicates;
[0052] Calculate molecular descriptors for each filtered molecular entry to form a molecular dataset.
[0053] In this embodiment, data filtering is performed according to molecular normalization to maintain consistency. Its complete process includes: desalting, electro-neutralization, deglycosylation, checking molecular validity, and removing duplicates.
[0054] In this embodiment, all natural products with stereochemical information are collected in the form of canonical SMILES strings. After data filtering is completed, molecular descriptors are calculated for each molecular entry, which are further used as constraints for model training and indicators for model evaluation.
[0055] In an alternative embodiment, in step S2, for a given molecular dataset, the first step is to perform fragment extraction, that is, to convert the molecular graph into an initial fragment sequence by using a special molecular fragmentation method; the second step is to standardize the fragment sequence and normalize the initial fragment sequence into a natural product fragment sequence that can be directly input into the model.
[0056] In an alternative embodiment, in step S2, the molecular dataset is subjected to molecular cleavage, including the following steps:
[0057] S210. Define a natural product structure as a directed graph S i =(V i , E i ) and use it as a subgraph of the target molecule G; where V i is the set of atoms of the i-th molecule in the molecular dataset, and E i is the set of chemical bonds of the i-th molecule in the molecular dataset;
[0058] S220. Based on a preset molecular cleavage rule, decompose the target natural product molecule G into molecular fragments; where the molecular fragments are initial fragments containing virtual atom IDs, and the bond identifiers in the target molecule G are used as the mapping identifiers of the molecular fragments;
[0059] S230. Based on the virtual atom IDs of the initial fragments, perform normalization and enhanced sampling processing on the generated molecular fragment sequence to obtain a normalized natural product fragment sequence containing rich semantic information.
[0060] Among them, in order to extract natural product fragments, the target molecule G is decomposed into fragments through a custom rule. This molecular cleavage method allows the model to capture key features in the molecular structure, and then generate molecules with expected physicochemical properties and biological activities.
[0061] Further, in an optional embodiment, the molecular cleavage rule includes:
[0062] Find all single bonds (μ, ν) and save them as the chemical bond set E; wherein, for the atoms in the ring where atom μ is located, if there is an atom ν that does not belong to this ring or is located in another ring, then save the single bond between atom μ and atom ν to the chemical bond set E;
[0063] Find all chemical bonds that conform to the substructure decomposition algorithm BRICS rule in reverse synthesis chemistry from the chemical bond set E, and break all the chemical bonds that conform to the BRICS rule to complete the molecular cleavage task.
[0064] In this embodiment, the result of molecular cleavage is the initial fragment containing virtual atoms, and the bond identifiers in the original target molecule G are retained as mapping identifiers for marking the positions where the two fragments can be connected. For example, [2*]=C([8*])O, where the mapping identifiers are 2 and 8.
[0065] Further optionally, divide the rings of the target molecule G skeleton as equally as possible into two parts. The separation of the rings will result in two fragments containing a common bond, and the carbon atoms on this common bond will be replaced by virtual atoms [*].
[0066] In this embodiment, by performing molecular cleavage according to the preset rules, the original molecule can be reconstructed without relying on virtual atom IDs.
[0067] In this embodiment, each normalized fragment containing rich semantic information is marked as a natural product fragment and used as the basic unit for training the model.
[0068] In an optional embodiment, in step S230, the normalization and enhanced sampling processing of the generated molecular fragment sequence includes the following steps:
[0069] S231. Generation of enhanced sampling sequence: Calculate the total length L of the initial natural product fragment sequence M 0 where Through cyclic shift operation, adjust the first fragment of the sequence M 0 to the end in turn, and generate L variant sequences M = {M 0 , M 1 , …, M L}, which constitute the enhanced sampling sequence set M;
[0070] S232. Verification of fragment adjacency relationship: Traverse each sequence M i in the set M, and check the first mapping identifiers of adjacent fragment pairs from left to right in turn; if the identifiers are different, execute step S233; if the identifiers are the same, execute step S234;
[0071] S233, Discontinuous fragment recombination: Search for fragments in the current sequence that are the same as the fragment at the first position Swap the positions of and , and go to step S234; if not, move the first fragment of the current sequence to the end to generate a new sequence M, and jump to step S232 for re-verification;
[0072] S234, Fragment iterative merging: Merge adjacent fragments and into an intermediate molecule, and continue to merge with until the current sequence traversal is completed, and then jump to step S235;
[0073] S235, Sequence normalized output: When all fragments of a certain sequence M i are merged; record the fragment merging order; remove the mapping identifier to generate a normalized fragment sequence;
[0074] S236, Termination condition determination: Check whether the set M has been traversed; if not, continue to process the next sequence and jump to step S232; if it has been completed, output all normalized fragment sequences as the final result.
[0075] In this embodiment, the finally generated natural product-inspired fragment sequence can reconstruct the molecule according to subsequent molecular reconstruction rules, without relying on virtual atom identifiers.
[0076] Exemplarily, as Figure 3 shown, it is the molecular fragment preprocessing flowchart of this embodiment.
[0077] In an optional embodiment, in step S2, molecular reconstruction verification is performed on the natural product fragment sequence, including the following steps:
[0078] S240, Connect the natural product fragment sequence from left to right in sequence according to the breakpoint identifier [*]; if the structure of the connected natural product fragment sequence is the same as the initial molecular structure before molecular cleavage, it is regarded as successful molecular reconstruction verification.
[0079] Wherein, the breakpoint identifier [*] is obtained by replacing the carbon atom on the broken chemical bond with a virtual atom [*] when performing molecular decomposition based on preset molecular cleavage rules.
[0080] This embodiment strictly examines all the heuristically generated natural product fragment sequences, ensuring an accuracy of 99.9% in molecular reconstruction and being unaffected by chiral differences. In contrast, HierVAE - decoder also uses natural product fragments as the building blocks for generation, but its accuracy is only about 80%. This indicates that the virtual atom omission attachment prediction process is retained in the natural product fragments and the loss during the generation process is reduced. This significant improvement shows that the fragment - based model has high reliability and interpretability.
[0081] In an optional embodiment, the constrained generation model is hierarchically stacked by several Transformer decoders. Specifically, the architecture of the constrained generation model is as follows:
[0082] (1) Hierarchical Transformer decoder stack: Composed of N identical Transformer decoders connected in series. Each layer contains an improved multi - head self - attention sub - layer, and its attention mechanism uses constraint - enhanced scaled dot - product attention:
[0083]
[0084] where Q, K, and V are the query vector, key vector, and value vector respectively; d k is the scaling factor;
[0085] (2) Chemical - aware positional encoding: Combining relative positional encoding to represent the distance between fragments;
[0086] (3) Constrained sampling algorithm: Using beam search for chemical validity pruning, that is, retaining candidates that satisfy the following constraints at each step:
[0087] under
[0088] x q ~p(x0,...,x Q- 1 ,c)
[0089] where p(·) is the probability function; this sampling problem is defined as searching for the most likely hypothesis y* according to the model and a set of constraint conditions c; its expression is:
[0090]
[0091] (4) Training optimization strategy using a multi - task loss function: The constrained generation model uses the negative log - likelihood function and the latent space regularization term as the loss function;
[0092] (5) Molecular structure screening and optimization based on multi - dimensional evaluation metrics: Using quantitative estimates of drug similarity metrics and synthetic accessibility scores, and allowing custom thresholds to achieve efficient candidate molecule selection.
[0093] Exemplarily, the constraint generation model is trained using the negative log-likelihood function as the loss function L(x); its expression is:
[0094]
[0095] where x q represents the output sequence of the q-th decoder, c represents the constraint condition, and Q is the number of decoders in the constraint generation model.
[0096] Exemplarily, the decoder generates subsequent samples by recursively adopting the beam search algorithm; wherein, the beam search algorithm retains K locally highest-probability candidates at each time step t to obtain N sequences, resulting in a natural product-inspired fragment sequence; where the hyperparameter K is the beam width.
[0097] In this embodiment, the multi-head self-attention sublayer is used to ensure that the prediction at the current position depends only on the sequence embedding information before that position; its self-attention mechanism adopts a scaled dot-product attention function, enabling the model to capture information from different subspaces at different positions. Among them, each processed natural product-inspired fragment sequence is fed into the embedding layer and added with positional encoding, and here standard sine positional encoding is used to enable the transformer to retain the relative position information of the words in the sentence.
[0098] Finally, the N natural product-inspired fragment sequences are transformed into the final molecular structure through a molecular reconstruction algorithm, which is the reverse process of molecular cleavage.
[0099] Exemplarily, as Figure 4 shown, it is a schematic diagram of the natural product fragment sequence enhancement sampling and normalization process. Among them, in this embodiment, the stereochemical information of the original natural product molecule is first eliminated, and 6 unordered initial fragment sequences are generated through fragment extraction, and then two natural product-inspired fragment sequences are generated through the enhancement sampling and normalization process.
[0100] Exemplarily, as Figure 5 shown, it is a schematic diagram of the molecular reconstruction process of the natural product-inspired fragment sequence. Among them, in this embodiment, the 6 natural product-inspired fragment sequences generated by molecular cleavage are gradually connected from left to right to finally form a complete molecular structure.
[0101] Exemplarily, as Figure 6As shown, this example application generates the property distribution of a heuristic molecular set of 5,000 natural products using the NIMO 2.0 sampling. From the distribution results, it can be seen that compared with the original chemical space of natural products, the chemical space reconstructed by NIMO 2.0 shows a trend of shrinking towards the ideal values in terms of target properties, including the quantitative evaluation of drug likeness (Qed), the octanol-water partition coefficient (logP), and the synthetic accessibility score (SAS). It should be noted that since the molecular weight is correlated with Qed, its distribution is also optimized. Statistical data shows that the molecular property distributions of hydrogen bond donors (HBD), hydrogen bond acceptors (HBA), and rotatable bonds (RB) exhibit similar and slightly concentrated characteristics. This result indicates that the natural product-inspired molecular structures generated by the present generation model NIMO 2.0 have a more drug-like distribution than the natural products in the training set.
[0102] Exemplarily, as Figure 7 shown, five common indicators for evaluating the molecular quality, such as molecular validity, internal diversity, novelty, nearest neighbor similarity (SNN), and fragment similarity (Frag), are used to evaluate the molecular quality of different generation models. Among them, the higher the values of the first three indicators, the better, and the lower the values of the last two indicators, the better. NIMO, MCMG, and QBMG are existing molecular generation models. Compared with them, the model NIMO 2.0 of the present invention performs optimally. Specifically, the model of the present invention achieves the best results in all five molecular quality evaluation indicators, showing particularly outstanding performance and significant progress in all of them.
[0103] Exemplarily, as Figure 8 shown, it is a structure-oriented evaluation of molecular generation performance. Among them, "Success" represents the success rate of generating terpene compound molecules; "RSs" represents the ring system; "FGs" represents the functional groups; "Coverage" represents the ratio of the number of unique RSs / FGs extracted from the generated set to the total number of RSs / FGs in the training set; "Recovery" represents the ratio of the number of unique RSs / FGs extracted from the generated set to the total number of unique RSs / FGs in the generated set.
[0104] To evaluate the accuracy of the model in generating molecules for specific structural regions, this embodiment specifically designs a generation task for complex polycyclic terpene compounds. All generated molecules are evaluated by NPClassifier to determine whether their structures belong to terpene compounds. As Figure 8 summarized, the model of the present invention performs best in the performance of effectively constructing terpene compounds, followed by NIMO. This indicates that the model of the present invention can generate more conforming molecules under the constraints of established structural rules.
[0105] Increasing evidence suggests that preserving specific substructures (such as the core scaffold) or general structural features (such as functional groups and ring systems) helps to inherit the biological relevance of natural products. Therefore, the present invention conducts a statistical analysis on the functional groups and ring systems of all model-generated molecules. According to the "Recovery" index, the model of the present invention shows the best replication ability in terms of functional groups and ring systems; while in the "Coverage" index, the model of the present invention ranks first in terms of the coverage rate of functional groups.
[0106] Obviously, the model proposed by the present invention performs excellently in replicating the substructure features of the original training set and demonstrates the best inheritance of the biological relevance of natural products among the benchmark models.
[0107] Exemplarily, as Figure 9 shown, it is the overall performance of the model of the present invention and the benchmark model in the task of generating anti-malaria activity-guided molecules. Among them, "EF[X]" represents the enrichment rate of active compounds.
[0108] From Figure 9 it can be seen that the proportion of anti-malaria active molecules in the original training set is 10.0%. The activity rate of the molecules generated by the benchmark model is increased to 52.1%, while the activity rate of the molecules generated by the model NIMO 2.0 proposed by the present invention is increased to 89.5%, showing the best performance. At the same time, the activity enrichment rate also gradually increases: the enrichment rate of the top 1% of active compounds is increased by 82%, the enrichment rate of the top 10% of active compounds is increased by 76%, and the enrichment rate of the top 50% of active compounds is increased by 68%, all significantly higher than the other two benchmark models. These results indicate that the molecular generation model proposed by the present invention performs excellently in conditional constrained generation and can significantly increase the generation proportion of potential anti-malaria active molecules.
[0109] Example 2
[0110] In this example, the method for generating the structure of natural product-like molecules proposed in Example 1 is applied to propose a system for generating the structure of natural product-like molecules. As Figure 10 shown, it is the architecture diagram of the system for generating the structure of natural product-like molecules in this example.
[0111] In the system for generating the structure of natural product-like molecules proposed in this example, it includes:
[0112] A data acquisition module, which is used to collect natural product structure data and clean and preprocess it to obtain a molecular data set;
[0113] A molecular cleavage module, which is used to cleave the molecular data set to obtain a natural product fragment sequence, and perform molecular enhanced sampling and reconstruction verification on the natural product fragment sequence;
[0114] A sequence generation module, on which a constraint generation model is mounted, is configured to input a normalized natural product fragment sequence verified by molecular reconstruction into the constraint generation model; the constraint generation model takes Transformer as the core and captures the internal semantic information of the natural product fragment sequence through the self-attention mechanism, and outputs a natural product fragment generation sequence;
[0115] A reconstruction module is configured to reconstruct the natural product fragment generation sequence based on molecular reconstruction rules to generate a natural product-like molecular structure.
[0116] It can be understood that the system of this embodiment corresponds to the method of Embodiment 1 above, and the optional items in Embodiment 1 above are equally applicable to this embodiment, so they will not be repeated here.
[0117] Embodiment 3
[0118] This embodiment provides a device, including a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor is caused to execute all or part of the steps of the natural product-like molecular structure generation method proposed in Embodiment 1.
[0119] Embodiment 4
[0120] This embodiment provides a storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, all or part of the steps of the natural product-like molecular structure generation method proposed in Embodiment 1 are implemented.
[0121] Exemplarily, the storage medium includes, but is not limited to, various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0122] Exemplarily, the instructions, programs, code sets, or instruction sets can be implemented using conventional programming languages.
[0123] Exemplarily, the processor includes, but is not limited to, smartphones, personal computers, servers, network devices, etc., and is configured to execute all or part of the steps of the natural product-like molecular structure generation method described in Embodiment 1.
[0124] The terms in the drawings are only for illustrative purposes and should not be construed as limiting the present invention;
[0125] Obviously, the above embodiments of the present invention are merely examples for clearly explaining the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for generating the molecular structure of a natural product-like compound, characterized in that, It includes the following steps: Collect natural product structure data, clean and preprocess it to obtain a molecular dataset; Perform molecular cleavage on the molecular dataset to obtain natural product fragment sequences, and perform molecular reconstruction verification on the natural product fragment sequences; Input the natural product fragment sequences that have passed the molecular reconstruction verification into a constrained generation model, which is used to capture the intrinsic semantic information of the natural product fragment sequences through the self-attention mechanism and output natural product fragment generation sequences; Reconstruct the natural product fragment generation sequences based on molecular reconstruction rules to generate molecular structures of natural product-like molecules.
2. The method for generating the molecular structure of the natural product-like compound according to claim 1, wherein The collection of natural product structure data and its cleaning and preprocessing include: Collect all natural products in the form of canonical SMILES strings; Eliminate the stereochemical information of natural products; Neutralize the charge; Remove glycosylation modifications, i.e., polysaccharide side chains; Perform data filtering according to molecular normalization and delete duplicates; Calculate molecular descriptors for each filtered natural product entry to form a molecular dataset.
3. The method for generating the molecular structure of the natural product-like compound according to claim 1, wherein, The molecular cleavage of the molecular dataset includes: Define a natural product structure as a directed graph S i =(V i , E i ), and use it as a subgraph of the target molecule G; where, V i is the atomic set of the i-th molecule in the molecular dataset, and E i is the chemical bond set of the i-th molecule in the molecular dataset; Based on preset molecular cleavage rules, decompose the target natural product molecule G into fragments; wherein, the natural product fragments are initial fragments containing virtual atom IDs, and the bond identifiers in the target molecule G are used as the mapping identifiers of the molecular fragments; Based on the virtual atom IDs of the initial fragments, perform normalization and enhanced sampling processing on the generated molecular fragment sequences to obtain normalized natural product fragment sequences containing rich semantic information.
4. The method for generating the molecular structure of a natural product-like compound according to claim 3, wherein The molecular cleavage rules include: Find all single bonds (μ, ν) and save them as a chemical bond set E, where μ and ν are atoms; Find all chemical bonds that conform to the substructure decomposition algorithm BRICS rule in reverse synthesis chemistry from the chemical bond set E, and break all chemical bonds that conform to the BRICS rule to complete the molecular cleavage task.
5. The method for generating the molecular structure of the natural product-like compound according to claim 3, wherein The normalization and enhanced sampling processing of the generated molecular fragment sequences include the following steps: Step 1. Enhanced sampling sequence generation: Calculate the total length L of the initial natural product fragment sequence M 0 where Through cyclic shift operations, sequentially adjust the first fragment of the sequence M 0 to the end to generate L variant sequences M = {M 0 , M 1 , …, M L}, forming the enhanced sampling sequence set M; Step 2, Fragment Adjacency Relationship Verification: Traverse each sequence M in set M i , and sequentially check the first mapping identifiers of adjacent fragment pairs from left to right; if the identifiers are different, execute Step 3; if the identifiers are the same, execute Step 4; Step 3, Discontinuous Fragment Recombination: Search for a fragment in the current sequence that is the same as the first identifier and swap it with and go to Step 4; if not, move the first fragment of the current sequence to the end to generate a new sequence M, and jump to Step 2 to recheck; Step 4, fragment iterative merging: Merge adjacent fragments and into intermediate molecules, and continue to merge with until the current sequence traversal is completed, then jump to Step 5; Step 5, Sequence Normalized Output: When all fragments of a certain sequence M i are merged; record the fragment merging order; remove the mapping identifier to generate a normalized fragment sequence; Step 6, termination condition determination: Check whether the set M has been traversed; if not, continue to process the next sequence and jump to step 2; if it has been completed, output all normalized fragment sequences as the final result.
6. The method for generating the molecular structure of a natural product-like compound according to claim 1, wherein The molecular reconstruction verification of the normalized natural product fragment sequences includes the following steps: Connect the normalized natural product fragment sequences in sequence from left to right according to the breakpoint identifier [*]; if the structure of the connected normalized natural product fragment sequence is the same as the initial molecular structure before molecular cleavage, it is regarded as successful molecular reconstruction verification.
7. The method for generating the molecular structure of a natural product-like substance according to any one of claims 1 to 6, characterized in that, The constrained generation model includes: (1) Hierarchical Transformer decoder stack: Composed of N identical Transformer decoder hierarchies cascaded, each layer contains an improved multi-head self-attention sublayer, and its attention mechanism adopts constrained enhanced scaled dot product attention: Among them, Q, K, and V are the query vector, key vector, and value vector respectively; d k is the scaling factor; (2) Chemical perception position encoding: Combine relative position encoding to represent the distance between fragments; (3) Constrained sampling algorithm: Use beam search for chemical validity pruning, that is, retain candidates that meet the following constraints at each step: x q ~p(x0,...,x Q- 1 ,c) where p(·) is a probability function; the sampling problem is defined as searching for the most probable hypothesis y* according to the model and a set of constraints c; its expression is: (4) Training and optimizing the strategy using a multi-task loss function: The constraint generation model uses the negative log-likelihood function and the latent space regularization term as the loss function; (5) Screening and optimizing the molecular structure based on multi-dimensional evaluation indicators: Using quantitative estimation of drug similarity indicators and synthetic accessibility scores, and allowing custom thresholds to achieve efficient candidate molecule selection.
8. A class of natural product molecular structure generation system, which applies the class of natural product molecular structure generation method described in any one of claims 1 to 7, characterized in that, Including: A data acquisition module, which is used to acquire natural product structure data, clean and preprocess it to obtain a molecular data set; A molecular cleavage module, which is used to cleave the molecular data set to obtain natural product fragment sequences, and perform molecular enhanced sampling and reconstruction verification on the natural product fragment sequences; A sequence generation module, on which a constraint generation model is mounted, and is used to input the normalized natural product fragment sequences that have passed the molecular reconstruction verification into the constraint generation model; the constraint generation model takes Transformer as the core, and captures the intrinsic semantic information of the natural product fragment sequences through the self-attention mechanism, and outputs the natural product fragment generation sequences; A reconstruction module, which is used to reconstruct the natural product fragment generation sequences based on the molecular reconstruction rules to generate natural product-like molecular structures.
9. An apparatus, comprising a memory and a processor, wherein computer-readable instructions are stored in the memory, characterized in that, When the computer-readable instructions are executed by the processor, the processor is caused to execute all or part of the steps of the natural product-like molecular structure generation method according to any one of claims 1 to 7.
10. A storage medium having computer-readable instructions stored thereon, characterized in that, When the computer-readable instructions are executed by the processor, all or part of the steps of the natural product-like molecular structure generation method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Compound optimization method based on deep learning connection fragment
CN114187978A
Natural language comparative learning optimization method, system and equipment for drug discovery
CN118398110A