Topology-Driven Chemical Data Completion
The data-driven molecular generation method addresses inefficiencies in materials science by using topology analysis and autoencoders to generate molecules with desired attributes, improving discovery efficiency and reducing resource waste.
Patent Information
- Application Number
- JP2023524956
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-23
- Filing Date
- 2021-11-11
- Publication Date
- 2025-11-05
- Estimated Expiration
- 2041-11-11
AI Technical Summary
Molecular discovery and synthesis in materials science is time-consuming and inefficient, often resulting in numerous candidates lacking desired attributes, leading to wasted resources and failure to explore viable materials.
A data-driven approach using topology data analysis and variational autoencoders to generate new molecules by filling gaps in molecular datasets, focusing on specific molecular backbones and scaffolds to enhance discovery efficiency.
Significantly reduces the number of infeasible candidates and enhances experimental efficiency by generating molecules with desired attributes, thereby optimizing the molecular discovery process.
Smart Images

Figure 0007764106000006 
Figure 0007764106000007 
Figure 0007764106000008
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD This disclosure relates generally to the field of materials science, and more specifically to the design, discovery, and synthesis of polymers. [Background technology]
[0002] Molecular discovery, design, synthesis, and testing often require significant time. This process can be accelerated through various computational tools. These tools can result in a vast number of generated molecular candidates, which often do not have the desired attributes. As a result, families of materials with the desired attributes cannot be adequately explored, and much time, money, and energy is wasted testing candidates that are not feasible or are feasible but do not meet the target requirements. Summary of the Invention
[0003] Embodiments of the present disclosure include systems, methods, and computer program products for generating new molecules.
[0004] In some embodiments, a processor may receive molecular data for a plurality of molecules. The processor may perform topology data analysis on the molecular data to generate a molecular topology map. The processor may identify one or more gaps in the molecular topology map. The processor may generate one or more additional molecules to fill at least one of the gaps.
[0005] In some embodiments of the present disclosure, the plurality of molecules share one or more common molecular properties.
[0006] In some embodiments of the present disclosure, a molecular scaffold may be generated for each of a plurality of molecules, hi some embodiments, a generability score may be generated for each scaffold.
[0007] In some embodiments of the present disclosure, multiple molecules may share a molecular backbone, and one or more additional molecules may contain the molecular backbone.
[0008] In some embodiments of the present disclosure, the one or more additional molecules are generated using a variational autoencoder. In some embodiments, the variational autoencoder can be adjusted by backbone adjustment so that the one or more additional molecules contain a specific molecular backbone. In some embodiments, the variational autoencoder has a variational autoencoder loss function, and the variational autoencoder loss function is modified to include the probability of generating the specific molecular backbone.
[0009] The above summary is not intended to describe each illustrated embodiment or every implementation of the present disclosure.
[0010] The drawings included in this disclosure are incorporated into and form a part of this specification. These drawings illustrate embodiments of the disclosure and, together with the description, serve to explain the principles of the disclosure. The drawings merely illustrate particular embodiments and are not intended to limit the disclosure. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 illustrates a pipeline for new molecule generation according to an embodiment of the present disclosure. [Figure 2A] FIG. 1 illustrates generating an identifier according to an embodiment of the present disclosure. [Figure 2B] FIG. 10 illustrates generating a bit vector using an identifier according to an embodiment of the present disclosure. [Figure 3] FIG. 1 illustrates topology data analysis according to an embodiment of the present disclosure. [Figure 4] FIG. 1 illustrates generating molecular candidates according to an embodiment of the present disclosure. [Figure 5] FIG. 1 illustrates backbone analysis molecule generation according to an embodiment of the present disclosure. [Figure 6]FIG. 1 illustrates a molecule generation pipeline according to an embodiment of the present disclosure. [Figure 7] FIG. 1 illustrates a molecule generation pipeline according to an embodiment of the present disclosure. [Figure 8] FIG. 1 illustrates a cloud computing environment according to an embodiment of the present disclosure. [Figure 9] FIG. 1 illustrates an abstraction model layer according to an embodiment of the present disclosure. [Figure 10] FIG. 1 is a high-level block diagram illustrating an example of a computer system that may be used to implement one or more of the methods, tools, and modules, and any associated functionality, described herein, in accordance with embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] While the embodiments described herein are susceptible to various modifications and alternative forms, specifics thereof have been shown by way of example in the drawings and will be described in detail. It is to be understood, however, that the specific embodiments described are not to be construed in a limiting sense. On the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure.
[0013] Aspects of the present disclosure relate generally to the field of materials science, and more specifically to polymer synthesis. It will be readily understood that the components herein, as generally described herein and illustrated in the drawings, could be arranged and designed in a wide variety of different configurations. Thus, the following detailed description of at least one embodiment of a method, apparatus, non-transitory computer-readable medium, and system, as illustrated in the accompanying drawings, is not intended to limit the scope of the present application as claimed, but merely represents selected embodiments.
[0014] These features, structures, or characteristics described throughout this specification may be combined in any suitable manner or omitted in one or more embodiments. For example, the phrases "example embodiment," "some embodiments," or other similar language used throughout this specification indicate the fact that particular features, structures, or characteristics described in connection with an embodiment may be included in at least one embodiment. Thus, the appearances of the phrases "example embodiment," "some embodiments," "other embodiments," or other similar language throughout this specification do not necessarily all refer to the same group of embodiments, and the described features, structures, or characteristics may be combined in any suitable manner or omitted in one or more embodiments. Furthermore, any connections between elements in the drawings, even if the connections are shown with one-way or two-way arrows, may enable one-way, two-way, or both communication. Additionally, any devices shown in the drawings may be different devices. For example, when a mobile device transmitting information is shown, a wired device may also be used to transmit that information.
[0015] Molecular discovery often takes a significant amount of time. For example, introducing a new polymeric material to the market can require more than 10 years of design, synthesis, and testing. This process can be accelerated through computational molecular design using tools such as combinatorial screening, inverse design, generative modeling, and reinforcement learning. These tools can generate a large number of computer-generated candidates, often on the scale of 10,000,000 candidates. However, the molecular candidates obtained through these tools often lack the desired attributes, such as synthetic feasibility, robust polymerization, and compliance with internal and external regulations. As a result, much time, money, and energy are wasted testing candidates that do not meet the requirements.
[0016] The historical data that drives the computational design and generation of new molecular candidates is incomplete and biased. These shortcomings can lead to the amplification of molecular candidates that lack novelty, lack necessary attributes, or have other undesirable characteristics. Furthermore, these shortcomings can lead to a failure to identify unexplored families of materials, reduced discovery efficiency, and unnecessary increased costs for conducting tests and experiments.
[0017] Experimental performance is critical to the development of new molecules in materials science, food science, preservation science, pharmacology, or any other field that benefits from the generation of new molecules. Data acquisition strategies that bridge the computational and experimental stages of molecular discovery can enhance the experimental phase by uncovering gaps in historical data and using guided candidate generation to complete the missing segments, increasing the percentage of successful results.
[0018] The present disclosure enables such data acquisition by filling gaps in the initial dataset and completing the dataset with previously overlooked molecules. Topological data analysis based on graphs (e.g., Reeb graphs) can reveal gaps in the available data that hinder efficient molecule discovery. The present disclosure can use these gaps to complete the data.
[0019] The present disclosure may use known molecular information as a starting point to find viable molecular candidates. The present disclosure can be described as exploring unknown areas on a map to discover topographical details and then filling in the map with the newly discovered details. The present disclosure uses what is known (e.g., the edges of the known world) to begin exploring what is known to be unknown, thereby filling in details about the unknown. The present disclosure may be likened to a gardener patching a hole in the garden, starting with a stable foothold and filling in the edges of the hole until the hole is completely sealed. The present disclosure may be likened to repairing clothing with a patch, where the patch is attached to a material known to be strong, stretched over the hole, and attached to another material known to be strong on the other side. The present disclosure can use known molecular data on the boundaries of what is known to discover new data.
[0020] Those skilled in the art will recognize that the present disclosure can be applied to generate candidates for any chemical dataset. For example, pharmaceuticals, polymeric materials, and other fields that may benefit from the generation of unknown organic or inorganic molecules will benefit from the present disclosure. Polymeric materials include, for example, products of ring-opening polymerization of cyclic lactones (including monomers and catalysts), polyimide block copolymers, and polyacrylic acids. For brevity and clarity, the present disclosure focuses on components of photoresists, such as photoacid generator (PAG) molecules. PAG molecules are frequently used in chemically amplified lithography, medicine, microfluidics, and three-dimensional (3D) printing.
[0021] The present disclosure provides constraints informed by the topological properties of the data to the generation procedure to avoid unimportant or undesirable molecular candidates, thereby reducing experimental costs by focusing on molecules likely to meet various requirements. Topological properties can be indicated by a topology data analysis graph. For example, loops and flares in a Reeve graph can indicate the presence of unknown molecular candidates worth pursuing. The present disclosure is compatible with other approaches for generating new molecules and promoting candidates with high viability or otherwise highly desirable. For example, subject matter experts (SMEs) can be included in the loop by showing the data to SMEs, who can evaluate the data and determine which datasets are most likely to produce the desired results.
[0022] The present disclosure may make active use of attributes assigned to molecules. Attributes may vary and may include location in a dataset, physical properties, National Fire Protection Association (NFPA) hazardous material identification or other indicators, and other attributes. These attributes may be used to construct a Reeb graph. In a Reeb graph, attributes (including scalar attributes) may be used as filter functions for initial data projection.
[0023] The present disclosure may use known molecules with similar attributes to generate candidate molecules with the same or similar attributes. This disclosure may be likened to a kitchen artist developing a new recipe. While a chocolate chip cookie recipe can be combined with recipes for other desserts (e.g., peanut butter cookies, cake, and brownies) to create a sweet treat, when the goal is to design a new dessert, it is unlikely that it would be combined with a main course recipe (e.g., steak, tofu, seitan, or macaroni). (E.g., the "skeleton" of a dessert group may be sugar, flour, and salt, and the "functional group" of the resulting dessert option may be a flavoring, such as vanilla, chocolate, or cinnamon, or a combination thereof.) In some embodiments, by limiting the discovery process to specific scaffolds (e.g., scaffolds known to have desired attributes), the search results can be narrowed to the most viable and desirable candidate molecules.
[0024] By limiting the molecular candidates for a condition-specific approach (e.g., by requiring candidates to have a specific scaffold), the number of infeasible candidates can be significantly reduced, and screening costs can be significantly reduced. Comparing the previous approach with the disclosed approach, the previous approach generated 44,000 molecular candidates, while the disclosed approach generated 137 molecular candidates, significantly increasing the efficiency of the experimental phase. In this example, the 44,000 candidates generated by the established approach failed to change the topology of the data in any meaningful way, whereas the 137 molecular candidates discovered by the disclosed approach changed the topology of the data by adding missing data to the dataset, using internal diagnostics of failure and success to elucidate the most viable molecular candidates.
[0025] Because the present disclosure uses a data-driven approach, input datasets in the present disclosure reflect output data. The present disclosure may use one dataset to generate one graph for one set of results, multiple datasets for multiple graphs for multiple sets of (possibly related) results, or multiple datasets combined into one graph for one set of results (e.g., exploring molecular candidates with hybrid qualities), or any combination thereof.
[0026] The initial dataset may include a set of molecules, such as a set of iconic PAGs from the sulfonium and iodonium families. New molecules may be PAG-like candidates believed to exhibit desirable photochemical behavior, currently unreachable levels of environmental friendliness, and other sought-after attributes. This disclosure serves as a control module in terms of improving the signal-to-noise ratio.
[0027] 1 illustrates a pipeline 100 for molecule generation according to an embodiment of the present disclosure. In some embodiments of the present disclosure, three main operations are used to develop new molecules: dataset generation 110, topology analysis 120, and scaffold-based VAE generation 130. A dataset 118 from dataset generation 110 may be submitted for topology analysis 120. Topology analysis results 126 may then be used for scaffold-based VAE molecule generation 130.
[0028] Dataset generation 110 can be accomplished in a variety of ways. Dataset 118 can be compiled 117 manually 112, using brute force 114, or using artificial intelligence (AI) 116, or a combination thereof. Dataset 118 can be provided 119 for topology analysis 120.
[0029] Topology analysis 120 can include obtaining a data set 118 and applying a kernel 121 to generate a molecular fingerprint 122. The molecular fingerprint 122 can be provided for topology data analysis 124 to produce 125 a topology analysis result 126. The topology analysis result 126 can be a compilation of data such as, for example, a topology graph or a Reeve graph.
[0030] The topology analysis results 126 may be provided 129 to a skeleton analysis 132 of a skeleton-based generation 130 operation. The skeleton analysis 132 may generate 131 skeletons 134 for the molecules in the molecular dataset 118. The skeletons 134 may be provided 133 to an encoder 136. The encoder 136 may be a skeleton-tuned one, such as a skeleton-tuned variational auto-encoder (VAE). The encoder 136 creates 135 new molecules 138. The new molecules 138 may be provided 139 to a topology analysis 120 for further learning and the addition of additional new molecules.
[0031] As will be further explained in the discussion of FIGS. 2A and 2B, the molecular dataset 118 may be provided 119 for topology analysis 120.
[0032] Figure 2A illustrates generating an identifier 230 from a molecule 210 according to an embodiment of the present disclosure, and Figure 2B illustrates generating a fingerprint 250 from the identifier 230 according to an embodiment of the present disclosure. Figures 2A and 2B may be considered illustrations of operation 121 of Figure 1.
[0033] 2A illustrates generating an identifier 230 from a molecule 210. A derivative 220 is derived from the molecule 210, which derivative 220 can be used to generate an identifier 230, which can be used to generate a binary representation 240 and a final molecular fingerprint 250 (FIG. 2B). Those skilled in the art will recognize that any method for generating the identifier 230 and molecular fingerprint 250 may be used in accordance with the present disclosure.
[0034] Molecular derivatives 220 can be derived from molecules 210. Molecules 210 can be derived to various diameters. The diameter of a fragment indicates the number of bonds from the center of the fragment. A 0-diameter fragment indicates a fragment with 0 bonds. In other words, a fragment with a 0-diameter describes only the central atom of the fragment. Larger fragments are built outward from the 0-diameter fragment. A fragment with a 2-diameter includes the central atom of the fragment and atoms directly bonded to it. A fragment with a 4-diameter includes the central atom of the fragment, atoms directly bonded to the central atom, and any atoms bonded to atoms directly bonded to the central atom.
[0035] Derivatives 222a, 222b, 222c, 222d, and 222e with a diameter of 0 are shown in a first derivative block 222. Derivatives 224a, 224b, 224c, 224d, 224e, and 224f with a diameter of 2 are shown in a second derivative block 224. Derivatives 226a, 226b, 226c, 226d, and 226e with a diameter of 4 are shown in a third derivative block 226. It may be advantageous to generate molecular derivatives 220 with varying diameters 222, 224, and 226.
[0036] Molecular derivatives 220 can be used to generate identifiers 230. Identifiers 232a, 232b, 232c, 232d, and 232e for diameter 0 derivatives are shown in a first identifier block 232. Identifiers 234a, 234b, 234c, 234d, 234e, and 234f for diameter 2 derivatives are shown in a second identifier block 234. Identifiers 236a, 236b, 236c, 236d, and 236e for diameter 4 derivatives are shown in a third identifier block 236.
[0037] 2B shows the identifier 230 being used to generate a bit vector 240 that can be used to generate a molecular fingerprint 250. The identifier 230 is hashed 239 to generate a fixed-length binary representation 240. Each identifier 234c is hashed 244c to generate a portion of the fixed-length binary representation 240. The fixed-length binary representation 240 is then used to generate the molecular fingerprint 250. The molecular fingerprint 250 is sometimes referred to as a bit vector fingerprint 250.
[0038] The molecular fingerprint 250 indicates the presence or absence of a structural motif. The kernel may extract molecular characteristics, hash 239 the molecular characteristics, and use the hash (e.g., binary representation 240) to determine the bits of the molecular fingerprint 250. In some embodiments, the fingerprint, molecular fingerprint 250, ranges in size from 1,000 to 4,000 bits. The molecular fingerprint 250 may be used to generate 123 the topology data analysis graph 124 shown in FIG. 1.
[0039] 3 illustrates a topology data analysis 300 according to an embodiment of the present disclosure. FIG. 3 may be considered an illustration of generating 125 a topology graph 126 using the topology data analysis graph 124 shown in FIG.
[0040] The Reeb graph or an approximation thereof, such as an adjacency graph 350, can be obtained from a three-dimensional or higher-dimensional model, such as a point cloud 312. An algorithm may be used for topological data analysis. The algorithm may be a mapper algorithm or any other method used to construct a Reeb graph or a Reeb graph approximation, or both. The algorithm may combine the construction of a Reeb graph approximation with pullback cover on the data.
[0041] Each molecule in the dataset may be represented by a bit vector and generated as a molecular topology fingerprint. For example, PAG may be represented by a Morgan fingerprint (MorganFP). The molecular dataset may be treated as a point cloud 312 with pairwise distances. The pairwise distances may be defined using any available cheminformatics approach. Each of the various dots in the matrix-structured point cloud 312 represents a molecular fingerprint. The molecular dataset may be treated as a point cloud 312 in a space 310 of bit vectors.
[0042] Dice similarity can be used on bit vectors to define pairwise distances in a set of molecules. Distance to a reference point can be used as a filter function in topology data analysis. For example, the reference point can be a PAG with the smallest number of heavy atoms in a data set, and distance to that PAG can be used to filter data in the topology data analysis graph. The filter function f320 can split the point cloud 312 horizontally with respect to the height of the point cloud 312. Alternative arrangements for splitting the point cloud 312 are suitable for use with the present disclosure, such as splitting the point cloud vertically or reorienting the point cloud 312 differently before splitting.
[0043] Overlapping range splitters 342, 344, 346, and 348 may be used to divide point cloud 312 into various overlapping segments 342a, 344a, 344b, 346a, and 348a. Numerators may be assigned to segment sets 340 based on the value of filter function f 320. Segments 342a, 344a, 344b, 346a, and 348a may also be referred to as levels 342a, 344a, 344b, 346a, and 348a, and segment sets 340 may also be referred to as level sets 340.
[0044] Algorithms can be used to generate simplified descriptions of data in the form of graphs. The algorithms can be computational methods (e.g., mappers) for extracting simple descriptions of high-dimensional datasets in the form of simplicial complexes. A graph can be described by the equation G = (C, E), where G is the graph, C is the set of clusters represented as nodes, and E is the set of all edges. Each node in the graph represents a cluster C of molecules, and the edges E between clusters indicate overlaps between clusters. Depending on the choice of approximation of the Reeb graph, other rules for establishing connections between nodes may be used.
[0045] Graphs generated by algorithms (e.g., mapper graphs) can directly visualize various aspects of the data shape. For example, loops (sometimes called holes) and flares (sometimes called bifurcations) can be seen in the graphical shape of the data. Loops and / or flares in the data indicate missing data and pinpoint where to look for new molecules, because new molecules that fill loops and close flares are considered desirable molecules.
[0046] To identify connected and disconnected components, segment set 340 may be clustered into disjoint sets using agglomerative clustering on pre-computed Dice distances. Clusters 342a, 344a, 344b, 346a, and 348a may be represented as nodes 352a, 354a, 354b, 356a, and 358a on a graph 350, such as mapper graph 350. Nodes 352a, 354a, 354b, 356a, and 358a may be connected to each other via links when the connected nodes have common members.
[0047] For example, due to the overlap range splitter 342 used to split the first segment cluster 342a from the second segment clusters 344a and 344b, the first segment cluster 342a has members in common with members of both the second segment clusters 344a and 344b. Thus, because these clusters share a molecular fingerprint, a link connects the first segment node 352a to each of the second segment nodes 354a and 354b. Note that because there is no overlap between the two second segment clusters 344a and 344b, the second segment clusters 344a and 344b do not share a common molecular fingerprint, and the second segment nodes 354a and 354b are not linked.
[0048] By mapping a molecular database in this manner, aspects of the shape of the dataset, such as loops and flares, are accurately captured. Flares are sometimes also called branches. Loops and flares indicate gaps. Gaps indicate gaps or holes in the data sufficient for molecule generation. A graph, such as mapper graph 350, may be submitted 129 for scaffold-based generation 130 of new molecules, as shown in FIG. 1.
[0049] 4 illustrates the generation of molecular candidates 400 according to an embodiment of the present disclosure. A topology data analysis graph 410 is provided for scaffold-based generation to create a more complete topology graph 420.
[0050] The topology data analysis graph 410 may have loops 412 and flares 414 and 416. The loops 412 and flares 414 and 416 may indicate that new molecules may be derived from the data set. The loops 412 in the topology graph 410 may be described as spaces that may allow for one or more additional unique links between nodes. The flares 414 and 416 in the topology graph 410 may be described as nodes with only one link, a free edge, or a space in the topology graph 410 from which a molecule point appears to dangle.
[0051] By providing input topology and data analysis graph 410 to scaffold-based molecule generation 130 (FIG. 1), output topology and data analysis graph 420 can be obtained. By performing scaffold-based molecule generation 130, additional molecules were added. Specifically, the addition of the scaffold-based molecule generation molecules narrowed loop 412 to smaller loop 422 and closed flares 414 and 416 to loops 424 and 426. Output topology and data analysis graph 420 can be provided to a further scaffold-based molecule generation 130 (shown in FIG. 1) so that loops 424 and 426 may be further narrowed and flares may be closed.
[0052] Adding a node to the topology data analysis graphs 410 and 420 represents adding a molecule to the dataset. In other words, an added node represents a new molecule. New molecules derived or discovered using the present disclosure improve the completeness of a chemical dataset and are considered relatively valuable for further exploration in the search for molecules with highly desirable attributes.
[0053] 5 illustrates a backbone analysis molecule generation 500 according to an embodiment of the present disclosure. The molecular backbone may represent the core of the molecule. The molecular core may be described as a molecule without functional groups. The backbone may be considered the primary constraint on the shape of the molecule. The backbone may be considered the primary constraint on the basic properties of the molecule.
[0054] Scaffolds allow for hierarchical representation of molecules. Scaffold hierarchies can be constructed to provide different levels of abstraction in the representation of molecules. Analysis of scaffolds uses definitions and existing hierarchies and their implementations. An example of an information database that may be useful for scaffold analysis is the Cheminformatics Toolkit.
[0055] 5, functional groups are removed 518 from molecule 510 to obtain a skeleton 520 of molecule 510. Skeleton 520 may then be used as a basis for generating 528 an assortment 530 of product molecules 532, 534, and 536.
[0056] The product molecules 532, 534, and 536 may have similar, different, more, or fewer functional groups than the molecule 510 from which the scaffold 520 was generated. Common to the product molecules 532, 534, and 536 is the scaffold 520 used to generate 528 the molecules 532, 534, and 536. Molecules generated from a given scaffold may have different functional groups attached to the same atom. For example, the two molecules 532 and 534 have different functional groups attached to the same position on the scaffold. Molecules generated from a given scaffold may have the same or different functional groups attached to one or more different atoms. For example, the first product molecule 532 and the third product molecule 536 have functional groups attached to different atoms in each molecule 532 and 536. The functional groups may be attached to any atom in the scaffold that supports the bond.
[0057] For skeletal analysis, an undirected graph can be written as G = (C,E). Given a data set S = {s1, s2, ..., s n} can be identified. Each skeleton s in the dataset is associated with one or more clusters C s ={c1,c2,…,c s}, and skeleton s can appear in the cluster C s The clusters may be analyzed into nodes, and either the clusters or the nodes may be used in the analysis. The shortest cycle for each cluster c in terms of hops: w is calculated by calculating the length l of the cycle w for the cluster c. lc When a shortest cycle exists, the first run of the first search breadth until c is reached will achieve the shortest cycle for that cluster. s can be generated.
[0058]
number
[0059] generation possibility g s is normalized between 0 and 1. High probability of generation g s indicates that the skeleton appears in small clusters with large cycle lengths. In other words, a high probability of generation g s indicates that the scaffold is part of a larger loop and flare in the topographical analysis graph and is therefore likely to generate new molecules when a scaffold-based generation is performed.
[0060] The VAE loss function is the probability of generating a molecular skeleton g s The standard VAE loss function can be expressed as:
[0061] L=L r +L KL
[0062] The loss l for a single data point (G;S) for a molecular graph G and corresponding scaffold S can be expressed as:
[0063]
number
[0064] Skeleton generation possibility g s Incorporating this into the loss function l, we get:
[0065]
number
[0066] where g s is the generability of the input skeleton, and g sn is a newly generated molecule in graph G=(C,E)
[0067]
number
[0068] is the probability of generating the skeleton of, and α is a hyperparameter [0,1].
[0069] Thus, the standard VAE loss function may become a modified loss function.
[0070] L = (1-g s )(L r +L KM +α(g s -g sn ))
[0071] where g s is the generability of the input skeleton, and g sn is a newly generated molecule in graph G=(C,E)
[0072]
number
[0073] is the probability of generating the skeleton of L, α is a hyperparameter [0,1], and L ris the input s and the generated skeleton s n is the reconstruction error with L KL is the Kullback-Leibler divergence between the prior and the approximate posterior distribution.
[0074] By using the modified loss function, the model can generate low probability g s By disfavoring the model when generating molecules with low generation probability g s The effect of the skeleton having the probability of generating a newly generated molecule g sn After each iteration, the newly generated molecules can be included in the graph to calculate
[0075] By using a modified loss function, the most promising skeletons in the dataset can be identified. Skeletons in small clusters along large loops in the graph (e.g., mapper graph) may be prioritized because they have the greatest generative probability g s The graph may be used to identify the minimum loop in terms of hops in various ways, such as using any variant of the Dijkstra algorithm. l We may then calculate the sum of the lengths of each edge in the loop, which is equivalent to: A larger cycle length indicates that the scaffold is part of a larger loop, thus increasing the likelihood of more than one desirable candidate. For each molecule in the cluster, we create a Bemis-Murko scaffold S = {s1, s2, ..., s n For each skeleton s, the generative probability g s may be calculated. s may be normalized between 0 and 1. High probability of generation g s indicates that the scaffolds appear in small clusters with large cycle lengths and are therefore more likely to generate new molecules.
[0076] The generation of new molecules to complete the loop of graphical data may use a graph generative model for scaffold-based molecular design adapted for the purpose of completing the loop of graphical data. In some embodiments, a VAE may be used in the generative modeling process.
[0077] The input may be extended by sequentially adding atoms and bonds. In this way, molecule generation is adjusted on the input skeleton to ensure that all generated molecules contain the input skeleton. The VAE loss function is the probability of generating the skeleton, g s may be modified to take into account
[0078] 6 illustrates a molecule generation pipeline 600 according to an embodiment of the present disclosure. A molecule database 610 can include data / information 612, 614, and 616 for molecules, molecular fragments, or some combination thereof. The molecule database 610 can originally include reference molecule information that was used and / or synthesized. The molecule database 610 may have started with molecule information, molecular fragment information, or some combination thereof. The molecule database 610 may include information about synthetic molecules, synthetic molecular fragments, naturally occurring molecules, molecular fragments, or some combination thereof.
[0079] Molecular fragments may be fragmented heuristically, randomly, or according to user (e.g., SME) decisions. Fragments may be combined according to constraints manually selected and set by the user, such as the number of fragments to combine, fragment compatibility, and fragment connectivity, among others. The user may also use algorithmic constraints on fragment combination. For example, metaheuristics such as particle swarm optimization and genetic algorithm optimization may be used. In some embodiments of the present disclosure, a convolutional neural network (CNN) may be used to replicate SME decisions based on a set confidence threshold.
[0080] Information for the main molecular database 610 may be provided by the literature (e.g., textbooks and chemical tables), subject matter experts, alternative sources, or some combination thereof. In some embodiments, the system provided by the present disclosure analyzes new molecules it creates to supplement the original molecular database 610 or to establish a distinct molecular database 610 consisting of molecular data for the newly generated molecules.
[0081] The molecular data 612, 614, and 616 may be provided 618 to an encoder 620 to generate 628 molecular fingerprint data 632, 634, and 636 for a molecular fingerprint database 630. The molecular fingerprint database 630 may include one or more molecular fingerprints 632, 634, and 636. The molecular fingerprints 632, 634, and 636 may be provided 638 to a scaffold-adjusted VAE generator 640 to create a candidate database 650 of molecular candidates 652a, 652b, 654a, 654b, 656a, and 656b from molecular scaffolds 652, 654, and 656.
[0082] The data for molecular scaffolds 652, 654, and 656 and new molecules 652a, 652b, 654a, 654b, 656a, and 656b may be included in candidate dataset 650, which may also be referred to as newly generated molecular dataset 650. New molecules 652a, 652b, 654a, 654b, 656a, and 656b may be derived from molecular scaffolds 652, 654, and 656. The molecular data for scaffolds 652, 654, and 656 and new molecules 652a, 652b, 654a, 654b, 656a, and 656b may be provided 608 directly to main molecular database 610 or provided 658 to analyzer 660 for analysis.
[0083] The molecular datasets 610 and 650 may be provided 658 to an analyzer 660. In some embodiments, the analyzer 660 may gather information about the data being provided to the molecular dataset 610. In some embodiments, the analyzer 660 may be a system with predetermined analysis thresholds. In some embodiments, the analyzer 660 may be an SME. The analyzer 660 may analyze the molecular datasets 610 and 650 and submit 668 the results of its analysis to the main molecular dataset 610 for provision to the molecule generation pipeline 600.
[0084] 7 illustrates a molecule generation pipeline 700 according to an embodiment of the present disclosure. A cylindrical shape 710 represents a corpus, pane-split rectangles 712, 714, 722, 724, 732, and 734 represent processes, methods, or functions, and wavy-bottom rectangles 720, 730, and 750 represent goals or results.
[0085] From a molecular dataset 710, the molecule generation pipeline 700 may generate new molecules 750 by generating a topology graph 720 and a scaffold 730 for the dataset 710. Data about the new molecules 750 may be incorporated into the molecular dataset 710. By adding each new set of data from each round of new molecules 750 to the molecular dataset 710, perhaps repeating the system 700 until all gaps have been identified 722 and filled, a particular molecular dataset 710 can be completed to the point where no additional viable candidates are likely to be generated from that dataset.
[0086] In some embodiments of the present disclosure, the generation of a new molecule may begin by submitting one or more molecular datasets 710 to a molecule generation system 700. The system 700 may convert 712 the molecule into a molecular fingerprint and perform topology data analysis 714 on the molecular fingerprint. The topology data analysis 714 may generate a molecular topology graph 720.
[0087] The molecular topology graph 720 may be analyzed to identify 722 gaps in the topology graph. By identifying 722 gaps, scaffolds to fill the gaps may be calculated 724. The result of calculating 724 scaffolds for the identified gaps 722 is a set 730 of scaffolds for the molecular dataset 710.
[0088] The dataset skeletons 730 may be used to calculate 732 a generativity score for each skeleton in the set of skeletons 730. The generativity scores 732 may be used to train 734 a skeleton-adjusted VAE. One or more new molecules 750 may be generated using the skeleton-adjusted VAE 734 trained by the generativity scores 732. The new molecules 750 may be added to the molecular dataset 710. This process may be repeated until all gaps in the molecular topology graph 720 have been identified 722 and the generation of new molecules 750 for the dataset has been completed.
[0089] Some embodiments of the present disclosure may use cloud computing. Accordingly, aspects of the present disclosure may relate to cloud computing. Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a service provider. The cloud model may include at least five characteristics, at least three service models, and at least four deployment models.
[0090] The characteristics are as follows:
[0091] On-demand self-service: Cloud consumers can unilaterally provision computing capacity, such as server time and network storage, as needed automatically and without human interaction with the provider of the service.
[0092] Broad network access. Functionality is available across the network and accessed through standard mechanisms that facilitate use by a variety of thin and thick client platforms (e.g., cell phones, laptops, and PDAs).
[0093] Resource Pooling. To serve multiple consumers using a multi-tenant model, a provider's computing resources are pooled, with different physical and virtual resources dynamically allocated and reallocated according to demand. The consumer generally has no control or knowledge over the exact portion of the resources provided, but there is a sense of portion independence in that they may be able to identify portions at a higher level of abstraction (e.g., country, state, or data center).
[0094] Rapid Elasticity. Capabilities can be rapidly and elastically provisioned, sometimes automatically, to quickly scale out, and rapidly released to quickly scale in. To the consumer, the capabilities available for provisioning often appear unlimited, and any amount can be purchased at any time.
[0095] Metered services. Cloud systems automatically control and optimize resource usage by utilizing metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services used.
[0096] The service model is as follows:
[0097] Software as a Service (SaaS). The functionality offered to the consumer is the use of the provider's applications running on a cloud infrastructure. The applications are accessible from a variety of client devices through thin-client interfaces, such as web browsers (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, storage, or even individual application functions, except possibly for limited user-specific application configuration settings.
[0098] Platform as a Service (PaaS). The functionality offered to the consumer is the deployment onto a cloud infrastructure of applications created or acquired by the consumer, written using programming languages and tools supported by the provider. While the consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, the consumer does have control over the deployed applications and possibly the application hosting environment configuration.
[0099] Infrastructure as a Service (IaaS). The functionality provided to the consumer is the provisioning of processing, storage, network, and other basic computing resources onto which the consumer can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but does have control over the operating system, storage, deployed applications, and perhaps limited control over the selection of networking components (e.g., host firewalls).
[0100] The deployment model is as follows:
[0101] Private cloud: This cloud infrastructure is operated solely for one organization. It may be managed by that organization or a third party and may reside on-premises or off-premises.
[0102] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with common concerns (e.g., mission, security requirements, policy or compliance considerations, or a combination thereof), and may be managed by those organizations or a third party and may reside on-premises or off-premises.
[0103] Public cloud: This cloud infrastructure is made available to the general public or large industry groups and is owned by an organization that sells cloud services.
[0104] Hybrid cloud: This cloud infrastructure is a composite of two or more clouds (private, community, or public) that remain their own entities but are bound together by standard or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).
[0105] Cloud computing environments are service-oriented and focus on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0106] 8 illustrates a cloud computing environment 810 according to an embodiment of the present disclosure. As shown, the cloud computing environment 810 includes one or more cloud computing nodes 800 with which local computing devices used by cloud consumers, such as a personal digital assistant (PDA) or mobile phone 800A, a desktop computer 800B, a laptop computer 800C, or an automobile computer system 800N, or any combination thereof, may communicate. The nodes 800 may also communicate with each other. These nodes may be physically or virtually grouped (not shown) in one or more networks, such as, for example, a private, community, public, or hybrid cloud, or any combination thereof, as described above.
[0107] This allows the cloud computing environment 810 to provide infrastructure, platform, or software, or a combination thereof, as a service for which cloud consumers are not required to maintain resources on their local computing devices. It is understood that the types of computing devices 800A-N shown in Figure 8 are intended to be exemplary only, and that the computing nodes 800 and the cloud computing environment 810 can communicate with any type of computing device over any type of network or network-addressable connection (e.g., using a web browser) or both.
[0108] 9 illustrates abstraction model layers 900 provided by cloud computing environment 810 (FIG. 8) according to an embodiment of the present disclosure. It should be understood in advance that the components, layers, and functions illustrated in FIG. 9 are intended to be merely exemplary, and embodiments of the present disclosure are not limited thereto. As shown below, the following layers and corresponding functions are provided:
[0109] The hardware and software layer 915 includes hardware and software components. Examples of hardware components include a mainframe 902, a RISC (Reduced Instruction Set Computer) architecture-based server 904, a server 906, a blade server 908, storage devices 911, and network and network forming components 912. In some embodiments, the software components include network application server software 914 and database software 916.
[0110] The virtualization layer 920 provides an abstraction layer from which the following examples of virtual entities may be provided: virtual servers 922, virtual storage 924, virtual networks including virtual private networks 926, virtual applications and operating systems 928, and virtual clients 930.
[0111] In one example, the management layer 940 may provide the following functions: Resource provisioning 942 provides dynamic procurement of computing and other resources used to perform tasks within the cloud computing environment. Metering and pricing 944 provides cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, as well as protection of data and other resources. User portal 946 provides access to the cloud computing environment for consumers and system administrators. Service level management 948 provides allocation and management of cloud computing resources to ensure required service levels are met. Service level agreement (SLA) planning and fulfillment 950 provides advance arrangements for and procurement of cloud computing resources where future demand is predicted by SLAs.
[0112] The workload layer 960 provides examples of functions for which a cloud computing environment may be used. Examples of workloads and functions that may be provided from this layer include mapping and navigation 962, software development and lifecycle management 964, virtual classroom instruction delivery 966, data analytics processing 968, transaction processing 970, and tools for generating new molecules 972.
[0113] 10 illustrates a high-level block diagram of an example computer system 1001 that may be used to implement one or more of the methods, tools, and modules described herein, and any associated functionality (e.g., using one or more processor circuits of a computer or computer processor), in accordance with embodiments of the present disclosure. In some embodiments, major components of computer system 1001 may include a processor 1002 having one or more central processing units (CPUs) 1002A, 1002B, 1002C, and 1002D, a memory subsystem 1004, a terminal interface 1012, a storage interface 1016, an I / O (Input / Output) device interface 1014, and a network interface 1018, all of which may be communicatively coupled, directly or indirectly, for inter-component communication via a memory bus 1003, an I / O bus 1008, and an I / O bus interface unit 1010.
[0114] The computer system 1001 may include one or more general-purpose programmable CPUs 1002A, 1002B, 1002C, and 1002D, collectively referred to herein as CPUs 1002. In some embodiments, the computer system 1001 may include multiple processors, as is typical of larger systems. However, in other embodiments, the computer system 1001 may alternatively be a single-CPU system. Each CPU 1002 may execute instructions stored in a memory subsystem 1004 and may include one or more levels of on-board cache.
[0115] System memory 1004 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 1022 or cache memory 1024. Computer system 1001 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 1036 may be provided for reading from and writing to non-removable, non-volatile magnetic media, such as a "hard drive." Although not shown, a magnetic disk drive may be provided for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), or an optical disk drive may be provided for reading from or writing to a removable, non-volatile optical disk, such as a CD-ROM, DVD-ROM, or other optical media. Additionally, memory 1004 may include flash memory, such as a flash memory stick drive or flash drive. Memory devices may be connected to memory bus 1003 by one or more data media interfaces. The memory 1004 may include at least one program product having a set (eg, at least one) of program modules configured to perform the functions of various embodiments.
[0116] The memory 1004 may store one or more programs / utilities 1028, each having at least one set of program modules 1030. The programs / utilities 1028 may include a hypervisor (also called a virtual machine monitor), one or more operating systems, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or any combination thereof, may comprise an implementation of a networked environment. The programs 1028 and / or program modules 1030 generally perform the functions or methods of various embodiments.
[0117] 10, memory bus 1003 is depicted as a single bus structure providing a direct communication path between CPU 1002, memory subsystem 1004, and I / O bus interface 1010, but in some embodiments, memory bus 1003 may include multiple different buses or communication paths, which may be arranged in any of a variety of configurations, such as point-to-point links in a hierarchical, star, or web configuration, multiple hierarchical buses, parallel and redundant paths, or any other suitable type of configuration. Additionally, while I / O bus interface 1010 and I / O bus 1008 are each depicted as single units, in some embodiments, computer system 1001 may include multiple I / O bus interface units 1010, multiple I / O buses 1008, or both. Additionally, although multiple I / O interface units 1010 are shown isolating the I / O bus 1008 from various communication paths to various I / O devices, in other embodiments, some or all of the I / O devices may be directly connected to one or more system I / O buses 1008.
[0118] In some embodiments, computer system 1001 may be a multi-user mainframe computer system, a single-user system, a server computer, or a similar device that has little or no direct user interface but receives requests from other computer systems (clients). Further, in some embodiments, computer system 1001 may be implemented as a desktop computer, a portable computer, a laptop or notebook computer, a tablet computer, a pocket computer, a telephone, a smartphone, a network switch or router, or any other suitable type of electronic device.
[0119] It should be noted that Figure 10 is intended to illustrate representative major components of exemplary computer system 1001. However, in some embodiments, individual components may have greater or less complexity than shown in Figure 10, there may be components other than or in addition to those shown in Figure 10, and the number, type, and configuration of such components may vary.
[0120] Although this disclosure includes a detailed description of cloud computing, it should be understood that implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of the present disclosure may be implemented in conjunction with any other type of computing environment now known or that may later be developed.
[0121] The present disclosure may be a system, method, or computer program product, or combination thereof, integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to perform aspects of the present disclosure.
[0122] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: Portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or raised structures in grooves with recorded instructions, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted through wires.
[0123] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium into each computing / processing device, or may be downloaded to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to a computer-readable storage medium within the respective computing / processing device for storage.
[0124] Computer-readable program instructions for carrying out the operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk or C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., over the Internet using an Internet Service Provider). In some embodiments, electronic circuitry, including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions by using state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects of the present disclosure.
[0125] The present disclosure contemplates methods, systems, and computer program products for generating new molecules, including receiving molecular data for a plurality of molecules, performing topology data analysis on the molecular data to generate a molecular topology map, identifying one or more gaps in the molecular topology map, and generating one or more additional molecules to fill at least one of the one or more gaps.
[0126] The present disclosure further contemplates that the plurality of molecules have one or more common molecular properties. The present disclosure further contemplates generating a molecular scaffold for each of the plurality of molecules. The present disclosure further contemplates generating a generativity score for each scaffold. The present disclosure further contemplates that the plurality of molecules share a molecular scaffold, and that one or more additional molecules contain the molecular scaffold.
[0127] The present disclosure further contemplates generating one or more additional molecules using a variational autoencoder. The present disclosure further contemplates adjusting the variational autoencoder to include a particular molecular skeleton via backbone adjustment. The present disclosure further contemplates having a variational autoencoder loss function, where the variational autoencoder loss function is modified to include the probability of generating a particular molecular skeleton.
[0128] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0129] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, whereby the instructions, executed by the processor of the computer or other programmable data processing apparatus, cause means for implementing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, whereby the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0130] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause the computer, other programmable apparatus, or other device to perform a series of operational steps to produce a computer-implemented process, where the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0131] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur in a different order than that shown in the figures. For example, two blocks shown in succession may actually be accomplished as a single step, may be executed concurrently, may be executed substantially concurrently in a partially or fully overlapping manner, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. In addition, it will be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.
[0132] The description of various embodiments of the present disclosure is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles of the embodiments, practical applications, or technical improvements to technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0133] While the present disclosure has been described in terms of specific embodiments, it is anticipated that alterations and modifications thereof will become apparent to those skilled in the art. It is therefore intended that the following claims be interpreted to embrace all such alterations and modifications as fall within the true spirit and scope of the present disclosure.
Claims
1. 1. A computer-implemented method for generating new molecules, comprising: receiving molecular data for a plurality of molecules; performing topological data analysis on the molecular data to generate a molecular topology map; identifying one or more defects in the molecular topology map; generating one or more additional molecules to fill at least one of the one or more gaps; A method comprising:
2. The method of claim 1 , wherein the plurality of molecules have one or more common molecular properties.
3. The method of claim 1, further comprising generating, by the processor, a molecular skeleton for each of the plurality of molecules.
4. The method of claim 3, further comprising generating a generatibility score for each skeleton by the processor.
5. The method of claim 1 , wherein the plurality of molecules share a molecular backbone and the one or more additional molecules contain the molecular backbone.
6. The method described in claim 1, wherein the processor generates the one or more additional molecules using a variational autoencoder.
7. The method of claim 6, further comprising adjusting the variational autoencoder by the processor via backbone adjustment so that the one or more additional molecules contain a particular molecular backbone.
8. The processor performs a step of: The method of claim 7 , further comprising modifying the variational autoencoder loss function to include the generability of the particular molecular scaffold.
9. 1. A system for generating new molecules, the system comprising: Memory and a processor in communication with the memory, the processor comprising: receiving molecular data for a plurality of molecules; performing topological data analysis on the molecular data to generate a molecular topology map; identifying one or more defects in the molecular topology map; generating one or more additional molecules to fill at least one of the one or more gaps; A system configured to perform operations including:
10. The system of claim 9 , further comprising generating a molecular scaffold for each of the plurality of molecules.
11. The system of claim 10 , further comprising generating a generatibility score for each skeleton.
12. The system of claim 9 , wherein the plurality of molecules share a molecular backbone and the one or more additional molecules contain the molecular backbone.
13. The system of claim 9 , wherein the one or more additional molecules are generated using a variational autoencoder.
14. 14. The system of claim 13, further comprising tuning the variational autoencoder so that the one or more additional molecules contain a particular molecular backbone via backbone tuning.
15. The variational autoencoder has a variational autoencoder loss function, and the operation is The system of claim 14 , further comprising modifying the variational autoencoder loss function to include the generability of the particular molecular scaffold.
16. A computer program for generating new molecules, comprising: A computer program product for causing a processor to execute the steps of any one of claims 1 to 8.
Citation Information
Patent Citations
Chemical compound generation device, chemical compound generation method, learning device, learning method, and program
JP2021068410A