Techniques for data-enabled drug discovery

CN116114024BActive Publication Date: 2026-08-11RECURSION PHARMACEUTICALS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-17
Publication Date
2026-08-11

AI Technical Summary

Benefits of technology

[0011] At least one technical advantage of the disclosed technique over the prior art is that it can be used to more effectively identify additional drug development candidates during the drug discovery process. In this regard, because the relevant editing heuristics implemented using the disclosed technique are tailored to a given drug discovery process, the likelihood that structural changes introduced by those editing heuristics will retain target biological activity is increased. Therefore, the proportion of structurally similar drug candidates ultimately relevant to a given drug discovery process is generally increased using the disclosed technique compared to prior art methods. Furthermore, unlike prior art, the computational complexity of the disclosed technique remains constant regardless of the total number of molecules searched, thus reducing the time and computational resources required to comprehensively search a molecule catalog. In particular, using the disclosed technique, all available molecules can be searched at a given interaction rate. These technical advantages provide one or more technical improvements over prior art methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116114024B_ABST
    Figure CN116114024B_ABST
Patent Text Reader

Abstract

In various implementations, the molecular exploration application identifies one or more potential drug candidates during the drug discovery process. The molecular exploration application generates derived molecular specifications based on querying molecular specifications and editing heuristics. Subsequently, the molecular exploration application performs one or more mapping operations on the derived molecular specifications using a mapping algorithm to generate mapped molecular specifications. The molecular exploration application then performs one or more search operations on the mapped molecular catalog based on the mapped molecular specifications to determine the one or more potential drug candidates. Advantageously, the molecular exploration application can be used to efficiently identify additional drug development candidates during the drug discovery process.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 067,025, filed August 18, 2020, entitled “Real-Time SAR Search Tool for Very Large Compound Databases,” and U.S. Patent Application No. 16 / 950,845, filed November 17, 2020, entitled “Techniques For Data-Enabled Drug Discovery.” Subject matter of these related applications is hereby incorporated by reference. Technical Field

[0003] The various implementation schemes generally involve computer science and biochemical analysis, and more specifically, technologies for data-supported drug discovery.

[0004] Description of related technologies

[0005] Drug discovery is the process of identifying molecules for the purpose of curing, alleviating, treating, and / or preventing diseases for further research. During the initial phase of a typical drug discovery process, many different molecules are tested to identify drug development candidates of interest that have the desired effect on cells (referred to as exhibiting “target biological activity”). For example, a drug development candidate of interest could be a drug that improves the health of cells afflicted with a specific disease. Based on the principle that structurally similar molecules often have some similar biological activity, after identifying a given drug development candidate of interest, molecules structurally similar to the drug development candidate of interest are searched in any number of available molecular catalogs. The molecules specified in the search results become potential drug candidates, and these potential drug candidates are then evaluated to identify additional drug development candidates, each of which has the same biological activity as the drug development candidate of interest. Generally, any of the additional drug development candidates identified using this process may have improved physical, chemical, biological, and / or pharmacological properties relative to the drug development candidate of interest.

[0006] In one method for identifying potential drug candidates, a similarity software application quantifies the structural similarity between a drug development candidate of interest and each of any number of molecules specified in a catalogue based on a similarity metric. In some embodiments, the similarity software application calculates a Tanimoto coefficient, which quantifies the structural similarity between two molecules based on a set of molecular structural fragments. In such embodiments, for each molecule specified in the search results, the similarity software application sets the Tanimoto coefficient value to be equal to the ratio of the number of molecular structural fragments shared between the molecule and the drug development candidate of interest to the number of molecular structural fragments present in either or both of the molecule and the drug development candidate of interest. The similarity software application then generates search results for molecules specifying Tanimoto coefficient values ​​exceeding a minimum similarity threshold. Any molecule with a Tanimoto coefficient exceeding the minimum similarity threshold is considered a potential drug candidate.

[0007] One drawback of identifying potential drug candidates based on similarity metrics is that different drugs with “similar” molecular structures according to the similarity metric do not necessarily share a given biological activity. Therefore, for a given drug discovery process, the proportion of potential drug candidates identified using a specific similarity metric that have the same target biological activity as the relevant drug development candidate of interest may be very low. Consequently, significant time and resources are often wasted evaluating potential drug candidates that ultimately become irrelevant to the given drug discovery process.

[0008] Another drawback of using similarity metrics to identify potential drug candidates is that the computational complexity of a catalog search is proportional to the total number of available molecules, meaning the time and computational resources required to fully search a molecular catalog can be extremely high. For example, evaluating hundreds of millions of available molecules in a current molecular catalog using a similarity metric based on three-dimensional shape could take decades of computation. Typically, due to time and computational resource constraints, only a small subset of the molecules specified in the current molecular catalog is searched during a given drug discovery process.

[0009] As previously explained, what is needed in the art is a more effective technique for identifying potential drug candidates during the drug discovery process. Summary of the Invention

[0010] One embodiment of the present invention describes a method for identifying one or more potential drug candidates during a drug discovery process. The method includes: generating a derived molecular specification based on a query molecular specification and an editing heuristic; performing one or more mapping operations on the derived molecular specification using a mapping algorithm to generate a mapped molecular specification; and performing one or more search operations on a mapped molecular catalog based on the mapped molecular specification to identify one or more potential drug candidates.

[0011] At least one technical advantage of the disclosed technique over the prior art is that it can be used to more effectively identify additional drug development candidates during the drug discovery process. In this regard, because the relevant editing heuristics implemented using the disclosed technique are tailored to a given drug discovery process, the likelihood that structural changes introduced by those editing heuristics will retain target biological activity is increased. Therefore, the proportion of structurally similar drug candidates ultimately relevant to a given drug discovery process is generally increased using the disclosed technique compared to prior art methods. Furthermore, unlike prior art, the computational complexity of the disclosed technique remains constant regardless of the total number of molecules searched, thus reducing the time and computational resources required to comprehensively search a molecule catalog. In particular, using the disclosed technique, all available molecules can be searched at a given interaction rate. These technical advantages provide one or more technical improvements over prior art methods. Attached Figure Description

[0012] Therefore, a more detailed understanding of the above-described features of the various embodiments can be obtained by referring to them, namely, a more specific description of the inventive concept briefly outlined above, some of which are shown in the accompanying drawings. However, it should be noted that the drawings only show typical embodiments of the inventive concept and should not be considered as limiting the scope in any way, and other equivalent embodiments exist.

[0013] Figure 1 It is a conceptual diagram of a system configured to implement one or more aspects of various implementation schemes;

[0014] Figure 2 It is based on various implementation plans. Figure 1 A more detailed illustration of the derivation engine;

[0015] Figure 3 It is based on various implementation plans. Figure 1 A more detailed illustration of the search results dataset; and

[0016] Figure 4 It is a flowchart of methodological steps for identifying potential drug candidates during the drug discovery process, according to various implementation schemes. Detailed Implementation

[0017] In the following description, numerous specific details are set forth to provide a more thorough understanding of the various embodiments. However, it will be apparent to those skilled in the art that the inventive concepts can be practiced without one or more of these specific details.

[0018] Exemplary Overview

[0019] For illustrative purposes only, this document describes an overview of exemplary embodiments. In some embodiments, the disclosed techniques can be used to efficiently derive potential drug candidates from drug development candidates of interest during a drug discovery process. The potential drug candidates can then be evaluated to determine which of them are further drug development candidates. For a given drug discovery process, each drug development candidate (e.g., a drug development candidate of interest or an additional drug development candidate) is a molecule that has a desired effect on cells consistent with the overall goals of the drug discovery process. The desired effect on cells is also referred to herein as “target biological activity.” For example, each drug development candidate may improve the health of cells afflicted with a specific disease.

[0020] During the initialization phase, the molecular exploration application applies a hash function to any number of molecular directories to generate mapped directories. Each of the molecular directories includes available molecular specifications representing any number of available molecules. The hash function maps molecular specifications (i.e., "keys") to array indices. Each mapped directory is a hash mapping of the associated molecular directory, which stores the available molecular specifications included in the molecular directory in buckets based on the associated array index.

[0021] During the subsequent search phase, the molecular exploration application receives any number of search requests, each associated with a query molecule specification and any number of editing heuristics. A query molecule specification is a representation of a molecule referred to herein as a “query molecule.” In some implementations, the query molecule is a drug development candidate of interest with a target biological activity. Each editing heuristic specifies a different type of modification to the structure of an existing molecule. Editing heuristics are typically designed to have some relevance to one or more drug discovery processes.

[0022] In particular, for each of at least one of the editing heuristics, empirical evidence suggests that, relative to random structural modifications, structural modifications specified by the editing heuristic are more likely to preserve the target biological activity associated with a typical drug discovery process. In other words, when applied to query molecular specifications associated with the target biological activity, at least one of the editing heuristics is designed to preferentially generate a derived molecular specification that is also associated with the target biological activity.

[0023] In response to a given search request, the molecular exploration application iteratively applies an associated edit heuristic to the associated query molecular specification to generate a derived molecular specification representing all possible combinations of the edit heuristic. The molecular exploration application then applies a hash function to the derived molecular specification to generate a corresponding array index. For each of the mapped directories, the molecular exploration application searches for each of the derived molecular specifications based on the corresponding array index to generate a matching subset. Each of the matching subsets specifies, but is not limited to, a derived molecular specification present in the associated molecular directories.

[0024] Subsequently, the molecular exploration application generates a search results dataset based on the matching subset. The search results dataset includes, but is not limited to, search reports specifying each of the derived molecular specifications present in at least one of the molecular directories and associated locations (i.e., associated molecular directories). Each of the derived molecular specifications included in the search reports represents a different potential drug candidate. The search results dataset may include any amount and / or type of additional data relevant to the drug discovery process. The molecular exploration application then stores any portion of the search results dataset and / or provides it to any number of other software applications and / or users for the purpose of identifying additional drug development candidates.

[0025] System Overview

[0026] Figure 1 This is a conceptual illustration of system 100 configured to implement one or more aspects of various implementation schemes. For illustrative purposes, multiple instances of similar objects are indicated by reference numerals that identify the objects and, where necessary, by alphanumeric characters in parentheses that identify the instances. As shown, system 100 includes, but is not limited to, computing instance 110, display device 108, and molecular catalogs 102(1)-102(M), where M can be any positive integer.

[0027] In some implementations, system 100 may include, but is not limited to, any number of computing instances 110 in any combination, any number (including zero) of display devices 108, and any number of sub-catalogs 102. Components of system 100 may be distributed across any number of shared geographic locations and / or any number of distinct geographic locations and / or in any combination within one or more cloud computing environments (i.e., encapsulated shared resources, software, data, etc.).

[0028] As shown in the figure, computing instance 110 includes, but is not limited to, processor 112 and memory 116. Computing instance 110 may be implemented in a cloud computing environment, as part of any other distributed computing environment, or in a standalone manner. In some embodiments, each of any number of computing instances 110 may include any number of processors 112 and any number of memories 116 in any combination. In the same or other embodiments, any number of computing instances 110 (including one) may provide a multiprocessing environment in any technically feasible manner. Computing instance 110 is also referred to herein as a “computing device”.

[0029] Processor 112 can be any instruction execution system, device, or apparatus capable of executing instructions. For example, processor 112 may include a central processing unit, graphics processing unit, controller, microcontroller, state machine, or any combination thereof. Memory 116 of computing instance 110 stores content used by processor 112 of computing instance 110, such as software applications and data. Memory 116 can be one or more readily available memories, such as random access memory, read-only memory, floppy disk, hard disk, or any other form of local or remote digital storage device.

[0030] In some embodiments, a storage device (not shown) may supplement or replace memory 116. The storage device may include any number and type of external memory accessible to processor 112. For example, but not limited to, the storage device may include a secure digital card, external flash memory, portable optical disc read-only memory, optical storage device, magnetic storage device, or any suitable combination of the foregoing.

[0031] In some implementations, computing instance 110 may be associated with any number (including zero) and / or type of input devices, output devices, and / or input / output devices in any combination. An input device is any device capable of receiving input from a user. Some examples of input devices include, but are not limited to, keyboards, mice, touchpads, microphones, cameras, etc. An output device is any device capable of providing output to a user. Some examples of output devices include, but are not limited to, display device 108, headphones, speakers, etc. An input / output device is any device capable of receiving input from a user and providing output to the user, such as a touchscreen.

[0032] As shown in the figure, in some embodiments, computing instance 110 is associated with display device 108. Display device 108 can be any device capable of displaying images and / or any other type of visual content. For example, display device 108 can be, but is not limited to, a liquid crystal display, a light-emitting diode display, a projection display, a plasma display panel, etc. In some embodiments, display device 108 is a touchscreen capable of displaying visual content and receiving input (e.g., from a user).

[0033] In some implementations, computing instance 110 may be integrated into a user device along with any number and / or type of other devices (e.g., other computing instances 110, input devices, output devices, input / output devices, etc.). Some examples of user devices include, but are not limited to, desktop computers, laptop computers, smartphones, smart TVs, game consoles, tablet computers, etc.

[0034] Generally, compute instance 110 is configured to implement one or more software applications. For illustrative purposes only, each software application is described as residing in the memory 116 of compute instance 110 and executing on the processor 112 of compute instance 110. However, in some embodiments, the functionality of each software application may be distributed across any number of other software applications residing in the memory 116 of any number of compute instances 110 and executing on any number of processors 112 of compute instances 110 in any combination. Furthermore, the functionality of any number of software applications may be combined into a single software application.

[0035] In some implementations, any number of software applications and / or portions of software applications are stored on one or more non-transitory computer-readable media. The term "non-transitory" as used herein refers to a limitation of the media itself (i.e., tangible, not tactile), not a limitation on the persistence of data storage. For example Random access memory (RAM) versus read-only memory (ROM). Non-transitory computer-readable media are also referred to herein as "computer-readable media". For example, in some embodiments, memory 116 is a computer-readable medium, and any number of software applications and / or portions of software applications are stored in memory 116.

[0036] In some embodiments, any number of software applications and / or portions of software applications are stored on one or more computer-readable media and then stored in memory 116. For example, in some embodiments, any number of software applications and / or portions of software applications are stored on a machine (e.g., a server machine), and any number of software applications and / or portions of software applications are downloaded from the machine to memory 116. In the same or other embodiments, any number of software applications and / or portions of software applications are stored on some form of portable computer-readable medium, and any number of applications and / or portions of applications are downloaded from the portable computer-readable medium to memory 116. Some examples of portable computer-readable media include, but are not limited to, digital video disks, storage disks, memory sticks, etc.

[0037] In some embodiments, aspects of this disclosure may take the form of a computer program product embodied in one or more computer-readable media having a computer-readable program codec embodied thereon. Any combination of one or more computer-readable media may be used. Each computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination of the foregoing. More specific examples (not an exhaustive list) of computer-readable storage media include electrical connectors having one or more wires, portable computer floppy disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, portable optical disc read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, device, or apparatus.

[0038] In some implementations, computational instance 110 is configured to identify potential drug candidates based on drug development candidates of interest during the drug discovery process, thereby facilitating the identification of additional drug development candidates. A “drug development candidate” is a molecule that has a target effect on cells, referred to herein as a “target biological activity.” An example of a target biological activity is improving the health of cells afflicted with a specific disease. Each drug development candidate of interest can be any drug development candidate that has already been identified in any technically feasible manner (e.g., by laboratory testing). For example, a drug development candidate of interest may be identified during an assay measuring cell proliferation or death. In another example, a drug development candidate of interest may be identified during an assay searching for the presence of a specific biomarker (e.g., the amount of a specific cytokine excreted by cells). In yet another example, a drug development candidate may be identified during an assay searching for positive effects in model organisms.

[0039] As referred to herein, “additional drug development candidates” means drug development candidates identified based on associated drug development candidates of interest. Each “potential drug candidate” is a molecule structurally similar to an associated drug development candidate of interest and may have the same target biological activity as the associated drug development candidate of interest. In some embodiments, a subset of potential drug candidates is subsequently identified as additional drug development candidates.

[0040] In some embodiments, each of the additional drug development candidates is specified in at least one of the molecular catalogues 102(1)-102(M). For purposes of explanation only, the molecular catalogues 102(1)-102(M) are also referred to herein individually as “Molecular Catalog 102” and collectively as “Molecular Catalog 102”. The molecular catalogue 102 is also referred to herein as “Catalog of Molecules”. In some embodiments, M equals 1, and system 100 includes, but is not limited to, molecular catalogue 102(1).

[0041] Each of the molecules in catalog 102 includes, but is not limited to, any number of available molecular specifications (not shown). Each available molecular specification is a structural representation of a molecule that can be obtained in any technically feasible manner (e.g., ordered, generated, etc.). As indicated herein, a structural representation of a molecule specifies, but is not limited to, which atoms are bonded to each other and optionally any amount of additional structural information. For example, some structural representations of molecules include additional structural information specifying the approximate spatial arrangement of atoms in the molecule and / or any lone pairs of electrons that may be present in the molecule.

[0042] In some implementations, the structural representation of a given molecule can be any of the following: Simplified Molecular Linear Input Specification (“SMILES”) string, International Union of Pure and Applied Chemistry (“IUPAC”) International Compound Identifier (“InChI”), skeletal structure, etc. SMILES is a linear notation system that uses American Standard Code for Information Interchange (USCIS) characters to designate the structure of a given molecule as a SMILES string. InChI is a textual identifier for a given molecule. A skeletal structure is a two-dimensional (“2D”) graphical representation or “structure” that depicts, but is not limited to, how the atoms of a molecule are arranged in 3D space.

[0043] As shown in italics, in some embodiments, molecular catalog 102(1) is an internal molecular catalog that includes, but is not limited to, any number of available molecular specifications representing molecules available within an organization associated with system 100. In the same or other embodiments, each of any number of molecular catalogs 102 includes, but is not limited to, any number of available molecular specifications representing molecules that can be ordered from an associated molecular provider. In some embodiments, each of any number of molecular catalogs 102 may specify any number and / or type of molecules and any number and / or type of description of how each molecule is obtained, associated with data.

[0044] As previously described, in some conventional methods for identifying potential drug candidates, conventional similarity software applications search a catalog of molecules based on a similarity metric to find potential drug candidates. Typically, for each of any number of molecules specified in any number of molecule catalogs, the conventional similarity software application calculates a value for the similarity metric based on the drug development candidate and the molecule of interest. The conventional software application then determines molecules associated with a similarity metric value exceeding a minimum similarity threshold as potential drug candidates.

[0045] One drawback of using similarity metrics to identify potential drug candidates is that, for a given drug discovery process, the proportion of potential drug candidates with the same target biological activity as the drug development candidate of interest, identified using a specific similarity metric, may be very low. Therefore, significant time and resources are often wasted evaluating potential drug candidates that ultimately become irrelevant to the given drug discovery process.

[0046] Another drawback of using similarity metrics to identify potential drug candidates is that the computational complexity of the operations performed by conventional similarity software applications is typically proportional to the total number of available molecules. Therefore, the time and computational resources required to perform a comprehensive search of a molecule catalog for potential drug candidates can be extremely high.

[0047] Identifying potential drug candidates based on editing heuristics

[0048] To address the above issues, computational instance 110 includes, but is not limited to, automatically determining potential drug candidate specifications based on querying molecular specifications 150, editing heuristic sets 152, and molecular catalogs 102. Figure 1 The molecular exploration application 120 (not shown in the figure) resides in the memory 116 of the computing instance 110 and executes on the processor 112 of the computing instance 110, as shown in the figure.

[0049] For the purpose of explanation only, Figure 1 In the depicted implementation, the molecular exploration application 120 initially operates in an initialization phase, during which it receives any number and / or type of load catalog requests 104 and, in response, acquires and preprocesses molecular catalogs 102(1)-102(M). Subsequently, the molecular exploration application 120 operates in a search phase, during which it receives search requests 128 and, in response, generates a search results dataset 198.

[0050] In other embodiments, the molecular exploration application 120 may acquire and preprocess any subset (including the empty set) of the molecular catalog 102 during the initialization phase and subsequently acquire and preprocess the remainder of the molecular catalog 102 on demand during the search phase. In the same or other embodiments, the molecular exploration application 120 receives any number of search requests 128 during the search phase and, in response, generates any number of search result datasets 198.

[0051] As shown in the figure, in some embodiments, the molecular exploration application 120 includes, but is not limited to, a workflow engine 122, a directory mapping engine 130(1)-130(M), a derivation engine 160, a molecular mapping engine 134, a search engine 180(1)-180(M), and a merging engine 190. The workflow engine 122 performs any number and / or type of input, output, translation, etc. operations and routes data to and from the directory mapping engine 130(1)-130(M), the derivation engine 160, the molecular mapping engine 134, the search engine 180(1)-180(M), and the merging engine 190, respectively. The workflow engine 122 can receive input and provide output in any technically feasible manner.

[0052] As shown in the figure, in some embodiments, the workflow engine 122 displays a graphical user interface (“GUI”) 106 via a display device 108. The workflow engine 122 can receive any amount and / or type of input via the GUI 106 and can display any amount and / or type of output via the GUI 106. In some embodiments, the workflow engine 122 receives a load directory request 104 and / or a search request 128 via the GUI 106. In the same or other embodiments, the workflow engine 122 displays a portion (including none or all) of the search results dataset 198 via the GUI 106.

[0053] In some implementations, in response to a load catalog request 104, workflow engine 122 acquires molecular catalog 102 in any technically feasible manner. To preprocess molecular catalog 102, workflow engine 122 inputs available molecular specifications (not shown) included in molecular catalogs 102(1)-102(M) into catalog mapping engines 130(1)-130(M), respectively. In response, catalog mapping engines 130(1)-130(M) output mapped catalogs 140(1)-140(M), respectively.

[0054] Directory mapping engines 130(0)-130(M) are different instances of a single directory mapping engine 130 (not explicitly shown). For purposes of explanation only, as used herein, “directory mapping engine 130” refers to any instance of directory mapping engine 130, regardless of whether the specific instance is depicted in any of the accompanying drawings. Mapped directories 140(1)-140(M) are also individually referred to herein as “mapped directories 140” and are collectively referred to herein as “mapped directories 140”. Mapped directories 140 are also referred to herein as “mapped subdirectories”.

[0055] In some other embodiments, the molecular exploration application 120 includes fewer than M instances of the directory mapping engine 130, and the workflow engine 122 inputs molecular directories 102(1)-102(M) sequentially, simultaneously, or in any combination thereof into any number of instances of the directory mapping engine 130. For example, in some embodiments, the workflow engine 122 sequentially inputs molecular directories 102(1)-102(M) into a single instance of the directory mapping engine 130. In response, the single instance of the directory mapping engine 130 sequentially outputs the mapped directories 140(1)-140(M).

[0056] As explicitly shown in the catalog mapping engine 130(1), each of the catalog mapping engines 130(1)-130(M) includes, but is not limited to, mapping algorithm 132. Catalog mapping engine 130(x) (integer x from 1 to M) performs a mapping operation on each of the available molecular specifications included in the molecular catalog 102(x) based on mapping algorithm 132 to generate a mapped catalog 140(x). The mapped catalog 140(x) includes, but is not limited to, mapped versions of each of the available molecular specifications included in the molecular catalog 102(x).

[0057] Mapping algorithm 132 can be any type of algorithm that generates a mapped version of the molecular specification when applied to it. For illustrative purposes only, a mapped version of the molecular specification (including one of the available molecular specifications) is also referred to herein as a “mapped molecular specification.” In some embodiments, mapping algorithm 132 is associated with any number of mapping operations that map the molecular specification to the mapped molecular specification when it is applied to the molecular specification. The mapped molecular specification is a compressed digital representation of the molecular specification that facilitates the search operation. Mapping algorithm 132 is compatible with at least one type of molecular specification (e.g., the string SMILES) and can generate a compressed digital representation of the molecular specification of at least one type.

[0058] In some implementations, the molecular catalog 102 is used to represent molecular specifications of available molecules that are compatible with the mapping algorithm 132. In some such implementations, the catalog mapping engine 130(x) applies the mapping algorithm 132 to each of the available molecular specifications included in the molecular catalog 102(x) to generate a mapped version of the available molecular specification.

[0059] In some other embodiments, the type of molecular specification used by molecular catalog 102 to represent available molecules is incompatible with mapping algorithm 132. In some such embodiments, workflow engine 122 and / or catalog mapping engine 130 perform any number and / or type of translation operations on each of the available molecular specifications included in molecular catalog 102 to generate a compatible version of molecular catalog 102. Catalog mapping engine 130(x) applies mapping algorithm 132 to each of the available molecular specifications included in the compatible version of molecular catalog 102(x) to generate a mapped version of the available molecular specification.

[0060] In some embodiments, mapping algorithm 132 performs a search for a specific type of molecular catalog 102. For example, in some embodiments, mapping algorithm 132 maps available molecular specifications to a vector space to perform a vector similarity search of molecular catalog 102 using mapping algorithm 132 and the mapped catalog 140. In the same or other embodiments, mapping algorithm 132 ensures that the time required to search for molecular specifications in molecular catalog 102 using mapping algorithm 132 and the mapped catalog 102 is independent of the size of the mapped catalog 102.

[0061] In some implementations, mapping algorithm 132 exhibits any number and / or type of properties typically required in a drug discovery process. For example, in some implementations, mapping algorithm 132 is tautomer-friendly. As indicated herein, mapping algorithm 132 is “tautomer-friendly” if it attempts to detect different possible tautomer forms of a molecular specification and map them to a single mapped version of the molecular specification.

[0062] As depicted in italics, in some embodiments, mapping algorithm 132 is a hash function. The hash function maps molecular specifications (i.e., "bonds") to array indices. Therefore, the mapped catalog 140(x) (where x is an integer from 1 to M) is a hash mapping of the molecular catalog 102(x). In some embodiments, the mapped catalog 140(x) stores available molecular specifications included in the molecular catalog 102(x) based on the associated array index. Advantageously, as those skilled in the art will recognize, the time to look up a key via the hash function and hash mapping is independent of the size of the hash mapping.

[0063] In some embodiments, molecular catalog 102 represents available molecules as the string SMILES, and mapping algorithm 132 is a tautomer-friendly hash function that maps InChI or SMILES strings to InChIKey. "InChIKey" encodes the molecular specification into a compressed, fixed-length format, which is often referred to as "hash InChI". Therefore, in some embodiments, the mapped catalog 140(x) (integer x from 1 to M) stores a hash mapping of the SMILES strings included in molecular catalog 102(x) based on the associated InChIKey.

[0064] Although not shown, in some embodiments, after generating each of the mapped directories 140, the workflow engine 122 stores the mapped directories 140 in any memory accessible to the molecular exploration application 120. In this way, one or more instances of the molecular exploration application 120 can generate the mapped directories 140, and other instances of the molecular exploration application 120 can reuse the mapped directories 140. In some embodiments, the workflow engine 122 can store the mapped directories 140 in memory in response to any number and / or type of requests and retrieve the mapped directories 140 from memory based on any number and / or type of requests.

[0065] During the search phase, workflow engine 122 receives search requests 128 associated with query molecule and, in response, determines query molecule specifications 150 and edits heuristic sets 152. The query molecule is the molecule that will become the starting point for searching the molecule catalog 102. For example, in some embodiments, the query molecule is a drug development candidate of interest in the drug discovery process, and therefore the query molecule has associated target biological activities. The query molecule can be described and associated with search request 128 in any technically feasible manner.

[0066] The query molecule specification 150 is a structural representation of the query molecule in any format supported by the mapping algorithm 132 previously described herein. For example, in some embodiments, the input to the mapping algorithm 132 is a string of SMILES, and therefore the query molecule specification 150 is a string of SMILES representing the query molecule. The workflow engine 122 performs any number and / or type of operations based on any description of the query molecule associated with the search request 128 to generate the query molecule specification 150.

[0067] For example, in some embodiments, search request 128 specifies a structural representation of the query molecule in a format supported by mapping algorithm 132, and workflow engine 122 sets query molecule specification 150 to be equal to the structural representation. In some other embodiments, search request 128 is associated with a skeletal structure of the query molecule graphically, and workflow engine 122 translates the skeletal structure into a structural representation of the query molecule in a format supported by mapping algorithm 132.

[0068] Editing heuristics set 152 includes, but is not limited to, any number and / or type of editing heuristics. Figure 1 (Not shown in the image). Each of the editing heuristics specifies a different type of modification to the structure of an existing molecule in any technically feasible manner. For example, in some embodiments, when applied to an existing molecule, one of the editing heuristics repositions any number of nitrogen atoms to generate any number (including zero) of isomers of the existing molecule that differ only in the nitrogen position.

[0069] In some embodiments, editing heuristics are designed to be correlated with one or more drug discovery processes. In the same or other embodiments, for each of at least one of the editing heuristics, empirical evidence suggests that, relative to random structural modifications, structural modifications specified by the editing heuristic are more likely to preserve the target biological activity associated with a typical drug discovery process. In some embodiments, one or more of the editing heuristics are designed to improve target biological activity, remove risk from the molecule, provide insight into whether a particular part of the molecule is important relative to the target biological activity, or any combination thereof.

[0070] Workflow engine 122 can determine the set of editing heuristics 152 associated with search request 128 in any technically feasible manner. In some embodiments, workflow engine 122 generates the set of editing heuristics 152 based on any number and / or type of commands received during the initialization phase in any technically feasible manner (e.g., from GUI 106). In the same or other embodiments, workflow engine 122 can add, delete, and / or modify any number of editing heuristics included in the set of editing heuristics 152 based on search request 128 and / or any number and / or type of commands associated with search request 128. For example, in some embodiments, search request 128 specifies any number and / or type of editing heuristics designed to increase the likelihood that applying editing heuristics to query molecule specification 150 will preserve the target biological activity associated with the query molecule.

[0071] As shown in the figure, in some implementations, workflow engine 122 inputs query molecular specifications 150 and edit heuristic set 152 into derivation engine 160. In response, derivation engine 160 generates derivation dataset 162. As shown in the figure, in some implementations, derivation dataset 162 includes, but is not limited to, the derivation molecular specifications 168(1)-168(N) and the applied edit lists 164(1)-164(N), where N can be any positive integer.

[0072] Each of the derived molecular specifications 168(1)-168(N) represents a different molecule derived based on the structure of the query molecule. For illustrative purposes only, the derived molecular specifications 168(1)-168(N) are also referred to individually as “derived molecular specifications 168” and collectively as “derived molecular specifications 168” in this document. Molecules represented by derived molecular specifications 168 are also referred to as “derived molecules” in this document.

[0073] The derivation engine 160 can generate the derivation of the molecular specification 168 based on the query molecular specification 150 and the editing heuristic set 152 in any technically feasible manner. In some embodiments, the derivation engine 160 may apply any number of editing heuristics included in the editing heuristic set 152 individually and / or in any number of combinations to the query molecular specification 150 to generate the derivation of the molecular specification 168(1)-168(N).

[0074] As those skilled in the art will recognize, in some embodiments, applying a given edit heuristic to a given molecular specification can produce any number (including zero) of derived molecular specifications 168. For example, if a given edit heuristic specifies the removal of a substituent and the given molecule does not contain said substituent, applying the edit heuristic to the associated molecular specification will not produce derived molecular specifications 168. In contrast, if a given edit heuristic specifies the addition to the list of substituents of a given molecule, applying the edit heuristic to the associated molecular specification can produce multiple derived molecular specifications 168.

[0075] For illustrative purposes only, each of the derived molecules is associated with the number of edits relative to the query molecule. The number of edits relative to the query molecule for a given derived molecule refers to the total number of heuristic-based edits made by the derivation engine 160 from the query molecule specification 150 to generate a derived molecule specification 168 representing the derived molecule. As referred to herein, "heuristic-based editing" is the application of one of the editing heuristics to either the query molecule specification 150 or one of the derived molecule specifications 168.

[0076] For example, in some implementations, the derivation engine 160 applies one of the editing heuristics to the query molecule specification 150 to generate a derived molecule specification 168(1), which represents a derived molecule edited one time from the query molecule. Subsequently, the derivation engine 160 applies the same or another editing heuristic to the derived molecule specification 168(1) to generate a different derived molecule specification 168, which represents a derived molecule edited two times from the query molecule.

[0077] As shown below Figure 2 In more detail, in some embodiments, the derivation engine 160 recursively applies the edit heuristics included in the edit heuristic set 152 to the query molecule specification 150 to generate the derived molecule specification 168. In some embodiments, during the first iteration, the derivation engine 160 applies each of the edit heuristics to the query molecule specification 150 to generate a first subset of the derived molecule specification 168. The first subset of the derived molecule specification 168 represents the molecule that has been edited once from the query molecule.

[0078] During the second iteration, the derivation engine 160 applies each of the edit heuristics to a first subset of the derived molecular specification 168 to generate a second subset of the derived molecular specification 168. The second subset of the derived molecular specification 168 represents molecules that have been edited twice from the query molecule. In some embodiments, the derivation engine 160 continues to apply the edit heuristics to the most recently generated subset of the derived molecular specification 168 until the derivation engine 160 has exhausted applying each of the edit heuristics and every possible combination of edit heuristics to the query molecular specification 150.

[0079] As depicted by dashed boxes and dashed arrows, in some embodiments, derivation engine 160 receives edit constraint 154. Edit constraint 154 can be any type of constraint that limits the total number of derived molecular specifications 168 generated by derivation engine 160 in any technically feasible manner. For example, in some embodiments, edit constraint 154 specifies the maximum number of iterations that derivation engine 160 can perform. In the same or other embodiments, edit constraint 154 specifies the maximum number of edits that any derived molecule can make relative to the query molecule. In some embodiments, derivation engine 160 implements a default value for edit constraint 154.

[0080] In some implementations, as the derivation engine 160 generates the derived molecular specifications 168(1)-168(N), the derivation engine 160 also generates an applied edit list 164(1)-164(N). The applied edit lists 164(1)-164(N) are also referred to herein separately as “applied edit list 164” and collectively as “applied edit list 164”. The applied edit list 164(y) (integer y from 1 to N) specifies, but is not limited to, the edits that the derivation engine 160 applies to the query molecular specification 150 to generate the derived molecular specification 168(y). The derivation engine 160 may specify the edits included in the applied edit list 164 at any level of detail and in any technically feasible manner. In some other implementations, the derivation engine 160 does not generate the applied edit list 164 and omits the applied edit list 164 from the derivation dataset 162.

[0081] As shown in the figure, in some implementations, workflow engine 122 inputs the derived molecular specifications 168(1)-168(N) into molecular mapping engine 134. In response, molecular mapping engine 134 generates mapped datasets 170(1)-170(N). Mapped datasets 170(1)-170(N) are also referred to herein as “mapped dataset 170” and collectively as “mapped dataset 170”.

[0082] Although not shown, the mapped dataset 170(y) (where y is an integer from 1 to N) includes, but is not limited to, the derived molecular specification 168(y) and the mapped version of the derived molecular specification 168(y). As mentioned earlier in this document, the mapped version of the molecular specification (including one of the derived molecular specifications 168) is also referred to herein as the “mapped molecular specification”.

[0083] Molecular mapping engine 134 can generate mapped versions of the derived molecular specifications 168(1)-168(N) in any technically feasible manner consistent with the mapping operations performed by catalog mapping engine 130. As shown, in some embodiments, molecular mapping engine 134 includes, but is not limited to, mapping algorithm 132, which is also included in catalog mapping engine 130. Molecular mapping engine 134 applies mapping algorithm 132 to each of the derived molecular specifications 168(1)-168(M) to generate mapped versions of the derived molecular specifications 168(1)-168(M) respectively.

[0084] As previously described herein, in some embodiments, mapping algorithm 132 is a hash function that maps the SMILES string to an InChIKey. In some such embodiments, each of the derived molecular specifications 168 is a SMILES string and each of the mapped versions of the derived molecular specifications 168 is an InChIKey. In the same or other embodiments, each of the mapped datasets 170 includes, but is not limited to, a SMILES string and an InChIKey that both represent the derived molecules.

[0085] As shown in the figures, in some implementations, workflow engine 122 inputs mapped datasets 170(1)-170(N) and mapped directories 140(1)-140(M) into search engines 180(1)-180(M), respectively. In response, search engines 180(1)-180(M) perform any number and / or type of search operations to generate matching subsets 188(1)-188(M), respectively. Search engines 180(1)-180(M) are different instances of a single search engine 180 (not explicitly shown). For purposes of explanation only, as used herein, "search engine 180" refers to any instance of search engine 180, regardless of whether the specific instance is depicted in any of the figures.

[0086] In some other embodiments, the molecular exploration application 120 includes fewer than M instances of the search engine 180, and the workflow engine 122 inputs mapped datasets 170(1)-170(N) and mapped directories 140(1)-140(M) sequentially, simultaneously, or in any combination thereof into any number of instances of the search engine 180. For example, in some embodiments, the molecular exploration application 120 inputs mapped datasets 170(1)-170(N) and sequentially inputs mapped directories 140(1)-140(M) into a single instance of the search engine 180. In response, the single instance of the search engine 180 sequentially outputs matching subsets 188(1)-188(M).

[0087] Matching subsets 188(1)-188(M) are also referred to individually as “matching subset 188” and collectively as “matching subset 188” in this document. Matching subset 188(x) (integers x from 1 to M) includes, but is not limited to, each of the derived molecular specifications 168(1)-168(N) that are also included in molecular catalog 102(x). As indicated herein, a derived molecular specification 168(y) is included in molecular catalog 102(x) if and only if it matches one of the available molecular specifications included in molecular catalog 102(x). Therefore, matching subset 188(x) is the intersection of the set of derived molecular specifications 168(1)-168(N) and the set of available molecular specifications included in molecular catalog 102(x).

[0088] Search engine 180(x) can perform any number and / or type of search operations (e.g., comparison operations, etc.) in any technically feasible manner to determine whether each of the derived molecular specifications 168 matches any available molecular specification included in the molecular catalog 102(x). Advantageously, to improve the efficiency of the search operations, search engine 180(x) performs the search operations at least in part based on the mapped version of the derived molecular specification 168 and the mapped catalog 140(x).

[0089] As those skilled in the art will recognize, in some embodiments, the computational complexity of performing a search operation based on the mapped version of the derived molecular specification 168 and the mapped catalog 140 is constant relative to the total number of available molecules included in the molecular catalog 102. In the same or other embodiments, because the computational complexity of searching for the derived molecular specification 168 in the molecular catalog 102 is independent of the total number of available molecules included in the molecular catalog 102, all available molecules can be searched at an interactive rate. As used herein, “interactive rate” refers to the response rate (e.g., less than one second) that typically does not interrupt the interactive flow between the user and the user-perceived application.

[0090] As previously described herein, in some embodiments, mapping algorithm 132 is a hash function, and search engine 180(x) can perform any technically feasible type of hash-based search to determine whether each of the derived molecular specifications 168(1)-168(N) is included in molecular catalog 102(x). In some such embodiments, the computational complexity of searching for the derived molecular specification 168 in molecular catalog 102 is approximately A^T. As used herein, the symbol A is proportional to the total number of atoms included in the query molecule, and the symbol T is the maximum number of edits any derived molecule has from the query molecule.

[0091] In some implementations, the mapped directory 140(x) is a hash mapping based on the mapped version of the available molecular specifications, storing the available molecular specifications in the molecular directory 102(x) in any number of buckets. To determine whether the derived molecular specification 168(y) (integer y from 1 to N) is included in the molecular directory 102(x), the search engine 180(x) performs a hash-based search based on the buckets.

[0092] In some implementations, to enable hash-based searching, search engine 180(x) identifies buckets included in the mapped directory 140(x) that are associated with the mapped version of the derived molecular specification 168(y). Search engine 180(x) then performs a search on the available molecular specifications stored in the identified buckets to determine whether any of the available molecular specifications matches the derived molecular specification 168(y). If the identified bucket is empty or none of the available molecular specifications stored in the identified bucket matches the derived molecular specification 168(y), search engine 180(x) does not add the derived molecular specification 168(y) to the matching subset 188(x). Otherwise, search engine 180(x) adds the derived molecular specification 168(y) to the matching subset 188(x).

[0093] As shown in the figure, in some implementations, workflow engine 122 inputs derivation dataset 162 and matching subset 188 into merging engine 190. In response, for each of the derivation molecular specifications 168 that are also included in at least one of the matching subset 188, merging engine 190 generates a potential drug candidate dataset ( Figure 1 (Not shown in the image). Each of the potential drug candidate datasets is associated with a different potential drug candidate. In some implementations, potential drug candidates are more likely to be associated with the drug discovery process associated with the query molecule than with other available molecules.

[0094] Each of the potential drug candidate datasets includes, but is not limited to, the potential drug candidate specification ( Figure 1 (not shown in the image), location list ( Figure 1 (Not shown in the table) and any other amount (including none) and / or type of additional data. A potential drug candidate specification is equal to the associated derived molecular specification 168. Each of the locations in the location list is specified, but not limited to, at least one of the molecular catalogs 102 that include the associated potential drug candidate specification. In some embodiments, the merging engine 190 generates the location list based on a matching subset 188 and any amount and / or type of preferences associated with the molecular catalog 102.

[0095] In some implementations, molecular directories 102(1)-102(M) are associated with a preference ranking from highest to lowest. If a molecule is included in molecular directory 102(1), the molecule is preferentially obtained from the associated provider. In some such implementations, for each potential drug candidate dataset, merging engine 190 determines an unordered list based on matching subset 188. Each unordered list is a subset of the associated potential drug candidate specification for each specified molecular directory 102. For each potential drug candidate dataset, merging engine 190 orders the associated unordered lists based on preference ranking to generate a list of positions.

[0096] In the same or other embodiments, instead of a list of locations or in addition to a list of locations, each of the potential drug candidate datasets also includes a preferred location. In some such embodiments, for each potential drug candidate dataset, the merging engine 190 specifies the molecular catalog 102 with the highest preference ranking of a subset of the molecular catalog 102 that includes the associated potential drug candidate specifications, by the preferred location.

[0097] Merging engine 190 generates a search results dataset 198 based on a dataset of potential drug candidates and any amount (including none) and / or type of additional data. Although not shown, in some embodiments, merging engine 190 interacts with workflow engine 122, inference engine 160, search engine 180, or any combination thereof to generate search results dataset 198. For example, in some embodiments, merging engine 190 receives query molecular specifications 150 and / or preference rankings associated with molecular catalog 102 from workflow engine 122.

[0098] In the same or other implementations, the merging engine 190 interacts with the workflow engine 122 to generate a skeletal structure that graphically represents the query molecule. Figure 1 (Not shown in the diagram). In the same or other embodiments, for each of the potential drug candidate specifications, workflow engine 122 interacts with derivation engine 160 to generate annotated skeletal structures that graphically represent the potential drug candidates and the structural differences between the potential drug candidates and the query molecule. Figure 1 (Not shown in the image).

[0099] As shown below Figure 3 In more detail, in some embodiments, the search results dataset 198 includes, but is not limited to, search reports, query datasets, and search summaries. In some embodiments, the search reports include, but are not limited to, datasets of potential drug candidates. In the same or other embodiments, the query dataset includes, but is not limited to, any amount and / or type of data associated with the query molecule. In some embodiments, the search summary includes, but is not limited to, any amount and / or type of data (e.g., statistics) associated with the search operation performed on the molecular catalog 102 through the mapped catalog 140.

[0100] As shown in the figure, in some embodiments, workflow engine 122 displays any portion (including all) of the search results dataset 198 via GUI 106. In the same or other embodiments, instead of displaying any portion of the search results dataset 198 via GUI 106 or otherwise, workflow engine 122 provides any portion of the search results dataset 198 to any number of users and / or any number and / or type of software applications. In the same or other embodiments, system 100 omits GUI 106, and workflow engine 122 can acquire input data and provide output data in any technically feasible manner.

[0101] Advantageously, the overall efficiency of the drug discovery process can be improved by prioritizing the evaluation of potential drug candidates when identifying additional drug development candidates. In particular, because the heuristic set 152 can be customized for a given drug discovery process, the relevance of potential drug candidates to the drug discovery process can be increased compared to conventional potential drug candidates identified based on similarity metrics. Therefore, in some embodiments, the amount of time and resources wasted evaluating molecules that do not possess the target biological activity when identifying additional drug development candidates can be reduced.

[0102] Note that the techniques described herein are illustrative and not limiting, and can be modified without departing from the broader spirit and scope of the invention. Many modifications and variations to the functionality provided by the molecular exploration application 120, workflow engine 122, directory mapping engine 130, derivation engine 160, molecular mapping engine 134, search engine 180, and merging engine 190 will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. For example, in some embodiments, the functionality provided by the molecular exploration application 120 as described herein is divided into an initialization application (not shown) stored in different memories 116 and executed on different processors 112, and a search application (not shown). In some embodiments, modifications may be made as needed. Figure 1 The connection topology between the various components.

[0103] Applying edit heuristics to query molecular norms

[0104] Figure 2 It is based on various implementation plans. Figure 1 A more detailed illustration of the derivation engine 160 is provided. As shown, in some embodiments, the derivation engine 160 generates a derivation dataset 162 based on a query molecule specification 150 and an edit heuristic set 152 including but not limited to edit heuristics 210(1)-210(4). Edit heuristics 210(1)-210(4) are also referred to herein individually as “edit heuristic 210” and collectively as “edit heuristic 210”. In other embodiments, the edit heuristic set 152 may include any number and / or type of edit heuristics 210.

[0105] As previously mentioned in this article Figure 1Each of the edit heuristics 210 specifies a different type of modification to the structure of an existing molecule. For illustrative purposes only, exemplary functions associated with edit heuristics 210(1)–210(4) are depicted in italics. As shown, the function of edit heuristic 210(1) will generate isomers differing only in the nitrogen position. The function of edit heuristic 210(2) will add or remove nitrogen from the ring. As shown, edit heuristic 210(3) includes, but is not limited to, a substituent list 212. The function of edit heuristic 210(3) will add, remove, or move (… For example (Repositioning) includes one of any number of substituents in the substituent list 212. Some examples of substituents include, but are not limited to, amino (“NH2”), carboxyl (“COOH”), etc. Editing the function of heuristic 210(4) will substitute sulfur with oxygen or vice versa.

[0106] As shown in the figure, the derivation engine 160 includes, but is not limited to, an iterative engine 220 that incrementally generates the derivation tree 240. For illustrative purposes only, the depiction of the derivation tree 240 is annotated with dashed boxes using skeleton structures 250(0)-250(N). Skeleton structures 250(0)-250(N) are also referred to herein as “skeleton structure 250” individually and collectively. Each of the skeleton structures 250 is a 2D representation of the associated molecule, which depicts, but is not limited to, how the atoms of the molecule are arranged in 3D space. Skeleton structure 250(0) graphically depicts the query molecule represented by the query molecule specification 150. Skeleton structures 250(1)-250(N) graphically depict the derivation molecules represented by the derivation molecule specifications 168(1)-168(N), respectively.

[0107] exist Figure 2 In the depicted embodiments, the skeleton structure 250 is merely illustrative and is not included in the derivation tree 240 generated by the iteration engine 220. In some other embodiments, the derivation engine 160 and / or the iteration engine 220 may acquire (e.g., receive, generate, etc.) the skeleton structures 250(0)-250(N) in any technically feasible manner. In some such embodiments, the derivation engine 160 and / or the iteration engine 220 add the skeleton structure 250 to the derivation tree 240 and / or the derivation dataset 162. In the same or other embodiments, the derivation engine 160, the iteration engine 220, the workflow engine 122, or any combination thereof may be displayed via the GUI 106 for any portion of the derivation tree 240 and / or any portion of the derivation dataset 162.

[0108] In some implementations, the iterative engine 220 generates derivation trees 240 of depths including, but not limited to, 242(0)-242(T), where T can be any positive integer. Depths 242(0)-242(T) are also referred to herein as "depth 242" and collectively as "depth 242". For illustrative purposes, and as described in detail below, Figure 2 In the described implementation scheme, T is an integer greater than 3.

[0109] When derivation engine 160 receives query molecule specification 150, iteration engine 220 generates an initial version of derivation tree 240, which includes, but is not limited to, a depth 252(0) corresponding to the root of derivation tree 240. Iteration engine 220 adds query molecule specification 150 to derivation tree 240 at depth 252(0). Therefore, depth 252(0) includes, but is not limited to, query molecule specification 150. For illustrative purposes only, the skeletal structure 250(0) of an exemplary query molecule represented by query molecule specification 150 is depicted by a dashed box.

[0110] During the first iteration, the iteration engine 220 will individually apply each of the edit heuristics 210 included in the edit heuristic set 152 to the query molecule specification 150. (As previously mentioned in this paper...) Figure 1 As described, in some implementations, applying each of the edit heuristics 210 to the query molecular specification 150 can produce any number (including zero) of the derived molecular specifications 168.

[0111] As shown in the figure, during the first iteration, the iteration engine 220 generates the derived molecular specifications 168(1)-168(a), where a can be any positive integer. The iteration engine 220 adds the derived molecular specifications 168(1)-168(a) to the derivation tree 240 at a depth of 252(1), which is associated with the derived molecule one edit away from the query molecule. For illustrative purposes only, the skeletal structures 250(1)-250(a) of the derived molecules, represented by the derived molecular specifications 168(1)-168(a), are depicted by dashed boxes.

[0112] During the second iteration, the iteration engine 220 individually applies each of the edit heuristics 210 to each of the derived molecular specifications 168(1)-168(a) at depth 252(1). More precisely, as shown, the iteration engine 220 individually applies each of the edit heuristics 210 to each of the derived molecular specifications 168(1) to generate the derived molecular specifications 168(a+1)-168(b), where b can be any integer greater than a. Although not shown, the iteration engine 220 individually applies each of the edit heuristics 210 to each of the derived molecular specifications 168(2)-168(a-1) to generate the derived molecular specifications 168(b+1)-168(c-1), where c can be any integer greater than b. As shown in the figure, the iteration engine 220 applies each of the editing heuristics 210 individually to the derived molecular specification 168(a) to generate the derived molecular specifications 168(c)-168(d), where d can be any integer greater than c.

[0113] Iteration engine 220 adds the derived molecular specifications 168(a+1)-168(d) to derivation tree 240 at a depth of 252(2), which is associated with the derived molecule at a distance of 2 edits from the query molecule. For illustrative purposes only, the skeletal structures 250(a+1), 250(c), and 250(d) of the derived molecules, represented by the derived molecular specifications 168(a+1), 168(c), and 168(d), are depicted by dashed boxes.

[0114] Subsequently, although not explicitly shown, the iterative engine 220 iteratively generates the derived molecular specification 168(d+1)-168(e-1), where e is any integer greater than d, and the derived molecular specification 168(d+1)-168(e-1) is distributed across depths 252(3)-252(T-1), where T is any positive integer greater than 3. Depths 252(3)-252(T-1) are associated with derived molecules that are 3-(T-1) edits away from the query molecule. Figure 2 In the described implementation, the iterative engine 220 generates the derived molecular specification 168 at depths 252(3)-252(T-1) based on the derived molecular specification at depths 252(2)-252(T-2).

[0115] Iteration engine 220 generates derived molecular specifications 168(e)-168(N) based on the derived molecular specification 168 at depth 252(T-1), where N is any integer greater than e. Iteration engine 220 adds the derived molecular specifications 168(e)-168(N) to the derivation tree 240 at depth 252(T), which is associated with the derived molecule T edits from the query molecule. Iteration engine 220 then determines that it has exhaustively applied each of the edit heuristics 210 and each possible combination of the edit heuristics 210 to the query molecule specification 150, and therefore stops iterating.

[0116] After the iteration engine 220 stops iterating, the derivation tree 240 includes, but is not limited to, the query molecular specification 150 at depth 252(0), and the derivation molecular specifications 168(1)-168(N) distributed across depths 252(1)-252(T). The derivation engine 160 generates a derivation dataset 162 based on the derivation tree 240. As previously described herein, in some embodiments, the derivation dataset 162 includes, but is not limited to, the derivation molecular specifications 168(1)-168(N) and the applied edit lists 164(1)-164(N).

[0117] As previously mentioned in this article Figure 2 The applied edit list 164(y) (integers y from 1 to N) specifies, but is not limited to, the edits that the iteration engine 220 applies to the query molecular specification 150 to generate the derived molecular specification 168(y). The derivation engine 160 may specify the edits included in the applied edit list 164 at any level of detail and in any technically feasible manner. As shown in italics for illustrative purposes only, in some embodiments, the derivation engine 160 sets the applied edit list 164(1) to be equal to “add nitrogen to the ring”.

[0118] Figure 3 It is based on various implementation plans. Figure 1 A more detailed illustration of the search results dataset 198. As previously combined in this article Figure 1 As stated, the merging engine 190 can generate the search results dataset 198 in any technically feasible manner. This is for illustrative purposes only. Figure 3 The description of the merge engine 190 is at least partially based on Figure 2 The example of the derivation dataset 162 depicted is an example of the search results dataset 198 generated. In some other embodiments, the search results dataset 198 can specify any number of potential drug candidates and any amount and / or type of associated data in any technically feasible manner.

[0119] As shown in the figure, in some implementations, the search results dataset 198 includes, but is not limited to, query dataset 302, search summary 310, and search report 380. Query dataset 302 describes the query molecule and includes, but is not limited to, query molecule specification 150 and skeleton structure 250(0). See again Figure 2 The skeletal structure 250(0) is a 2D representation of how the atoms of the query molecule are arranged in 3D space, but not limited to a description of the molecule.

[0120] Search summary 310 is a 2D table associated with the search operation performed by search engine 180 on molecular directory 102 through mapped directory 140. As shown, the search summary includes, but is not limited to, rows 312(1)-312(M), columns 314(1)-314(T), and match counts 320(1,1)-320(M,T). Rows 312(1)-312(M) correspond to molecular directories 102(1)-102(M) respectively, and are referred to herein as “row 312”. Columns 314(1)-314(T) correspond to depths 252(1)-252(T) respectively, and are referred to herein as “column 314”.

[0121] Match counts 320(1,1)-320(M,T) are also referred to as "match count 320" in this document and are collectively referred to as "match count 320". See again Figure 1 and Figure 2 The match count 320(i,j) (where i is an integer between 1 and M and j is an integer between 1 and T) is the total number of deduced molecular specifications 168 associated with depth 252(j) included in the match subset 188(i). Therefore, the match count 320(i,j) specifies the total number of deduced molecules edited j times from the query molecule specified in the molecular catalog 102(i).

[0122] For illustrative purposes only, some exemplary values ​​of match count 320 are depicted in italics. As shown, match count 320(1,1) is 4, indicating that molecule directory 102(1) includes a representation of 4 derived molecules from the molecule one edited from the query molecule. Match count 320(2,1) is 17, indicating that molecule directory 102(2) includes a representation of 17 derived molecules from the molecule one edited from the query molecule. Match count 320(M,1) is 0, indicating that molecule directory 102(M) includes a representation of zero derived molecules from the molecule one edited from the query molecule.

[0123] As shown in the figure, match count 320(1,T) is 10, indicating that molecule directory 102(1) includes the representation of 10 derived molecules from the derived molecules edited T times from the query molecule. Match count 320(2,1) is 167, indicating that molecule directory 102(2) includes the representation of 167 derived molecules from the derived molecules edited T times from the query molecule. Match count 320(M,T) is 2, indicating that molecule directory 102(M) includes the representation of 2 derived molecules from the derived molecules edited T times from the query molecule.

[0124] As shown in the figure, in some implementations, search report 380 includes, but is not limited to, potential drug candidate datasets 390(1)-390(P), where P is an integer less than or equal to N (the total number of derived molecular specifications 168). For illustrative purposes only, potential drug candidate datasets 390(1)-390(P) are also referred to herein as “potential drug candidate dataset 390” and collectively as “potential drug candidate dataset 390”. Each of the potential drug candidate datasets 390 describes a different potential drug candidate.

[0125] As shown in the figure, the potential drug candidate dataset 390(1) includes, but is not limited to, potential drug candidate specifications 392(1), a position list 394(1), an annotated skeleton structure 396(1), and a modification level 398(1). Also as shown in the figure, the potential drug candidate dataset 390(P) includes, but is not limited to, potential drug candidate specifications 392(P), a position list 394(P), an annotated skeleton structure 396(P), and a modification level 398(P). Although not explicitly shown, the potential drug candidate dataset 390(k) (where k is an integer between 2 and P-1) includes, but is not limited to, potential drug candidate specifications 392(k), a position list 394(k), an annotated skeleton structure 396(k), and a modification level 398(k).

[0126] Each of the potential drug candidate specifications 392(1)-392(P) is equal to a different derived molecular specification in the derived molecular specification 168 included in at least one of the molecular catalogs 102. Each of the position lists 394(1)-394(P) specifies, but is not limited to, a subset of the molecular catalogs 102 that respectively include potential drug candidate specifications 392(1)-393(P). In some embodiments, each of the position lists 394(1)-394(P) is ordered based on a preference order associated with the molecular catalogs 102.

[0127] The annotated skeletal structures 396(1)-396(P) are skeletal structures 250 representing associated potential molecules, which are annotated in any technically feasible manner (e.g., by coloring scheme) to graphically depict the structural differences between the associated potential drug candidate and the query molecule. The merging engine 190 can obtain the annotated skeletal structures 396(1)-396(P) in any technically feasible manner. In some embodiments, the workflow engine 122, the derivation engine 160, the merging engine 190, or any combination thereof can generate the annotated skeletal structures 396(1)-396(P).

[0128] Modification levels 398(1)-398(P) specify that derivation engine 160 is applied to query molecule specification 150 to generate the total number of editing heuristics 210 for potential drug candidate specifications 392(1)-392(P). See again Figure 2 In some implementations, each of the modification levels 398(1)-398(P) is a different depth 252, and thus indicates the total number of edits from the query molecule.

[0129] For illustrative purposes only, some exemplary values ​​of the potential drug candidate datasets 390(1) and 390(P) are depicted in italics. As shown, potential drug candidate specification 392(1) is the derived molecular specification 168(3) (not explicitly depicted). The location list 394(1) indicates that potential drug candidate specification 392(1) is included in the molecular catalog 102(2). The modification level 398(1) indicates that the potential drug candidate described by potential drug candidate specification 392(1) is one edit away from the query molecule.

[0130] In some implementations, potential drug candidate specification 392(P) is the derived molecular specification 168(e) (in Figure 2 (Depicted in the middle). Location list 394(P) indicates that the potential drug candidate specification 392(P) is included in the molecular catalog 102(1) and 102(M). Modification level 398(1) indicates that the potential drug candidate described in the potential drug candidate specification 392(1) is T edits away from the queried molecule.

[0131] Figure 4 This is a flowchart of methodological steps for identifying potential drug candidates during the drug discovery process, according to various implementation schemes. (Although references are available...) Figures 1 to 3 The system described herein describes the method steps, but those skilled in the art will understand that any system configured to implement the method steps in any order is within the scope of this invention.

[0132] As shown in the figure, method 400 begins with step 402, where any number of instances of the directory mapping engine 130 generate mapped directories 140(1)-140(M) based on molecular directories 102(1)-102(M) and mapping algorithm 132. The workflow engine 122 then waits for search requests 128. In some embodiments, the workflow engine 122 stores any number of mapped directories 140(1)-140(M) in any memory accessible to the molecular exploration application 120.

[0133] In step 404, workflow engine 122 determines the query molecular specification 150 and edit heuristic set 152 associated with search request 128. In step 406, derivation engine 160 calculates the derivation molecular specifications 168(1)-168(N) and optionally the applied edit lists 164(1)-164(N) based on the query molecular specification 150 and edit heuristic set 152.

[0134] In step 408, the molecular mapping engine 134 calculates the mapped versions of the derived molecular specifications 168(1)-168(N) based on the mapping algorithm 132. In step 410, any number of instances of the search engine 180 search for each of the derived molecular specifications 168(1)-168(N) in each of the mapped directories 140(1)-140(M) and the mapped versions of the derived molecular specifications 168(1)-168(N) to generate a matching subset 188(1)-188(M).

[0135] In step 412, the merging engine 190 generates a potential drug candidate dataset 390 based on the matching subset 188. In step 414, the merging engine 190 generates a search results dataset 198 based on the potential drug candidate dataset 390, and optionally any number of query molecular specifications 150, derived molecular specifications 168, and applied edit lists 164.

[0136] In step 416, workflow engine 122 stores any portion of the search results dataset 198 and / or provides it to any number of users and / or software application types for identifying additional drug development candidates. In step 418, workflow engine 122 determines whether it has received a new search request 128. If, in step 418, workflow engine 122 determines that it has not received a new search request 128, method 400 terminates.

[0137] However, if in step 418, workflow engine 122 determines that it has received a new search request 128, then method 400 returns to step 404, where workflow engine 122 determines the query molecule specification 150 and the edit heuristic set 152 associated with search request 128. Method 400 continuously loops through steps 404-418 to generate a new search results dataset 198 until workflow engine 122 determines in step 418 that it has not received a new search request 128. Then method 400 terminates.

[0138] In summary, the disclosed techniques can be used to infer potential drug candidates for the drug discovery process based on query molecules, edit heuristics tailored to the drug discovery process, and any number of molecular catalogs. In some implementations, the molecular exploration application includes, but is not limited to, workflow engines, catalog mapping engines, inference engines, molecular mapping engines, search engines, and merging engines.

[0139] During the initialization phase, the workflow engine inputs molecular directories into any number of instances of the directory mapping engine to generate mapped directories. Each of the molecular directories includes, but is not limited to, any number of available molecular specifications, where each of the available molecular specifications represents a different existing molecule. Each of the mapped directories includes, but is not limited to, mapped versions of the available molecular specifications included in the associated molecular directories. To generate a mapped directory corresponding to a given molecular directory, the directory mapping engine applies a mapping algorithm (e.g., a hash function) to each of the available molecular specifications included in the molecular directories.

[0140] Subsequently, during the search phase, the workflow engine receives any number of search requests, each associated with a query molecule. In some implementations, the query molecule is a drug development candidate of interest in the drug discovery process. In response to a given search request, the workflow engine determines the query molecule specification and edits the heuristic set. The query molecule specification represents the query molecule associated with the search request.

[0141] The set of editing heuristics includes, but is not limited to, any number and / or type of editing heuristics, each specifying a different type of modification to the structure of an existing molecule. Editing heuristics are typically designed to have some relevance to one or more drug discovery processes. In particular, for each of at least one of the editing heuristics, empirical evidence suggests that, relative to random structural modifications, structural modifications specified by an editing heuristic are more likely to preserve the target biological activity associated with a typical drug discovery process.

[0142] The derivation engine iteratively applies the edit heuristics from the edit heuristic set to the associated query molecular specifications to generate derived molecular specifications corresponding to all possible combinations of the edit heuristics. The molecular mapping engine then applies a mapping algorithm to each of the derived molecular specifications to generate a mapped version of the derived molecular specification. Importantly, the molecular mapping engine and the catalog mapping engine implement the same mapping algorithm.

[0143] The search engine performs any number and / or type of search operation on each of the molecular directories based on the associated mapped directories and the mapped versions of the derived molecular specifications to generate an associated subset of matches for the derived molecular specifications. The subset of matches associated with a given molecular directory specifies, but is not limited to, the derived molecular specifications that match the available molecular specifications included in the molecular directories.

[0144] Based on a matching subset, the workflow engine generates a search results dataset that specifies, but is not limited to, any number of potential drug candidate specifications and an associated list of locations. Each of the potential drug candidate specifications is a distinct, derived molecular specification that matches at least one of the available molecular specifications. For each potential drug candidate specification, the associated list of locations specifies at least one of the molecular directories that include the potential drug candidate specification. The workflow engine then provides any portion of the search results dataset to any number of users. For example (via GUI) and / or transfer any part of the search results dataset to any number of other software applications.

[0145] At least one technical advantage of the disclosed technique over existing techniques is that the molecular exploration application can be used to more efficiently identify additional drug development candidates during the drug discovery process. In particular, because the editing heuristic is tailored to the drug discovery process, the likelihood that each of the derived molecular specifications represents a molecule with the target biological activity is increased. Therefore, the proportion of potential drug candidates ultimately identified as additional drug development candidates is generally increased compared to existing techniques. Furthermore, unlike existing techniques, the computational complexity of the operations performed by the search engine remains constant regardless of the total number of molecules searched, thus reducing the time and computational resources required to comprehensively search the molecular catalog. Notably, using the disclosed technique, all available molecules can be searched at a given interaction rate. These technical advantages provide one or more technical improvements over existing methods.

[0146] 1. In some embodiments, a computer-implemented method for identifying one or more potential drug candidates during a drug discovery process includes: generating a plurality of derived molecular specifications based on query molecular specifications and a plurality of editing heuristics; performing one or more mapping operations on the plurality of derived molecular specifications using a mapping algorithm to generate a plurality of mapped molecular specifications; and performing one or more search operations on a mapped molecular catalog based on the plurality of mapped molecular specifications to identify the one or more potential drug candidates.

[0147] 2. The computer-implemented method as described in Clause 1, further comprising performing one or more mapping operations on a plurality of molecular specifications associated with a molecular catalog using the mapping algorithm to generate the mapped molecular catalog.

[0148] 3. The computer-implemented method as described in Clause 1 or 2, wherein generating the plurality of derived molecular specifications comprises recursively applying the plurality of edit heuristics to the query molecular specifications to generate a derivation tree comprising the plurality of derived molecular specifications.

[0149] 4. A computer-implemented method as described in any one of Clauses 1-3, wherein generating the plurality of derived molecular specifications comprises: applying a first editing heuristic, included in the plurality of editing heuristics, to the query molecular specification to generate a first derived molecular specification; and applying a second editing heuristic, included in the plurality of editing heuristics, to the first derived molecular specification to generate a second derived molecular specification.

[0150] 5. A computer-implemented method as described in any one of Clauses 1-4, wherein the plurality of editing heuristics includes at least one editing heuristic, which, when applied to the query molecular specification, adds a nitrogen or substituent to the query molecular specification, removes a nitrogen or substituent from the query molecular specification, or repositions a substituent included in the query molecular specification to generate the derived molecular specification.

[0151] 6. A computer-implemented method as described in any one of Clauses 1-5, wherein the plurality of editing heuristics includes at least one editing heuristic, which, when applied to the query molecular specification, repositions nitrogen included in the query molecular specification to generate a derived molecular specification representing an isomer of a query molecule corresponding to the query molecular specification.

[0152] 7. A computer-implemented method as described in any one of Clauses 1-6, wherein the mapping algorithm includes a hash function and the mapped molecular directory includes a hash mapping.

[0153] 8. A computer-implemented method as described in any one of Clauses 1-7, wherein performing the one or more search operations comprises performing a hash-based search on the mapped molecular catalog based on a first mapped molecular specification included in the plurality of mapped molecular specifications to determine that the first derived molecular specification included in the plurality of derived molecular specifications matches a first molecular specification included in the mapped molecular catalog.

[0154] 9. A computer-implemented method as described in any one of claims 1-8, further comprising: performing another hash-based search on another mapped molecular directory to determine that the first derived molecular specification matches a second molecular specification included in the other mapped molecular directory; determining that the first molecular directory corresponding to the mapped molecular directory has a first preference ranking lower than that of the second molecular directory corresponding to the other mapped molecular directory; and displaying on a computing device via a graphical user interface that the first derived molecule corresponding to the first derived molecular specification is a first potential drug candidate and is located in the second molecular directory.

[0155] 10. A computer-implemented method as described in any one of Clauses 1-9, wherein the query molecular specification represents a drug development candidate of interest associated with the drug discovery process.

[0156] 11. In some embodiments, one or more non-transitory computer-readable media include instructions that, when executed by one or more processors, cause the one or more processors to determine one or more potential drug candidates during a drug discovery process by performing the following steps: generating a plurality of derived molecular specifications based on query molecular specifications and a plurality of editing heuristics; performing one or more mapping operations on the plurality of derived molecular specifications using a mapping algorithm to generate a plurality of mapped molecular specifications; and searching a molecular catalog for each derived molecular specification included in the plurality of derived molecular specifications based on the plurality of mapped molecular specifications and a mapped molecular catalog to determine the one or more potential drug candidates.

[0157] 12. One or more non-transitory computer-readable media as described in Clause 11, further comprising performing one or more mapping operations on a plurality of molecular specifications associated with a molecular catalog using the mapping algorithm to generate the mapped molecular catalog.

[0158] 13. One or more non-transitory computer-readable media as described in clause 11 or 12, wherein generating the plurality of derived molecular specifications comprises recursively applying the plurality of edit heuristics to the query molecular specifications to generate a derivation tree including the plurality of derived molecular specifications.

[0159] 14. One or more non-transitory computer-readable media as described in any one of clauses 11-13, wherein generating a plurality of derived molecular specifications comprises: applying a first editing heuristic, included in the plurality of editing heuristics, to the query molecular specification to generate a first derived molecular specification; and applying a second editing heuristic, included in the plurality of editing heuristics, to the query molecular specification to generate a second derived molecular specification.

[0160] 15. One or more non-transitory computer-readable media as described in any one of Clauses 11-14, wherein the plurality of editing heuristics includes at least one editing heuristic that, when applied to the query molecular specification, adds a nitrogen or substituent to the query molecular specification, removes a nitrogen or substituent from the query molecular specification, or repositions a substituent included in the query molecular specification to generate the derived molecular specification.

[0161] 16. One or more non-transitory computer-readable media as described in any one of Clauses 11-15, wherein the plurality of editing heuristics includes at least one editing heuristic that, when applied to the query molecular specification, replaces oxygen in the query molecular specification with sulfur or replaces sulfur in the query molecular specification with oxygen.

[0162] 17. One or more non-transitory computer-readable media as described in any one of clauses 11-16, wherein the one or more mapping operations map the plurality of derived molecular specifications to a vector space to generate the plurality of mapped molecular specifications.

[0163] 18. One or more non-transitory computer-readable media as described in any one of clauses 11-17, wherein searching the molecular catalog comprises performing a hash-based search on the mapped molecular catalog based on a first mapped molecular specification included in the plurality of mapped molecular specifications to determine that the first derived molecular specification included in the plurality of derived molecular specifications matches a first molecular specification included in the mapped molecular catalog.

[0164] 19. One or more non-transitory computer-readable media as described in any one of Clauses 11-18, wherein the query molecular specification represents a drug development candidate of interest associated with the drug discovery process.

[0165] 20. In some embodiments, a system includes: one or more memories storing instructions and one or more processors coupled to the one or more memories, the one or more processors performing the following steps when executing the instructions: generating a plurality of derived molecular specifications based on query molecular specifications and a plurality of editing heuristics; applying a mapping algorithm to each of the plurality of derived molecular specifications to generate a plurality of mapped molecular specifications; and performing one or more search operations on a mapped molecular catalog based on the plurality of mapped molecular specifications to determine one or more potential drug candidates.

[0166] Any and all combinations of any claim element set forth in any claim and / or any element described in this application, in any manner, fall within the scope of the implementation and protection contemplated.

[0167] Various embodiments have been described for illustrative purposes; however, these descriptions are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

[0168] Various aspects of the embodiments of this invention may be embodied as systems, methods, or computer program products. Therefore, aspects of this disclosure may take the form of entirely hardware implementations, entirely software implementations (including firmware, resident software, microcode, etc.), or implementations combining software and hardware aspects, all of which are generally referred to herein as “modules,” “systems,” or “computers.” Furthermore, any hardware and / or software techniques, processes, functions, components, engines, modules, or systems described in this disclosure may be implemented as circuits or groups of circuits.

[0169] As previously described herein, aspects of this disclosure may take the form of a computer program product embodied in one or more computer-readable media having a computer-readable program codec embodied thereon. Any combination of one or more computer-readable media may be used. Each computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination of the foregoing. More specific examples (not an exhaustive list) of computer-readable storage media will include electrical connections having one or more wires, portable computer floppy disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory or flash memory, optical fiber, portable optical disc read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, device, or apparatus.

[0170] The foregoing description of aspects of this disclosure is based on flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks of the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine. When executed by a processor of a computer or other programmable data processing apparatus, these instructions enable the apparatus to perform the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such processors can be, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.

[0171] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code comprising one or more executable instructions for implementing a specified logical function. It should also be noted that in some embodiments, the functions indicated in the blocks may not occur in the order shown in the drawings. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a system based on dedicated hardware or a combination of dedicated hardware and computer instructions that performs the specified function or action.

[0172] While the foregoing relates to embodiments of this disclosure, other and additional embodiments of this disclosure may be contemplated without departing from the essential scope of this disclosure, the scope of which is defined by the appended claims.

Claims

1. A computer-implemented method for identifying one or more potential drug candidates during a drug discovery process, the method comprising: For each editing heuristic included in a plurality of editing heuristics, determine the likelihood that the target biological activity associated with the query molecular specification will be preserved when the editing heuristic is applied to the query molecular specification; Generate a subset of the plurality of editing heuristics that increase the likelihood of preserving the target biological activity; The subset of the plurality of editing heuristics is applied to the query molecular specification to generate a plurality of derived molecular specifications based on the query molecular specification and the plurality of editing heuristics; One or more mapping operations are performed on the multiple derived molecular specifications using a mapping algorithm to generate multiple mapped molecular specifications. as well as One or more search operations are performed on the mapped molecular catalog based on the multiple mapped molecular specifications to determine the one or more potential drug candidates.

2. The computer-implemented method of claim 1, further comprising performing one or more mapping operations on a plurality of molecular specifications associated with a molecular catalog using the mapping algorithm to generate the mapped molecular catalog.

3. The computer-implemented method of claim 1, wherein generating the plurality of derived molecular specifications comprises recursively applying the plurality of edit heuristics to the query molecular specifications to generate a derivation tree comprising the plurality of derived molecular specifications.

4. The computer-implemented method of claim 1, wherein generating the plurality of derived molecular specifications comprises: The first editing heuristic, included in the plurality of editing heuristics, is applied to the query molecular specification to generate the first derived molecular specification; as well as The second editing heuristic, included in the plurality of editing heuristics, is applied to the first derived molecular specification to generate the second derived molecular specification.

5. The computer-implemented method of claim 1, wherein the plurality of editing heuristics includes at least one editing heuristic, which, when applied to the query molecular specification, adds a nitrogen or substituent to the query molecular specification, removes a nitrogen or substituent from the query molecular specification, or repositions a substituent included in the query molecular specification to generate the derived molecular specification.

6. The computer-implemented method of claim 1, wherein the plurality of editing heuristics includes at least one editing heuristic, which, when applied to the query molecular specification, repositions nitrogen included in the query molecular specification to generate a derived molecular specification representing an isomer of a query molecule corresponding to the query molecular specification.

7. The computer-implemented method of claim 1, wherein the mapping algorithm includes a hash function, and the mapped molecular directory includes a hash mapping.

8. The computer-implemented method of claim 1, wherein performing the one or more search operations comprises performing a hash-based search on the mapped molecular catalog based on a first mapped molecular specification included in the plurality of mapped molecular specifications to determine that the first derived molecular specification included in the plurality of derived molecular specifications matches a first molecular specification included in the mapped molecular catalog.

9. The computer-implemented method of claim 8, further comprising: Perform another hash-based search on another mapped molecular catalog to determine if the first derived molecular specification matches a second molecular specification included in the other mapped molecular catalog; A first preference ranking is determined to be lower than that of a second molecular directory corresponding to the other mapped molecular directory; as well as The first derived molecule, corresponding to the first derived molecular specification, is displayed on a computing device via a graphical user interface. It is a first potential drug candidate and is located in the second molecule catalog.

10. The computer-implemented method of claim 1, wherein the query molecular specification represents a drug development candidate of interest associated with the drug discovery process.

11. One or more non-transitory computer-readable media, comprising instructions that, when executed by one or more processors, cause the one or more processors to identify one or more potential drug candidates during a drug discovery process by performing the following steps: For each editing heuristic included in a plurality of editing heuristics, determine the likelihood that the target biological activity associated with the query molecular specification will be preserved when the editing heuristic is applied to the query molecular specification; Generate a subset of the plurality of editing heuristics that increase the likelihood of preserving the target biological activity; The subset of the plurality of editing heuristics is applied to the query molecular specification to generate a plurality of derived molecular specifications based on the query molecular specification and the plurality of editing heuristics; One or more mapping operations are performed on the multiple derived molecular specifications using a mapping algorithm to generate multiple mapped molecular specifications. as well as Based on the plurality of mapped molecular specifications and the mapped molecular catalog, a search is performed in the molecular catalog for each of the plurality of derived molecular specifications included in the plurality of derived molecular specifications to identify the one or more potential drug candidates.

12. The one or more non-transitory computer-readable media of claim 11, further comprising performing one or more mapping operations on a plurality of molecular specifications associated with a molecular catalog using the mapping algorithm to generate the mapped molecular catalog.

13. One or more non-transitory computer-readable media as claimed in claim 11, wherein generating the plurality of derived molecular specifications comprises recursively applying the plurality of editing heuristics to the query molecular specifications to generate a derivation tree including the plurality of derived molecular specifications.

14. The one or more non-transitory computer-readable media of claim 11, wherein generating the plurality of derived molecular specifications comprises: The first editing heuristic, included in the plurality of editing heuristics, is applied to the query molecular specification to generate the first derived molecular specification; as well as The second editing heuristic, included in the plurality of editing heuristics, is applied to the query molecular specification to generate the second derived molecular specification.

15. One or more non-transitory computer-readable media as claimed in claim 11, wherein the plurality of editing heuristics includes at least one editing heuristic that, when applied to the query molecular specification, adds a nitrogen or substituent to the query molecular specification, removes a nitrogen or substituent from the query molecular specification, or repositions a substituent included in the query molecular specification to generate the derived molecular specification.

16. The one or more non-transitory computer-readable media of claim 11, wherein the plurality of editing heuristics includes at least one editing heuristic that, when applied to the query molecular specification, replaces oxygen in the query molecular specification with sulfur or replaces sulfur in the query molecular specification with oxygen.

17. One or more non-transitory computer-readable media as claimed in claim 11, wherein the one or more mapping operations map the plurality of derived molecular specifications to a vector space to generate the plurality of mapped molecular specifications.

18. One or more non-transitory computer-readable media as claimed in claim 11, wherein searching the molecular catalog comprises performing a hash-based search on the mapped molecular catalog based on a first mapped molecular specification included in the plurality of mapped molecular specifications to determine that the first derived molecular specification included in the plurality of derived molecular specifications matches a first molecular specification included in the mapped molecular catalog.

19. One or more non-transitory computer-readable media as claimed in claim 11, wherein the query molecular specification represents a drug development candidate of interest associated with the drug discovery process.

20. A system for identifying one or more potential drug candidates during a drug discovery process, comprising: One or more memories that store instructions; and One or more processors coupled to the one or more memories, wherein the one or more processors perform the following steps when executing the instructions: For each editing heuristic included in a plurality of editing heuristics, determine the likelihood that the target biological activity associated with the query molecular specification will be preserved when the editing heuristic is applied to the query molecular specification; Generate a subset of the plurality of editing heuristics that increase the likelihood of preserving the target biological activity; The subset of the plurality of editing heuristics is applied to the query molecular specification to generate a plurality of derived molecular specifications based on the query molecular specification and the plurality of editing heuristics; The mapping algorithm is applied to each of the multiple derived molecular specifications to generate multiple mapped molecular specifications. as well as One or more search operations are performed on the mapped molecular catalog based on the multiple mapped molecular specifications to identify one or more potential drug candidates.