Systems and Methods for In-Silico Screening of Compounds

The method addresses the computational complexity of current in-silico drug discovery by iteratively updating a prediction model based on target results from a subset of subjects, enabling efficient evaluation of large compound databases and improving drug discovery efficiency.

JP7691417B2Active Publication Date: 2025-06-11ATOMWISE INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022519999
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-03
Filing Date
2020-09-30
Publication Date
2025-06-11
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

Current in-silico drug discovery methods are limited by their computational complexity, making them inefficient for evaluating large libraries of compounds and discovering drugs for new diseases.

Method used

A method is provided that involves obtaining a subject dataset, applying a target model to a subset of subjects to obtain target results, training a prediction model using these results, and iteratively updating the model to reduce the number of subjects based on predefined criteria, thereby reducing computational complexity.

Benefits of technology

This approach enables the efficient evaluation of large chemical compound databases, reducing the number of subjects while maintaining accuracy, thus overcoming the limitations of existing methods in drug discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691417000005
    Figure 0007691417000005
  • Figure 0007691417000006
    Figure 0007691417000006
  • Figure 0007691417000007
    Figure 0007691417000007
Patent Text Reader

Abstract

A system and method are provided for reducing the number of subjects in a subject dataset. A target model having a first computational complexity is applied to a subset of subjects from the subject dataset and the target subjects, thereby obtaining a subset of target outcomes. A predictive model having a second computational complexity is trained using the subset of subjects and the subset of target outcomes. The predictive model is applied to a plurality of subjects, thereby obtaining a plurality of predicted outcomes. A portion of the subjects is excluded from the plurality of subjects based at least in part on the plurality of predicted outcomes. The method determines whether one or more predefined reduction criteria are met. If the predefined reduction criteria are not met, an additional subset of subjects and target outcomes is obtained, and the method is repeated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 62 / 910,068, entitled "Systems and Methods for Screening Compounds In Silico", filed on October 3, 2019, which is hereby incorporated by reference in its entirety.

[0002] This specification generally relates to techniques for dataset reduction by using multiple computational models having different computational complexities.

Background Art

[0003] The need to diversify molecular scaffolds to increase the likelihood of drug discovery success is referred to as moving away from the "flatlands", i.e., relying on synthetic methods that build flat molecules. Another approach to exploring the untapped potential of the molecular universe is to find ways to reveal what is hidden. According to some estimates, there are at least 10 60 different drug - like molecules: the potential of the unknown billion. One approach to opening up this uncharted chemical space is to study very large virtual libraries, i.e., libraries of compounds that do not need to be synthesized but whose molecular properties can be inferred from their calculated molecular structures.

[0004] New insights can be generated from large amounts of data such as these virtual libraries using the application of classifiers such as deep learning neural networks. In fact, lead identification and optimization in drug discovery, support for patient recruitment in clinical trials, medical image analysis, biomarker identification, drug efficacy analysis, drug adherence evaluation, sequencing data analysis, virtual screening, molecular profiling, metabolic data analysis, electronic health record analysis and medical device data evaluation, off-target side effect prediction, toxicity prediction, efficacy optimization, drug repurposing, drug resistance prediction, personalized medicine, drug trial design, pesticide design, materials science and simulation are all examples of applications where the use of classifiers such as deep learning-based solutions is being explored. Specifically, in healthcare, the American Recovery and Reinvestment Act of 2009 and the Precision Medicine Initiative of 2015 have widely supported the value of healthcare data in medicine. Thanks to several such initiatives, the volume of healthcare big data is expected to increase approximately 50-fold by 2020, reaching 25,000 petabytes. See, for example, Roots Analysis, February 22, 2017, “Deep Learning in Drug Discovery and Diagnostics, 2017 - 2035,” available at rootsanalysis.com on the Internet.

[0005] With the progress of drug repurposing and preclinical research, the application of classifiers to drug discovery has significantly improved the drug discovery process, thus creating opportunities to improve patient outcomes across the healthcare system. See, for example, Rifaioglu et al., 2018, “Recent applications of deep learning and machine intelligence on in silico drug discovery: methods, tools and databases,” Briefings in Bioinform 1-35, and Lavecchia, 2015, “Machine-learning approaches in drug discovery: methods and applications,” Drug Discovery Today 20(3), 318-331. Methods of in silico drug discovery are a particularly valuable use of classifiers because of their potential to reduce the time and cost of pharmaceutical development. Currently, the average cost of developing a new drug for human use is estimated to far exceed $2 billion. See, for example, DiMasi et al., 2016, J Health Econ 47, 20-33. In addition, the United States federal government spent over $100 billion, mostly through NIH funding, on the basic research that contributed to all 210 new drugs approved by the FDA between 2010 and 2016. See Cleary et al., 2018, “Contributions of NIH funding to new drug approvals 2010-2016,” PNAS 115(10), 2329-2334. Therefore, computational methods for discovering or at least screening for lead compounds (e.g., in databases of known and / or FDA-approved chemicals) have the potential to revolutionize drug discovery and pharmaceutical development.

[0006] There are many computational approaches that assist in drug discovery. The discovery of polypharmacology (e.g., the understanding that many drugs can bind to two or more molecular targets and actually do so) has opened up the field of repurposing already approved drugs for diseases lacking treatment. See, for example, Hopkins, 2009, “Predicting promiscuity,” Nature 462, 167-168 and Keiser et al., 2007, “Relating protein pharmacology by ligand chemistry,” Nat Biotechnol 25(2), 197-206. In silico drug discovery has already generated potential treatments for diseases ranging from Chagas disease to Zika virus. See, for example, Ramarack et al., 2017, “Zika virus NS5 protein potential inhibitors: an enhanced in silico approach in drug discovery,” J Biomol Structure and Dynamics 36(5), 1118-1133, Castillo-Garit et al., 2012, “Identification in silico and in vitro of Novel Trypanosomicidal Drug-Like Compounds,” Chem Biol and Drug Des 80, 38-45, and Raj et al. 2015 “Flavonoids as Multi-target Inhibitors for Proteins associated with Ebola Virus,” Interdisip Sci Comput Life Sci 7, 1-10. However, one drawback of many of the methods currently used for drug discovery, including the evaluation of virtual libraries, is their computational complexity.

[0007] In particular, many of the in-silico drug discovery methods are mainly applicable to pre-filtered and size-restricted molecular databases. See, for example, Macalino et al., 2018, “Evolution of in Silico Strategies for Protein-Protein Interaction Drug Discovery,” Molecules 23, 1963, and Lionata et al., 2014, “Structure-Based Virtual Screening for Drug Discovery: Principles, Applications and Recent Advances,” Curr Top Med Chem 14(16):1923-1938. In particular, the datasets are typically restricted to at least a few million compounds. See Ramsundar et al., 2015, “Massively Multitask Networks for Drug Discovery,” arXiv:1502.02072. The limitation on database size imposes a limitation on the ability to discover or screen for potential drugs to treat new diseases.

[0008] Considering the importance of identifying promising lead compounds, there is a need in the art for improved computational methods for drug discovery that enable the evaluation of large libraries of compounds.

Prior Art Documents

Non-Patent Documents

[0009]

Non-Patent Document 1

Non-Patent Document 2

Non - Patent Document 3

Non - Patent Document 4

Non - Patent Document 5

Non - Patent Document 6

Non - Patent Document 7

Non - Patent Document 8

Non - Patent Document 9

[0010] The present disclosure addresses the drawbacks identified in the background by providing a method for the evaluation of large-scale chemical compound databases.

[0011] In one aspect of the present disclosure, a method for reducing the number of subjects among a plurality of subjects in a subject dataset is provided. The method includes obtaining, in electronic form, a subject dataset.

[0012] The method further includes, for each respective subject in a subset of subjects from the plurality of subjects, applying a target model to each respective subject and at least one target subject to obtain a corresponding target result, thereby obtaining a corresponding subset of target results.

[0013] The method further includes training an initially trained state of a prediction model by using at least i) the subset of subjects as independent variables of the prediction model and ii) the corresponding subset of target results as dependent variables of the prediction model, thereby updating the prediction model to an updated trained state.

[0014] The method further includes applying the prediction model in the updated trained state to the plurality of subjects, thereby obtaining instances of a plurality of prediction results.

[0015] The method further excludes a portion of the subjects from the plurality of subjects, at least in part based on the instances of the plurality of prediction results.

[0016] The method further includes determining whether one or more predefined reduction criteria are met. If the one or more predefined reduction criteria are not met, the method further includes (i) for each respective subject of an additional subset of subjects from the plurality of subjects, applying a target model to each respective subject and at least one target subject to obtain corresponding target results, thereby obtaining an additional subset of target results. The additional subset of subjects is selected at least in part on instances of the plurality of prediction results. The method includes (ii) updating the subset of subjects by incorporating the additional subset of subjects into the subset of subjects, (iii) updating the subset of target results by incorporating the additional subset of target results into the subset of target results, and (iv) after updating (ii) and (iii), modifying the prediction model by applying the prediction model to at least 1) the subset of subjects as independent variables and 2) the corresponding subset of target results as corresponding dependent variables, thereby providing an updated trained state of the prediction model. The method then repeats the application of the updated trained state of the prediction model to the plurality of subjects, thereby obtaining instances of the plurality of prediction results. The method further excludes a portion of the subjects from the plurality of subjects based at least in part on the instances of the plurality of prediction results until the one or more predefined reduction criteria are met.

[0017] In some embodiments, the target model exhibits a first computational complexity when evaluating a subject, the prediction model exhibits a second computational complexity when evaluating a subject, and the second computational complexity is less than the first computational complexity. In some embodiments, the target model is at least 3 times, at least 5 times, or at least 100 times more computationally complex than the prediction model.

[0018] In some embodiments, the subject dataset includes a plurality of feature vectors (e.g., protein fingerprints, computational properties, and / or graph descriptors). In some embodiments, each feature vector is for each subject among the plurality of subjects, and the size of each feature vector among the plurality of feature vectors is the same. In some embodiments, each feature vector among the plurality of feature vectors is a one-dimensional vector.

[0019] In some embodiments, for each respective subject of a subset of subjects from the plurality of subjects, applying a target model to each respective subject and at least one target subject to obtain corresponding target results, thereby obtaining a corresponding subset of target results, further includes randomly selecting one or more subjects from the plurality of subjects to form a subset of subjects.

[0020] In some embodiments, for each respective subject of a subset of subjects from the plurality of subjects, applying a target model to each respective subject and at least one target subject to obtain corresponding target results, thereby obtaining a corresponding subset of target results, further includes selecting one or more subjects from the plurality of subjects of the subset of subjects based on the evaluation of one or more features selected from the plurality of feature vectors. In some embodiments, the selection is based on clustering (e.g., of the plurality of subjects).

[0021] In some embodiments, meeting one or more predefined reduction criteria includes comparing each prediction result among the plurality of prediction results with the corresponding target result from the subset of target results. In some embodiments, one or more predefined reduction criteria are met when the difference between the training result and the target result is below a predetermined threshold.

[0022] In some embodiments, meeting one or more predefined reduction criteria includes determining that the number of subjects among a plurality of subjects is below a threshold number of subjects.

[0023] In some embodiments, the target model is a convolutional neural network.

[0024] In some embodiments, the prediction model includes a random forest tree, a random forest including a plurality of multi-additive decision trees, a neural network, a graph neural network, a dense neural network, principal component analysis, nearest neighbor analysis, linear discriminant analysis, quadratic discriminant analysis, support vector machine, evolutionary method, projection pursuit, linear regression, naive Bayes algorithm, multi-categorical logistic regression algorithm, or an ensemble thereof.

[0025] In some embodiments, at least one target object is a single object, and the single object is a polymer. In some embodiments, the polymer includes an active site. In some embodiments, the polymer is an assembly of a protein, polypeptide, polynucleotide, polyribonucleic acid, polysaccharide, or any combination thereof.

[0026] In some embodiments, the plurality of test subjects includes at least 1 billion test subjects, at least 5 billion test subjects, at least 10 billion test subjects, at least 20 billion test subjects, at least 30 billion test subjects, at least 40 billion test subjects, at least 50 billion test subjects, at least 60 billion test subjects, at least 70 billion test subjects, at least 80 billion test subjects, at least 90 billion test subjects, at least 100 billion test subjects, at least 110 billion test subjects, at least 150 billion test subjects, at least 200 billion test subjects, at least 300 billion test subjects, at least 400 billion test subjects, at least 500 billion test subjects, at least 600 billion test subjects, at least 700 billion test subjects, at least 800 billion test subjects, at least 900 billion test subjects, at least 1000 billion test subjects, or at least 1100 billion test subjects before the application of an instance that excludes a portion of the test subjects from the plurality of test subjects.

[0027] In some embodiments, one or more predefined reduction criteria require that the plurality of test subjects have 30 or fewer test subjects, 40 or fewer test subjects, 50 or fewer test subjects, 60 or fewer test subjects, 70 or fewer test subjects, 90 or fewer test subjects, 100 or fewer test subjects, 200 or fewer test subjects, 300 or fewer test subjects, 400 or fewer test subjects, 500 or fewer test subjects, 600 or fewer test subjects, 700 or fewer test subjects, 800 or fewer test subjects, 900 or fewer test subjects, or 1000 or fewer test subjects (e.g., after one or more instances that exclude a portion of the test subjects from the plurality of test subjects).

[0028] In some embodiments, each test subject among the plurality of test subjects is a chemical compound.

[0029] In some embodiments, the initial trained state prediction model includes a classifier that is untrained or partially trained. In some embodiments, the updated trained state prediction model includes a classifier that is untrained or partially trained and is different from the initial trained state prediction model.

[0030] In some embodiments, a subset of subjects and / or an additional subset of subjects includes at least 1,000 subjects, at least 5,000 subjects, at least 10,000 subjects, at least 25,000 subjects, at least 50,000 subjects, at least 75,000 subjects, at least 100,000 subjects, at least 250,000 subjects, at least 500,000 subjects, at least 750,000 subjects, at least 1 million subjects, at least 2 million subjects, at least 3 million subjects, at least 4 million subjects, at least 5 million subjects, at least 6 million subjects, at least 7 million subjects, at least 8 million subjects, at least 9 million subjects, or at least 10 million subjects. In some embodiments, the additional subset of subjects is different from the subset of subjects.

[0031] In some embodiments, training the initial trained state prediction model by using at least i) a subset of subjects as a plurality of independent variables (of the prediction model) and ii) a corresponding subset of target results as a plurality of dependent variables (of the prediction model) further includes using at least one target object as an independent variable of the prediction model.

[0032] In some embodiments, the at least one target object includes at least two target objects, at least three target objects, at least four target objects, at least five target objects, or at least six target objects.

[0033] In some embodiments, modifying the prediction model by applying the prediction model (iv) after updating (ii) and updating (iii) further includes using at least 1) a subset of the subjects as independent variables and 2) the corresponding subset of the target results as the corresponding dependent variables, and in addition, 3) using at least one target as an independent variable.

[0034] In some embodiments, when one or more pre-defined reduction criteria are met, the method further includes clustering a plurality of subjects, thereby assigning each subject among the plurality of subjects to a cluster among the plurality of clusters, and excluding one or more subjects from the plurality of subjects based at least in part on the redundancy of the subjects in the individual clusters among the plurality of clusters.

[0035] In some embodiments, the method further includes selecting a subset of subjects from the plurality of subjects by clustering the plurality of subjects, thereby assigning each subject among the plurality of subjects to each of the clusters among the plurality of clusters, and selecting a subset of subjects from the plurality of subjects based at least in part on the redundancy of the subjects in the individual clusters among the plurality of clusters.

[0036] In some embodiments, when one or more pre-defined reduction criteria are met, the method further includes applying a plurality of subjects and at least one target to the prediction model, thereby causing the prediction model to provide respective prediction results for each subject among the plurality of subjects. In some embodiments, each respective prediction result corresponds to a prediction of the interaction between each subject and at least one target (e.g., IC 50 EC 50 , Kd, or KI). In some embodiments, each respective prediction score is used to characterize at least one target.

[0037] In some embodiments, excluding a portion of the subjects from the plurality of subjects, at least in part based on instances of a plurality of prediction results, includes: i) clustering the plurality of subjects, thereby assigning each subject among the plurality of subjects to a respective cluster among the plurality of clusters; and ii) excluding a subset of the subjects from the plurality of subjects, at least in part based on the redundancy of the subjects in the individual clusters among the plurality of clusters.

[0038] In some embodiments, the clustering of the plurality of subjects is performed using a density-based spatial clustering algorithm, a divisive clustering algorithm, an agglomerative clustering algorithm, a k-means clustering algorithm, a supervised clustering algorithm, or an ensemble thereof.

[0039] In some embodiments, excluding a portion of the subjects from the plurality of subjects, at least in part based on instances of a plurality of prediction results, includes: i) ranking the plurality of subjects based on the instances of the plurality of prediction results; and ii) removing from the plurality of subjects those subjects among the plurality of subjects that do not have a corresponding interaction score that meets a threshold cutoff.

[0040] In some embodiments, the threshold cutoff is an upper threshold percentage. In some embodiments, the upper threshold percentage is the upper 90 percent, upper 80 percent, upper 75 percent, upper 60 percent, or upper 50 percent of the plurality of prediction results.

[0041] In some embodiments, each instance of excluding a portion of the subjects from the plurality of subjects, at least in part based on instances of a plurality of prediction results, excludes from 1 / 10 to 9 / 10 of the subjects among the plurality of subjects. In some embodiments, each instance of excluding excludes from 1 / 4 to 3 / 4 of the subjects among the plurality of subjects.

[0042] Another aspect of the present disclosure provides a computing system including at least one processor and a memory storing at least one program executed by the at least one processor, where the at least one program includes instructions for reducing the number of subjects among a plurality of subjects in a subject dataset by any of the methods disclosed above.

[0043] Yet another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing at least one program for reducing the number of subjects among a plurality of subjects in a subject dataset. The at least one program is configured to be executed by a computer and includes instructions for performing any of the methods disclosed above.

[0044] As disclosed herein, any embodiment disclosed herein may be applied to any other aspect where applicable. Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, which shows and describes only exemplary embodiments of the present disclosure. As will be appreciated, the present disclosure is capable of other and different embodiments and some of the details thereof are capable of modification in various obvious respects without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.

[0045] Incorporation by Reference All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In case of conflict between the terms herein and those of the incorporated references, the terms herein shall control. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The implementations disclosed herein are illustrated by way of example in the accompanying drawings and not by way of limitation. The description and drawings are for illustrative purposes only and are not intended to define the limitations of the systems and methods of the present disclosure. Like reference numerals refer to corresponding parts throughout the drawings.

[0047]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

[0048] The computational effort required for drug discovery has been increasing in tandem with the growth in the size and complexity of drug datasets. In particular, very accurate models of target molecules have enabled the detection of additional test compounds (e.g., potential lead compounds) that may not have been considered using traditional drug discovery methods. The use of computational compound discovery scrutinizes the search space of potential drug databases (e.g., by determining which test compounds are most likely to have the desired effects given a particular target molecule) and simplifies the very laborious and time-consuming downstream processes of conducting clinical trials to validate good test compounds.

[0049] Here, embodiments are referred to in detail, and examples of these embodiments are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks are not described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0050] The implementation modes described herein provide various technical solutions for training a reference model for determining a target tumor fraction.

[0051] Definition. As used herein, the term "clustering" refers to various methods for optimizing the grouping of one or more sets of data points (e.g., clusters), where each data point in each set contains a higher degree of similarity to every other data point in that set than to data points not in that set. There are a wide variety of clustering algorithms suitable for evaluating different types of data. These algorithms include hierarchical models, centroid models, distribution models, density-based models, subspace models, graph-based models, and neural models. Each of these different models has its own separate computational requirements (e.g., complexity) and is suitable for different data types. Applying two separate clustering models to the same data set often results in two different groupings of the data. In some embodiments, repeated application of a clustering model to a data set results in a different grouping of the data each time.

[0052] As used herein, the term "feature vector" or "vector" is an enumerated list of elements, such as an array of elements, where each element has an assigned meaning. Thus, the term "feature vector" as used in this disclosure is interchangeable with the term "tensor". For ease of presentation, in some cases, a vector may be described as being one-dimensional. However, this disclosure is not so limited. A feature vector of any dimension may be used in this disclosure, provided that a description of what each element of the vector represents is defined.

[0053] As used herein, the term "polypeptide" means two or more amino acids or residues linked by peptide bonds. The terms "polypeptide" and "protein" are used interchangeably herein and include oligopeptides and peptides. "Amino acid", "residue", or "peptide" refers to any of the 20 standard structural units of proteins known in the art and includes amino acids such as proline and hydroxyproline. The designation of amino acid isomers may include D, L, R, and S. The definition of amino acids includes non-natural amino acids. Thus, selenocysteine, pyrroline, lanthionine, 2-aminoisobutyric acid, γ-aminobutyric acid, dehydroalanine, ornithine, citrulline, and homocysteine are all considered amino acids. Other variants or analogs of amino acids are known in the art. Thus, polypeptides can include synthetic peptide mimetic structures such as peptides. See Simon et al., 1992, Proceedings of the National Academy of Sciences USA, 89, 9367, which is hereby incorporated by reference in its entirety. See also Chin et al., 2003, Science 301, 964, and Chin et al., 2003, Chemistry & Biology 10, 511, each of which is hereby incorporated by reference in its entirety.

[0054] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the invention. As used in the detailed description of the invention and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. As used herein, the term "and / or" refers to any and all possible combinations of one or more of the associated listed items and is to be construed as inclusive. As used herein, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Further, as long as the terms "including", "includes", "having", "has", "with" or variations thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in the same manner as the term "comprising".

[0055] Some aspects are described below with reference to exemplary applications for illustration. It should be understood that numerous specific details, relationships, and methods are set forth in order to provide a complete understanding of the features described herein. However, one of ordinary skill in the art will readily recognize that the features described herein may be practiced without one or more of the specific details or in other ways. The features described herein are not limited by the order of acts or events recited as the acts or events may occur in different orders and / or concurrently with other acts or events. Further, not all acts or events illustrated are required to implement the methodology in accordance with the features described herein.

[0056] Exemplary system embodiments Here, the details of an exemplary system are described in conjunction with FIG. 1. FIG. 1 is a block diagram illustrating a system 100 according to some implementations. The system 100 in some implementations includes at least one or more processing units CPU102 (also referred to as processors), one or more network interfaces 104, an optional user interface 108 (e.g., having a display 106, an input device 110, etc.), a memory 111, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 optionally include circuitry (sometimes referred to as a chipset) for interconnecting and controlling communications between system components.

[0057] In some embodiments, each processing unit in the one or more processing units 102 is a single-core processor or a multi-core processor. In some embodiments, the one or more processing units 102 are multi-core processors that enable parallel processing. In some embodiments, the one or more processing units 102 are multiple processors (single-core or multi-core) that enable parallel processing. In some embodiments, each of the one or more processing units 102 is configured to execute a series of machine-readable instructions that can be embodied in a program or software. The instructions can be stored in a memory location such as the memory 111. The instructions are directed to the one or more processing units 102 and can subsequently program or otherwise configure the one or more processing units 102 to implement the methods of the present disclosure. Examples of operations performed by the one or more processing units 102 can include fetch, decode, execute, and write-back. The one or more processing units 102 can be part of a circuit such as an integrated circuit. One or more other components of the system 100 can be included in the circuit. In some embodiments, the circuit is an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA) architecture.

[0058] In some embodiments, display 106 is a touch-sensing display, such as a touch-sensing surface. In some embodiments, user interface 106 includes one or more soft keyboard embodiments. In some implementations, the soft keyboard embodiments include standard (QWERTY) and / or non-standard configurations of symbols on the displayed icons. User interface 106 may be configured to provide the user, for example, with a graphic display of an interaction score, or a prediction result, as a result of reducing the number of subjects among a plurality of subjects in a subject dataset. The user interface may enable user interaction with a particular task (e.g., reviewing and adjusting predefined reduction criteria).

[0059] Memory 111 may be non-persistent memory, persistent memory, or any combination thereof. Non-persistent memory typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, etc., while persistent memory typically includes CD-ROM, digital versatile disk (DVD) or other optical storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage device devices, magnetic disk storage device devices, optical disk storage device devices, flash memory devices, or other non-volatile solid state storage device devices. Memory 111 optionally includes one or more storage device devices located remotely from CPU 102. Memory 111, and the non-volatile memory devices within memory 111, include non-transitory computer-readable storage media. In some embodiments, memory 111 includes at least one non-transitory computer-readable storage medium and stores computer-executable executable instructions that can be in the form of programs, modules, and data structures.

[0060] In some embodiments, as shown in FIG. 1, memory 111 stores the following programs, modules, and data structures, or subsets thereof: ● Instructions, programs, data, or information associated with an operating system 116 (e.g., an embedded operating system such as iOS, ANDROID, DARWIN, RTXC, LINUX, UNIX, OS X, WINDOWS, or VxWorks), including various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management), and facilitating communication between various hardware components and software components, ● Instructions, programs, data, or information associated with an optional network communication module (or instructions) 118 for connecting system 100 to other devices and / or to a communication network, ● At least one target 122, which in some embodiments includes at least one target 122 that includes a polymer, ● A subject database 122 including a plurality of subjects 124 (e.g., subjects 124-1,..., 124-X), from which a subset 130 of subjects (e.g., subjects 124-A,..., 124-B) is selected for analysis by a target model 150, and optionally one or more additional subsets of subjects (e.g., 140-1,..., 140-Y) are selected from the plurality of subjects 124 and then added to subset 130, each subject 124 in subset 130 having a corresponding target result 132 and a corresponding predicted result 134, ● A target model 150 having a first computational complexity 152, the application of the target model to a subset 130 of subjects resulting in respective target results 132 for each subject 124 in the subject subset 130, and ● A prediction model 160 having a second computational complexity 162, which applies the prediction model in either an initial untrained state 164 or an updated untrained state 166 to a subset of subjects 130 to obtain respective prediction results 136 for each subject 132 in the subset of subjects 130.

[0061] In various implementations, one or more of the elements identified above are stored in one or more of the aforementioned memory devices and correspond to a set of instructions for performing the functions described above. The identified modules, data, or programs (e.g., instruction sets) above need not be implemented as separate software programs, procedures, data sets, or modules, and thus, various subsets of these modules and data can be combined in various implementations or rearranged otherwise. In some implementations, the memory 111 optionally stores a subset of the modules and data structures identified above. Further, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the elements identified above are stored in a computer system other than the computer system of the system 100, are addressable by the system 100, and the system 100 can retrieve all or a portion of such data when needed.

[0062] FIG. 1 depicts "System 100", but the figure is intended more as a functional explanation of various features that may exist in a computer system than as a structural schematic of the implementations described herein. In fact, as will be recognized by those skilled in the art, the separately shown items can be combined and some items may be separate. Moreover, although FIG. 1 depicts specific data and modules (which may be non-persistent memory or persistent memory) of memory 111, it should be recognized that these data and modules, or portions thereof, may be stored in two or more memories. For example, in some embodiments, at least the first data set 122, the second data set 124, the reference module 120, and the reference model 140 may be stored in a remote storage device that may be part of a cloud-based infrastructure. In some embodiments, at least the first data set 122 and the second data set 124 are stored in a cloud-based infrastructure. In some embodiments, the reference module 120 and the reference model 140 may also be stored in a remote storage device.

[0063] A system for training a prediction model according to the present disclosure is disclosed with reference to FIG. 1, where a method for performing such training according to the present disclosure is detailed with reference to FIG. 2.

[0064] Block 202. Referring to block 202 of FIG. 2A, a method is provided for reducing the number of subjects among a plurality of subjects in a subject data set.

[0065] Blocks 204-206. Referring to block 204 in FIG. 2A, the method proceeds by obtaining a subject data set in electronic form. An example of such a subject data set is ZINC15. See Sterling and Irwin, 2005, J. Chem. Inf. Model 45(1), p. 177-182. Zinc 15 is a commercially available database of compounds for virtual screening. ZINC 15 contains over 230 million purchasable compounds in 3D format that can be docked immediately. ZINC 15 also contains over 750 million purchasable compounds. Other examples of subject data sets include, but are not limited to, MASSIV, AZ Space with Enamine BBs, EVOspace, PGVL, BICLAIM, Lilly, GDB-17, SAVI, CHIPMUNK, REAL‘Space’, SCUBIDOO 2.1, REAL‘Database’, WuXi Virtual, PubChem Compounds, Sigma Aldrich‘in-stock’, eMolecules Plus, and WuXi Chemistry Services, which are summarized in Hoffmann and Gastreich, 2019, “The next level in chemical space navigation: going far beyond enumerable compound libraries,” Drug Discovery Today 24(5), pp. 1148, which is incorporated herein by reference.

[0066] In some embodiments, the plurality of subjects includes at least 100 million subjects, at least 500 million subjects, at least 1 billion subjects, at least 2 billion subjects, at least 3 billion subjects, at least 4 billion subjects, at least 5 billion subjects, at least 6 billion subjects, at least 7 billion subjects, at least 8 billion subjects, at least 9 billion subjects, at least 10 billion subjects, at least 11 billion subjects, at least 15 billion subjects, at least 20 billion subjects, at least 30 billion subjects, at least 40 billion subjects, at least 50 billion subjects, at least 60 billion subjects, at least 70 billion subjects, at least 80 billion subjects, at least 90 billion subjects, at least 100 billion subjects, or at least 110 billion subjects (e.g., before application of an instance that excludes a portion of the subjects from the plurality of subjects as described below with respect to blocks 232 - 234). In some embodiments, the plurality of subjects includes from 100 million to 500 million subjects, from 100 million to 1 billion subjects, from 1 billion to 2 billion subjects, from 1 billion to 5 billion subjects, from 1 billion to 10 billion subjects, from 1 billion to 15 billion subjects, from 5 billion to 10 billion subjects, from 5 billion to 15 billion subjects, or from 10 billion to 15 billion subjects. In some embodiments, the plurality of subjects includes 10 6 、10 7 、10 8 、10 9 、10 10 、10 11 、10 12 、10 13 、10 14 、10 15 、10 16 、10 17 、10 18 、10 19 、10 20 、10 21 、10 22 、10 23 、10 24 、10 25 、10 26 、10 27 、1028 , 10 29 , 10 30 , 10 31 , 10 32 , 10 33 , 10 34 , 10 35 , 10 36 , 10 37 , 10 38 , 10 39 , 10 40 , 10 41 , 10 42 , 10 43 , 10 44 , 10 45 , 10 46 , 10 47 , 10 48 , 10 49 , 10 50 , 10 51 , 10 52 , 10 53 , 10 54 , 10 55 , 10 56 , 10 57 , 10 58 , 10 59 , or 10 60 and is about 10 such compounds.

[0067] In some embodiments, the size of the subject data set is at least 100 kilobytes, at least 1 megabyte, at least 2 megabytes, at least 3 megabytes, at least 4 megabytes, at least 10 megabytes, at least 20 megabytes, at least 100 megabytes, at least 1 gigabyte, at least 10 gigabytes, or at least 1 terabyte. In some embodiments, the subject data set is a collection of files or data sets (e.g., 2 or more, 3 or more, 4 or more, 100 or more, 1000 or more, or 1 million or more) having a total file size of at least 100 kilobytes, at least 1 megabyte, at least 2 megabytes, at least 3 megabytes, at least 4 megabytes, at least 10 megabytes, at least 20 megabytes, at least 100 megabytes, at least 1 gigabyte, at least 10 gigabytes, or at least 1 terabyte.

[0068] Regarding block 206, in some embodiments, each subject among a plurality of subjects represents a respective chemical compound. In some embodiments, each subject represents a chemical compound that satisfies the five Lipinski rules. In some embodiments, each subject is an organic compound that satisfies two or more rules, three or more rules, or all four of Lipinski's Rule of Five rules: (i) no more than 5 hydrogen bond donors (e.g., OH and NH groups), (ii) no more than 10 hydrogen bond acceptors (e.g., N and O), (iii) a molecular weight of less than 500 Daltons, and (iv) a LogP of less than 5. The "Rule of Five" is so named because three of the four criteria involve the number 5. See Lipinski, 1997, Adv. Drug Del. Rev. 23, 3, which is hereby incorporated by reference in its entirety. In some embodiments, each subject satisfies one or more criteria in addition to Lipinski's Rule of Five. For example, in some embodiments, each subject has no more than 5 aromatic rings, no more than 4 aromatic rings, no more than 3 aromatic rings, or no more than 2 aromatic rings. In some embodiments, each subject describes a chemical compound, and the description of the chemical compound includes the modeled atomic coordinates of the chemical compound. In some embodiments, each subject among a plurality of subjects represents a different chemical compound.

[0069] In some embodiments, each subject represents an organic compound having a molecular weight of less than 2000 Daltons, less than 4000 Daltons, less than 6000 Daltons, less than 8000 Daltons, less than 10000 Daltons, or less than 20000 Daltons.

[0070] In some embodiments, at least one subject among the plurality of subjects represents a corresponding pharmaceutical compound. In some embodiments, at least one subject among the plurality of subjects represents a corresponding bioactive compound. As used herein, the term "bioactive compound" refers to a compound that has a physiological effect on humans (e.g., through interaction with proteins). A subset of bioactive compounds can be developed into pharmaceuticals. See, for example, Gu et al. 2013 "Use of Natural Products as Chemical Library for Drug Discovery and Network Pharmacology" PLoS One 8(4), e62839. Bioactive compounds can be naturally occurring or synthetic. Various definitions of bioactivity have been proposed. See, for example, Lagunin et al. 2000 "PASS: Prediction of activity spectra for biologically active substances" Bioinform 16, 747-748.

[0071] In some embodiments, the subjects in the subject dataset represent chemical compounds having an "alkyl" group. The term "alkyl" means, by itself or as part of another substituent of a chemical compound, a straight-chain, branched-chain, or cyclic hydrocarbon radical, or combinations thereof, unless otherwise specified, which can be fully saturated, monounsaturated, or polyunsaturated, and can include divalent, trivalent, and polyvalent radicals having the specified number of carbon atoms (i.e., C 1 ~C 10(which means 1 to 10 carbons). Examples of saturated hydrocarbon radicals include, but are not limited to, methyl, ethyl, n-propyl, isopropyl, n-butyl, t-butyl, isobutyl, secondary butyl, cyclohexyl, (cyclohexyl)methyl, cyclopropylmethyl, and homologs and isomers such as, for example, n-pentyl, n-hexyl, n-heptyl, n-octyl, etc. An unsaturated alkyl group is a group having one or more double or triple bonds. Examples of unsaturated alkyl groups include, but are not limited to, vinyl, 2-propenyl, crotyl, 2-isopentenyl, 2-(butadienyl), 2,4-pentadienyl, 3-(1,4-pentadienyl), ethynyl, 1- and 3-propynyl, 3-butynyl, and higher homologs and isomers. The term "alkyl" also means, optionally, including those derivatives of alkyl, such as "heteroalkyl", which are defined in more detail below, unless otherwise specified. An alkyl group limited to hydrocarbon groups is referred to as "homoalkyl". Exemplary alkyl groups include monounsaturated C 9-10 , an oleoyl chain, or diunsaturated C 9-10,12-13 a linoleyl chain. The term "alkylene", by itself or as part of another substituent, means a divalent radical derived from an alkane, exemplified by -CH 2 CH 2 CH 2 CH 2 - and further includes groups as described below as "heteroalkylene". Typically, an alkyl (or alkylene) group has 1 to 24 carbon atoms, and in the present invention, those groups preferably have 10 or fewer carbon atoms. "Lower alkyl" or "lower alkylene" is a shorter chain alkyl or alkylene group generally having 8 or fewer carbon atoms.

[0072] In some embodiments, the subject in the subject dataset represents a chemical compound having "alkoxy", "alkylamino", and "alkylthio" groups. The terms "alkoxy", "alkylamino", and "alkylthio" (or thioalkoxy) are used in their conventional meanings and refer to an alkyl group bonded to the remainder of the molecule via an oxygen atom, an amino group, or a sulfur atom, respectively.

[0073] In some embodiments, the subject in the subject dataset represents a chemical compound having "aryloxy" and "heteroaryloxy" groups. The terms "aryloxy" and "heteroaryloxy" are used in their conventional meanings and refer to an aryl or heteroaryl group bonded to the remainder of the molecule via an oxygen atom.

[0074] In some embodiments, the subject in the subject dataset represents a chemical compound having a "heteroalkyl" group. The term "heteroalkyl", by itself or in combination with another term, unless otherwise stated, means a stable straight-chain, branched-chain, or cyclic hydrocarbon radical, or combinations thereof, consisting of the stated number of carbon atoms and at least one heteroatom selected from the group consisting of O, N, Si, and S, where the nitrogen and sulfur atoms may optionally be oxidized and the nitrogen heteroatom may optionally be quaternized. The heteroatoms O, N, S, and Si may be placed at any internal position of the heteroalkyl group or at the position where the alkyl group is bonded to the remainder of the molecule. Examples include, but are not limited to, -CH 2 -CH 2 -O-CH 3 、-CH 2 -CH 2 -NH-CH 3 、-CH 2 -CH 2 -N(CH 3 )-CH 3 、-CH 2 -S-CH 2 -CH 3 、-CH2 -CH 2 、 -S(O)-CH 3 、 -CH 2 -CH 2 -S(O) 2 -CH 3 、 -CH=CH-O-CH 3 、 -Si(CH 3 ) 3 、 -CH 2 -CH=N-OCH 3 、 and -CH=CH-N(CH 3 )-CH 3 are included. Up to two heteroatoms may be consecutive, for example -CH 2 -NH-OCH 3 and -CH 2 -O-Si(CH 3 ) 3 . Similarly, the term "heteroalkylene" is not limited in itself or as part of another substituent, but is exemplified by -CH 2 -CH 2 -S-CH 2 -CH 2 - and -CH 2 -S-CH 2 -CH 2 -NH-CH 2 and means a divalent radical derived from heteroalkyl. For a heteroalkylene group, the heteroatom(s) can also occupy either or both of the chain ends (e.g., alkyleneoxy, alkylenedioxy, alkyleneamino, alkylenediamino, etc.). Furthermore, for alkylene and heteroalkylene linking groups, the orientation of the linking group is not implied by the direction in which the formula of the linking group is written. For example, the formula -CO 2 R’- represents both -C(O)OR’ and -OC(O)R’.

[0075] In some embodiments, the subject in the subject data set represents a chemical compound having a "cycloalkyl" and a "heterocycloalkyl" group. The terms "cycloalkyl" and "heterocycloalkyl", by themselves or in combination with other terms, unless otherwise defined, each represent a cyclic version of "alkyl" and "heteroalkyl", respectively. In addition, for heterocycloalkyl, the heteroatom can occupy the position where the heterocycle is attached to the rest of the molecule. Examples of cycloalkyl include, but are not limited to, cyclopentyl, cyclohexyl, 1-cyclohexenyl, 3-cyclohexenyl, cycloheptyl, etc. Further exemplary cycloalkyl groups include steroids such as cholesterol and its derivatives. Examples of heterocycloalkyl include, but are not limited to, 1-(1,2,5,6-tetrahydropyridyl), 1-piperidinyl, 2-piperidinyl, 3-piperidinyl, 4-morpholinyl, 3-morpholinyl, tetrahydrofuran-2-yl, tetrahydrofuran-3-yl, tetrahydrothien-2-yl, tetrahydrothien-3-yl, 1-piperazinyl, 2-piperazinyl, etc.

[0076] In some embodiments, the subject in the subject data set represents a chemical compound having a "halo" or "halogen". The terms "halo" or "halogen", by themselves or as part of another substituent, unless otherwise defined, mean a fluorine, chlorine, bromine, or iodine atom. Further, terms such as "haloalkyl" mean including monohaloalkyl and polyhaloalkyl. For example, the term "halo(C 1 ~C 4 )alkyl" means including, but not limited to, trifluoromethyl, 2,2,2-trifluoroethyl, 4-chlorobutyl, 3-bromopropyl, etc.

[0077] In some embodiments, the subject in the subject dataset represents a chemical compound having an "aryl" group. The term "aryl" means a polyunsaturated aromatic substituent that can be a monocyclic or polycyclic (preferably 1 to 3 rings) that are either fused together or covalently bonded, unless otherwise defined.

[0078] In some embodiments, the subject in the subject dataset represents a chemical compound having a "heteroaryl" group. The term "heteroaryl" refers to an aryl substituent (or ring) containing 1 to 4 heteroatoms selected from N, O, S, Si, and B, where nitrogen and sulfur atoms are optionally oxidized and nitrogen atoms are optionally quaternized. Exemplary heteroaryl groups are six-membered azines such as pyridinyl, diazinyl, and triazinyl. The heteroaryl group can be attached to the rest of the molecule through a heteroatom. Non-limiting examples of aryl and heteroaryl groups include phenyl, 1-naphthyl, 2-naphthyl, 4-biphenyl, 1-pyrrolyl, 2-pyrrolyl, 3-pyrrolyl, 3-pyrazolyl, 2-imidazolyl, 4-imidazolyl, pyrazinyl, 2-oxazolyl, 4-oxazolyl, 2-phenyl-4-oxazolyl, 5-oxazolyl, 3-isoxazolyl, 4-isoxazolyl, 5-isoxazolyl, 2-thiazolyl, 4-thiazolyl, 5-thiazolyl, 2-furyl, 3-furyl, 2-thienyl, 3-thienyl, 2-pyridyl, 3-pyridyl, 4-pyrimidyl, 4-pyrimidyl, 5-benzothiazolyl, purinyl, 2-benzimidazolyl, 5-indyl, 1-isoquinolyl, 5-isoquinolyl, 2-quinoxynyl, 5-quinoxynyl, 3-quinolyl, and 6-quinolyl. Each substituent of the aryl and heteroaryl ring systems listed above is selected from the group of acceptable substituents described below.

[0079] Briefly, the term "aryl" when used in combination with other terms (e.g., aryloxy, arylthioxy, arylalkyl) includes aryl, heteroaryl, and heteroarene rings as defined above. Thus, the term "arylalkyl" means a radical in which an aryl group is attached to an alkyl group (e.g., benzyl, phenethyl, pyridylmethyl, etc.) that includes an alkyl group in which a carbon atom (e.g., a methylene group) is replaced by an oxygen atom (e.g., phenoxymethyl, 2-pyridyloxymethyl, 3-(1-naphthyloxy)propyl, etc.).

[0080] Each of the above terms (e.g., "alkyl", "heteroalkyl", "aryl", and "heteroaryl") is meant to optionally include both substituted and unsubstituted forms of the species shown. Exemplary substituents of these species are provided below.

[0081] The substituents of the alkyl and heteroalkyl radicals of the chemical compounds represented by the subject data set (often including groups such as those often called alkylene, alkenyl, heteroalkylene, heteroalkenyl, alkynyl, cycloalkyl, heterocycloalkyl, cycloalkenyl, and heterocycloalkenyl) are generally referred to as "alkyl group substituents", and they can be one or more of a variety of groups selected from, but not limited to: a number in the range of zero to (2m'+1) of H, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted heterocycloalkyl, -OR', =O, =NR', =N-OR', -NR'R'', SR', halogen, SiR'R''R''', OC(O)R', C(O)R', CO 2 R', CONR'R'', OC(O)NR'R'', NR''C(O)R', NR'C(O)NR''R''', NR''C(O)2R', NR C(NR'R''R''')=NR'''', NR C(NR'R'')=NR''', -S(O)R', -S(O) 2 R', -S(O) 2NR’R’’, NRSO2R’, -CN, and -NO 2 wherein m is the total number of carbon atoms in such a radical. Each of R’, R’’, R’’’, and R’’’’ is preferably, independently, hydrogen, substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, e.g., aryl substituted with 1 to 3 halogens, substituted or unsubstituted alkyl, alkoxy or thioalkoxy groups, or arylalkyl groups. When the compounds of the present invention contain two or more R groups, for example, when two or more of the R groups are present, each of the R groups is independently selected as each being an R’, R’’, R’’’, and R’’’’ group. When R’ and R’’ are attached to the same nitrogen atom, they can be combined with the nitrogen atom to form a five-, six-, or seven-membered ring. For example, -NR’R’’ means, without limitation, including 1-pyrrolidinyl and 4-morpholinyl. From the above discussion of substituents, one of ordinary skill in the art will understand that the term “alkyl” includes groups containing carbon atoms bonded to groups other than hydrogen groups such as haloalkyl (e.g., -CF 3 and -CH 2 CF 3 ) and acyl (e.g., -C(O)CH 3 , -C(O)CF 3 , -C(O)CH 2 OCH 3 etc.). These terms include groups that are considered exemplary “alkyl group substituents” which are components of the exemplary “substituted alkyl” and “substituted heteroalkyl” moieties.

[0082] Similar to the substituents described for alkyl radicals, the substituents of aryl, heteroaryl, and heteroarene groups are generally referred to as "aryl group substituents". The substituents are, for example, without limitation, a number in the range from zero to the total number of open valences on the aromatic ring system, of substituted or unsubstituted alkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted heterocycloalkyl, OR', =O, =NR', =N-OR', -NR'R'', -SR', -halogen, -SiR'R''R''', -OC(O)R', -C(O)R', -CO 2 R', -CONR'R'', -OC(O)NR'R'', -NR''C(O)R', -NR'-C(O)NR''R''', -NR''C(O) 2 R', -NR-C(NR'R''R''')=NR'''', -NR-C(NR'R'')=NR''', -S(O)R', -S(O) 2 R', -S(O) 2 NR'R'', -NRSO 2 R', -CN and -NO 2 , -R', -N 3 , -CH(Ph) 2 , fluoro(C 1 ~C 4 )alkoxy, and fluoro(C 1 ~C 4 )alkyl bonded to the heteroaryl or heteroarene nucleus via a carbon or heteroatom (e.g., P, N, O, S, Si, or B). Each of the groups named above is bonded directly to the heteroarene or heteroaryl nucleus, or via a heteroatom (e.g., P, N, O, S, Si, or B), where R', R", R''', and R'''' are preferably independently selected from hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, and substituted or unsubstituted heteroaryl. When the compounds of the present invention contain two or more R groups, for example, when two or more of these R groups are present, each of the R groups is independently selected as each being an R', R'', R''', and R'''' group.

[0083] Two of the substituents on adjacent atoms of the aryl ring, heteroarene ring or heteroaryl ring may optionally be replaced by a substituent of the formula -T-C(O)-(CRR’) q -U-, wherein T and U are independently -NR-, -O-, -CRR’- or a single bond, and q is an integer from 0 to 3. Alternatively, two of the substituents on adjacent atoms of the aryl or heteroaryl ring may optionally be replaced by a substituent of the formula -A-(CH 2 ) r -B-, wherein A and B are independently -CRR’-, -O-, -NR-, -S-, -S(O)-, -S(O) 2 -, -S(O) 2 NR’-, or a single bond, and r is an integer from 1 to 4. One of the single bonds of the newly formed ring may optionally be replaced by a double bond. Alternatively, two of the substituents on adjacent atoms of the aryl, heteroarene, or heteroaryl ring may optionally be replaced by a substituent of the formula -(CRR’) s -X-(CR’’R’’’) d -, wherein s and d are independently integers from 0 to 3, and X is -O-, -NR’-, -S-, -S(O)-, -S(O) 2 ]-, or -S(O) 2 NR’-. The substituents R, R’, R’’, and R’’’ are preferably independently hydrogen or substituted or unsubstituted (C 1 ~C 6 )alkyl. These terms include groups that are considered exemplary “aryl group substituents” which are components of the exemplary “substituted aryl”, “substituted heteroarene” and “substituted heteroaryl” moieties.

[0084] In some embodiments, the subject in the subject data set represents a chemical compound having an "acyl" group. As used herein, the term "acyl" describes a carbonyl residue, a substituent containing C(O)R. Exemplary species of R include H, halogen, substituted or unsubstituted alkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, and substituted or unsubstituted heterocycloalkyl.

[0085] In some embodiments, the subject in the subject data set represents a chemical compound having a "condensed ring system". As used herein, the term "condensed ring system" means at least two rings, each ring having at least two atoms in common with another ring. A "condensed ring system" can include aromatic rings as well as non-aromatic rings. Examples of "condensed ring systems" are naphthalene, indole, quinoline, chromene, and the like.

[0086] As used herein, the term "heteroatom" includes oxygen (O), nitrogen (N), sulfur (S), and silicon (Si), boron (B), and phosphorus (P).

[0087] The symbol "R" is a general abbreviation representing a substituent selected from H, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, and substituted or unsubstituted heterocycloalkyl groups.

[0088] Block 208. Referring to block 208 of FIG. 2A, in some embodiments, the subject data set includes a plurality of feature vectors (e.g., each feature vector corresponds to an individual subject in the subject data set and includes one or more features). In some embodiments, each respective feature vector among the plurality of feature vectors includes a chemical fingerprint, a molecular fingerprint, one or more computed properties, and / or a graph descriptor of the respective chemical compound represented by the corresponding subject. Exemplary molecular fingerprints include, but are not limited to, Daylight fingerprint, BCI fingerprint, ECFP fingerprint, ECFC fingerprint, MDL fingerprint, APFP fingerprint, TTFP fingerprint, UNITY 2D fingerprint, and the like.

[0089] In some embodiments, some of the features in the vector include any combination of molecular properties of the corresponding subject, such as molecular weight, number of rotatable bonds, computed LogP (e.g., computed octanol-water partition coefficient or other method), number of hydrogen bond donors, number of hydrogen bond acceptors, number of chiral centers, number of chiral double bonds (E / Z isomers), polar and nonpolar desolvation energies (in kcal / mol units), net charge, and number of rigid fragments. In some embodiments, one or more subjects in the subject data set are annotated with a function or activity. In some such embodiments, the features in the vector include such function or activity.

[0090] In some embodiments, the subject dataset includes the chemical structure of each subject. For example, in some embodiments, the chemical structure is a SMILES string. In some embodiments, a canonical representation of the subject is calculated to represent the chemical structure of the subject (see, e.g., OpenEye's OEchem library, OpenEye.com on the Internet). In some embodiments, the initial 3D model is generated from the unambiguous isomer SMILES of the subject (using, for example, OpenEye's Omega program). In some embodiments, the relevant correctly protonated form of the subject at pH 5-9.5 is then created (using, for example, Schrodinger's ligprep program available from Schrodinger, Inc. at schrodinger.com on the Internet). This includes, for example, the deprotonation of carboxylic acids and tetrazoles, and the protonation of most aliphatic amines. In some embodiments, the partial atomic charges and atomic desolvation penalties of a single 3D conformation of each protonation state, stereoisomer, and tautomer are calculated (using, for example, the semi-empirical quantum mechanical program AMSOL16). In some embodiments, OpenEye's program Omega is used to generate 3D conformations. See, e.g., Sterling and Irwin, 2005, J. Chem. Inf. Model 45(1), p. 177-182. In some embodiments, the subjects in the subject dataset are represented by the subject dataset having at least in part a data structure in the form of SMILES, mol2, 3D SDF, DOCK flexibase, or equivalent.

[0091] In an embodiment of a subject dataset in which a subject is represented by a feature vector, each feature vector is for each subject among a plurality of subjects. In some embodiments, the size (e.g., the number of features) of each feature vector among the plurality of feature vectors is the same. In some embodiments, the size (e.g., the number of features) of each feature vector among the plurality of feature vectors is not the same. That is, in some embodiments, at least one of the feature vectors among the plurality of feature vectors is of a different size. In some embodiments, each feature vector is of any length (e.g., each feature vector can be of any size). In some embodiments, the number of dimensions of each feature vector among the plurality of feature vectors can vary (e.g., a feature vector can have any number of dimensions). In some embodiments, each feature vector among the plurality of feature vectors is a one-dimensional vector. In some embodiments, one or more of the feature vectors among the plurality of feature vectors are two-dimensional vectors. In some embodiments, one or more of the feature vectors among the plurality of feature vectors are three-dimensional vectors. In some embodiments, the number of dimensions of each feature vector among the plurality of feature vectors is the same (e.g., each feature vector has the same number of dimensions). In some embodiments, each feature vector among the plurality of feature vectors is at least a two-dimensional vector. In some embodiments, each feature vector among the plurality of feature vectors is at least an N-dimensional vector, where N is a positive integer greater than or equal to 2 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or greater).

[0092] In some embodiments, each respective subject among a plurality of subjects includes a corresponding chemical fingerprint of a chemical compound represented by each respective subject. In some embodiments, the chemical fingerprint of a subject is represented by a corresponding feature vector of the subject. As used herein, the term "chemical fingerprint" refers to a unique pattern (e.g., a unique vector or matrix) corresponding to a particular molecule. In some embodiments, each chemical fingerprint is of a fixed size. In some embodiments, one or more chemical fingerprints are variably sized. In some embodiments, the chemical fingerprint of each respective subject among a plurality of subjects can be determined directly (e.g., through a mass spectrometry method such as MALDI-TOF). In some embodiments, the chemical fingerprint of each respective subject among a plurality of subjects can be obtained via a computational method. See, for example, Daina et al. (2017) "SwissADME: a free web tool to evaluate pharmacokinetics, drug-likeness and medicinal chemistry friendliness of small molecules" Sci Reports 7, 42717, O'Boyle et al. 2011 "Open Babel: An open chemical toolbox" J Cheminforma 3, 33, Cereto-Massague et al. 2015 "Molecular fingerprint similarity search in virtual screening" Methods 71, 58-63, and Mitchell 2014 "Machine learning methods in cheminformatics" WIREs Comput Mol Sci. 4:468-481, each of which is hereby incorporated by reference herein.

[0093] Many different ways of representing chemical compounds in a computational space are known in the art.

[0094] In some embodiments, each chemical fingerprint includes information regarding the interaction between a respective chemical compound and one or more additional chemical compounds and / or biological macromolecules. In some embodiments, the chemical fingerprint includes information regarding protein-ligand binding affinity. See Wojcikowski et al. 2018 “Development of a protein-ligand extended connectivity (PLEC) fingerprint and its application for binding affinity predictions” Bioinformatics 35(8), 1334-1341, which is incorporated herein by reference. In some embodiments, a neural network is used to determine one or more chemical properties (and / or chemical fingerprints) of at least one subject in a subject database.

[0095] In some embodiments, each subject in a subject database corresponds to a known chemical compound having one or more known chemical properties. In some embodiments, the same number of chemical properties are provided for each subject within a plurality of subjects in a subject dataset. In some embodiments, different numbers of chemical properties are provided for one or more subjects in a subject dataset. In some embodiments, one or more subjects in a subject dataset are synthetic (e.g., the chemical structure of the subject can be determined even though the subject has not been analyzed in a laboratory). See, for example, Gomez-Bombarelli et al. 2017 “Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules” arXiv:1610.02415v3, which is incorporated herein by reference.

[0096] In some embodiments, graph comparison is used to compare the three-dimensional structures of molecules represented by a subject data set (e.g., to determine clusters or sets of similar molecules). The concept of graph comparison relies on comparing graph descriptors and yields measures of difference or similarity that can be used for pattern recognition. See, for example, Czech 2011 “Graph Descriptors form B -Matrix Representation” Graph-Based Representations in Patter Recognition, LNCS 6658, 12-21, which is hereby incorporated by reference. In some embodiments, measures such as clustering coefficient, efficiency, or betweenness centrality can be utilized to capture relevant structural properties within the graph (e.g., of a set of subjects). See, for example, Costa et al. 2007 “Characterization of complex networks: A survey of measurements” Advances Phys 56(1), 198-200, which is hereby incorporated by reference.

[0097] Block 210. Referring to block 210 of FIG. 2A, for each respective subject of a subset of subjects from a plurality of subjects, a target model is applied to each respective subject and at least one target subject to obtain corresponding target results, thereby obtaining a corresponding subset of target results. In an exemplary embodiment, each respective subject is docked to each target subject of at least one target subject. In some embodiments, only a single target subject exists.

[0098] In some embodiments, the target is a polymer. Examples of polymers include, but are not limited to, assemblies of proteins, polypeptides, polynucleic acids, polyribonucleic acids, polysaccharides, or any combination thereof. For example, polymers such as those studied using some embodiments of the disclosed systems and methods are large molecules composed of repeating residues. In some embodiments, the polymer is a natural material. In some embodiments, the polymer is a synthetic material. In some embodiments, the polymer is an elastomer, shellac, amber, natural or synthetic rubber, cellulose, bakelite, nylon, polystyrene, polyethylene, polypropylene, polyacrylonitrile, polyethylene glycol, or a polysaccharide.

[0099] In some embodiments, the target is a heteropolymer (copolymer). A copolymer is a polymer derived from two (or more) monomer species, as opposed to a homopolymer in which only one monomer is used. Copolymerization refers to the method used to chemically synthesize a copolymer. Examples of copolymers include, but are not limited to, ABS resin, SBR, nitrile rubber, styrene-acrylonitrile, styrene-isoprene-styrene (SIS), and ethylene-vinyl acetate. Since a copolymer consists of at least two types of constitutional units (structural units, or particles), the copolymer can be classified based on how these units are arranged along the chain. These include alternating copolymers having regular alternating A and B units. See, for example, Jenkins, 1996, “Glossary of Basic Terms in Polymer Science,” Pure Appl. Chem. 68(12):2287-2311, which is incorporated herein by reference in its entirety. Additional examples of copolymers are repeating sequences (e.g., (A-B-A-B-B-A-A-A-A-A-B-B-B) n) It is a periodic copolymer having A units and B units arranged therein. Additional examples of copolymers are statistical copolymers in which the sequence of monomer residues in the copolymer follows statistical rules. For example, see Painter, 1997, Fandamentals of Polymer Science, CRC Press, 1997, p14, which is hereby incorporated by reference in its entirety. Still other examples of copolymers that can be evaluated using the disclosed systems and methods are block copolymers that include two or more homopolymer subunits linked by covalent bonds. The bonding of the homopolymer subunits may require an intermediate non-repeating subunit known as a junction block. Block copolymers having two or three separate blocks are called diblock copolymers and triblock copolymers, respectively.

[0100] In some embodiments, the target is actually a plurality of polymers, and each polymer in the plurality of polymers does not all have the same molecular weight. In some such embodiments, the polymers in the plurality of polymers fall within a weight range having a corresponding distribution of chain lengths. In some embodiments, the polymer is a branched polymer molecule that includes a main chain having one or more substituent side chains or branches. Types of branched polymers include, but are not limited to, star polymers, comb polymers, brush polymers, dendritic polymers, ladders, and dendrimers. For example, see Rubinstein et al., 2003, Polymer physics, Oxford; New York: Oxford University Press.p.6, which is hereby incorporated by reference in its entirety.

[0101] In some embodiments, the target is a polypeptide. As used herein, the term "polypeptide" means two or more amino acids or residues linked by peptide bonds. The terms "polypeptide" and "protein" are used interchangeably herein and include oligopeptides and peptides. "Amino acid", "residue", or "peptide" refers to any of the 20 standard structural units of proteins known in the art and includes amino acids such as proline and hydroxyproline. The nomenclature of amino acid isomers can include D, L, R, and S. The definition of amino acids includes non-natural amino acids. Thus, selenocysteine, pyrroline, lanthionine, 2-aminoisobutyric acid, γ-aminobutyric acid, dehydroalanine, ornithine, citrulline, and homocysteine are all considered amino acids. Other variants or analogs of amino acids are known in the art. Thus, polypeptides can include synthetic peptide mimetic structures such as peptides. See Simon et al., 1992, Proceedings of the National Academy of Sciences USA, 89, 9367, which is incorporated herein by reference in its entirety. See also Chin et al., 2003, Science 301, 964, and Chin et al., 2003, Chemistry & Biology 10, 511, each of which is incorporated herein by reference in its entirety.

[0102] In some embodiments, the target being evaluated according to some embodiments of the disclosed systems and methods may also have any number of post-translational modifications. Thus, targets include polymers modified by acylation, alkylation, amidation, biotinylation, formylation, γ-carboxylation, glutamylation, glycosylation, glycylation, hydroxylation, iodination, isoprenylation, lipoylation, cofactor addition (such as heme, flavin, metal, etc.), addition of nucleosides and their derivatives, oxidation, reduction, pegylation, phosphatidylinositol addition, phosphopantetheinylation, phosphorylation, pyroglutamic acid formation, racemization, addition of amino acids by tRNA (such as arginylation), sulfation, selenoylation, ISGylation, SUMOylation, ubiquitination, chemical modification (such as citrullination and deamidation), and treatment by other enzymes (such as proteases, phosphotases, and kinases). Other types of post-translational modifications are known in the art and are also included.

[0103] In some embodiments, the target is an organometallic complex. An organometallic complex is a chemical compound containing a bond between carbon and a metal. In some cases, organometallic compounds are distinguished by the prefix "organo", for example, organopalladium compounds.

[0104] In some embodiments, the target is a surfactant. A surfactant is a compound that reduces the surface tension of a liquid, the interfacial tension between two liquids, or the interfacial tension between a liquid and a solid. Surfactants can act as detergents, wetting agents, emulsifiers, foaming agents, and dispersants. Surfactants are typically organic compounds that are amphiphilic, meaning that these organic compounds contain both a hydrophobic group (their tails) and a hydrophilic group (their heads). Thus, surfactant molecules contain both a water-insoluble (or oil-soluble) component and a water-soluble component. Surfactant molecules diffuse in water and adsorb at the interface between air and water or, when water is mixed with oil, at the interface between oil and water. The insoluble hydrophobic group can extend from the bulk aqueous phase into the air or into the oil phase while the water-soluble head group remains in the aqueous phase. This orientation of surfactant molecules at the surface modifies the surface properties of water at the water / air or water / oil interface.

[0105] Examples of ionic surfactants include ionic surfactants such as anionic surfactants, cationic surfactants, or zwitterionic (amphoteric) surfactants. In some embodiments, the target is an inverse micelle or a liposome.

[0106] In some embodiments, the target is a fullerene. A fullerene is any molecule composed entirely of carbon in the form of a hollow sphere, ellipsoid, or tube. Spherical fullerenes are also called buckyballs, and they resemble the balls used in association football. Cylindrical ones are called carbon nanotubes or buckytubes. Fullerenes have a structure similar to graphite, which is composed of stacked graphene sheets of connected hexagonal rings, but they can also contain pentagonal (or sometimes heptagonal) rings.

[0107] In some embodiments, the target is a polymer, and the spatial coordinates are a set of three-dimensional coordinates {x 1 ,...,x N} and (208), N is an integer of 2 or more (for example, 10 or more, 20 or more, etc.). In some embodiments, the target is a polymer, and the spatial coordinates are a set {x1,..., xN} of three-dimensional coordinates of the crystal structure of the polymer resolved with a resolution of 3.3 Å or more (210). In some embodiments, the target is a polymer, and the spatial coordinates are resolved with a resolution of 3.3 Å or more, 3.2 Å or more, 3.1 Å or more, 3.0 Å or more, 2.5 Å or more, 2.2 Å or more, 2.0 Å or more, 1.9 Å or more, 1.85 Å or more, 1.80 Å or more, 1.75 Å or more, or 1.70 Å or more (for example, by X-ray crystallographic techniques) and are a set of three-dimensional coordinates of the crystal structure of the polymer {x 1 ,..., x N}.

[0108] In some embodiments, the target is a polymer, and the spatial coordinates are an ensemble of 10 or more, 20 or more, 30 or more three-dimensional coordinates of the polymer determined by nuclear magnetic resonance, and the ensemble has a backbone RMSD of 1.0 Å or more, 0.9 Å or more, 0.8 Å or more, 0.7 Å or more, 0.6 Å or more, 0.5 Å or more, 0.4 Å or more, 0.3 Å or more, or 0.2 Å or more. In some embodiments, the spatial coordinates are determined by neutron diffraction or cryo-electron microscopy.

[0109] In some embodiments, the target includes two different types of polymers, such as a nucleic acid bound to a polypeptide. In some embodiments, the native polymer includes two polypeptides bound to each other. In some embodiments, the native polymer under study includes one or more metal ions (for example, a metalloprotease having one or more zinc atoms). In such cases, the metal ions and / or small organic molecules may be included in the spatial coordinates of the target.

[0110] In some embodiments, the target is a polymer, and the polymer has 10 or more, 20 or more, 30 or more, 50 or more, 100 or more, 100 to 1000, or less than 500 residues.

[0111] In some embodiments, the spatial coordinates of the target are determined using a modeling method such as an ab initio method, a density functional method, semi-empirical and empirical methods, molecular mechanics, chemical kinetics, or molecular dynamics.

[0112] In an embodiment, the spatial coordinates are represented by the Cartesian coordinates of the centers of the atoms including the target. In some alternative embodiments, the spatial coordinates of the target are represented by the electron density of the target, which is measured, for example, by X-ray crystallography. For example, in some embodiments, the spatial coordinates are 2F calculated using the calculated atomic coordinates of the target. observed -F calculated including an electron density map, where F observed is the observed structure factor amplitude of the target, and Fc is the structure factor amplitude calculated from the calculated atomic coordinates of the target.

[0113] Thus, the spatial coordinates of the target can be received as input data from a variety of sources including, but not limited to, a structural ensemble generated by solution NMR, X-ray crystallography, neutron diffraction, or a complex interpreted from cryo-electron microscopy, sampling from a computational simulation, homology modeling or rotamer library sampling, and combinations of these techniques.

[0114] In some embodiments, block 210 includes obtaining the spatial coordinates of the target. Further, block 210 includes modeling each subject with the target at each of a plurality of different poses, thereby creating a plurality of voxel maps, where each respective voxel map of the plurality of voxel maps includes the subject at each respective pose of the plurality of different poses.

[0115] In some embodiments, the target is a polymer having an active site, each test subject is a chemical compound, and modeling each test subject at the target in each of a plurality of different poses involves docking the test subject to the active site of the target. In some embodiments, each test subject is docked onto the target multiple times to form a plurality of poses (e.g., each docking represents a different pose). In some embodiments, the test subject is docked onto the target 2, 3, 4, 5 or more, 10 or more, 50 or more, 100 or more, or 1000 or more times. Each such docking represents a different pose of the respective test subject docked on the target. In some embodiments, each target is a polymer having an active site, and the test subject is docked to the active site in each of a plurality of different ways, with each such approach representing a different pose. It is envisioned that many of these poses are incorrect, meaning that such poses do not represent the true interaction between the respective test subject and the actual target that occurs. Without intending to be limited to any particular theory, it is envisioned that the object-to-object (e.g., intermolecular) interactions observed among the incorrect poses will cancel each other out like white noise, whereas the object-to-object interactions formed by the correct poses formed by the test subject will reinforce each other. In some embodiments, the test subject is docked either by a random pose generation technique or by biased pose generation. In some embodiments, the test subject is docked by Markov Chain Monte Carlo sampling. In some embodiments, such sampling enables full flexibility of the test subject in the docking calculation, a scoring function that is the sum of the interaction energy between the test subject and the target, and the conformational energy of the test subject.For example, reference is made to Liu and Wang, 1999, “MCDOCK: A Monte Carlo simulation approach to the molecular docking problem,” Journal of Computer - Aided Molecular Design 13, 435 - 451, which is incorporated herein by reference.

[0116] In some embodiments, algorithms such as DOCK (Shoichet, Bodian, and Kuntz, 1992, “Molecular docking using shape descriptors,” Journal of Computational Chemistry 13(3), pp. 380 - 397, and Knegtel, Kuntz, and Oshiro, 1997”Molecular docking to ensembles of protein structures,” Journal of Molecular Biology 266, pp. 424 - 440, each of which is incorporated herein by reference) are used to find multiple poses for each respective subject relative to each of the target subjects. Such algorithms model the target subject and the subject as rigid bodies. The docked conformations are explored using complementary surfaces to find the poses.

[0117] In some embodiments, algorithms such as AutoDOCK (Morris et al., 2009, “AutoDock4 and AutoDockTools4: Automated Docking with Selective Receptor Flexibility,” J. Comput. Chem. 30(16), pp. 2785-2791, Sotriffer et al., 2000, “Automated docking of ligands to antibodies: methods and applications,” Methods: A Companion to Methods in Enzymology 20, pp. 280-291, and Morris et al., 1998, “Automated Docking Using a Lamarckian Genetic Algorithm and Empirical Binding Free Energy Function,” Journal of Computational Chemistry 19: pp. 1639-1662, each of which is incorporated herein by reference) are used to find multiple poses for each respective test subject against each of the target subjects. AutoDOCK uses a kinematic model of the ligand and supports Monte Carlo, simulated annealing, Lamarckian genetic algorithm, and genetic algorithm. Thus, in some embodiments, multiple different poses (for a given test subject-target subject pair) are obtained by Markov chain Monte Carlo sampling, simulated annealing, Lamarckian genetic algorithm, or genetic algorithm using a docking scoring function.

[0118] In some embodiments, algorithms such as FlexX (Rarey et al., 1996, “A Fast Flexible Docking Method Using an Incremental Construction Algorithm,” Journal of Molecular Biology 261, pp. 470 - 489, which is incorporated herein by reference) are used to find multiple poses for each of a subset of test subjects with respect to each of the target subjects. FlexX uses a greedy algorithm to perform sequential construction of the test subject at the active site of the target subject. Thus, in some embodiments, multiple different poses (for a given test subject - target subject pair) are obtained by the greedy algorithm.

[0119] In some embodiments, algorithms such as GOLD (Jones et al., 1997, “Development and Validation of a Genetic Algorithm for flexible Docking,” Journal Molecular Biology 267, pp. 727 - 748, which is incorporated herein by reference) are used to find multiple poses for each of a subset of test subjects with respect to each of the target subjects. GOLD is short for Genetic Optimization for Ligand Docking. GOLD constructs a genetically optimized hydrogen - bonding network between the test subject and the target subject.

[0120] In some embodiments, modeling includes performing molecular dynamics runs of the target and the subject. During the molecular dynamics runs, the atoms of the target and the subject interact for a fixed period, enabling a view of the dynamic evolution of the system. The trajectories of the atoms of the target and the subject are determined by numerically solving Newton's equations of motion for a system of interacting particles, and the forces between the particles and their potential energy are calculated using an interatomic potential or a molecular mechanics force field. See Alder and Wainwright, 1959, “Studies in Molecular Dynamics. I. General Method,” J. Chem. Phys. 31(2):459, and Bibcode, 1959, J.Ch.Ph. 31, 459A, doi:10.1063 / 1.1730376, each of which is incorporated herein by reference. Thus, in this way, the molecular dynamics runs generate the trajectories of the target and the subject over time. These trajectories include the trajectories of the atoms of the target and the subject. In some embodiments, a subset of multiple different poses is obtained by taking snapshots of this trajectory over a period of time. In some embodiments, the poses are obtained from snapshots of several different trajectories, each trajectory including a different molecular dynamics run of the target interacting with the subject. In some embodiments, prior to the molecular dynamics run, the subject is first docked to the active site of the target using docking techniques.

[0121] Regardless of what modeling method is used, what is achieved for any given subject-target pair is a set of diverse poses of the subject with respect to the target, and one or more of the poses are assumed to be close enough to native poses to illustrate some of the relevant intermolecular interactions between the given subject / target pair.

[0122] In some embodiments, the initial pose of the subject at the active site of the target is generated using any of the techniques described above, and additional poses are generated through the application of some combination of rotation, translation, and mirroring operators in any combination of three X, Y, and Z planes. The rotation and translation of the subject may be randomly selected (within a range, e.g., plus or minus 5 Å from the origin) or uniformly generated in a pre-specified increment (e.g., in 5-degree increments around the circumference). FIG. 4 provides a sample illustration of the subject 122 at two different poses (402-1 and 402-2) at the active site of the target 124.

[0123] After generating each pose for each of the target and / or the subject, in some embodiments, a voxel map of each pose is created, thereby creating a plurality of voxel maps for each given target with respect to the target. In some embodiments, each respective voxel map among the plurality of voxel maps is created by a method comprising: (i) sampling the subject at each respective pose among the plurality of different poses and sampling the target on a three-dimensional grid basis, thereby forming a corresponding three-dimensional uniform space-filling honeycomb comprising a corresponding plurality of space-filling (three-dimensional) polyhedral cells; and (ii) filling the voxels (an individual set of regularly spaced polyhedral cells) of each respective voxel map based on the attributes (e.g., chemical attributes) of each respective three-dimensional polyhedral cell among the corresponding plurality of three-dimensional cells. Thus, in such embodiments, if a particular subject has 10 poses with respect to a target, 10 corresponding voxel maps are created, if a particular subject has 100 poses with respect to a target, 100 corresponding voxel maps are created, and so on. Examples of space-filling honeycombs include a cubic honeycomb with parallelogram cells, a hexagonal prism honeycomb with hexagonal prism cells, a rhombic dodecahedron with rhombic dodecahedron cells, an elongated dodecahedron with elongated dodecahedron cells, and a truncated octahedron with truncated octahedron cells.

[0124] In some embodiments, the space-filling honeycomb is a cubic honeycomb having cubic cells, and the dimensions of such voxels determine their resolution. For example, a resolution of 1 Å may be selected, which means that in such embodiments, each voxel represents a corresponding cube of geometric data having dimensions of 1 Å (e.g., 1 Å×1 Å×1 Å in the height, width, and depth of each cell). However, in some embodiments, a finer grid interval (e.g., 0.1 Å, or even 0.01 Å) or a coarser grid interval (e.g., 4 Å) is used, and this interval results in an integer number of voxels to cover the input geometric data. In some embodiments, the sampling is performed at a resolution that is from 0.1 Å to 10 Å. By way of illustration, for a 40 Å input cube, a resolution of 1 Å would result in such an arrangement having 40*40*40 = 64,000 input voxels.

[0125] In some embodiments, each subject is a first compound, the target is a second compound, and the properties of the atoms resulting from sampling (i) are assigned to a single voxel of each voxel map by filling (ii), and each voxel among the plurality of voxels represents the properties of at most one atom. In some embodiments, the properties of the atoms consist of a listing of the types of atoms. As an example, for biological data, some embodiments of the disclosed systems and methods are configured to represent the presence of any atom in a given voxel of the voxel map as a different number for its entry, e.g., if carbon is in the voxel, since the atomic number of carbon is 6, a value of 6 is assigned to that voxel. However, in such encoding, it may imply that atoms with close atomic numbers behave similarly, which may not be particularly useful depending on the application. Further, the behavior of elements may be more similar within a group (a column in the periodic table), and thus such encoding poses additional work for the convolutional neural network to decode.

[0126] In some embodiments, the properties of the atoms are encoded into the voxels as binary categorical variables. In such embodiments, the atom types are encoded into what is referred to as "one-hot" encoding: each atom type has a separate channel. Thus, in such embodiments, each voxel has a plurality of channels, and at least a subset of the plurality of channels represents the atom type. For example, one channel within each voxel may represent carbon, while another channel within each voxel may represent oxygen. When a given atom type is found in the three-dimensional grid element corresponding to a given voxel, the channel of that atom type within the given voxel is assigned a first value of a binary categorical variable, such as "1", and when the atom type is not found in the three-dimensional grid element corresponding to the given voxel, the channel of that atom type within the given voxel is assigned a second value of the binary categorical variable, such as "0", within the given voxel.

[0127] There are over 100 elements, but most are not encountered in biology. However, even for those representing the most common biological elements (e.g., H, C, N, O, F, P, S, Cl, Br, I, Li, Na, Mg, K, Ca, Mn, Fe, Co, Zn), 18 channels per voxel, or 10,483 * 18 = 188,694 inputs to the receptor field can result. Thus, in some embodiments, each respective voxel in the voxel map in the plurality of voxel maps includes a plurality of channels, and each channel in the plurality of channels represents a different attribute that can occur in the three-dimensional space-filling polyhedral cell corresponding to the respective voxel. The number of possible channels for a given voxel is even greater in those embodiments where additional properties of the atoms (e.g., partial charge, presence at a ligand-to-protein target, electronegativity, or SYBYL atom type) are additionally presented as independent channels for each voxel, and more input channels are required to distinguish other equivalent atoms.

[0128] In some embodiments, each voxel has five or more input channels. In some embodiments, each voxel has fifteen or more input channels. In some embodiments, each voxel has twenty or more input channels, twenty-five or more input channels, thirty or more input channels, fifty or more input channels, or one hundred or more input channels. In some embodiments, each voxel has five or more input channels selected from the descriptors found in Table 1 below. For example, in some embodiments, each voxel has five or more channels, and each channel is encoded as a binary categorical variable, where each channel represents a SYBYL atom type selected from Table 1 below. For example, in some embodiments, each voxel of the voxel map includes a channel of the C.3 (sp3 carbon) atom type, which means that when the grid in the space of a given subject-target complex represented by each voxel includes sp3 carbon, the channel adopts a first value (e.g., "1"), and a second value (e.g., "0") otherwise.

Table 1-1

Table 1-2

[0129] In some embodiments, each voxel includes ten or more input channels, fifteen or more input channels, or twenty or more input channels selected from the descriptors found in Table 1 above. In some embodiments, each voxel includes a channel for a halogen.

[0130] In some embodiments, a Structural Protein-Ligand Interaction Fingerprint (SPLIF) score is generated for each pose of each subject with respect to a target, and this SPLIF score is used as an additional input to a target model or individually encoded in a voxel map. For an explanation of SPLIF, see Da and Kireev, 2014, J. Chem. Inf. Model. 54, pp. 2555-2561, “Structural Protein-Ligand Interaction Fingerprints (SPLIF) for Structure-Based Virtual Screening: Method and Benchmark Study”, which is incorporated herein by reference. SPLIF implicitly encodes all possible interaction types (e.g., π-π, CH-π, etc.) that can occur between the interacting fragments of a subject and a target. In the first step, a subject-target complex (pose) is examined for intermolecular contacts. Two atoms are considered to be in contact if the distance between them is within a specified threshold (e.g., within 4.5 Å). For each such intermolecular atom pair, each subject atom and target atom are extended into circular fragments, e.g., fragments that include the atom in question and their successive neighborhoods up to a certain distance. An identifier is assigned to each type of circular fragment. In some embodiments, such identifiers are encoded in individual channels of each voxel. In some embodiments, an extended-connectivity fingerprint up to the first nearest neighbor (ECFP2), as defined in Pipeline Pilot software, can be used. See Pipeline Pilot, ver8.5, Accelrys Software Inc., 2009, which is incorporated herein by reference. ECFP holds information about all atom / bond types and uses a single unique integer identifier to represent one substructure (e.g., a cyclic fragment). The SPLIF fingerprint encodes all of the circular fragment identifiers found.In some embodiments, the SPLIF fingerprint functions as a separate independent input into the target model, rather than being encoded individual voxels.

[0131] In some embodiments, instead of or in addition to SPLIF, a Structural Interaction Fingerprint for Target (SIFt) is calculated for each pose of a given subject with respect to a target object and provided independently as an input into the target model or encoded into a voxel map. For the calculation of SIFt, see Deng et al., 2003, “Structural Interaction Fingerprint (SIFt): A Novel Method for Analyzing Three-Dimensional Protein-Ligand Binding Interactions,” J. Med. Chem. 47(2), pp. 337-344, which is incorporated herein by reference.

[0132] In some embodiments, instead of or in addition to SPLIF and SIFT, an Atom Pair-based Interaction Fragment (APIF) is calculated for each pose of a given subject with respect to a target object and provided independently as an input into the target model or individually encoded into a voxel map. For the calculation of APIF, see Perez-Nueno et al., 2009, “APIF: a new interaction fingerprint based on atom pairs and its application to virtual screening,” J. Chem. Inf. Model. 49(5), pp. 1245-1260, which is incorporated herein by reference.

[0133] Data representation may be encoded with biological data, for example, by means that enable expression of various structural relationships associated with molecules / proteins. Geometric representations may be implemented in various ways and topographies according to various embodiments. Geometric representations are used for visualization and analysis of data. For example, in an embodiment, geometric shapes may be represented using voxels laid out on various topographies such as 2D, 3D Cartesian / Euclidean space, 3D non-Euclidean space, manifolds, etc. For example, FIG. 5 illustrates a sample three-dimensional grid structure 500 including a series of sub-containers according to an embodiment. Each sub-container 502 may correspond to a voxel. A coordinate system may be defined for the grid such that each sub-container has an identifier. In some embodiments of the disclosed systems and methods, the coordinate system is a Cartesian system of 3D space, but in other embodiments of the system, the coordinate system may be any other type of coordinate system, such as, inter alia, an oblate spheroid, a cylindrical coordinate system or a spherical coordinate system, a polar coordinate system, other coordinate systems designed for various manifolds and vector spaces. In some embodiments, the voxels may have specific values associated with these voxels, which may be represented, for example, by, inter alia, applying a label and / or determining the orientation of these voxels.

[0134] In some embodiments, block 210 includes expanding each voxel map in a plurality of voxel maps into a corresponding vector, thereby creating a plurality of vectors, where each vector in the plurality of vectors is of the same size. In some embodiments, each respective vector in the plurality of vectors is input into a target model. In some embodiments, the target model includes (i) an input layer for sequentially receiving the plurality of vectors, (ii) a plurality of convolutional layers, and (iii) a scorer, the plurality of convolutional layers including an initial convolutional layer and a final convolutional layer, where each layer in the plurality of convolutional layers is associated with a different set of weights. In such embodiments, in response to the input of each respective vector in the plurality of vectors, the input layer supplies a first plurality of values to the initial convolutional layer as a first function of the values of the respective vector, and each respective convolutional layer other than the final convolutional layer supplies intermediate values to another convolutional layer in the plurality of convolutional layers as a function of (i) a different set of weights associated with the respective convolutional layer and (ii) a second function of the input values received by the respective convolutional layer, and the final convolutional layer supplies final values to the scorer as a function of (i) a different set of weights associated with the final convolutional layer and (ii) a third function of the input values received by the final convolutional layer. In this way, a plurality of scores are obtained from the scorer, where each score in the plurality of scores corresponds to the input of the vector in the plurality of vectors to the input layer. The plurality of scores are then used to provide corresponding target results for each respective subject. In some embodiments, the target result is a weighted average of the plurality of scores. In some embodiments, the target result is a measure of the central tendency of the plurality of scores. Examples of measures of central tendency include the arithmetic mean, weighted average, midrange, midhinge, trimean, Winsorized mean, median, or mode of the plurality of scores.

[0135] In some embodiments, the scorer includes a plurality of fully connected layers and an evaluation layer supplied by a fully connected layer among the plurality of fully connected layers to the evaluation layer. In some embodiments, the scorer includes a decision tree, a multiple additive regression tree, a clustering algorithm, principal component analysis, nearest neighbor analysis, linear discriminant analysis, quadratic discriminant analysis, a support vector machine, an evolutionary approach, projection pursuit, and an ensemble thereof. In some embodiments, each vector among the plurality of vectors is a one-dimensional vector. In some embodiments, the plurality of different poses includes two or more poses, ten or more poses, one hundred or more poses, or one thousand or more poses. In some embodiments, the plurality of different poses is obtained using a docking scoring function in one of a markup chain Monte Carlo sampling, simulated annealing, a Lamarckian genetic algorithm, or a genetic algorithm. In some embodiments, the plurality of different poses is obtained by sequential search using a greedy algorithm.

[0136] Blocks 212 and 214. In some embodiments, the target model has a higher computational complexity than the prediction model. In some such embodiments, applying the target model to all subjects in the subject dataset is computationally prohibited. For this reason, the target model is typically applied not to every subject in the subject dataset, but to a subset of the subjects. In some embodiments, some diversity in the subset of subjects (e.g., a subset of subjects including subjects having a certain range of structural or functional qualities) is desired. In some embodiments, the subset of subjects includes at least 1,000 subjects, at least 5,000 subjects, at least 10,000 subjects, at least 25,000 subjects, at least 50,000 subjects, at least 75,000 subjects, at least 100,000 subjects, at least 250,000 subjects, at least 500,000 subjects, at least 750,000 subjects, at least 1,000,000 subjects, at least 2,000,000 subjects, at least 3,000,000 subjects, at least 4,000,000 subjects, at least 5,000,000 subjects, at least 6,000,000 subjects, at least 7,000,000 subjects, at least 8,000,000 subjects, at least 9,000,000 subjects, or at least 10,000,000 subjects.

[0137] To ensure this, referring to block 212 of FIG. 2A, in some embodiments, the subset of subjects is selected from the subject dataset on a randomized basis (e.g., the subset of subjects is selected from the subject dataset using any random method known in the art).

[0138] Referring to block 214 of FIG. 2A, in other embodiments, a subset of the subjects is selected from the subject dataset based on an evaluation of one or more features of the subject's feature vector. In some such embodiments, the evaluation of the features includes selecting subjects from a plurality of subjects based on clustering (e.g., selecting subjects from a plurality of clusters when forming each subset of subjects). The subset of subjects is then selected based at least in part on the redundancy of the subjects in the individual clusters among the plurality of clusters (e.g., to obtain a subset of subjects representing different types of chemical compounds). For example, consider a case where the subjects in the subject dataset are clustered into 100 different clusters based on their feature vectors. One approach to selecting a subset of subjects is to select a fixed number of subjects (e.g., 10, 100, 1000, etc.) from each of the different clusters to form the subset of subjects. Within each cluster, the selection of subjects can be done in a random manner. Alternatively, within each cluster, the subjects that are closest to the center of each cluster are selected based on the fact that such subjects best represent the characteristics of their respective clusters. In some embodiments, the form of clustering used is unsupervised clustering. The benefit of clustering a plurality of subjects from the subject dataset is that this provides more accurate training of the prediction model. For example, if all or most of the subjects in the subset of subjects are similar chemical compounds (e.g., contain the same chemical group, have a similar structure, etc.), there is a risk that the prediction model is biased towards that particular type of chemical compound or is overfitting. This can, in some cases, have an adverse effect on downstream training (e.g., it may be difficult to efficiently retrain the prediction model to accurately analyze subjects from different types of chemical compounds).

[0139] To illustrate how the feature vectors of the subjects are used in clustering, consider the case where a set of 10 common features (the same 10 features) within each feature vector are used for clustering. In some embodiments, each subject in the subject dataset can have a value for each of the 10 features. In some embodiments, each subject in the subject dataset has measured values for some of the features, and missing values are filled using interpolation techniques or ignored (underestimated). In some embodiments, each subject in the subject dataset has values for some of the features, and missing values are filled using constraints. The values from the feature vectors of the subjects in the subject dataset define the vector: X 1 , X 2 , X 3 , X 4 , X 5 , X 6 , X 7 , X 8 , X 9 , X 10 , where X i is the value of the i-th feature in the feature vector of a particular subject. If there are Q subjects in the subject dataset, the selection of 10 features can define Q vectors. In clustering, members of the subject dataset that exhibit similar measurement patterns across their respective feature vectors tend to be clustered together.

[0140] Specific exemplary clustering techniques that may be used include, but are not limited to, hierarchical clustering (agglomerative clustering using the nearest neighbor algorithm, farthest neighbor algorithm, average linkage algorithm, centroid algorithm, or sum of squares algorithm), k-means clustering, fuzzy k-means clustering algorithm, Jarvis-Patrick clustering, density-based spatial clustering algorithm, divisive clustering algorithm, supervised clustering algorithm, or an ensemble thereof. Such clustering may relate to features within the feature vector of each subject, or principal components (or other forms of reduced components) derived therefrom. In some embodiments, the clustering includes unsupervised clustering where no preconception is imposed as to what clusters may be formed when the subject data set is clustered.

[0141] Data clustering is an unsupervised process that requires effective optimization. For example, describing a data set using either too few or too many clusters can result in loss of information. See, e.g., Jain et al. 1999 “Data Clustering: A review” ACM Computing Surveys 31(3), 264-323, and Berkin 2002”Survey of clustering datamining techniques” Tech Report, Accrue Software, San Jose, CA, each of which is incorporated herein by reference. In some embodiments, to improve the clustering process, multiple subjects are normalized prior to clustering (e.g., one or more dimensions of each feature vector among multiple feature vectors are normalized (e.g., to the respective average value of the corresponding dimension determined from the multiple feature vectors).

[0142] In some embodiments, a centroid-based clustering algorithm is used to perform clustering of a plurality of subjects. Centroid-based clustering organizes data into non-hierarchical clusters and represents all of the subjects from the perspective of a centroid vector (where the vector itself may not be part of the data set). The algorithm then calculates a distance measurement between each subject and the centroid vector and clusters the subjects based on their proximity to one of the centroid vectors. In some embodiments, a Euclidean distance measurement, Manhattan distance measurement, or Minkowski distance measurement is used to calculate the distance measurement between each subject and the centroid vector. In some embodiments, a k-means, k-medoid, CLARA, or CLARANS clustering algorithm is used to cluster the plurality of subjects. An example of the k-means algorithm is described in Uppada 2014 “Centroid Based Clustering Algorithms - A Clarion Study” Int J Comp Sci and Inform Technol 5(6),7309-7313, which is incorporated herein by reference.

[0143] In some embodiments, a density-based clustering algorithm is used to perform clustering of a plurality of subjects. A density-based spatial clustering algorithm identifies clusters as regions (e.g., a plurality of feature vectors) of a dataset with a higher density (e.g., a high-density region of subjects). In some embodiments, density-based spatial clustering can be performed as described in Ester et al. 1996 “A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise” KDD’96: Proceedings of the Second International Conference on Knowledge Discovery and Data Mining, 226-231, which is incorporated herein by reference. In such embodiments, the algorithm allows for arbitrarily shaped distributions and does not assign outliers (e.g., subjects outside the density of other subjects) to clusters.

[0144] In some embodiments, a hierarchical clustering (e.g., connectivity-based clustering) algorithm is used to perform clustering of a plurality of subjects. Generally, hierarchical clustering is used to construct a series of clusters and can be agglomerative or divisive, as further described below (e.g., there are agglomerative or divisive subsets of hierarchical clustering methods). For example, Rokach et al., incorporated herein by reference, describes various versions of agglomerative clustering methods (“Clustering Methods” 2005 Data Mining and Knowledge Discovery Handbook, 321-352).

[0145] In some embodiments, hierarchical clustering includes divisive clustering. Divisive clustering first groups a plurality of subjects into one cluster and then divides the plurality of subjects into more clusters (e.g., it is a recursive process) until a specific threshold (e.g., the number of clusters) is reached. Examples of different methods of divisive clustering are described, for example, in Chavent et al. 2007 “DIVCLUS-T: a monothetic divisive hierarchical clustering method” Comp Stats Data Anal 52(2), 687-701, Sharma et al. 2017”Divisive hierarchical maximum likelihood clustering” BMC Bioinform 18(Suppl 16):546, and Xiong et al. 2011”DHCC: Divisive hierarchical clustering of categorical data” Data Min Knowl Disc doi 10.1007 / s10618-011-0221-2, which are each incorporated herein by reference.

[0146] In some embodiments, hierarchical clustering includes agglomerative clustering. Agglomerative clustering generally first separates a plurality of subjects into a number of distinct clusters (e.g., in some cases starting with each individual subject defining a cluster), and continuously and repeatedly merges pairs of clusters. Ward's method is an example of agglomerative clustering that uses the sum of squares to reduce the variance among the members of each cluster (e.g., it is a minimum variance agglomerative clustering technique). See Murtagh and Legendre 2014 “Ward’s Hierarchical Agglomerative Clustering Method” J.Class 31,274-295, which is incorporated herein by reference. A drawback of many agglomerative clustering methods is their high computational requirements. In some embodiments, the agglomerative clustering algorithm can be combined with a k-means clustering algorithm. Non-limiting examples of agglomerative and k-means clustering are described in Karthikeyan et al.2020 “A comparative study of k-means clustering and agglomerative hierarchical clustering” Int J Emer Trends Eng Res 8(5),1600-1604, which is incorporated herein by reference. As an example, the k-means clustering algorithm divides a plurality of subjects into individual sets of k clusters within the data space (e.g., initial k partitions). In some embodiments, k-means clustering is repeatedly applied to a plurality of subjects (e.g., k-means clustering is applied to a plurality of subjects a number of times, e.g., consecutively). In some embodiments, the use of a combination of agglomerative and k-means clustering requires less computation than either agglomerative clustering or k-means clustering alone.

[0147] Block 216. Referring to block 216, in some embodiments, the target model is a convolutional neural network.

[0148] In some embodiments (e.g., when at least one target is a polymer having an active site and the subject is a chemical composition), the description of the subject presented for each target is obtained by docking the atomic representation of the subject to the atomic representation of the active site of the polymer. Non-limiting examples of such docking are Liu and Wang, 1999, “MCDOCK: A Monte Carlo simulation approach to the molecular docking problem,” Journal of Computer-Aided Molecular Design 13, 435-451, Shoichet et al., 1992, “Molecular docking using shape descriptors,” Journal of Computational Chemistry 13(3), 380-397, Knegtel et al., 1997 “Molecular docking to ensembles of protein structures,” Journal of Molecular Biology 266, 424-440, Morris et al., 2009, “AutoDock4 and AutoDockTools4: Automated Docking with Selective Receptor Flexibility,” J Comput Chem 30(16), 2785-2791, Sotriffer et al., 2000, “Automated docking of ligands to antibodies: methods and applications,” Methods: A Companion to Methods in Enzymology 20, 280-291, Morris et al., 1998, “Automated Docking Using a Lamarckian Genetic Algorithm and Empirical Binding Free Energy Function,” Journal of Computational Chemistry 19:1639-1662, and Rarey et al., 1996, “A Fast Flexible Docking Method Using an Incremental Construction Algorithm,” Journal of Molecular Biology 261, 470 - 489, each of which is incorporated herein by reference. Then, the description of this pose of each test subject with respect to at least one target object is applied to the target model. In some such embodiments, the test subject is a chemical compound, each target object includes a polymer having a binding pocket, and presenting a description of the test subject with respect to each target object includes docking the atom coordinates modeled for the chemical compound to the atom coordinates for the binding pocket..

[0149] In some embodiments, each test subject is presented with respect to one or more target objects and is a chemical compound presented to the target model using any of the techniques disclosed in U.S. Patent Nos. 10,546,237, 10,482,355, 10,002,312, and 9,373,059, each of which is incorporated herein by reference.

[0150] In some embodiments, the convolutional neural network includes an input layer, a plurality of individually weighted convolutional layers, and an output score layer, as described in U.S. Patent No. 10,002,312, entitled "Systems and Methods for Applying a Convolutional Network to Spatial Data," issued on June 19, 2018, the entirety of which is incorporated herein by reference. For example, in some such embodiments, the convolutional layers of the target model include an initial layer and a final layer. In some embodiments, the final layer may include gating using a threshold function or activation function f, which can be a linear or non-linear function. The activation function can be, for example, a rectified linear unit (ReLU) activation function, a leaky ReLU activation function, or other functions such as saturated hyperbolic tangent, identity, binary step, logistic, arctangent, softsign, parametric rectified linear unit, exponential linear unit, softplus, bent identity, softExponential, sinusoid, sine, Gaussian, or sigmoid function, or any combination thereof.

[0151] In response to the input, in some embodiments, the input layer supplies values to an initial convolutional layer. In some embodiments, each convolutional layer other than the final convolutional layer supplies to another one of the convolutional layers an intermediate value as a function of the weights of the respective convolutional layer and the input values of the respective convolutional layer. In some embodiments, the final convolutional layer supplies values to a scorer as a function of the final layer weights and the input values. In this way, the scorer may score each of the feature vectors (e.g., the input vectors described in U.S. Patent No. 10,002,312) that describe each subject, and use these scores together to provide the corresponding target results (e.g., the classifications described in U.S. Patent No. 10,002,312) for each respective subject. In some embodiments, the scorer provides a respective single score for each of the feature vectors and uses a weighted average of these scores to provide the corresponding target results for each respective subject.

[0152] In some embodiments, the total number of layers (including the input layer and the output layer) used in the convolutional neural network ranges from about 3 to about 200. In some embodiments, the total number of layers is at least 3, at least 4, at least 5, at least 10, at least 15, or at least 20. In some embodiments, the total number of layers is at most 20, at most 15, at most 10, at most 5, at most 4, or at most 3. One of ordinary skill in the art will recognize that the total number of layers used in the convolutional neural network may have any value within this range, e.g., 8 layers.

[0153] In some embodiments, the total number of trainable or learnable parameters used in a convolutional neural network, such as weighting coefficients, biases, or thresholds, ranges from about 1 to about 10,000. In some embodiments, the total number of learnable parameters is at least 1, at least 10, at least 100, at least 500, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000, or at least 10,000. Alternatively, the total number of learnable parameters is any number less than 100, any number from 100 to 10,000, or a number greater than 10,000. In some embodiments, the total number of learnable parameters is at most 10,000, at most 9,000, at most 8,000, at most 7,000, at most 6,000, at most 5,000, at most 4,000, at most 3,000, at most 2,000, at most 1,000, at most 500, at most 100, at most 10, or at most 1. Those skilled in the art will recognize that the total number of learnable parameters used can have any value within this range.

[0154] Since convolutional neural networks require a fixed input size, some embodiments of the disclosed systems and methods that utilize a convolutional neural network for a target model crop geometric data (target - subject complex) to fit within an appropriate bounding box. For example, a cube with sides of 25 - 40 Å may be used. In some embodiments where the target and / or subject is docked to the active site of the target, the center of the active site functions as the center of the cube.

[0155] In some embodiments, a square cube of a fixed size centered on the active site of the target is used to divide the space into a voxel grid, but the disclosed system is not so limited. In some embodiments, any of a variety of shapes is used to divide the space into a voxel grid. In some embodiments, polyhedra such as rectangular prisms, polyhedral shapes, etc. are used to divide the space.

[0156] In embodiments, the grid structure may be configured to be similar to the arrangement of voxels. For example, each sub-structure may be associated with a channel for each atom being analyzed. Also, a coding method for numerically representing each atom may be provided.

[0157] In some embodiments, the voxel map that describes the interface between the subject and the target object may take into account the factor of time and thus be four-dimensional (X, Y, Z, and time).

[0158] In some embodiments, instead of voxels, other implementation manners such as pixels, points, polygonal shapes, polyhedra, or any other type of shape in multiple dimensions (e.g., shapes in 3D, 4D, etc.) may be used.

[0159] In some embodiments, the geometric data is normalized by selecting the origin of the X, Y, and Z coordinates to be the center of mass of the binding site of the target object, as determined by a cavity flooding algorithm. For representative details of such algorithms, see Ho and Marshall, 1990, “Cavity search: An algorithm for the isolation and display of cavity-like binding regions,” Journal of Computer-Aided Molecular Design 4, pp. 337-354, and Hendlich et al., 1997, “Ligsite: automatic and efficient detection of potential small molecule-binding sites in proteins,” J. Mol. Graph. Model 15, no. 6, which are hereby incorporated by reference. Alternatively, in some embodiments, the origin of the voxel map is centered on the center of mass of the entire complex (of the subject bound to the target object, of the target object only, or of the subject only). The basis vectors may optionally be selected to be the principal moments of the entire complex, of the target object only, or of the subject only. In some embodiments, the target object is a polymer having an active site, and the sampling is in a three-dimensional grid format with the center of mass of the active site as the origin, sampling the subject at each of the respective poses of the plurality of different poses described above for the subject and the active site, and the corresponding three-dimensional uniform honeycomb for sampling represents a portion of the polymer and the subject centered on the center of mass. In some embodiments, the uniform honeycomb is a regular cubic honeycomb, and the portions of the polymer and the subject are cubes of a predetermined fixed dimension. The use of cubes of a predetermined fixed dimension in such embodiments ensures that the relevant portions of the geometric data are used and that each voxel map is of the same size.In some embodiments, a predetermined fixed dimension of the cube is N Å × N Å × N Å, where N is an integer or real value from 5 to 100, an integer from 8 to 50, or an integer from 15 to 40. In some embodiments, the uniform honeycomb is a rectangular prism honeycomb, which is a part of the polymer, and the test subject has a predetermined fixed dimension of Q Å x R Å x S Å of the rectangular prism, where Q is a first integer from 5 to 100, R is a second integer from 5 to 100, S is a third integer or real value from 5 to 100, and at least one number in the set {Q, R, S} is not equal to another value in the set {Q, R, S}.

[0160] In some embodiments, each voxel has one or more input channels, and the one or more input channels can have various values associated therewith, which, in one implementation, can be on / off and can be configured to encode the type of atom. The atom type may represent the element of the atom, or the atom type may be further refined to distinguish other atomic features. Then, the atoms present may be encoded in each voxel. Various types of encoding may be utilized using various techniques and / or methodologies. As an exemplary encoding method, the atomic number of the atom may be utilized, and one value per voxel is obtained, ranging from 1 for hydrogen to 118 for ununoctium (or any other element).

[0161] However, as discussed above, other encoding methods such as "one-hot encoding" may be utilized, in which case each voxel has multiple parallel input channels, each of which encodes whether a certain type of atom is either on or off. The atom type may represent the element of the atom and may be further refined to distinguish other atomic features. For example, SYBYL atom types distinguish single-bonded carbon from double-bonded carbon, triple-bonded carbon, or aromatic carbon. For SYBYL atom types, see Clark et al., 1989, "Validation of the General Purpose Tripos Force Field, 1989, J. Comput. Chem. 10, pp. 982-1012, which is incorporated herein by reference.

[0162] In some embodiments, each voxel further includes one or more channels for distinguishing cofactors for atoms that are part of the target object or for parts of the subject. For example, in one embodiment, each voxel further includes a first channel for the target object and a second channel for the subject. If the atoms in a portion of the space represented by the voxel are from the target object, the first channel is set to a value such as "1" and zero otherwise (e.g., because this portion of the space represented by the voxel does not contain atoms from the subject or contains one or more atoms). Further, if the atoms in a portion of the space represented by the voxel are from the subject, the second channel is set to a value such as "1" and zero otherwise (e.g., because this portion of the space represented by the voxel does not contain atoms from the target object or contains one or more atoms). Similarly, other channels can additionally (or alternatively) specify further information such as partial charge, polarization, electronegativity, solvent accessible space, and electron density. For example, in some embodiments, the electron density map of the target object is overlaid with a set of three-dimensional coordinates, and creating the voxel map further samples the electron density map. Examples of suitable electron density maps include, but are not limited to, multiple isomorphous replacement maps, single isomorphous replacement with anomalous signal maps, single wavelength anomalous dispersion maps, multiple wavelength anomalous dispersion maps, and 2F observable -F calculated maps. See McRee, 1993, Practical Protein Crystallography, Academic Press, which is hereby incorporated by reference herein.

[0163] In some embodiments, voxel encoding according to the disclosed systems and methods can include additional optional encoding refinements. Two of the following are provided as examples.

[0164] In the first encoding refinement, the required memory can be reduced by reducing the set of atoms represented by voxels (e.g., by reducing the number of channels represented by voxels), based on the fact that most elements occur rarely in biological systems. The atoms can be mapped to share the same channel in a voxel either by combining rare atoms (thus potentially rarely affecting the performance of the system) or by combining atoms with similar properties (thus minimizing inaccuracies from the combination).

[0165] Another encoding refinement is to have voxels represent atomic positions by partially activating adjacent voxels. This results in partial activation of adjacent neurons in subsequent neural networks, moving from one-hot encoding to "somewhat warm" encoding. For example, this could be a van der Waals diameter of 3.5 Å, and thus a volume of 22.4 Å when the grid is placed at 1 Å 3 when the grid is placed 3Considering the chlorine atoms involved, it can be exemplified that the voxels inside the chlorine atoms are completely filled, while only the voxels at the atom's edges are partially filled. Thus, the channels representing chlorine in the partially filled voxels will turn on in proportion to the amount of such voxels that fall within the chlorine atom. For example, if 50% of the voxel volume is within the chlorine atom, 50% of the channels in the voxels representing chlorine will be activated. This can result in a "smoothed" and more accurate representation compared to individual one-hot encodings. Thus, in some embodiments, the subject is a first compound, the target is a second compound, and the atomic features resulting from the sampling spread across a subset of the voxels of each voxel map, where the subset of voxels includes two or more voxels, three or more voxels, five or more voxels, ten or more voxels, or twenty-five or more voxels. In some embodiments, the atomic properties consist of a listing of atomic types (e.g., one of the SYBYL atomic types).

[0166] Thus, the voxelization (rasterization) of the encoded geometric data (docking of the target onto the subject) is based on various rules applied to the input data.

[0167] Figures 6 and 7 provide diagrams of two subjects 602 encoded on a two-dimensional grid 600 of voxels, according to some embodiments. Figure 6 provides two subjects superimposed on a two-dimensional grid. Figure 7 provides a one-hot encoding that uses different shading patterns to encode the presence of oxygen, nitrogen, carbon, and empty space, respectively. As noted above, such an encoding may be referred to as a "one-hot" encoding. Figure 7 shows grid 500 of Figure 6 with subject 502 omitted. Figure 8 shows a diagram of the two-dimensional grid of voxels of Figure 7 with the voxels numbered.

[0168] In some embodiments, the characteristic geometric shape is represented in a form other than voxels. FIG. 9 provides diagrams of various representations in which features (e.g., atom centers) are represented as 0-D points (representation 902), 1-D points (representation 904), 2-D points (representation 906), or 3-D points (representation 908). First, the spacing between points can be randomly selected. However, when training the target model, the points can move closer to or farther from each other. FIG. 10 illustrates the range of possible positions for each point.

[0169] In embodiments where the interaction between the subject and the target is encoded as a voxel map, each voxel map is optionally expanded into a corresponding vector, thereby creating a plurality of vectors, and each vector in the plurality of vectors is of the same size. In some embodiments, each vector in the plurality of vectors is a one-dimensional vector. For example, in some embodiments, a 20 Å cube on each side is centered on the active site of the target and sampled at a 1 Å three-dimensional fixed grid interval to hold the basis of voxel structure features such as atom types in each channel in the corresponding voxels of the voxel map, and optionally, more complex subject-target descriptors are formed. In some embodiments, the voxels of this three-dimensional voxel map are expanded into one-dimensional floating-point vectors. In some embodiments where the target model is a convolutional neural network, the vectorized representation of the voxel map is provided to the convolutional network.

[0170] In some embodiments, the convolutional layers in a plurality of convolutional layers include a set of filters (also referred to as kernels). Each filter has a fixed three-dimensional size that is convolved (stepped at a predetermined step rate) across the depth, height, and width of the input volume of the convolutional layer, calculates a dot product (or other function) between the filter and the entries (weights) of the input, thereby creating a multi-dimensional activation map for that filter. In some embodiments, the filter step rate is one element, two elements, three elements, four elements, five elements, six elements, seven elements, eight elements, nine elements, ten elements, or more than ten elements of the input space. Thus, consider the case where the filter is size 5 3 . In some embodiments, this filter calculates a dot product (or other mathematical function) between contiguous cubes of the input space having a depth of five elements, a width of five elements, and a height of five elements, for a total count of 125 values of the input space per voxel channel.

[0171] The input space to the initial convolutional layer (e.g., the output from the input layer) is formed from either a voxel map or a vectorized representation of the voxel map. In some embodiments, the vectorized representation of the voxel map is a one-dimensional vectorized representation of the voxel map that functions as the input space to the initial convolutional layer. Nevertheless, when the filter convolves its input space and the input space is a one-dimensional vectorized representation of a voxel map, the filter still obtains those elements that represent the corresponding contiguous cubes of the fixed space within the target - subject complex from the one-dimensional vectorized representation. In some embodiments, the filter selects those elements from within the one-dimensional vectorized representation that forms the corresponding contiguous cubes of the fixed space of the target - subject complex using standard bookkeeping techniques. Thus, in some cases, this necessarily involves taking a non - contiguous subset of the elements of the one-dimensional vectorized representation to obtain the element values of the corresponding contiguous cubes of the fixed space of the target - subject complex.

[0172] In some embodiments, the filter is initialized (e.g., to Gaussian noise) or trained to have corresponding weights of 125 (per input channel) and performs a dot product (or some other form of mathematical operation, such as a function of 125 input space values, to compute a first single value (or set of values) of the activation layer corresponding to the filter). In some embodiments, the value computed by the filter is summed, weighted, and / or biased. To compute additional values of the activation layer corresponding to the filter, the filter is then stepped (convolved) in one of the three dimensions of the input volume by a step rate (stride) associated with the filter, at which point a dot product or some other form of mathematical operation between the filter weights and 125 input space values (per channel) is performed at the new location of the input volume. This stepping (convolution) is repeated until the filter samples the entire input space according to the step rate. In some embodiments, the boundaries of the input space are zero-padded to control the spatial volume of the output space generated by the convolutional layer. In a typical embodiment, each of the filters of the convolutional layer thus canvases the entire three-dimensional input volume, thereby forming the corresponding activation map. The collection of activation maps from the filters of the convolutional layer collectively forms the three-dimensional output volume of one convolutional layer, thereby functioning as the three-dimensional (three spatial dimensions) input to a subsequent convolutional layer. Thus, any entry in the output volume can also be interpreted as the output of a single neuron (or set of neurons) that looks at a small region of the input space to the convolutional layer and shares the neurons and parameters of the same activation map. Thus, in some embodiments, the convolutional layer among a plurality of convolutional layers has a plurality of filters, and each filter among the plurality of filters convolves a cubic input space of N 3 with a stride Y in (three spatial dimensions), where N is an integer greater than or equal to 2 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or greater than 10), and Y is a positive integer (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or greater than 10).

[0173] Each layer in the plurality of convolutional layers is associated with a different set of weights. More specifically, each layer in the plurality of convolutional layers includes a plurality of filters, and each filter includes a plurality of independent weights. In some embodiments, the convolutional layer has 128 filters of size 5 3 and thus the convolutional layer has 128 × 5 × 5 or 16,000 weights per channel of the voxel map. Thus, if there are 5 channels in the voxel map, the convolutional layer will have 16,000 × 5 weights, or 80,000 weights. In some embodiments, some or all of such weights (and optionally, biases) of any filter of a given convolutional layer may be tied together, e.g., constrained to be identical.

[0174] In response to the input of each vector in the plurality of vectors, the input layer supplies a first plurality of values to the initial convolutional layer as a first function of the values of each vector.

[0175] Each convolutional layer other than the final convolutional layer supplies intermediate values as a second function of (i) a different set of weights associated with each convolutional layer and (ii) the input values received by each convolutional layer to another convolutional layer in the plurality of convolutional layers. For example, each filter of each convolutional layer canvases the input volume to the convolutional layer at each respective filter position (in three spatial dimensions) according to the characteristic three-dimensional stride of the convolutional layer, takes the dot product of the filter weights of each filter (or some other mathematical function) and the values of the input volume (a concatenated cube that is a subset of the total input space) at each respective filter position, thereby generating a calculated point (or set of points) on the activation layer corresponding to each filter position. The activation layer of the filters of each convolutional layer collectively represents the intermediate values of each convolutional layer.

[0176] The final convolutional layer supplies the final value as a third function of (i) a different set of weights associated with the final convolutional layer and (ii) the input value received by the final convolutional layer to the score layer. For example, each respective filter of the final convolutional layer canvases the input volume (in three spatial dimensions) up to the final convolutional layer at each respective filter position according to the characteristic three-dimensional stride of the convolutional layer, takes the dot product of the filter weights of the filter (or some other mathematical function) and the values of the input volume at each position, thereby calculating a point (or set of points) on the activation layer corresponding to each respective filter position. The activation layers of the filters of the final convolutional layer collectively represent the final values supplied to the score layer.

[0177] In some embodiments, the convolutional neural network has one or more activation layers. In some embodiments, the activation layer is a layer of neurons that applies the non-saturating activation function f(x)=max(0,x). This increases the decision function and the non-linear characteristics of the network as a whole without affecting the receptive field of the convolutional layer. In other embodiments, the activation layer is another function for increasing non-linearity, such as the saturating hyperbolic tangent function f(x)=tanh, f(x)=│tanh(x)│, and the sigmoid function f(x)=(1+e -x ) -1 has. In some embodiments of the neural network, non-limiting examples of other activation functions found in other activation layers include, but are not limited to, logistic (or sigmoid), softmax, Gaussian, Boltzmann weighted averaging, absolute value, linear, rectified linear, bounded rectified linear, soft rectified linear, parameterized rectified linear, mean, max, min, some vector norms LP (for p = 1, 2, 3,..., ∞), square, square root, polynomial, inverse quadratic, inverse polynomial, multi-harmonic spline, thin plate spline.

[0178] In some embodiments, zero or more of the layers of the target model (in embodiments where the target model is a convolutional neural network) may consist of pooling layers. Like convolutional layers, pooling layers are a set of function calculations that apply the same function to different spatially local patches of the input. For pooling layers, the output is given by a pooling operator, such as some vector norms LP for p = 1, 2, 3, ..., ∞, over some voxels. Pooling is typically done per channel rather than between channels. Pooling divides the input space into a set of three-dimensional boxes and outputs the maximum value for each such sub-region. The pooling operation provides a form of translational invariance. The function of the pooling layer is to gradually reduce the spatial size of the representation, reducing the amount of parameters and calculations within the network, and thus also controlling overfitting. In some embodiments, the pooling layer is inserted between consecutive convolutional layers within a target model in the form of a convolutional neural network. Such a pooling layer operates independently on every depth slice of the input and changes its size spatially. The pooling unit can also perform other functions such as average pooling or even L2 norm pooling in addition to max pooling.

[0179] In some embodiments, zero or more of the layers of the target model (in embodiments where the target model is a convolutional neural network) may consist of normalization layers, such as local response normalization or local contrast normalization, that can be applied across channels at the same position or to specific channels across several positions. These normalization layers can promote diversity in the responses of several function calculations to the same input.

[0180] In some embodiments, the scorer (in embodiments where the target model is a convolutional neural network) includes a plurality of fully connected layers and an evaluation layer to which the fully connected layers in the plurality of fully connected layers supply. The neurons in the fully connected layers have a full connection to all activations of the previous layer, as seen in a normal neural network. Thus, those activations can be calculated with matrix multiplication followed by a bias offset. In some embodiments, each fully connected layer has 512 hidden units, 1024 hidden units, or 2048 hidden units. In some embodiments, the scorer has no fully connected layer, one fully connected layer, two fully connected layers, three fully connected layers, four fully connected layers, five fully connected layers, six or more fully connected layers, or ten or more fully connected layers.

[0181] In some embodiments, the evaluation layer discriminates a plurality of active classes. In some embodiments, the evaluation layer includes a logistic regression cost layer over two active classes, three active classes, four active classes, five active classes, or six or more active classes.

[0182] In some embodiments, the evaluation layer includes a logistic regression cost layer over a plurality of active classes. In some embodiments, the evaluation layer includes a logistic regression cost layer over two active classes, three active classes, four active classes, five active classes, or six or more active classes.

[0183] In some embodiments, the evaluation layer discriminates two active classes, and the first active class (the first classification) is the subject's IC for the target object that exceeds the first binding value 50 , EC 50 , Kd, or KI, and the second active class (the second classification) is the subject's IC for the target object that is below the first binding value 50 , EC 50, Kd, or KI. In some such embodiments, the target result is an indication that the subject has a first activity or a second activity. In some embodiments, the first binding value is 1 nanomolar, 10 nanomolars, 100 nanomolars, 1 micromolar, 10 micromolars, 100 micromolars, or 1 millimolar.

[0184] In some embodiments, the evaluation layer includes a logistic regression cost layer across two activity classes, and the first activity class (first classification) is the subject's IC for a target that exceeds the first binding value 50 , EC 50 , Kd, or KI, and the second activity class (second classification) is the subject's IC for a target that is below the first binding value 50 , EC 50 , Kd, or KI. In some such embodiments, the target result is an indication that the subject has a first activity or a second activity. In some embodiments, the first binding value is 1 nanomolar, 10 nanomolars, 100 nanomolars, 1 micromolar, 10 micromolars, 100 micromolars, or millimolar.

[0185] In some embodiments, the evaluation layer discriminates three activity classes, and the first activity class (first classification) is the subject's IC for a target that exceeds the first binding value 50 , EC 50 , Kd, or KI, and the second activity class (second classification) is the subject's IC for a target between the first binding value and the second binding value 50 , EC 50 , Kd, or KI, and the third activity class (third classification) is the subject's IC for a target that is below the second binding value 50 , EC 50 , Kd, or KI, and the first binding value is other than the second binding value. In some such embodiments, the target result is an indication that the subject has a first activity, a second activity, or a third activity.

[0186] In some embodiments, the evaluation layer includes a logistic regression cost layer across three active classes, where the first active class (first classification) is the subject's IC for a target object that exceeds a first binding value 50 , EC 50 , Kd, or KI, the second active class (second classification) is the subject's IC for a target object between the first binding value and a second binding value 50 , EC 50 , Kd, or KI, and the third active class (third classification) is the subject's IC for a target object that is below the second binding value 50 , EC 50 , Kd, or KI, where the first binding value is other than the second binding value. In some such embodiments, the target result is an indication that the subject has first activity, second activity, or third activity

[0187] In some embodiments, the scorer (in embodiments where the target model is a convolutional neural network) includes a fully connected single or multi-layer perceptron. In some embodiments, the scorer includes a support vector machine, random forest, nearest neighbor. In some embodiments, the scorer assigns a numerical score indicating the strength (or confidence or probability) of classifying the input into various output categories. In some cases, the categories are binder and non-binder, or alternatively, potency levels (e.g., <1 molar, <1 millimolar, <100 micromolar, <10 micromolar, <1 micromolar, <100 nanomolar, <10 nanomolar, <1 nanomolar IC 50 , EC 50 or KI potency). In some such embodiments, the target result is that the indication is an identification of one of these categories for the subject

[0188] Details for obtaining target results of a target model of a complex of a subject under test and a target object have been described above. As discussed above, in some embodiments, each subject is docked to the target object in a plurality of poses. To present all such poses to the target model at once, a very large input field (e.g., an input field of a size equal to the number of voxels * number of channels * number of poses if the target model is a convolutional neural network) may be required. In some embodiments, all poses are presented to the target model simultaneously, while in other embodiments, each such pose is processed into a voxel map, vectorized, and functions as a sequential input to the target model (e.g., if the target model is a convolutional neural network). In this way, a plurality of scores are obtained from the target model, and each score among the plurality of scores corresponds to the input of a vector among the plurality of vectors to the input layer of the score layer of the target model. In some embodiments, the scores for each of the poses of a given subject having a given target object are combined together (e.g., as a weighted average of the scores, as a measure of the central tendency of the scores, etc.) to generate a final target result for each subject.

[0189] In some embodiments where the score layer output of the target model is a numerical value, the output may be combined using any of the activation functions described herein, or any activation function known or developed. Examples include, but are not limited to, the non-saturating activation function f(x) = max(0, x), the saturating hyperbolic tangent function f(x) = tanh, f(x) = |tanh(x)|, the sigmoid function f(x) = (1 + e -x ) -1 , logistic (or sigmoid), softmax, Gaussian, Boltzmann averaging, absolute value, linear, rectified linear, bounded rectified linear, soft rectified linear, parameterized rectified linear, mean, max, min, some vector norms LP (for p = 1, 2, 3,..., ∞), square, square root, polynomial, inverse quadratic, inverse polynomial, multi-harmonic spline, thin plate spline.

[0190] In some embodiments of the present disclosure, the target model may be configured to utilize the Boltzmann distribution to combine outputs, since if the output is interpreted as indicating the binding energy, this is consistent with the physical probability of the pose. In other embodiments of the present disclosure, the max() function may also provide a reasonable approximation to Boltzmann and is computationally efficient.

[0191] In some embodiments where the score output of the target model is not numerical, the scorer may be configured to combine the outputs using various ensemble voting schemes, which may include, by way of non-limiting illustrative examples, majority vote, weighted average, Condorcet method, Borda count, among others, to form the corresponding target result.

[0192] In some embodiments, the system may be configured to apply an ensemble of scorers to generate, for example, an indicator of binding affinity.

[0193] In some embodiments, the subject is a chemical compound, and characterizing (e.g., determining classification) the subject using multiple scores (from multiple poses of the subject) involves taking a measure of the central tendency of the multiple scores. When the measure of central tendency meets a predetermined threshold or a predetermined threshold range, the subject is considered to have a first classification. If the measure of central tendency does not reach the predetermined threshold or the predetermined threshold range, the subject is considered to have a second classification. In some such embodiments, the target result output by the target model for each subject is an indication of one of these classifications.

[0194] In some embodiments, using a plurality of scores to characterize a subject includes taking a weighted average of the plurality of scores (from a plurality of poses of the subject). When the weighted average meets a predetermined threshold or a predetermined threshold range, the subject is considered to have a first classification. If the weighted average does not reach the predetermined threshold or the predetermined threshold range, the subject is considered to have a second classification. In some embodiments, the weighted average is the Boltzmann average of the plurality of scores. In some embodiments, the first classification is the subject's IC 50 、EC 50 、Kd, or KI for a target object that exceeds a first binding value (e.g., 1 nanomolar, 10 nanomolar, 100 nanomolar, 1 micromolar, 10 micromolar, 100 micromolar, or 1 millimolar), and the second classification is the subject's IC 50 、EC 50 、Kd, or KI for a target object that is below the first binding value. In some such embodiments, the target result output by the target model for each subject is a display of one of these classifications.

[0195] In some embodiments, providing a target result of a subject using a plurality of scores includes taking a weighted average of the plurality of scores (from a plurality of poses of the subject). When the weighted average meets each of the threshold ranges among the plurality of threshold ranges, the subject is considered to have each of the classifications among the plurality of classifications that uniquely corresponds to each of the threshold ranges. In some embodiments, each of the classifications among the plurality of classifications is the subject's IC 50 、EC 50 、Kd, or KI range (e.g., 1 micromolar to 10 micromolar, 1 nanomolar to 100 nanomolar).

[0196] In some embodiments, a single pose of each subject for a given target object is run through a target model, and the subject is classified using each score assigned by each subject's respective target model based thereon.

[0197] In some embodiments, a weighted average of target model scores for one or more poses of a subject with respect to each of a plurality of target objects evaluated by a target model using the techniques disclosed herein is used to provide a target result for the subject. For example, in some embodiments, the plurality of target objects are derived from a molecular dynamics run, and each target object in the plurality of target objects represents the same polymer at different time steps in the molecular dynamics run. Each voxel map of one or more poses of the subject with respect to each of these target objects is evaluated by the target model to obtain a score for each independent pose-target object pair and a weighted average of these scores, or some other measure of central tendency of these scores is used to provide a target result for the target object.

[0198] Block 218. Referring to block 218 of FIG. 2A, in some embodiments, at least one target object is a single object (e.g., each target object is a respective single object). In some embodiments, the single object is a polymer. In some embodiments, the polymer includes an active site (e.g., the polymer is an enzyme having an active site). In some embodiments, the polymer is an assembly of proteins, polypeptides, polynucleic acids, polyribonucleic acids, polysaccharides, or any combination thereof. In some embodiments, the single object is an organometallic complex. In some embodiments, the single object is a surfactant, reverse micelle, or liposome.

[0199] In some embodiments, each subject among the plurality of subjects includes a respective chemical compound that may or may not bind to the active site of at least one target object having a corresponding affinity (e.g., an affinity to form a chemical bond to at least one target object).

[0200] In some embodiments, at least one target object includes at least two target objects, at least three target objects, at least four target objects, at least five target objects, or at least six target objects. In some embodiments, each target object is, as described above, a respective single object (e.g., a single protein, a single polypeptide, etc.). In some embodiments, one or more of the at least one target object includes a plurality of objects (e.g., a protein complex and / or an enzyme having a plurality of subunits such as a ribosome).

[0201] Block 220. Referring to block 220 of FIG. 2B, the method proceeds by training an initial state prediction model using at least i) a subset of the test subjects as independent variables and ii) a corresponding subset of the target results as dependent variables, thereby updating the prediction model to an updated trained state. That is, the prediction model is trained to predict what the target result (target model score) for a given test compound will be without incurring the computational cost of the target model. Moreover, in some embodiments, the prediction model does not utilize at least one target object. In such embodiments, the prediction model attempts to predict the score of the target model based solely on the information provided to the test subjects in the test subject dataset (e.g., the chemical structure of the test subject) rather than on the interaction between the test subject and one or more target objects.

[0202] Referring to block 222, in some embodiments, the target model exhibits a first computational complexity when evaluating each test subject, and the prediction model exhibits a second computational complexity when evaluating each test subject, and the second computational complexity is less than the first computational complexity (e.g., the prediction model requires less time and / or less computational effort to provide respective prediction results for the test subjects than the target model requires to provide corresponding target results for the same test subjects).

[0203] As used herein, the term "computational complexity" is interchangeable with the term "time complexity" and relates to the time required to obtain results when applying a model to a test subject and at least one target subject with a given number of processors, and also relates to the required number of processors necessary to obtain results when applying a model to a test subject and at least one target subject within a given time when each processor has a given amount of processing power. Thus, as used herein, computational complexity refers to the predictive complexity of a model. However, in some embodiments, the target model exhibits a first training computational complexity, the predictive model exhibits a second training computational complexity, and the second training computational complexity is less than the first training computational complexity. Table 2 below lists some exemplary predictive models for making predictions and their estimated computational complexities (predictive complexities). [Table 2]

[0204] In Table 2, p is the number of features of the test subject evaluated by the classifier when providing the result of the classifier, and n trees is the number of trees (in the case of methods based on various trees), and O refers to the Bachmann-Landau notation that indicates the upper bound of the growth rate of a function. See, for example, Arora and Barak, 2009, Computational Complexity: Arora and Barak, 2009, Cumutatucation Complexity: A Modern Approach, Cambridge University Press, Cambridge England. In contrast, one estimate of the total time complexity of a convolutional neural network, which is a form of training model, is [Equation] where l is the index of the convolutional layer, d is the depth (number of convolutional layers), and n l is the number of filters in the l-th layer (also known as "width") (n l-1(also known as the number of input channels of the l-th layer), s l is the spatial size (length) of the filter, m l is the spatial size of the output feature map. This time complexity applies to both training time and test time, but the scales are different. The training time per subject is approximately 3 times the test time per subject (once for forward propagation and twice for backpropagation). See Hi and Sun, 2014, “Convolutional Neural Networks at Constrained Time Cost,” arXiv:1412.1710v1 [cs.CV] 4 Dec 2014, which is incorporated herein by reference. Thus, clearly, the time complexity of the convolutional neural network is greater than that of the exemplary prediction model provided in Table 1.

[0205] Block 224. Referring to block 224 in FIG. 2B, in some embodiments, the initially trained state prediction model includes an untrained or partially trained classifier. For example, in some embodiments, the prediction model is trained partially with other forms of data, such as assay data that is distinct from the data provided by a subject or from multiple subjects in a subject dataset, e.g., using transfer learning techniques. In one example, the prediction model is partially trained with binding affinity data for a set of compounds, and such compounds may or may not be in the subject dataset using transfer learning techniques.

[0206] Referring to block 226, in some embodiments, the updated trained state prediction model includes an untrained or partially trained classifier that is different from the initially trained state prediction model (e.g., one or more weights of the prediction model have been changed). The ability to retrain or update an existing classifier is particularly useful when the training dataset is subject to change (e.g., when the training dataset increases in class size and / or number).

[0207] In some embodiments, a boosting algorithm is used to update (train) the prediction model. The boosting algorithm is generally described by Dai et al. 2007 “Boosting for transfer learning” in Proc 24th Int Conf on Mach Learn, which is incorporated herein by reference. The boosting algorithm can include re-weighting the data (e.g., a subset of subjects) previously used to train the prediction model when new data (e.g., an additional subset of subjects) is added to the data set used to retrain or update the prediction model. See, e.g., Freund et al. 1997 “A decision-theoretic generalization of on-line learning and an application to boosting” J Computer and System Sciences 55(1), 119-139, which is incorporated herein by reference.

[0208] In some embodiments, as discussed above, depending on the type of algorithm used for the initial trained state of the prediction model (e.g., if the prediction model is not a single decision tree), a transfer learning method is used to update the prediction model to an updated trained state (e.g., with each successive iteration of the method). Transfer learning generally involves the transfer of knowledge from a first model to a second model (e.g., any knowledge from a first set of tasks or from a first dataset to a second set of tasks or a second dataset). An additional review of transfer learning methods can be found in Torrey et al. 2009 “Transfer Learning” in the Handbook of Research on Machine Learning Applications, Pan et al. 2009 “A Survey on Transfer Learning” IEEE Transactions on Knowledge and Data Engineering doi:10.1109 / TKDE.2009.191, and Molochanov et al. 2016” Pruning Convolutional Neural Networks for Resource Efficient Transfer Learning” arXiv:1611.06440v1, each of which is hereby incorporated by reference herein. In some embodiments, a variant of random forest can be used with the dynamic training dataset. See Ristin et al. 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3654-3661, which is hereby incorporated by reference herein.

[0209] In some embodiments, the prediction model includes a random forest tree, a random forest including a plurality of multiple additive decision trees, a neural network, a graph neural network, a dense neural network, principal component analysis, nearest neighbor analysis, linear discriminant analysis, quadratic discriminant analysis, a support vector machine, an evolutionary approach, projection pursuit, regression, a naive Bayes algorithm, or an ensemble thereof.

[0210] Random forests, decision trees, and boosted tree algorithms. Decision trees are generally described by Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, 395-396, which is hereby incorporated by reference. Random forests are generally defined as a collection of decision trees. Tree-based methods partition the feature space into a set of rectangles and fit a model (such as a constant) to each rectangle. In some embodiments, the decision tree includes random forest regression. One particular algorithm that can be used for a predictive model is classification and regression trees (CART). Other particular decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and random forests. CART, ID3, and C4.5 are described by Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, 396-408 and 411-412, which is hereby incorporated by reference. CART, MART, and C4.5 are described by Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is hereby incorporated by reference in its entirety. Random forests in general are described by Breiman, 1999, Technical Report 567, Statistics Department, U.C. Berkeley, September 1999, which is hereby incorporated by reference in its entirety.

[0211] Neural networks, graph neural networks, dense neural networks. Various neural networks may be employed as either or both the target model and / or the prediction model, provided that the prediction model has a lower computational complexity than the target model. Neural network algorithms, including convolutional neural network (CNN) algorithms, are disclosed, for example, in Vincent et al., 2010, J Mach Learn Res 11, 3371-3408, Larochelle et al., 2009, J Mach Learn Res 10, 1-40, and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference. In some embodiments, but not limited to, graph neural networks (GNNs) and dense neural networks (DNNs) are included, but other variants of neural network algorithms are used for the prediction model. Graph neural networks are useful for data represented in non-Euclidean spaces (e.g., particularly complex datasets). An overview of GNNs is provided by Wu et al. 2019 “A Comprehensive Survey on Graph Neural Networks” arVix:1901.00596, and Zhou et al 2018 “Graph Neural Networks: A Review of Methods and Applications” arVix:1812.08434. Combining GNNs with other data analysis methods can enable drug discovery. See, for example, Altre-Tran et al. 2017 “Low Data Drug Discovery with One-Shot Learning” ACS Cent Sci 3, 283-293.Dense neural networks generally include a large number of neurons in each layer and are described in Montavon et al. 2018 “Methods for interpreting and understanding deep neural networks” Digit Signal Process 73, 1-15, and Finnegan et al. 2017 “Maximum entropy methods for extracting the learned features of deep neural networks” PLoS Comput Biol. 13(10), 1005836, each of which is incorporated herein by reference.

[0212] Principal component analysis. Principal component analysis is one of several methods often used for dimensionality reduction of complex data (e.g., to reduce the number of subjects under consideration). An example of using PCA for data clustering is provided, for example, by Yeung and Ruzzo 2001 “Principal component analysis for clustering gene expression data” Bioinformat 17(9), 763-774, which is incorporated herein by reference. The principal components are typically ordered by the range of variance present (e.g., it is thought that only the first n components convey signal instead of noise) and are uncorrelated (e.g., each component is orthogonal to the other components).

[0213] Nearest neighbor analysis. Nearest neighbor analysis is typically performed using Euclidean distance. An example of nearest neighbor analysis is provided by Weinberger et al. 2006 “Distance metric learning for large margin nearest neighbor classification” in NIPS MIT Press 2, 3. Nearest neighbor analysis is beneficial because it is effective in settings with large training datasets in some embodiments. See Sonawane 2015 “A Review on Nearest Neighbour Techniques for Large Data” International Journal of Advances Research in Computer and Communication Engineering 4(11), 459 - 461, which is incorporated herein by reference.

[0214] Linear discriminant analysis. Linear discriminant analysis (LDA) is typically performed to identify a linear combination of features that characterize classes of subjects or distinguish between separate classes. Examples of LDA are provided by Ye et al. 2004 “Two-Dimensional Linear Discriminant Analysis” Advances in Neural Information Processing Systems 17, 1569-1576, Prince et al. 2007 “Probabilistic Linear Discriminant Analysis for Inferences about Identity” 11th International Conference on Computer Vision, 1-8. LDA is beneficial because it can be applied to both large and small sample sizes and can be used in high dimensions. See Kaipatnen 1997 “Utilizing Geometric Anomalies of High Dimension: When Complexity Makes Computation Easier” Computer-Intensive Methods in Control and Signal Processing, 283-294.

[0215] Quadratic discriminant analysis. Quadratic discriminant analysis (QDA) is closely related to LDA, but in QDA, individual covariance matrices are estimated for every class of interest. See Wu et al. 1996 “Comparison of regularized discriminant analysis, linear discriminant analysis and quadratic discriminant analysis, applied to NIR data” Analytica Chimica Acta 329, 257-265. Examples of QDA are provided by Zhang 1997 “Identification of protein coding regions in the human genome by quadratic discriminant analysis” PNAS 94, 565-568, Zhang et al. 2003 “Splice site prediction with quadratic discriminant analysis using diversity measure” Nuc Acids Res 31(21), 6124-6220, each of which is hereby incorporated by reference. QDA is beneficial as it provides more effective parameters than LDA, as described in Wu et al. 1996 “Comparison of regularized discriminant analysis, linear discriminant analysis and quadratic discriminant analysis, applied to NIR data” Analytica Chimica Acta 329, 257-265, which is hereby incorporated by reference.

[0216] Support Vector Machine. Non-limiting examples of the Support Vector Machine (SVM) algorithm are described in Cristianini and Shawe-Taylor, 2000 “An Introduction to Support Vector Machines,” Cambridge University Press; Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.; Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., 259, 262-265; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is hereby incorporated by reference in its entirety. When used for classification, the SVM separates a given binary-labeled data training set using a hyperplane that is maximally distant from the labeled data. If linear separation is not possible, the SVM can operate in combination with a “kernel” technique that automatically implements a non-linear mapping into the feature space. The hyperplane found by the SVM in the feature space corresponds to a non-linear decision boundary in the input space.

[0217] Linear regression. As used herein, linear regression can include simple, multivariate, and / or multiple linear regression analysis. Linear regression uses a linear approach to model the relationship between a dependent variable (also known as a scalar response) and one or more independent variables (also known as explanatory variables), and thus can be used as a predictive model in the present disclosure. See Altman et al. 2015 “Simple Linear Regression” Nature Methods 12, 999 - 1000, which is incorporated herein by reference. The relationship is predicted using a linear predictor function, and its parameters are estimated from the data using a linear model. In some embodiments, simple linear regression is used to model the relationship between a dependent variable and a single independent variable. An example of simple linear regression can be found in Altman et al. 2015 “Simple Linear Regression” Nature Methods 12, 999 - 1000, which is incorporated herein by reference.

[0218] In some embodiments, multiple linear regression is used to model the relationship between a dependent variable and multiple independent variables and can thus be used as a prediction model in the present disclosure. An example of multiple linear regression can be found in Sousa et al. 2007 “Multiple linear regression and artificial neural networks based on principal components to predict ozone concentration” Environ Model & Soft 22(1), 97 - 103, which is incorporated herein by reference. In some embodiments, multivariate linear regression is used to model the relationship between multiple dependent variables and any number of independent variables. A non - limiting example of multivariate linear regression can be found in Wang et al. 2016 “Discriminative Feature Extraction via Multivariate Linear Regression for SSVEP - Based BCI” IEEE Transactions on Neural Systems and Rehabilitation Engineering 24(5), 532 - 541, which is incorporated herein by reference.

[0219] Naive Bayes algorithm. The Naive Bayes classifier (algorithm) is a family of “probabilistic classifiers” based on applying Bayes' theorem with a strong (naive) independence assumption between features. In some embodiments, they are combined with kernel density estimation. See Hastie, Trevor, 2001, The elements of statistical learning: data mining, inference, and prediction, Tibshirani, Robert, Friedman, J.H. (Jerome H.), New York: Springer, which is incorporated herein by reference.

[0220] In some embodiments, training an initially trained prediction model using at least i) a subset of subjects as independent variables of the prediction model and ii) a corresponding subset of target outcomes as dependent variables of the prediction model further includes iii) using at least one target subject as an independent variable of the prediction model to update the prediction model to an updated trained state.

[0221] Blocks 228-230. Referring to block 228 of FIG. 2B, the method proceeds by applying an updated trained state prediction model (e.g., a retrained prediction model) to a complete plurality of subjects, thereby obtaining instances of a plurality of prediction results. Referring to block 230, in some embodiments, the instances of the plurality of prediction results include the respective prediction results of each subject among the plurality of subjects. In this way, a balance is achieved between the high computational burden of the target model and the corresponding performance improvement, and the low computational burden of the prediction model and the corresponding inferior performance. Use the target model to obtain target results for only a subset of the subjects, thereby forming a training set for training the prediction model. This training set is likely to be more accurate due to the performance of the more computationally intensive target model, as well as the fact that it exploits the interaction between at least one target and the subject. For example, in some embodiments, the target is an enzyme having an active site, and the target model scores the interaction between each subject of the subset of subjects and the target. Then, use the training set to train the prediction model. Thus, in a typical embodiment, the prediction model is trained using the training set, the training set includes the target model scores of each subject of the subset of subjects, and the chemical data is provided for each such subject in the subject dataset, whereby the prediction model can predict the scores of the target model without using the target (e.g., without docking the subject to the target). Then, apply the prediction model thus trained to a complete plurality of subjects to obtain instances of a plurality of prediction results. The instances of the prediction results include the scores predicted by the trained prediction model to be the target model scores of each target among the complete plurality of targets. In this way, it helps to reduce the number of subjects in the subject dataset by fully utilizing the performance of the more computationally intensive target model where docking occurs simultaneously. Moreover, to fully utilize the efficiency of the prediction model to reduce the number of subjects in the subject dataset, obtain the test result of each subject.

[0222] Blocks 232-234. Referring to block 232 of FIG. 2B, the method proceeds by excluding a portion of the subjects from a plurality of subjects based at least in part on a plurality of instances of prediction results (e.g., according to any of the exclusion criteria described below). In some embodiments, for each respective subject of a subset of subjects from the plurality of subjects, the target model is applied to each respective subject and at least one target subject to obtain a corresponding target result, thereby obtaining a corresponding subset of target results (block 210), training a prediction model in an initial trained state (block 220), applying the updated trained state of the prediction model to the plurality of subjects, thereby obtaining a plurality of instances of prediction results (block 228), and excluding a portion of the subjects from the plurality of subjects based at least in part on the plurality of instances of prediction results (block 232) is an iterative process that is repeated several times (e.g., 2, 3, more than 3, more than 10, more than 15, etc.) as the subject of evaluation described in block 236 below. Each time the process is repeated (in each iteration), a portion of the subjects remaining in the plurality of subjects is removed from the plurality of subjects based at least in part on the latest instance of the plurality of prediction results from block 228.

[0223] Referring to block 234, in some embodiments, excluding involves: i) clustering a plurality of subjects, thereby assigning each subject among the plurality of subjects to a respective cluster among the plurality of clusters; and ii) excluding a subset of subjects from the plurality of subjects, at least in part based on the redundancy of the subjects in the individual clusters among the plurality of clusters (e.g., to ensure a diverse variety of chemical compounds among the plurality of subjects). In other words, in such embodiments, in each iteration of block 232, the remaining plurality of subjects are clustered. In some embodiments, this clustering is based on the feature vectors of the subjects as described above. In some embodiments, any of the clustering described in block 214 may be used to perform the clustering of block 234. In block 214, such clustering was performed to select a subset of subjects to use for the target model, whereas in block 234, the clustering is performed to permanently exclude subjects from the plurality of subjects. Considering an example where the clustering of block 234 clusters the subjects remaining in the plurality of subjects into Q clusters, Q is a positive integer greater than or equal to 2 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, greater than 10, greater than 20, greater than 30, greater than 100, etc.). In some such embodiments, the same number of subjects in each of these clusters are retained in the plurality of subjects, and all other subjects are deleted from the plurality of subjects. In this way, the subjects remaining in the plurality of subjects are balanced across all clusters.

[0224] The plurality of prediction results generated in step 232 represent scores predicted by the prediction model as to what the target model would call for the plurality of subjects.

[0225] When scoring is performed in a scheme where compounds with lower scores have better affinity for one or more target objects, it is interesting to remove those test subjects with high scores. Thus, in some alternative embodiments, clustering is not used and the exclusion of block 232 involves: i) ranking a plurality of test subjects based on instances of a plurality of prediction results; and ii) removing from the plurality of test subjects those test subjects among the plurality of test subjects that do not have corresponding prediction scores that meet a threshold cut-off (e.g., to ensure that the test subjects remaining among the plurality of test subjects have high prediction scores). In some embodiments, the threshold cut-off is an upper threshold percentage (e.g., the percentage of the plurality of test subjects ranked highest based on the plurality of prediction results). In some such embodiments, the upper threshold percentage represents the test subjects among the plurality of test subjects where the prediction results are in the top 90 percent, top 80 percent, top 75 percent, top 60 percent, top 50 percent, top 40 percent, top 30 percent, top 25 percent, top 20 percent, top 10 percent, or top 5 percent of the plurality of prediction results. In such embodiments, the corresponding lower percentage of test subjects is excluded from the plurality of test subjects for further consideration (e.g., thereby reducing the number of test subjects among the plurality of test subjects).

[0226] When scoring is performed in a scheme where compounds with higher scores have better affinity for one or more target objects, it is interesting to remove those test subjects with low scores. Thus, in some alternative embodiments, clustering is not used, and the exclusion of block 232 involves: i) ranking a plurality of test subjects based on instances of a plurality of prediction results; and ii) removing from the plurality of test subjects those test subjects among the plurality of test subjects that do not have corresponding prediction scores that meet a threshold cut-off (e.g., to ensure that the test subjects remaining among the plurality of test subjects have low prediction scores). In some such embodiments, the threshold cut-off is a lower threshold percentage (e.g., the percentage of the plurality of test subjects that are ranked lowest based on the plurality of prediction results). In some embodiments, the lower threshold percentage represents test subjects among the plurality of test subjects where the prediction results are in the lower 90 percent, lower 80 percent, lower 75 percent, lower 60 percent, lower 50 percent, lower 40 percent, lower 30 percent, lower 25 percent, lower 20 percent, lower 10 percent, or lower 5 percent of the plurality of prediction results. In such embodiments, the corresponding upper percentage of test subjects is excluded from the plurality of test subjects for further consideration (e.g., thereby reducing the number of test subjects among the plurality of test subjects).

[0227] In some embodiments, each instance of exclusion (e.g., in embodiments where the method repeatedly excludes a portion of test subjects from a plurality of test subjects) excludes from 1 / 10 to 9 / 10 of the test subjects among the plurality of test subjects in a particular iteration of block 232. In some embodiments, each instance of exclusion excludes more than 5 percent, more than 10 percent, more than 15 percent, more than 20 percent, or more than 25 percent of the test subjects present among the plurality of test subjects in a particular iteration of block 232.

[0228] In some embodiments, each instance of exclusion excludes 5 percent to 30 percent, 10 percent to 40 percent, 15 percent to 70 percent, 20 percent to 50 percent, 25 percent to 90 percent of the plurality of subjects in a particular iteration of block 232. In some embodiments, each instance of exclusion excludes one quarter to three quarters of the subjects among the plurality of subjects in a particular iteration of block 232. In some embodiments, each instance of exclusion excludes one quarter to one half of the subjects among the plurality of subjects in a particular iteration of block 232.

[0229] In some embodiments, each instance of exclusion (block 232) excludes a predetermined number (or portion) of subjects from a plurality of subjects. For example, in some embodiments, each instance of exclusion (block 232) excludes 5 percent of the subjects among the plurality of subjects in each instance of exclusion. In some embodiments, one or more instances of exclusion exclude different numbers (or portions) of subjects. For example, an initial instance of exclusion (block 232) may exclude a higher percentage of the plurality of subjects among the plurality of subjects in these initial instances of exclusion 232, while a subsequent instance of exclusion may exclude a lower percentage of the plurality of subjects among the plurality of subjects in these subsequent instances of exclusion 232. For example, the initial instance excludes 10 percent of the plurality of test compounds, while the subsequent instance excludes 5 percent of the plurality of test compounds. In another example, an initial instance of exclusion (block 232) may exclude a lower percentage of the plurality of subjects among the plurality of subjects in these initial instances of exclusion, while a subsequent instance of exclusion may exclude a higher percentage of the plurality of subjects among the plurality of subjects in these subsequent instances of exclusion 232. For example, in the initial instance of exclusion, 5 percent of the plurality of test compounds are excluded, while in the subsequent instance of exclusion 232, 10 percent of the plurality of test compounds are excluded.

[0230] Block 236. Referring to block 236 of FIG. 2C, the method proceeds by determining whether one or more predefined reduction criteria are met. If one or more predefined reduction criteria are not met, the method further includes the following. For each respective subject in an additional subset of subjects among a plurality of subjects, apply a target model to each respective subject and at least one target object to obtain a corresponding target result, thereby obtaining an additional subset of target results (i). The additional subset of subjects is selected at least in part on instances of a plurality of prediction results. Update the subset of subjects by incorporating the additional subset of subjects into a subset of subjects (e.g., a previous subset of subjects) (ii). Update the subset of target results by incorporating the additional subset of target results into the subset of target results (iii). Thus, as the method progressively repeats performing the target model, training the prediction model, and executing the prediction model, the subset of target results grows. After update (ii) and update (iii), modify the prediction model (iv) by applying the prediction model to at least 1) a subset of subjects as an independent variable and a corresponding subset of target results as a corresponding dependent variable, thereby providing an updated trained state of the prediction model. Applying (block 228), excluding (block 232), and determining (block 236) are repeated until one or more predefined reduction criteria are met.

[0231] In some embodiments, modifying the prediction model (iv) includes either retraining or training a new partially trained prediction model.

[0232] In some embodiments, when one or more predefined reduction criteria are met, the method further includes: i) clustering a plurality of subjects, thereby assigning each subject among the plurality of subjects to a cluster among the plurality of clusters; and ii) excluding one or more subjects from the plurality of subjects based at least in part on the redundancy of the subjects in an individual cluster among the plurality of clusters.

[0233] In some embodiments, clustering the plurality of subjects is performed as described with respect to block 212.

[0234] Referring to block 238, in some embodiments, applying (i) further includes forming an additional subset of subjects by selecting one or more subjects from the plurality of subjects (e.g., by selecting subjects from diverse clusters) based on an evaluation of one or more features selected from the plurality of feature vectors, as described above.

[0235] In some embodiments, the additional subset of subjects is the same size as, or a similar size to, the subset of subjects. In some embodiments, the additional subset of subjects is a different size than the subset of subjects. In some embodiments, the additional subset of subjects is a separate subset from the subset of subjects.

[0236] In some embodiments, an additional subset of the subjects includes at least 1,000 subjects, at least 5,000 subjects, at least 10,000 subjects, at least 25,000 subjects, at least 50,000 subjects, at least 75,000 subjects, at least 100,000 subjects, at least 250,000 subjects, at least 500,000 subjects, at least 750,000 subjects, at least 1 million subjects, at least 2 million subjects, at least 3 million subjects, at least 4 million subjects, at least 5 million subjects, at least 6 million subjects, at least 7 million subjects, at least 8 million subjects, at least 9 million subjects, or at least 10 million subjects.

[0237] In some embodiments, modifying the prediction model (iv) includes retraining the prediction model (e.g., re-running the training process with an updated subset of the subjects and potentially changing some parameters or hyperparameters of the prediction model). In some embodiments, modifying the prediction model (iv) includes training a new prediction model (e.g., replacing the previous prediction model).

[0238] In some embodiments, modifying (iv) further includes using at least 1) a subset of the subjects as an independent variable, 2) a corresponding subset of the target results as a corresponding dependent variable, and 3) at least one target object as an independent variable. In other words, in some embodiments, the prediction model docks the subjects to the target object in order to generate a prediction result trained against the target result of the target model, provided that the prediction model with docking is actually less computationally burdensome than the target model with concurrent associations.

[0239] Referring to block 240, in some embodiments, meeting one or more predefined reduction criteria includes correlating a plurality of prediction results with corresponding target results from a subset of the target results. For example, in some embodiments, one or more predefined reduction criteria are met when the correlation between the plurality of prediction results and the corresponding target results is 0.60 or greater, 0.65 or greater, 0.70 or greater, 0.75 or greater, 0.80 or greater, 0.85 or greater, or 0.90 or greater.

[0240] Referring to block 240, in some embodiments, meeting one or more predefined reduction criteria includes determining the average difference between a plurality of prediction results and the corresponding target results on an absolute scale or a normalized scale, and the one or more predefined reduction criteria are met when this average difference is less than a threshold amount. In such embodiments, the threshold amount is application-dependent.

[0241] In some embodiments, meeting one or more predefined reduction criteria includes determining that the number of subjects among a plurality of subjects is less than a threshold number of subjects. In some embodiments, one or more predefined reduction criteria require that the plurality of subjects have 30 or fewer subjects, 40 or fewer subjects, 50 or fewer subjects, 60 or fewer subjects, 70 or fewer subjects, 90 or fewer subjects, 100 or fewer subjects, 200 or fewer subjects, 300 or fewer subjects, 400 or fewer subjects, 500 or fewer subjects, 600 or fewer subjects, 700 or fewer subjects, 800 or fewer subjects, 900 or fewer subjects, or 1000 or fewer subjects.

[0242] In some embodiments, one or more predefined reduction criteria require that a plurality of subjects have 2 to 30 subjects, 4 to 40 subjects, 5 to 50 subjects, 6 to 60 subjects, 5 to 70 subjects, 10 to 90 subjects, 5 to 100 subjects, 20 to 200 subjects, 30 to 300 subjects, 40 to 400 subjects, 40 to 500 subjects, 40 to 600 subjects, or 50 to 700 subjects.

[0243] In some embodiments, meeting one or more predefined reduction criteria includes determining that the number of subjects among a plurality of subjects has been reduced by a threshold percentage of the number of subjects in the subject database. In some embodiments, one or more predefined reduction criteria require reducing a plurality of subjects by at least 10% of the subject database, at least 20% of the subject database, at least 30% of the subject database, at least 40% of the subject database, at least 50% of the subject database, at least 60% of the subject database, at least 70% of the subject database, at least 80% of the subject database, at least 90% of the subject database, at least 95% of the subject database, or at least 99% of the subject database.

[0244] In some embodiments, one or more predefined reduction criteria is a single reduction criterion. In some embodiments, one or more predefined reduction criteria is a single reduction criterion, and this single reduction criterion is any one of the reduction criteria described in this disclosure.

[0245] In some embodiments, one or more predefined reduction criteria is a combination of reduction criteria. In some embodiments, this combination of reduction criteria is any combination of the reduction criteria described in this disclosure.

[0246] Referring to block 242, in some embodiments, if one or more pre-defined reduction criteria are met, the method further includes applying a predictive model to a plurality of subjects and at least one target subject, thereby causing the predictive model to provide a respective score for each subject among the plurality of subjects (e.g., each score is for a respective subject and target subject). In some such embodiments, each respective score corresponds to an interaction between a respective subject and at least one target subject. In some embodiments, each score is used to characterize at least one target subject. In some embodiments, the score refers to a binding affinity (e.g., between one or more target subjects and respective subjects) described in U.S. Patent No. 10,002,312, entitled "Systems and Methods for Applying a Convolutional Network to Spatial Data", the entirety of which is incorporated herein by reference. In some embodiments, the interaction between a subject and a target subject is affected by distance, angle, atomic type, molecular charge and / or polarization, and surrounding stabilizing or destabilizing environmental factors.

[0247] In some alternative embodiments, if one or more pre-defined reduction criteria are met, the method applies the target model to the remaining plurality of test subjects and at least one target subject, thereby causing the target model to provide a respective target score for each of the remaining test subjects among the plurality of test subjects (e.g., each target score is for each test subject and target subject among one or more target subjects). In some such embodiments, each respective target score corresponds to an interaction between a respective test subject and at least one target subject. In some embodiments, each target score is used to characterize at least one target subject. In some embodiments, the target score refers to a binding affinity (e.g., between respective test subjects having one or more target subjects) as described in U.S. Patent No. 10,002,312, titled "Systems and Methods for Applying a Convolutional Network to Spatial Data", which is hereby incorporated by reference in its entirety. In some embodiments, the interaction between a test subject and a target subject is affected by distance, angle, atom type, molecular charge and / or polarization, and ambient stabilizing or destabilizing environmental factors.

[0248] Example 1 - Usage example. The following are sample usage examples provided for illustrative purposes only regarding some uses of some embodiments of the present invention. Other uses may be contemplated, and the examples provided below are non-limiting and may be subject to variations, omissions, or may include additional elements.

[0249] The following examples illustrate binding affinity predictions, although these examples may differ in whether the prediction is made for a single molecule, a set of iteratively modified molecules, or a series of iteratively modified molecules, whether the prediction is made for a single target or multiple targets, whether activity against the target is desired or avoided, and whether the critical amount is absolute activity or relative activity, or whether the molecule set or target set is specifically selected (e.g., for molecules, as existing drugs or pesticides, for proteins, as having known toxicity or side effects).

[0250] Hit discovery. Pharmaceutical companies spend millions of dollars screening compounds to discover new promising drug leads. They test large compound collections to find a few compounds that have any interaction with an interesting disease target. Unfortunately, wet-lab screening is subject to experimental error in addition to the cost and time to perform the assay experiments, and the collection of large screening collections poses significant challenges through storage constraints, storage stability, or chemical cost. Even the largest pharmaceutical companies have only hundreds of thousands to millions of compounds against hundreds of millions of commercially available molecules and billions of simulatable molecules.

[0251] A potentially more efficient alternative to physical experiments is virtual high-throughput screening. Just as physical simulations can be useful for aerospace engineers to evaluate possible wing designs before the models are physically tested, computational screening of molecules can focus on experimental testing of a small subset of likely molecules. This can reduce screening costs and time, reduce false negatives, improve success rates, and / or cover a wide chemical space.

[0252] In this application, a protein target may function as a target of interest. A large set of molecules may also be provided in the form of a subject dataset. For each subject remaining upon application of the disclosed method, the binding affinity for the protein target is predicted. The resulting scores can be used to rank the remaining molecules, and the molecules with the best scores are most likely to bind to the target protein. Optionally, the ranked list of molecules may be analyzed for clusters of similar molecules, and large clusters may be used as stronger predictors of molecular binding or molecules may be selected between clusters to ensure diversity in confirmatory experiments.

[0253] Off-target side effect prediction. Many drugs may be found to have side effects. In many cases, these side effects are due to interactions with biological pathways other than those responsible for the therapeutic effect of the drug. These off-target side effects can be unpleasant or dangerous and limit the patient population for which the drug can be used safely. Thus, off-target side effects are an important criterion for evaluating which drug candidates to further develop. Characterizing the interactions of a drug with many alternative biological targets is important, but such tests can be expensive and time-consuming to develop and perform. Computational prediction can make this process more efficient.

[0254] When applying embodiments of the present invention, a panel of biological targets associated with a significant biological response and / or side effect may be constructed. In that case, the system may be configured to predict binding to each protein in the panel by processing such proteins as targets of interest in turn. Strong activity against a particular target (i.e., as strong as a compound known to activate an off-target protein) can involve the molecule in side effects due to off-target effects.

[0255] Toxicity prediction. Toxicity prediction is a particularly important special case of off-target side effect prediction. Approximately half of drug candidates in late-stage clinical trials fail due to unacceptable toxicity. As part of the new drug approval process (and before a drug candidate becomes testable in humans), the FDA requires toxicity test data for a set of targets including cytochrome P450 liver enzymes (inhibition of which can lead to toxicity from drug-drug interactions) or hERG channels (binding to which can lead to QT prolongation leading to ventricular arrhythmias and other adverse cardiac effects).

[0256] In toxicity prediction, the system may be configured to bind off-target proteins to primary anti-targets (e.g., CYP450, hERG, or 5-HT 2B receptors). The binding affinity for drug candidates can then be predicted for each of these proteins by treating each of these proteins as a target (e.g., in separate independent runs). Optionally, the molecule may be analyzed to predict a set of metabolites (subsequent molecules produced by the body during metabolism / degradation of the original molecule), which can also be analyzed for binding to anti-targets. Molecules in question can be identified and modified to avoid toxicity, or the development of the molecular series can be stopped to avoid wasting additional resources.

[0257] Pesticide design. In addition to pharmaceutical applications, the pesticide industry uses binding prediction in the design of new pesticides. For example, one requirement for a pesticide is that it stops a single species of interest without adversely affecting other species. For ecological safety, one may desire to kill cockroaches without killing honey bees.

[0258] For this use, a user may input into the system a set of protein structures as one or more target objects from different species under consideration. A subset of the proteins can be designated as the proteins to be activated, while the remainder can be designated as the proteins for which the molecule should be inactive. As in previous usage examples, several sets of molecules (regardless of whether from an existing database or newly generated) are considered as test subjects for each target object, and the system returns molecules that have maximum effectiveness against the proteins in the first group while avoiding the second group.

[0259] Conclusion For the components, operations, or structures described herein as a single instance, multiple instances may be provided. Ultimately, the boundaries between various components, operations, and data stores are somewhat arbitrary, and particular operations are illustrated in the context of a particular exemplary configuration. Other assignments of functionality are envisioned and may be within the scope of the implementation. In general, the structures and functionality presented as separate components of an exemplary configuration may be implemented as a combined structure or component. Similarly, the structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements are within the scope of the implementation.

[0260] As used herein, the term "if" may be construed to mean "when" or "upon" or "in response to determining" or "in response to detecting", depending on the context. Similarly, depending on the context, the phrases "if it is determined" or "if [stated state or event] is detected" may be construed to mean "upon determining" or "in response to determining" or "(upon detecting the stated state or event)" or "(in response to detecting the stated state or event)".

[0261] The terms first, second, and the like may be used herein to describe various elements, but it should be understood that these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the present disclosure, a first object may be referred to as a second object, and similarly, a second object may be referred to as a first object. The first object and the second object are both objects, but they are not the same object.

[0262] The foregoing description has included example systems, methods, techniques, instruction sequences, and computing machine program products that embody example implementations. For purposes of explanation, numerous specific details have been set forth in order to provide an understanding of the various implementations of the inventive subject matter. It will be apparent to one skilled in the art, however, that the inventive subject matter may be practiced without these specific details. In general, well-known instruction instances, protocols, structures, and techniques have not been shown in detail.

[0263] The foregoing description has been presented for purposes of illustration and is described with respect to specific implementations. However, the above exemplary considerations are not intended to be exhaustive or to limit the implementations to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The implementations were chosen and described in order to best explain the principles and their practical application, thereby enabling one skilled in the art to best utilize the various implementations with various modifications as are suited to the particular use contemplated.

Claims

**Claim 1** A method for reducing the number of compounds in a plurality of compounds in a compound dataset to reduce the screening cost and time for screening a plurality of compounds for binding affinity to a protein disease target, which is an enzyme having an active site, comprising: A) obtaining the compound dataset in electronic form; B) for each compound in a subset of compounds from the plurality of compounds, applying a target model to the three-dimensional spatial coordinates of the respective compound and the protein disease target to obtain a corresponding target model score for the interaction between the respective compound and the protein disease target, thereby obtaining a corresponding subset of target model scores; C) training a prediction model in an initial trained state using at least i) the subset of compounds as an independent variable and ii) the corresponding subset of target model scores as a dependent variable, thereby updating the prediction model to an updated trained state; D) applying the prediction model in the updated trained state to the plurality of compounds, thereby obtaining a plurality of instances of prediction results, wherein the plurality of instances of prediction results include respective prediction scores for the interaction between each compound in the plurality of compounds and the protein disease target; E) excluding a portion of the compounds from the plurality of compounds based at least in part on the instances of the plurality of prediction scores; F) for each compound in an additional subset of compounds from the plurality of compounds, applying the target model to the three-dimensional spatial coordinates of the respective compound and the protein disease target to obtain a corresponding target model score for the interaction between the respective compound and the protein disease target, thereby obtaining an additional subset of target model scores, wherein the additional subset of compounds is selected at least in part on the instances of the plurality of prediction scores; G) updating the subset of compounds by incorporating the additional subset of compounds into the subset of compounds. H) updating a subset of the target model scores by incorporating an additional subset of the target model scores into the subset of the target model scores; I) after said updating G) and said updating H), applying the prediction model to at least 1) a subset of the compounds as independent variables of the prediction model and 2) a corresponding subset of the target model scores as corresponding dependent variables of the prediction model to modify the prediction model, thereby providing the updated and trained prediction model; J) repeating said applying D), excluding E), applying F), updating G), updating H), and modifying I) until one or more predefined reduction criteria are met, wherein the plurality of compounds includes at least 100 million compounds prior to the application of the instance of said excluding E); K) upon completion of said determining J), experimentally testing the selection of the plurality of compounds using a wet-lab binding affinity assay to determine which of the selections of compounds have binding affinity to the protein disease target. A method comprising.

2. wherein the target model exhibits a first computational complexity; wherein the prediction model exhibits a second computational complexity; The method according to claim 1, wherein the second computational complexity is less than the first computational complexity.

3. The method according to claim 1 or 2, wherein the compound dataset includes a plurality of feature vectors, each feature vector being for a respective compound in the plurality of compounds.

4. The method according to any one of claims 1 to 3, wherein said applying B) further comprises randomly selecting one or more compounds from the plurality of compounds to form the subset of compounds.

5. The method according to claim 3, wherein said applying B) further comprises selecting one or more compounds from the plurality of compounds of the subset of compounds based on an evaluation of one or more features selected from the plurality of feature vectors, or each feature vector in the plurality of feature vectors is a one-dimensional vector.

6. The method according to claim 3 or 4, wherein the applying F) further comprises forming an additional subset of the compounds by selecting one or more compounds from the plurality of compounds based on an evaluation of one or more features selected from the plurality of feature vectors.

7. Satisfying the one or more predefined reduction criteria includes (i) comparing each prediction result among the plurality of prediction results with a corresponding target model score from a subset of the target model scores, or (ii) determining that the number of the compounds among the plurality of compounds is less than a threshold number of compounds. The method according to any one of claims 1 to 6.

8. The target model is a convolutional neural network, or the prediction model includes a random forest tree, a random forest including a plurality of multiple additive decision trees, a neural network, a graph neural network, a dense neural network, principal component analysis, nearest neighbor analysis, linear discriminant analysis, quadratic discriminant analysis, support vector machine, evolutionary method, projection pursuit, linear regression, naive Bayes algorithm, multinomial logistic regression algorithm, or an ensemble thereof. The method according to any one of claims 1 to 7.

9.

10. Before applying the instance of the excluding E), the plurality of compounds include at least 500 million compounds, at least 1 billion compounds, at least 2 billion compounds, at least 3 billion compounds, at least 4 billion compounds, at least 5 billion compounds, at least 6 billion compounds, at least 7 billion compounds, at least 8 billion compounds, at least 9 billion compounds, at least 10 billion compounds, at least 11 billion compounds, at least 15 billion compounds, at least 20 billion compounds, at least 30 billion compounds, at least 40 billion compounds, at least 50 billion compounds, at least 60 billion compounds, at least 70 billion compounds, at least 80 billion compounds, at least 90 billion compounds, at least 100 billion compounds, or at least 110 billion compounds. The method according to any one of claims 1 to 9.

11. ​ ​ ​ The target model is a set of three-dimensional coordinates of the crystal structure of the protein disease target resolved at a resolution of 2.5 Å or higher or the crystal structure of the protein disease target resolved at a resolution of 3.3 Å or higher {x 1 ,..., x N}, which is applied to the three-dimensional spatial coordinates of each of the compounds and the protein disease target, or the target model is based on the spatial coordinates of an ensemble of the three-dimensional coordinates of the protein disease target determined by nuclear magnetic resonance, neutron diffraction, or cryo-electron microscopy, and is applied to the protein disease target. The method according to any one of claims 1 to 8. ​ ​ ​ A subset of said compounds comprises at least 1,000 compounds, at least 5,000 compounds, at least 10,000 compounds, at least 25,000 compounds, at least 50,000 compounds, at least 75,000 compounds, at least 100,000 compounds, at least 250,000 compounds, at least 500,000 compounds, at least 750,000 compounds, at least 1 million compounds, at least 2 million compounds, at least 3 million compounds, at least 4 million compounds, at least 5 million compounds, at least 6 million compounds, at least 7 million compounds, at least 8 million compounds, at least 9 million compounds, or at least 10 million compounds, and optionally, an additional subset of said compounds comprises at least 1,000 compounds, at least 5,000 compounds, at least 10,000 compounds, at least 25,000 compounds, at least 50,000 compounds, at least 75,000 compounds, at least 100,000 compounds, at least 250,000 compounds, at least 500,000 compounds, at least 750,000 compounds, at least 1 million compounds, at least 2 million compounds, at least 3 million compounds, at least 4 million compounds, at least 5 million compounds, at least 6 million compounds, at least 7 million compounds, at least 8 million compounds, at least 9 million compounds, or at least 10 million compounds, and optionally, the additional subset of said compounds is distinct from the subset of said compounds. The method according to any one of claims 1 to 10.

12. Modifying the prediction model (I) includes retraining the prediction model, or training (C) includes, in addition to using at least i) a subset of the compounds as a plurality of independent variables of the prediction model and ii) a corresponding subset of the target model scores as a plurality of dependent variables of the prediction model, further including iii) using the protein disease target as an independent variable of the prediction model. The method according to claim 1.

13. The method according to any one of claims 1 to 12, wherein the modifying (I) further comprises, in addition to using at least 1) a subset of the compounds as independent variables and 2) a corresponding subset of the target model scores as the corresponding dependent variables of the prediction model, 3) using the protein disease target as an independent variable.

14. When the one or more predefined reduction criteria are met, the method comprises i) clustering the plurality of compounds, thereby assigning each compound in the plurality of compounds to a cluster in a plurality of clusters; and ii) excluding one or more compounds from the plurality of compounds, at least in part based on the redundancy of the compounds in the individual clusters among the plurality of clusters, or the method comprises i) clustering the plurality of compounds, thereby assigning each compound in the plurality of compounds to a respective cluster in a plurality of clusters; and ii) selecting a subset of the compounds from the plurality of compounds by selecting a subset of the compounds from the plurality of compounds, at least in part based on the redundancy of the compounds in the individual clusters among the plurality of clusters. The method according to any one of claims 1 to 13.

15. When the one or more predefined reduction criteria are met, the method further comprises applying the prediction model to the plurality of compounds and the protein disease target, thereby causing the prediction model to provide respective interaction scores for each compound in the plurality of compounds, optionally, each respective interaction score corresponding to an interaction between a respective compound and the protein disease target, and optionally, using each respective interaction score to characterize the protein disease target. The method according to any one of claims 1 to 14.

16. The excluding (E) comprises i) clustering the plurality of compounds, thereby assigning each compound in the plurality of compounds to a respective cluster in a plurality of clusters; and ii) excluding a subset of compounds from the plurality of compounds, at least in part, based on the redundancy of the compounds of the individual clusters among the plurality of clusters, and optionally, clustering the plurality of compounds is performed using a density-based spatial clustering algorithm, a divisive clustering algorithm, an agglomerative clustering algorithm, a k-means clustering algorithm, a supervised clustering algorithm, or an ensemble thereof, the method according to claim 1.

17. wherein said excluding (E) comprises ranking the plurality of compounds based on the instances of the plurality of prediction results; deleting from the plurality of compounds those compounds among the plurality of compounds that do not have corresponding prediction results that meet a threshold cut-off, and optionally, the threshold cut-off is a top threshold percentage, and optionally, the top threshold percentage is the top 90 percent, top 80 percent, top 75 percent, top 60 percent, or top 50 percent of the plurality of prediction results, the method according to claim 1.

18. each instance of said excluding (E) excludes from 1 / 10 to 9 / 10, or from 1 / 4 to 3 / 4 of the compounds among the plurality of compounds, the method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Data set selecting device and experiment designing system

    JP2007304782A

  • Systems and methods for applying convolutional networks to spatial data

    JP2019501433A

  • Method for the synthesis of DNA conjugates by micellar catalysis

    WO2018149863A1

  • Mass spectrometry distinguishable synthetic compounds, libraries, and methods thereof

    WO2019050504A1