System and method for screening compound in silico

By iteratively applying a target and predictive model with varying complexities to a large chemical compound dataset, the method addresses the limitations of current in silico drug discovery methods, effectively reducing subjects and enhancing drug discovery efficiency.

JP2025143257APending Publication Date: 2025-10-01ATOMWISE INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025090729
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-03
Filing Date
2025-05-30
Publication Date
2025-10-01

AI Technical Summary

Technical Problem

Current in silico drug discovery methods are limited by computational complexity and database size, restricting the ability to discover potential drugs for new diseases due to the evaluation of small, pre-filtered molecular databases.

Method used

A method involving a target model and a predictive model with varying computational complexities is applied to a large chemical compound dataset, iteratively reducing the number of subjects based on predicted outcomes until predefined criteria are met, using techniques like clustering and model updating to refine the dataset.

Benefits of technology

This approach enables the evaluation of large chemical compound databases, significantly reducing the number of subjects while maintaining accuracy, thereby enhancing the discovery of potential drugs and reducing computational effort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025143257000001_ABST
    Figure 2025143257000001_ABST
Patent Text Reader

Abstract

To provide a system and method for reducing the number of test objects in a test dataset with different computational complexities.SOLUTION: A system 100 performs the following: a target model with a first computational complexity is applied to a subset of test objects from a test object dataset and a target object, thereby obtaining a subset of target results; and a predictive model with a second computational complexity is trained using the subset of test objects and the subset of target results. The predictive model is applied to the plurality of test objects, thereby obtaining a plurality of predictive results. A portion of the test objects are eliminated from the plurality of test objects based at least in part on the plurality of predictive results. The system also determines whether predefined reduction criteria are satisfied. When the predefined reduction criteria are not satisfied, an additional subset of test objects and target results are obtained, and the method is repeated.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 910,068, entitled "Systems and Methods for Screening Compounds In Silico," filed October 3, 2019, which is incorporated herein by reference.

[0002] This document generally relates to techniques for dataset reduction by using multiple computational models with different computational complexities. [Background technology]

[0003] The need to diversify molecular scaffolds to increase the chances of success in drug discovery has been called a move away from "flatland," or the reliance on synthetic methods to build flat molecules. Another way to explore the unexplored potential of the molecular universe is to find ways to reveal what lies in the shadows. By some estimates, at least 10 60 It is said that there is a possibility of developing different drug-like molecules: novandecylions. One approach to unlocking this unexplored chemical space is to explore ultra-large virtual libraries, i.e., libraries of compounds that do not need to be synthesized but whose molecular properties can be inferred from their calculated molecular structures.

[0004] The application of classifiers, such as deep learning neural networks, can be used to generate novel insights from large amounts of data, such as these virtual libraries. Indeed, lead identification and optimization in drug discovery, support for patient recruitment in clinical trials, medical image analysis, biomarker identification, drug efficacy analysis, drug adherence assessment, sequencing data analysis, virtual screening, molecular profiling, metabolic data analysis, electronic medical record analysis and medical device data evaluation, off-target side effect prediction, toxicity prediction, efficacy optimization, drug repurposing, drug resistance prediction, personalized medicine, drug trial design, pesticide design, materials science, and simulation are all examples of applications in which the use of classifiers, such as deep learning-based solutions, is being explored. Specifically, in healthcare, the American Recovery and Reinvestment Act of 2009 and the Precision Medicine Initiative of 2015 have broadly endorsed the value of medical data in healthcare. Thanks to several such initiatives, the volume of medical big data is expected to grow approximately 50-fold by 2020, reaching 25,000 petabytes. See, for example, Roots Analysis, February 22, 2017, "Deep Learning in Drug Discovery and Diagnostics, 2017-2035," available online at rootsanalysis.com.

[0005] With advances in drug repurposing and preclinical research, the application of classifiers to drug discovery presents an opportunity to significantly improve the drug discovery process and, therefore, patient outcomes throughout the healthcare system. See, for example, Rifaioglu et al., 2018, “Recent applications of deep learning and machine intelligence on in silico drug discovery: methods, tools, and databases,” Briefings in Bioinform 1-35, and Lavecchia, 2015, “Machine-learning approaches in drug discovery: methods and applications,” Drug Discovery Today 20(3), 318-331. In silico drug discovery methods are a particularly valuable application of classifiers because they have the potential to reduce the time and cost of drug development. Currently, the average cost of developing a new drug for human use is estimated to be well over $2 billion. See, for example, DiMasi et al., 2016, J Health Econ 47, 20-33. In addition, the U.S. federal government, largely through NIH funding, spent over $100 billion on primary basic research that contributed to all 210 new drugs approved by the FDA between 2010 and 2016. See Cleary et al., 2018, “Contributions of NIH funding to new drug approvals 2010-2016,” PNAS 115(10), 2329-2334. Thus, computational methods for discovering, or at least screening, lead compounds (e.g., in databases of known and / or FDA-approved chemicals) have the potential to revolutionize drug discovery and development.

[0006] There are many examples of computational approaches aiding drug discovery. The discovery of polypharmacology (e.g., the understanding that many drugs can and do bind to more than one molecular target) has opened up the field of repurposing already approved drugs for diseases that lack treatments. See, for example, Hopkins, 2009, “Predicting promiscuity,” Nature 462, 167-168 and Keiser et al., 2007, “Relating protein pharmacology by ligand chemistry,” Nat Biotechnol 25(2), 197-206. In silico drug discovery has already yielded potential treatments for diseases ranging from Zika to Chagas disease. See, for example, Ramarack et al., 2017, "Zika virus NS5 protein potential inhibitors: an enhanced in silico approach in drug discovery," J Biomol Structure and Dynamics 36(5), 1118-1133; Castillo-Garit et al., 2012, "Identification in silico and in vitro of Novel Trypanosomicidal Drug-Like Compounds," Chem Biol and Drug Des 80, 38-45; and Raj et al., 2015, "Flavonoids as Multi-target Inhibitors for Proteins associated with Ebola Virus," Interdisip Sci Comput Life Sci 7, 1-10. However, one drawback of many of the methods currently used for drug discovery, including virtual library evaluation, is their computational complexity.

[0007] In particular, many in silico drug discovery methods are primarily applicable to pre-filtered, size-limited molecular databases. See, for example, Macalino et al., 2018, “Evolution of in Silico Strategies for Protein-Protein Interaction Drug Discovery,” Molecules 23, 1963, and Lionata et al., 2014, “Structure-Based Virtual Screening for Drug Discovery: Principles, Applications and Recent Advances,” Curr Top Med Chem 14(16):1923-1938. In particular, datasets are typically limited to at least a few million compounds. See Ramsundar et al., 2015, “Massively Multitask Networks for Drug Discovery,” arXiv:1502.02072. Database size limitations impose corresponding limitations on the ability to discover or screen potential drugs to treat new diseases.

[0008] Given the importance of identifying promising lead compounds, there is a need in the art for improved computational methods of drug discovery that allow for the evaluation of large libraries of compounds. [Prior art documents] [Non-patent literature]

[0009] [Non-Patent Document 1] Roots Analysis,February 22,2017,“Deep Learning in Drug Discovery and Diagnostics,2017-2035” [Non-patent document 2] Rifaioglu et al.,2018,“Recent applications of deep learning and machine intelligence on in silico drug discovery:methods, tools and databases,”Briefings in Bioinform 1-35 [Non-patent document 3] Lavecchia, 2015, “Machine-learning approaches in drug discovery:methods and applications,” Drug Discovery Today 20(3), 318-331 [Non-patent document 4] DiMasi et al.,2016,J Health Econ 47,20-33 [Non-Patent Document 5] Cleary et al., 2018, “Contributions of NIH funding to new drug approvals 2010-2016,” PNAS 115(10), 2329-2334 [Non-patent document 6] Hopkins, 2009, “Predicting promiscuity,” Nature 462, 167-168 [Non-Patent Document 7] Keiser et al., 2007, “Relating protein pharmacology by ligand chemistry,” Nat Biotechnol 25(2), 197-206 [Non-patent document 8] Ramarack et al., 2017, “Zika virus NS5 protein potential inhibitors: an enhanced in silico approach in drug discovery,” J Biomol Structure and Dynamics 36(5), 1118-1133 [Non-Patent Document 9] Castillo-Garit et al., 2012, “Identification in silico and in vitro of Novel Trypanosomicidal Drug-Like Compounds,” Chem Biol and Drug Des 80, 38-45 [Non-Patent Document 10] Raj et al.2015 “Flavonoids as Multi-target Inhibitors for Proteins associated with Ebola Virus,” Interdisip Sci Comput Life Sci 7,1-10 [Non-Patent Document 11] Macalino et al., 2018, “Evolution of in Silico Strategies for Protein-Protein Interaction Drug Discovery,” Molecules 23, 1963 [Non-Patent Document 12] Lionata et al., 2014, “Structure-Based Virtual Screening for Drug Discovery: Principles, Applications and Recent Advances,” Curr Top Med Chem 14(16):1923-1938 [Non-Patent Document 13] Ramsundar et al., 2015, “Massively Multitask Networks for Drug Discovery,” arXiv:1502.02072 Summary of the Invention

[0010] The present disclosure addresses the deficiencies identified in the background by providing methods for the evaluation of large chemical compound databases.

[0011] In one aspect of the present disclosure, a method for reducing the number of subjects among a plurality of subjects in a subject dataset is provided, the method including obtaining the subject dataset in electronic form.

[0012] The method further includes, for each respective subject of a subset of subjects from the plurality of subjects, applying a target model to the respective subject and at least one target object to obtain a corresponding target result, thereby obtaining a corresponding subset of target results.

[0013] The method further trains the predictive model of the initial trained state using at least i) a subset of the test subjects as independent variables of the predictive model, and ii) a corresponding subset of the target outcomes as dependent variables of the predictive model, thereby updating the predictive model to an updated trained state.

[0014] The method further applies the updated trained state predictive model to a plurality of test subjects, thereby obtaining a plurality of instances of predicted results.

[0015] The method further excludes a portion of the subjects from the plurality of subjects based at least in part on instances of the plurality of predicted outcomes.

[0016] The method further includes determining whether one or more predefined reduction criteria are met. If the one or more predefined reduction criteria are not met, the method further includes (i) for each respective subject of an additional subset of subjects from the plurality of subjects, applying the target model to each subject and at least one target subject to obtain a corresponding target outcome, thereby obtaining an additional subset of target outcomes. The additional subset of subjects is selected at least in part based on the instances of the plurality of predicted outcomes. The method further includes (ii) updating the subset of subjects by incorporating the additional subset of subjects into the subset of subject subjects; (iii) updating the subset of target outcomes by incorporating the additional subset of target outcomes into the subset of target outcomes; and (iv) after updating (ii) and (iii), modifying the predictive model by applying it to at least 1) the subset of subjects as an independent variable and 2) the corresponding subset of target outcomes as a corresponding dependent variable, thereby providing an updated trained state predictive model. The method then iteratively applies the updated trained predictive model to a plurality of subjects, thereby obtaining a plurality of instances of predicted outcomes, and further eliminates a portion of the subjects from the plurality of subjects based at least in part on the plurality of instances of predicted outcomes until one or more predefined reduction criteria are met.

[0017] In some embodiments, the target model exhibits a first computational complexity when evaluating the subject, and the predictive model exhibits a second computational complexity when evaluating the subject, the second computational complexity being less than the first computational complexity. In some embodiments, the target model is at least 3 times, at least 5 times, or at least 100 times more computationally complex than the predictive model.

[0018] In some embodiments, the subject dataset includes a plurality of feature vectors (e.g., protein fingerprints, computational properties, and / or graph descriptors). In some embodiments, each feature vector is for a respective subject in the plurality of subjects, and each feature vector in the plurality of feature vectors has the same size. In some embodiments, each feature vector in the plurality of feature vectors is a one-dimensional vector.

[0019] In some embodiments, for each respective subject of a subset of subjects from the plurality of subjects, applying a target model to the respective subject and at least one target subject to obtain a corresponding target outcome, thereby obtaining a corresponding subset of target outcomes, further comprises randomly selecting one or more subjects from the plurality of subjects to form the subset of subjects.

[0020] In some embodiments, for each respective subject of a subset of subjects from the plurality of subjects, applying a target model to the respective subject and at least one target object to obtain a corresponding target outcome, thereby obtaining a corresponding subset of target outcomes, further comprises selecting one or more subjects from the plurality of subjects of the subset of subjects based on evaluation of one or more features selected from the plurality of feature vectors. In some embodiments, the selection is based on clustering (e.g., of the plurality of subjects).

[0021] In some embodiments, satisfying the one or more predefined reduction criteria includes comparing each predicted outcome in the plurality of predicted outcomes to a corresponding target outcome from a subset of target outcomes, hi some embodiments, the one or more predefined reduction criteria are satisfied when a difference between the training outcome and the target outcome is below a predetermined threshold.

[0022] In some embodiments, meeting one or more predefined reduction criteria comprises determining that the number of subjects in the plurality of subjects has fallen below a threshold number of subjects.

[0023] In some embodiments, the target model is a convolutional neural network.

[0024] In some embodiments, the predictive model comprises a random forest tree, a random forest including multiple multi-additive decision trees, a neural network, a graph neural network, a dense neural network, principal component analysis, nearest neighbor analysis, linear discriminant analysis, quadratic discriminant analysis, support vector machine, evolutionary method, projection pursuit, linear regression, a naive Bayes algorithm, a multi-category logistic regression algorithm, or an ensemble thereof.

[0025] In some embodiments, the at least one target entity is a single entity, and the single entity is a polymer. In some embodiments, the polymer comprises an active site. In some embodiments, the polymer is an assembly of a protein, a polypeptide, a polynucleic acid, a polyribonucleic acid, a polysaccharide, or any combination thereof.

[0026] In some embodiments, the plurality of subjects comprises at least 100 million subjects, at least 500 million subjects, at least 1 billion subjects, at least 2 billion subjects, at least 3 billion subjects, at least 4 billion subjects, at least 5 billion subjects, at least 6 billion subjects, at least 7 billion subjects, at least 8 billion subjects, at least 9 billion subjects, or at least 100 million subjects, prior to application of the instance of excluding a portion of the subjects from the plurality of subjects. at least 10 billion subjects, at least 11 billion subjects, at least 15 billion subjects, at least 20 billion subjects, at least 30 billion subjects, at least 40 billion subjects, at least 50 billion subjects, at least 60 billion subjects, at least 70 billion subjects, at least 80 billion subjects, at least 90 billion subjects, at least 100 billion subjects, or at least 110 billion subjects.

[0027] In some embodiments, the one or more predefined reduction criteria require that the plurality of subjects (e.g., after one or more instances of excluding a portion of subjects from the plurality of subjects) have 30 or fewer subjects, 40 or fewer subjects, 50 or fewer subjects, 60 or fewer subjects, 70 or fewer subjects, 90 or fewer subjects, 100 or fewer subjects, 200 or fewer subjects, 300 or fewer subjects, 400 or fewer subjects, 500 or fewer subjects, 600 or fewer subjects, 700 or fewer subjects, 800 or fewer subjects, 900 or fewer subjects, or 1000 or fewer subjects.

[0028] In some embodiments, each subject in the plurality of subjects is a chemical compound.

[0029] In some embodiments, the initial trained state predictive model includes an untrained or partially trained classifier, hi some embodiments, the updated trained state predictive model includes an untrained or partially trained classifier that is distinct from the initial trained state predictive model.

[0030] In some embodiments, the subset of subjects and / or the additional subset of subjects comprises at least 1,000 subjects, at least 5,000 subjects, at least 10,000 subjects, at least 25,000 subjects, at least 50,000 subjects, at least 75,000 subjects, at least 100,000 subjects, at least 250,000 subjects, at least 500,000 subjects, at least 750,000 subjects, at least 1 million subjects, at least 2 million subjects, at least 3 million subjects, at least 4 million subjects, at least 5 million subjects, at least 6 million subjects, at least 7 million subjects, at least 8 million subjects, at least 9 million subjects, or at least 10 million subjects. In some embodiments, the additional subset of subjects is distinct from the subset of subjects.

[0031] In some embodiments, training the initial, trained state predictive model using at least i) a subset of the subject subjects as a plurality of independent variables (of the predictive model) and ii) a corresponding subset of the target outcomes as a plurality of dependent variables (of the predictive model) further comprises iii) using at least one target subject as an independent variable (of the predictive model).

[0032] In some embodiments, the at least one target object comprises at least two target objects, at least three target objects, at least four target objects, at least five target objects, or at least six target objects.

[0033] In some embodiments, after updating (ii) and updating (iii), modifying the predictive model by applying the predictive model (iv) further comprises, in addition to using at least 1) a subset of the subject subjects as independent variables and 2) a corresponding subset of the target outcomes as corresponding dependent variables, 3) using at least one target subject as an independent variable.

[0034] In some embodiments, if one or more predefined reduction criteria are met, the method further includes clustering the plurality of subjects, thereby assigning each subject in the plurality of subjects to a cluster in the plurality of clusters, and eliminating one or more subjects from the plurality of subjects based at least in part on redundancy of the subjects in individual clusters in the plurality of clusters.

[0035] In some embodiments, the method further includes selecting a subset of subjects from the plurality of subjects by clustering the plurality of subjects, thereby assigning each subject in the plurality of subjects to a respective cluster in the plurality of clusters, and selecting the subset of subjects from the plurality of subjects based at least in part on redundancy of subjects in individual clusters in the plurality of clusters.

[0036] In some embodiments, if one or more predefined reduction criteria are met, the method further includes applying the plurality of subject subjects and at least one target subject to the predictive model, thereby causing the predictive model to provide a respective predicted outcome for each subject in the plurality of subject subjects. In some embodiments, each respective predicted outcome is a respective predicted outcome for each subject and at least one target subject (e.g., IC 50 , E.C. 50 , Kd, ​​or KI). In some embodiments, each respective prediction score is used to characterize at least one target object.

[0037] In some embodiments, eliminating a portion of the subjects from the plurality of subjects based at least in part on instances of the plurality of predicted outcomes comprises: i) clustering the plurality of subjects, thereby assigning each subject in the plurality of subjects to a respective cluster in the plurality of clusters; and ii) eliminating a subset of the subjects from the plurality of subjects based at least in part on redundancy of subjects in individual clusters in the plurality of clusters.

[0038] In some embodiments, clustering of the plurality of subjects is performed using a density-based spatial clustering algorithm, a partitional clustering algorithm, an agglomerative clustering algorithm, a k-means clustering algorithm, a supervised clustering algorithm, or an ensemble thereof.

[0039] In some embodiments, eliminating a portion of the subjects from the plurality of subjects based at least in part on instances of the plurality of predicted outcomes comprises: i) ranking the plurality of subjects based on instances of the plurality of predicted outcomes; and ii) removing from the plurality of subjects those subjects in the plurality of subjects that do not have a corresponding interaction score that meets a threshold cutoff.

[0040] In some embodiments, the threshold cutoff is an upper threshold percentage, hi some embodiments, the upper threshold percentage is the top 90 percent, top 80 percent, top 75 percent, top 60 percent, or top 50 percent of the plurality of predicted outcomes.

[0041] In some embodiments, each instance of excluding a portion of the subjects from the plurality of subjects based at least in part on the instances of the plurality of predicted outcomes excludes between one-tenth and nine-tenths of the subjects in the plurality of subjects. In some embodiments, each instance of excluding excludes between one-quarter and three-quarters of the subjects in the plurality of subjects.

[0042] Another aspect of the present disclosure provides a computing system including at least one processor and a memory storing at least one program executed by the at least one processor, the at least one program including instructions for reducing a number of subjects among a plurality of subjects in a subject dataset by any of the methods disclosed above.

[0043] Yet another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing at least one program for reducing the number of subjects among a plurality of subjects in a subject dataset, the at least one program being configured to be executed by a computer, the at least one program including instructions for performing any of the methods disclosed above.

[0044] As disclosed herein, any embodiment disclosed herein may be applied to any other aspect, where applicable. Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only exemplary embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.

[0045] Incorporation by Reference All publications, patents, and patent applications mentioned herein are incorporated by reference in their entirety as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In the event of a conflict between a term in this specification and a term in an incorporated reference, the term in this specification will control. [Brief explanation of the drawings]

[0046] Implementations disclosed herein are illustrated by way of example, and not by way of limitation, in the accompanying drawings. The description and drawings are for illustrative purposes and as an aid to understanding only, and are not intended as a definition of the limits of the disclosed systems and methods. Like reference numerals refer to corresponding parts throughout the drawings.

[0047] [Figure 1] FIG. 1 is a block diagram illustrating an example of a computing system according to some embodiments of the present disclosure. [Figure 2A] 1 illustrates generally an example flowchart of a method for reducing the number of subjects among a plurality of subjects in a subject dataset, according to some embodiments of the present disclosure. [Figure 2B] 1 illustrates generally an example flowchart of a method for reducing the number of subjects among a plurality of subjects in a subject dataset, according to some embodiments of the present disclosure. [Figure 2C] 1 illustrates generally an example flowchart of a method for reducing the number of subjects among a plurality of subjects in a subject dataset, according to some embodiments of the present disclosure. [Figure 3] 1 illustrates an example of evaluating a compound library according to some embodiments of the present disclosure. [Figure 4] 1A-1C are schematic diagrams of an exemplary subject in two different poses relative to a target object, according to an embodiment of the present disclosure. [Figure 5] FIG. 1 is a schematic diagram of a geometric representation of input features in the form of a three-dimensional grid of voxels, according to an embodiment of the present disclosure. [Figure 6] FIG. 1 is a diagram of two subjects encoded onto a two-dimensional grid of voxels, according to an embodiment of the present disclosure. [Figure 7] FIG. 1 is a diagram of two subjects encoded onto a two-dimensional grid of voxels, according to an embodiment of the present disclosure. [Figure 8] FIG. 8 is a diagram of the visualization of FIG. 7 with voxels numbered, according to an embodiment of the present disclosure. [Figure 9]FIG. 1 is a schematic diagram of a geometric representation of input features in the form of coordinate locations of atom centers, according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a schematic diagram of the coordinate locations of FIG. 9 with a range of positions, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0048] The computational effort required for drug discovery has increased simultaneously with the expansion of drug datasets in size and complexity. In particular, highly accurate models of target molecules have enabled the detection of additional test compounds (e.g., potential lead compounds) that may not have been considered using traditional drug discovery methods. The use of computational compound discovery further simplifies the laborious and time-consuming downstream process of combing the search space of potential drug databases (e.g., by determining which test compounds are most likely to have the desired effect given a particular target molecule) and conducting clinical trials to validate successful test compounds.

[0049] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0050] The implementations described herein provide various technical solutions for training a reference model for determining a subject's tumor fraction.

[0051] Definition. As used herein, the term "clustering" refers to various methods of optimizing the grouping of data points into one or more sets (e.g., clusters), where each data point in each set contains a higher degree of similarity to every other data point in the respective set than to data points not in the respective set. There are a wide variety of clustering algorithms suitable for evaluating different types of data. These algorithms include hierarchical models, centroid models, distributional models, density-based models, subspace models, graph-based models, and neural models. Each of these different models has different computational requirements (e.g., complexity) and is suitable for different data types. Applying two separate clustering models to the same data set often results in two different data groupings. In some embodiments, repeated application of a clustering model to a data set results in different data groupings each time.

[0052] As used herein, the term "feature vector" or "vector" is an enumerated list of elements, such as an array of elements, with each element having an assigned meaning. Thus, the term "feature vector" as used in this disclosure is interchangeable with the term "tensor." For ease of presentation, in some instances, vectors may be described as being one-dimensional, although the present disclosure is not so limited. Feature vectors of any dimension may be used in this disclosure, provided a description of what each element of the vector represents is defined.

[0053] As used herein, the term "polypeptide" refers to two or more amino acids or residues linked by a peptide bond. The terms "polypeptide" and "protein" are used interchangeably herein and include oligopeptides and peptides. "Amino acid," "residue," or "peptide" refers to any of the 20 standard structural units of proteins known in the art, including imino acids such as proline and hydroxyproline. Designations for amino acid isomers may include D, L, R, and S. The definition of amino acid includes unnatural amino acids. Thus, selenocysteine, pyrrolysine, lanthionine, 2-aminoisobutyric acid, γ-aminobutyric acid, dehydroalanine, ornithine, citrulline, and homocysteine ​​are all considered amino acids. Other variants or analogs of amino acids are known in the art. Thus, polypeptides can include synthetic peptidomimetic structures such as peptides. See Simon et al., 1992, Proceedings of the National Academy of Sciences USA, 89, 9367, which is incorporated herein by reference in its entirety. See also Chin et al., 2003, Science 301, 964, and Chin et al., 2003, Chemistry & Biology 10, 511, each of which is incorporated herein by reference in its entirety.

[0054] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in the detailed description of the invention and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. As used herein, the term "and / or" will also be understood to refer to and include any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "comprises" and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof, are used in either the detailed description and / or claims, such terms are intended to be inclusive in the same manner as the term "comprising."

[0055] Some aspects are described below with reference to example applications for illustration. It should be understood that numerous specific details, relationships, and methods are set forth to provide a thorough understanding of the features described herein. However, those skilled in the art will readily recognize that the features described herein can be implemented without one or more of the specific details or using other methods. The features described herein are not limited by the illustrated order of acts or events, as some acts may occur in different orders and / or concurrently with other acts or events. Furthermore, not all illustrated acts or events are required to implement a methodology in accordance with the features described herein.

[0056] Exemplary System Embodiments Details of an exemplary system will now be described in conjunction with Figure 1. Figure 1 is a block diagram illustrating a system 100 according to some implementations. In some implementations, the system 100 includes at least one or more processing units CPU 102 (also referred to as a processor), one or more network interfaces 104, an optional user interface 108 (e.g., having a display 106, input devices 110, etc.), memory 111, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 optionally include circuitry (sometimes referred to as a chipset) that interconnects and controls communications between the system components.

[0057] In some embodiments, each processing unit in the one or more processing units 102 is a single-core processor or a multi-core processor. In some embodiments, the one or more processing units 102 is a multi-core processor enabling parallel processing. In some embodiments, the one or more processing units 102 are multiple processors (single-core or multi-core) enabling parallel processing. In some embodiments, each of the one or more processing units 102 is configured to execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as memory 111. The instructions may be directed to the one or more processing units 102, which may subsequently program or otherwise configure the one or more processing units 102 to implement the methods of the present disclosure. Examples of operations performed by the one or more processing units 102 may include fetch, decode, execute, and writeback. The one or more processing units 102 may be part of a circuit, such as an integrated circuit. One or more other components of the system 100 may be included in the circuit. In some embodiments, the circuit is an application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA) architecture.

[0058] In some embodiments, display 106 is a touch-sensitive display, such as a touch-sensitive surface. In some embodiments, user interface 106 includes one or more soft keyboard embodiments. In some implementations, the soft keyboard embodiments include standard (QWERTY) and / or non-standard configurations of symbols on displayed icons. User interface 106 may be configured to provide a user with a graphical display of, for example, the results of reducing the number of subjects among multiple subjects in a subject dataset, an interaction score, or a predicted outcome. The user interface may allow user interaction with a particular task (e.g., reviewing and adjusting predefined reduction criteria).

[0059] Memory 111 may be non-persistent memory, persistent memory, or any combination thereof. Non-persistent memory typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, and flash memory, while persistent memory typically includes CD-ROM, digital versatile disk (DVD) or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, magnetic disk storage device, optical disk storage device, flash memory device, or other non-volatile solid-state storage device. Memory 111 optionally includes one or more storage devices located remotely from CPU 102. Memory 111 and the non-volatile memory devices within memory 111 include non-transitory computer-readable storage media. In some embodiments, memory 111 includes at least one non-transitory computer-readable storage medium for carrying and storing computer-executable instructions, which may be in the form of programs, modules, and data structures.

[0060] In some embodiments, as shown in FIG. 1, memory 111 stores the following programs, modules, and data structures, or a subset thereof: ● instructions, programs, data, or information associated with an operating system 116 (e.g., an embedded operating system such as iOS, ANDROID, DARWIN, RTXC, LINUX, UNIX, OS X, WINDOWS, or VxWorks), including various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management), and facilitating communication between various hardware and software components; • Instructions, programs, data, or information associated with an optional network communication module (or instructions) 118 for connecting the system 100 with other devices and / or a communication network; at least one target object 122, where in some embodiments the target object comprises a polymer; a subject database 122 including a plurality of subjects 124 (e.g., subjects 124-1, ..., 124-X), from which a subset 130 of subjects (e.g., subjects 124-A, ..., 124-B) is selected for analysis by a target model 150, and from which, optionally, one or more additional subsets of subjects (e.g., 140-1, ..., 140-Y) are selected and subsequently added to the subset 130, each subject 124 of the subset 130 having a corresponding target outcome 132 and a corresponding predicted outcome 134; a target model 150 having a first computational complexity 152, where application of the target model to a subset 130 of subjects results in a respective target outcome 132 for each subject 124 of the subject subset 130; and ●A predictive model 160 having a second computational complexity 162, which applies either an initial untrained state 164 or an updated untrained state 166 of the predictive model to a subject subset 130 to obtain respective prediction results 136 for each subject 132 in the subject subset 130.

[0061] In various implementations, one or more of the above-identified elements are stored in one or more of the aforementioned memory devices and correspond to sets of instructions for performing the functions described above. The above-identified modules, data, or programs (e.g., instruction sets) need not be implemented as separate software programs, procedures, data sets, or modules; thus, various subsets of these modules and data may be combined or otherwise rearranged in various implementations. In some implementations, memory 111 optionally stores a subset of the above-identified modules and data structures. Furthermore, in some embodiments, memory stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system other than that of system 100 and are addressable by system 100, and system 100 may retrieve all or a portion of such data when needed.

[0062] While FIG. 1 depicts “system 100,” the diagram is intended more as a functional description of various features that may be present in a computer system than as a structural schematic of the implementations described herein. In practice, items shown separately can be combined, and some items may be separate, as will be recognized by those skilled in the art. Moreover, while FIG. 1 depicts certain data and modules in memory 111 (which may be non-persistent or persistent), it should be appreciated that these data and modules, or portions thereof, may be stored in two or more memories. For example, in some embodiments, at least first dataset 122, second dataset 124, reference module 120, and reference model 140 are stored in a remote storage device that may be part of a cloud-based infrastructure. In some embodiments, at least first dataset 122 and second dataset 124 are stored in a cloud-based infrastructure. In some embodiments, reference model 120 and reference model 140 may also be stored in a remote storage device.

[0063] Having disclosed a system for training a predictive model according to the present disclosure with reference to FIG. 1, a method for performing such training according to the present disclosure will now be detailed with reference to FIG.

[0064] Block 202. Referring to block 202 of Figure 2A, a method for reducing the number of subjects among a plurality of subjects in a subject dataset is provided.

[0065] Blocks 204-206. Referring to block 204 of FIG. 2A, the method proceeds by obtaining a subject dataset in electronic form. An example of such a subject dataset is ZINC15. See Sterling and Irwin, 2005, J. Chem. Inf. Model 45(1), pp. 177-182. ZINC15 is a database of commercially available compounds for virtual screening. ZINC15 contains over 230 million commercially available compounds in a 3D format that can be readily docked. ZINC15 also contains over 750 million commercially available compounds. Other examples of subject datasets include, but are not limited to, MASSIV, AZ Space with Enamine BBs, EVOspace, PGVL, BICLAIM, Lilly, GDB-17, SAVI, CHIPMUNK, REAL'Space', SCUBIDOO 2.1, REAL'Database', WuXi Virtual, PubChem Compounds, Sigma Aldrich'in-stock', eMolecules Plus, and WuXi Chemistry Services, which are summarized in Hoffmann and Gastreich, 2019, "The next level in chemical space navigation: going far beyond enumerable compound libraries," Drug Discovery Today 24(5), pp. 1148, which is incorporated herein by reference.

[0066] In some embodiments, the plurality of subjects comprises (e.g., prior to application of instances of excluding a portion of subjects from the plurality of subjects, as described below with respect to blocks 232-234) at least 100 million subjects, at least 500 million subjects, at least 1 billion subjects, at least 2 billion subjects, at least 3 billion subjects, at least 4 billion subjects, at least 5 billion subjects, at least 6 billion subjects, at least 7 billion subjects, at least 8 billion subjects, or The subject may comprise at least 9 billion subjects, at least 10 billion subjects, at least 11 billion subjects, at least 15 billion subjects, at least 20 billion subjects, at least 30 billion subjects, at least 40 billion subjects, at least 50 billion subjects, at least 60 billion subjects, at least 70 billion subjects, at least 80 billion subjects, at least 90 billion subjects, at least 100 billion subjects, or at least 110 billion subjects. In some embodiments, the plurality of subjects comprises 100 million to 500 million subjects, 100 million to 1 billion subjects, 1 billion to 2 billion subjects, 1 billion to 5 billion subjects, 1 billion to 10 billion subjects, 1 billion to 15 billion subjects, 5 billion to 10 billion subjects, 5 billion to 15 billion subjects, or 10 billion to 15 billion subjects. 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 16 , 10 17 , 10 18 , 10 19 , 10 20 , 10 21 , 10 22 , 10 23 , 10 24 , 10 25 , 10 26 , 10 27 , 1028 , 10 29 , 10 30 , 10 31 , 10 32 , 10 33 , 10 34 , 10 35 , 10 36 , 10 37 , 10 38 , 10 39 , 10 40 , 10 41 , 10 42 , 10 43 , 10 44 , 10 45 , 10 46 , 10 47 , 10 48 , 10 49 , 10 50 , 10 51 , 10 52 , 10 53 , 10 54 , 10 55 , 10 56 , 10 57 , 10 58 , 10 59 , or 10 60 It is a compound of about 1000.

[0067] In some embodiments, the size of the subject dataset is at least 100 kilobytes, at least 1 megabyte, at least 2 megabytes, at least 3 megabytes, at least 4 megabytes, at least 10 megabytes, at least 20 megabytes, at least 100 megabytes, at least 1 gigabyte, at least 10 gigabytes, or at least 1 terabyte in size. In some embodiments, the subject dataset is a collection of files or datasets (e.g., 2 or more, 3 or more, 4 or more, 100 or more, 1000 or more, or 1 million or more) that collectively have a file size of at least 100 kilobytes, at least 1 megabyte, at least 2 megabytes, at least 3 megabytes, at least 4 megabytes, at least 10 megabytes, at least 20 megabytes, at least 100 megabytes, at least 1 gigabyte, at least 10 gigabytes, or at least 1 terabyte.

[0068] With respect to block 206, in some embodiments, each subject in the plurality of subjects represents a respective chemical compound. In some embodiments, each subject represents a chemical compound that satisfies the five criteria of Lipinski's Rule of Five. In some embodiments, each subject is an organic compound that satisfies two or more rules, three or more rules, or all four of Lipinski's Rule of Five: (i) five or fewer hydrogen bond donors (e.g., OH and NH groups), (ii) ten or fewer hydrogen bond acceptors (e.g., N and O), (iii) a molecular weight of less than 500 daltons, and (iv) a LogP of less than 5. The "Rule of Five" is so named because three of the four criteria involve the number five. See Lipinski, 1997, Adv. Drug Del. Rev. 23, 3, which is incorporated herein by reference in its entirety. In some embodiments, each subject satisfies one or more criteria in addition to Lipinski's Rule of Five. For example, in some embodiments, each subject has five or fewer aromatic rings, four or fewer aromatic rings, three or fewer aromatic rings, or two or fewer aromatic rings. In some embodiments, each subject describes a chemical compound, and the description of the chemical compound includes modeled atomic coordinates of the chemical compound. In some embodiments, each subject of the plurality of subjects represents a different chemical compound.

[0069] In some embodiments, each subject represents an organic compound having a molecular weight of less than 2000 daltons, less than 4000 daltons, less than 6000 daltons, less than 8000 daltons, less than 10000 daltons, or less than 20000 daltons.

[0070] In some embodiments, at least one subject in the plurality of subjects exhibits a corresponding pharmaceutical compound. In some embodiments, at least one subject in the plurality of subjects exhibits a corresponding biologically active compound. As used herein, the term "biologically active compound" refers to a compound that has a physiological effect on humans (e.g., through interaction with a protein). A subset of biologically active compounds can be developed into pharmaceuticals. See, for example, Gu et al. 2013 "Use of Natural Products as Chemical Library for Drug Discovery and Network Pharmacology" PLoS One 8(4), e62839. Biologically active compounds can be naturally occurring or synthetic. Various definitions of biological activity have been proposed. See, for example, Lagunin et al. 2000 "PASS: Prediction of activity spectra for biologically active substances" Bioinform 16, 747-748.

[0071] In some embodiments, the subjects in the subject dataset represent chemical compounds having an "alkyl" group. The term "alkyl," by itself or as part of another substituent of a chemical compound, means, unless otherwise specified, a straight-chain, branched-chain, or cyclic hydrocarbon radical, or combinations thereof, which may be fully saturated, monounsaturated, or polyunsaturated, and may include divalent, trivalent, and polyvalent radicals having the specified number of carbon atoms (i.e., C1-C6). 10means 1 to 10 carbons). Examples of saturated hydrocarbon radicals include, but are not limited to, groups such as methyl, ethyl, n-propyl, isopropyl, n-butyl, t-butyl, isobutyl, sec-butyl, cyclohexyl, (cyclohexyl)methyl, cyclopropylmethyl, homologs and isomers of, for example, n-pentyl, n-hexyl, n-heptyl, n-octyl, etc. Unsaturated alkyl groups are groups with one or more double or triple bonds. Examples of unsaturated alkyl groups include, but are not limited to, vinyl, 2-propenyl, crotyl, 2-isopentenyl, 2-(butadienyl), 2,4-pentadienyl, 3-(1,4-pentadienyl), ethynyl, 1- and 3-propynyl, 3-butynyl, and higher homologs and isomers. The term "alkyl," unless otherwise specified, is also meant to optionally include those derivatives of alkyl defined in more detail below, such as "heteroalkyl." Alkyl groups that are limited to hydrocarbon groups are referred to as "homoalkyl." Exemplary alkyl groups include monounsaturated C 9-10 , oleoyl chain, or diunsaturated C 9-10,12-13 Examples include linoleyl chains. The term "alkylene," by itself or as part of another substituent, means a divalent radical derived from an alkane, exemplified by, but not limited to, -CH2CH2CH2CH2-, and further includes groups such as those described below as "heteroalkylene." Typically, an alkyl (or alkylene) group will have 1 to 24 carbon atoms, with those groups having 10 or fewer carbon atoms being preferred for the present invention. A "lower alkyl" or "lower alkylene" is a shorter chain alkyl or alkylene group, generally having 8 or fewer carbon atoms.

[0072] In some embodiments, the subjects in the subject dataset represent chemical compounds having "alkoxy," "alkylamino," and "alkylthio" groups. The terms "alkoxy," "alkylamino," and "alkylthio" (or thioalkoxy) are used in their conventional sense to refer to such alkyl groups attached to the remainder of the molecule via an oxygen atom, an amino group, or a sulfur atom, respectively.

[0073] In some embodiments, the subjects in the subject dataset represent chemical compounds having "aryloxy" and "heteroaryloxy" groups. The terms "aryloxy" and "heteroaryloxy" are used in their conventional sense to refer to an aryl or heteroaryl group attached to the remainder of the molecule via an oxygen atom.

[0074] In some embodiments, the subjects in the subject dataset represent chemical compounds having a "heteroalkyl" group. The term "heteroalkyl," by itself or in combination with another term, means, unless otherwise specified, a stable linear or branched chain, or cyclic hydrocarbon radical, or combination thereof, consisting of the stated number of carbon atoms and at least one heteroatom selected from the group consisting of O, N, Si, and S, where the nitrogen and sulfur atoms can be optionally oxidized and the nitrogen heteroatom can be optionally quaternized. The heteroatoms O, N, S, and Si can be located at any interior position of the heteroalkyl group or at the position at which the alkyl group is attached to the remainder of the molecule. Examples include, but are not limited to, -CH-CH-O-CH, -CH-CH-NH-CH, -CH-CH-N(CH)-CH, -CH-S-CH-CH, -CH-CH, -S(O)-CH, -CH-CH-S(O)-CH, -CH=CH-O-CH, -Si(CH), -CH-CH=N-OCH, and -CH=CH-N(CH)-CH. Up to two heteroatoms may be consecutive, such as, for example, -CH-NH-OCH and -CH-O-Si(CH). Similarly, the term "heteroalkylene," by itself or as part of another substituent, means a divalent radical derived from a heteroalkyl, exemplified by, but not limited to, -CH-CH-S-CH-CH- and -CH-S-CH-CH-NH-CH. For heteroalkylene groups, heteroatoms can also occupy either or both of the chain termini (e.g., alkyleneoxy, alkylenedioxy, alkyleneamino, alkylenediamino, etc.). Furthermore, for alkylene and heteroalkylene linking groups, no orientation of the linking group is implied by the direction in which the formula of the linking group is written. For example, the formula -COR'- represents both -C(O)OR' and -OC(O)R'.

[0075] In some embodiments, the subjects in the subject dataset represent chemical compounds having "cycloalkyl" and "heterocycloalkyl" groups. The terms "cycloalkyl" and "heterocycloalkyl," by themselves or in combination with other terms, represent cyclic versions of "alkyl" and "heteroalkyl," respectively, unless otherwise specified. Additionally, for heterocycloalkyl, a heteroatom can occupy the position where the heterocycle is attached to the remainder of the molecule. Examples of cycloalkyl include, but are not limited to, cyclopentyl, cyclohexyl, 1-cyclohexenyl, 3-cyclohexenyl, cycloheptyl, and the like. Further exemplary cycloalkyl groups include steroids, such as cholesterol and its derivatives. Examples of heterocycloalkyl include, but are not limited to, 1-(1,2,5,6-tetrahydropyridyl), 1-piperidinyl, 2-piperidinyl, 3-piperidinyl, 4-morpholinyl, 3-morpholinyl, tetrahydrofuran-2-yl, tetrahydrofuran-3-yl, tetrahydrothien-2-yl, tetrahydrothien-3-yl, 1-piperazinyl, 2-piperazinyl, and the like.

[0076] In some embodiments, the subjects in the subject dataset represent chemical compounds having a "halo" or "halogen." The terms "halo" or "halogen," by themselves or as part of another substituent, mean, unless otherwise specified, a fluorine, chlorine, bromine, or iodine atom. Additionally, terms such as "haloalkyl" are meant to include monohaloalkyl and polyhaloalkyl. For example, the term "halo(C1-C4)alkyl" is meant to include, but is not limited to, trifluoromethyl, 2,2,2-trifluoroethyl, 4-chlorobutyl, 3-bromopropyl, and the like.

[0077] In some embodiments, the subjects in the subject dataset represent chemical compounds having an "aryl" group. The term "aryl," unless otherwise specified, refers to a polyunsaturated aromatic substituent, which may be a single ring or multiple rings (preferably 1-3 rings) that are fused or covalently linked together.

[0078] In some embodiments, the subjects in the subject dataset represent chemical compounds having a "heteroaryl" group. The term "heteroaryl" refers to an aryl substituent (or ring) containing 1 to 4 heteroatoms selected from N, O, S, Si, and B, where the nitrogen and sulfur atoms are optionally oxidized and the nitrogen atoms are optionally quaternized. Exemplary heteroaryl groups are six-membered azines, such as pyridinyl, diazinyl, and triazinyl. Heteroaryl groups can be attached to the remainder of the molecule via a heteroatom. Non-limiting examples of aryl and heteroaryl groups include phenyl, 1-naphthyl, 2-naphthyl, 4-biphenyl, 1-pyrrolyl, 2-pyrrolyl, 3-pyrrolyl, 3-pyrazolyl, 2-imidazolyl, 4-imidazolyl, pyrazinyl, 2-oxazolyl, 4-oxazolyl, 2-phenyl-4-oxazolyl, 5-oxazolyl, 3-isoxazolyl, 4-isoxazolyl, 5-isoxazolyl, and 5-isoxazolyl.

[0033] Examples of thiazolyl include thiazolyl, 2-thiazolyl, 4-thiazolyl, 5-thiazolyl, 2-furyl, 3-furyl, 2-thienyl, 3-thienyl, 2-pyridyl, 3-pyridyl, 4-pyrimidyl, 4-pyrimidyl, 5-benzothiazolyl, purinyl, 2-benzimidazolyl, 5-indyl, 1-isoquinolyl, 5-isoquinolyl, 2-quinoxynyl, 5-quinoxynyl, 3-quinolyl, and 6-quinolyl. Substituents for each of the above noted aryl and heteroaryl ring systems are selected from the group of acceptable substituents described below.

[0079] Briefly, the term "aryl" when used in combination with other terms (e.g., aryloxy, arylthioxy, arylalkyl) includes aryl, heteroaryl, and heteroarene rings as defined above. Thus, the term "arylalkyl" is meant to include those radicals in which an aryl group is attached to an alkyl group (e.g., benzyl, phenethyl, pyridylmethyl, etc.), including those alkyl groups in which a carbon atom (e.g., a methylene group) has been replaced with an oxygen atom (e.g., phenoxymethyl, 2-pyridyloxymethyl, 3-(1-naphthyloxy)propyl, etc.).

[0080] Each of the above terms (e.g., "alkyl," "heteroalkyl," "aryl," and "heteroaryl") is meant to include both optionally substituted and unsubstituted forms of the indicated species. Exemplary substituents for these species are provided below.

[0081] Substituents on the alkyl and heteroalkyl radicals (including those groups often referred to as alkylene, alkenyl, heteroalkylene, heteroalkenyl, alkynyl, cycloalkyl, heterocycloalkyl, cycloalkenyl, and heterocycloalkenyl) of the chemical compounds represented by the subject dataset are generally referred to as "alkyl group substituents," and can be one or more of a variety of groups selected from, but not limited to, the following: H, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted heterocycloalkyl, -OR', ═O, ═NR', ═N-OR', -NR'R'', SR', halogen, SiR'R''R''', OC(O)R', C(O)R', COR', CONR'R'', OC(O)NR'R'', NR''C(O)R', NR'C(O)NR''R''', NR''C(O)R', NR C(NR'R''R''')=NR'''', NR C(NR'R'')=NR'''', -S(O)R', -S(O)R', -S(O)NR'R'', NRS02R', -CN, and -NO2, where m is the total number of carbon atoms in such radical. R', R'', R'', and R'''' each preferably independently represent hydrogen, substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, e.g., aryl substituted with 1 to 3 halogens, substituted or unsubstituted alkyl, alkoxy or thioalkoxy groups, or arylalkyl groups. When a compound of the invention includes more than one R group, e.g., when two or more of these groups are present, each of the R groups is independently selected as being an R', R'', R'', and R'''' group, respectively. When R' and R'' are attached to the same nitrogen atom, they can be combined with the nitrogen atom to form a five-, six-, or seven-membered ring. For example, --NR'R'' is meant to include, but not be limited to, 1-pyrrolidinyl and 4-morpholinyl.From the above discussion of substituents, one of skill in the art will understand that the term "alkyl" is meant to include groups that contain carbon atoms bonded to groups other than hydrogen groups, such as haloalkyl (e.g., -CF and -CHCF) and acyl (e.g., -C(O)CH, -C(O)CF, -C(O)CHOCH, etc.). These terms encompass groups that are considered exemplary "alkyl group substituents," which are components of exemplary "substituted alkyl" and "substituted heteroalkyl" moieties.

[0082] Similar to the substituents described for the alkyl radical, substituents for the aryl, heteroaryl, and heteroarene groups are generally referred to as "aryl group substituents." Substituents include, for example, but are not limited to, substituted or unsubstituted alkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted heterocycloalkyl, OR', ═O, ═NR', ═N—OR', —NR'R'', —SR', -halogen, —SiR'R''R''', —OC(O)R', —C(O)R', —COR', —CONR'R'', —OC(O)NR'R'', —NR''C(O)R', —NR'—C(O)NR''R'''', —NR''C(O )2R', -NR-C(NR'R''R''')=NR'''', -NR-C(NR'R'')=NR''', -S(O)R', -S(O)2R', -S(O)2NR'R'', -NRSO2R', -CN and -NO2, -R', -N3, -CH(Ph)2, fluoro(C1-C4)alkoxy, and fluoro(C1-C4)alkyl, and are selected from groups attached to a heteroaryl or heteroarene nucleus through a carbon or heteroatom (e.g., P, N, O, S, Si, or B), including: Each of the above-named groups is attached to the heteroarene or heteroaryl nucleus directly or through a heteroatom (e.g., P, N, O, S, Si, or B), where R', R", R'", and R"" are preferably independently selected from hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, and substituted or unsubstituted heteroaryl. When a compound of the invention includes more than one R group, e.g., when more than one of these groups is present, each of the R groups is independently selected as are each R', R", R'", and R"" groups.

[0083] Two of the substituents on adjacent atoms of the aryl, heteroarene, or heteroaryl ring are optionally represented by the formula -TC(O)-(CRR') q-U-, where T and U are independently -NR-, -O-, -CRR'- or a single bond, and q is an integer from 0 to 3. Alternatively, two of the substituents on adjacent atoms of the aryl or heteroaryl ring may optionally be replaced by a substituent of the formula -A-(CH2) r wherein A and B are independently -CRR'-, -O-, -NR-, -S-, -S(O)-, -S(O)2-, -S(O)2NR'-, or a single bond, and r is an integer from 1 to 4. One of the single bonds in the new ring so formed may optionally be replaced with a double bond. Alternatively, two of the substituents on adjacent atoms of the aryl, heteroarene, or heteroaryl ring may optionally be replaced with a substituent of the formula -(CRR') s -X-(CR''R''') d where s and d are independently integers from 0 to 3, and X is -O-, -NR'-, -S-, -S(O)-, -S(O)2-, or -S(O)2NR'-. The substituents R, R', R", and R'" are preferably independently selected from hydrogen or substituted or unsubstituted (C1-C6) alkyl. These terms encompass groups considered exemplary "aryl group substituents," which are components of exemplary "substituted aryl," "substituted heteroarene," and "substituted heteroaryl" moieties.

[0084] In some embodiments, the subjects in the subject dataset represent chemical compounds having an "acyl" group. As used herein, the term "acyl" describes a substituent that includes a carbonyl residue, C(O)R. Exemplary species of R include H, halogen, substituted or unsubstituted alkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, and substituted or unsubstituted heterocycloalkyl.

[0085] In some embodiments, the subjects in the subject dataset represent chemical compounds having a "fused ring system." As used herein, the term "fused ring system" means at least two rings, each ring having at least two atoms in common with another ring. A "fused ring system" can include aromatic rings as well as non-aromatic rings. Examples of "fused ring systems" are naphthalene, indole, quinoline, chromene, etc.

[0086] As used herein, the term "heteroatom" includes oxygen (O), nitrogen (N), sulfur (S), and silicon (Si), boron (B), and phosphorus (P).

[0087] The symbol "R" is a general abbreviation representing a substituent selected from H, substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, and substituted or unsubstituted heterocycloalkyl groups.

[0088] Block 208. Referring to block 208 of FIG. 2A, in some embodiments, the subject dataset includes a plurality of feature vectors (e.g., each feature vector corresponds to an individual subject in the subject dataset and includes one or more features). In some embodiments, each respective feature vector in the plurality of feature vectors includes a chemical fingerprint, a molecular fingerprint, one or more computational properties, and / or a graph descriptor of a respective chemical compound represented by the corresponding subject. Exemplary molecular fingerprints include, but are not limited to, a Daylight fingerprint, a BCI fingerprint, an ECFP fingerprint, an ECFC fingerprint, an MDL fingerprint, an APFP fingerprint, a TTFP fingerprint, a UNITY 2D fingerprint, etc.

[0089] In some embodiments, some of the features in the vector include molecular properties of the corresponding subjects, such as any combination of molecular weight, number of rotatable bonds, calculated LogP (e.g., calculated octanol-water partition coefficient or other method), number of hydrogen bond donors, number of hydrogen bond acceptors, number of chiral centers, number of chiral double bonds (E / Z isomers), polar and nonpolar desolvation energies (in kcal / mol), net charge, and number of rigid fragments. In some embodiments, one or more subjects in the subject dataset are annotated with a function or activity. In some such embodiments, the features in the vector include such a function or activity.

[0090] In some embodiments, the subject dataset includes a chemical structure for each subject. For example, in some embodiments, the chemical structure is a SMILES string. In some embodiments, a canonical representation of the subject is calculated to represent the subject's chemical structure (see, e.g., OpenEye's OEchem library, available online at OpenEye.com). In some embodiments, an initial 3D model is generated from the subject's unambiguous isomer SMILES (e.g., using OpenEye's Omega program). In some embodiments, the relevant, correctly protonated form of the subject at pH 5-9.5 is then created (e.g., using Schrodinger's ligprep program, available from Schrodinger, Inc., available online at schrodinger.com). This includes, for example, deprotonation of carboxylic acids and tetrazoles, and protonation of most aliphatic amines. In some embodiments, partial atomic charges and atomic desolvation penalties for a single 3D conformation of each protonation state, stereoisomer, and tautomer are calculated (e.g., using the semi-empirical quantum mechanical program AMSOL16). In some embodiments, the 3D conformations are generated using OpenEye's program Omega. See, e.g., Sterling and Irwin, 2005, J. Chem. Inf. Model 45(1), pp. 177-182. In some embodiments, the subjects in the subject dataset are represented, at least in part, by a subject dataset having a data structure in SMILES, mol2, 3D SDF, DOCK flexibase, or an equivalent format.

[0091] In embodiments of a subject dataset in which subjects are represented by feature vectors, each feature vector is for a respective subject in the plurality of subjects. In some embodiments, the size (e.g., number of features) of each feature vector in the plurality of feature vectors is the same. In some embodiments, the size (e.g., number of features) of each feature vector in the plurality of feature vectors is not the same. That is, in some embodiments, at least one of the feature vectors in the plurality of feature vectors is a different size. In some embodiments, each feature vector is of any length (e.g., each feature vector can be of any size). In some embodiments, the number of dimensions of each feature vector in the plurality of feature vectors can vary (e.g., a feature vector can have any number of dimensions). In some embodiments, each feature vector in the plurality of feature vectors is a one-dimensional vector. In some embodiments, one or more feature vectors in the plurality of feature vectors are two-dimensional vectors. In some embodiments, one or more feature vectors in the plurality of feature vectors are three-dimensional vectors. In some embodiments, the number of dimensions of each feature vector in the plurality of feature vectors is the same (e.g., each feature vector has the same number of dimensions). In some embodiments, each feature vector in the plurality of feature vectors is a vector of at least two dimensions, hi some embodiments, each feature vector in the plurality of feature vectors is a vector of at least N dimensions, where N is a positive integer greater than or equal to 2 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more).

[0092] In some embodiments, each respective subject in the plurality of subjects comprises a corresponding chemical fingerprint of the chemical compound represented by the respective subject. In some embodiments, the chemical fingerprint of a subject is represented by the subject's corresponding feature vector. As used herein, the term "chemical fingerprint" refers to a unique pattern (e.g., a unique vector or matrix) corresponding to a particular molecule. In some embodiments, each chemical fingerprint is a fixed size. In some embodiments, one or more chemical fingerprints are variably sized. In some embodiments, the chemical fingerprint of each subject in the plurality of subjects can be determined directly (e.g., via mass spectrometry such as MALDI-TOF). In some embodiments, the chemical fingerprint of each subject in the plurality of subjects can be obtained via computational methods. See, for example, Daina et al. (2017) "SwissADME: a free web tool to evaluate pharmacokinetics, drug-likeness, and medicinal chemistry friendliness of small molecules," Sci Reports 7, 42717; O'Boyle et al. 2011 "Open Babel: An open chemical toolbox," J Cheminforma 3, 33; Cereto-Massague et al. 2015 "Molecular fingerprint similarity search in virtual screening," Methods 71, 58-63; and Mitchell 2014 "Machine learning methods in cheminformatics," WIREs Comput Mol Sci. 4:468-481, each of which is incorporated herein by reference.

[0093] Many different methods are known in the art for representing chemical compounds in computational space.

[0094] In some embodiments, each chemical fingerprint includes information about interactions between the respective chemical compound and one or more additional chemical compounds and / or biological macromolecules. In some embodiments, the chemical fingerprint includes information about protein-ligand binding indeterminacy. See Wojcikowski et al. 2018, "Development of a protein-ligand extended connectivity (PLEC) fingerprint and its application for binding affinity predictions," Bioinformatics 35(8), 1334-1341, which is incorporated herein by reference. In some embodiments, a neural network is used to determine one or more chemical properties (and / or chemical fingerprints) of at least one subject in a subject database.

[0095] In some embodiments, each subject in the subject database corresponds to a known chemical compound having one or more known chemical properties. In some embodiments, the same number of chemical properties are provided for each subject in the plurality of subjects in the subject dataset. In some embodiments, a different number of chemical properties are provided for one or more subjects in the subject dataset. In some embodiments, one or more subjects in the subject dataset are synthetic (e.g., the chemical structure of the subject can be determined even though the subject has not been analyzed in a laboratory). See, e.g., Gomez-Bombarelli et al. 2017 "Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules" arXiv:1610.02415v3, which is incorporated herein by reference.

[0096] In some embodiments, graph comparison is used to compare the three-dimensional structures of molecules represented by the subject datasets (e.g., to determine clusters or sets of similar molecules). The concept of graph comparison relies on comparing graph descriptors, resulting in a measure of dissimilarity or similarity that can be used for pattern recognition. See, for example, Czech 2011 "Graph Descriptors form B -Matrix Representation," Graph-Based Representations in Patter Recognition, LNCS 6658, 12-21, which is incorporated herein by reference. In some embodiments, measures such as clustering coefficient, efficiency, or betweenness centrality can be utilized to capture relevant structural properties within the graphs (e.g., of the subject set). See, for example, Costa et al. 2007 "Characterization of complex networks: A survey of measurements," Advances Phys 56(1), 198-200, which is incorporated herein by reference.

[0097] Block 210. Referring to block 210 of FIG. 2A, for each respective subject of a subset of subject subjects from the plurality of subject subjects, a target model is applied to the respective subject subject and at least one target object to obtain corresponding target results, thereby obtaining a corresponding subset of target results. In typical embodiments, each subject subject is docked to each target object of the at least one target object. In some embodiments, there is only a single target object.

[0098] In some embodiments, the target entity is a polymer. Examples of polymers include, but are not limited to, assemblies of proteins, polypeptides, polynucleic acids, polyribonucleic acids, polysaccharides, or any combination thereof. Polymers, such as those studied using some embodiments of the disclosed systems and methods, are large molecules composed of repeating residues. In some embodiments, the polymer is a natural material. In some embodiments, the polymer is a synthetic material. In some embodiments, the polymer is an elastomer, shellac, amber, natural or synthetic rubber, cellulose, bakelite, nylon, polystyrene, polyethylene, polypropylene, polyacrylonitrile, polyethylene glycol, or a polysaccharide.

[0099] In some embodiments, the target entity is a heteropolymer (copolymer). A copolymer is a polymer derived from two (or more) monomer species, as opposed to a homopolymer, in which only one monomer is used. Copolymerization refers to the method used to chemically synthesize the copolymer. Examples of copolymers include, but are not limited to, ABS resin, SBR, nitrile rubber, styrene-acrylonitrile, styrene-isoprene-styrene (SIS), and ethylene-vinyl acetate. Because copolymers are composed of at least two types of building blocks (structural units, or particles), copolymers can be classified based on how these units are arranged along the chain. These include alternating copolymers, which have regularly alternating A and B units. See, e.g., Jenkins, 1996, "Glossary of Basic Terms in Polymer Science," Pure Appl. Chem. 68(12):2287-2311, which is incorporated herein by reference in its entirety. Additional examples of copolymers include copolymers with repeating sequences (e.g., (ABABBAAAAABBB) n) are periodic copolymers having A and B units arranged in a cyclic manner. Additional examples of copolymers are statistical copolymers, in which the arrangement of monomer residues in the copolymer follows statistical rules. See, for example, Painter, 1997, Fundamentals of Polymer Science, CRC Press, 1997, p. 14, which is incorporated herein by reference in its entirety. Yet another example of a copolymer that can be evaluated using the disclosed systems and methods is a block copolymer containing two or more homopolymer subunits linked by covalent bonds. The linkage of the homopolymer subunits may require an intermediate non-repeating subunit known as a junction block. Block copolymers with two or three separate blocks are called diblock and triblock copolymers, respectively.

[0100] In some embodiments, the target object is actually a plurality of polymers, and the polymers in the plurality do not all have the same molecular weight. In some such embodiments, the polymers in the plurality of polymers fall into weight ranges with a corresponding distribution of chain lengths. In some embodiments, the polymers are branched polymer molecules comprising a main chain with one or more substituent side chains or branches. Types of branched polymers include, but are not limited to, star polymers, comb polymers, brush polymers, dendritic polymers, ladder polymers, and dendrimers. See, e.g., Rubinstein et al., 2003, Polymer Physics, Oxford; New York: Oxford University Press, p. 6, which is incorporated herein by reference in its entirety.

[0101] In some embodiments, the target entity is a polypeptide. As used herein, the term "polypeptide" refers to two or more amino acids or residues linked by a peptide bond. The terms "polypeptide" and "protein" are used interchangeably herein and include oligopeptides and peptides. "Amino acid," "residue," or "peptide" refers to any of the 20 standard structural units of proteins known in the art, including imino acids such as proline and hydroxyproline. Designations for amino acid isomers can include D, L, R, and S. The definition of amino acid includes unnatural amino acids. Thus, selenocysteine, pyrrolysine, lanthionine, 2-aminoisobutyric acid, γ-aminobutyric acid, dehydroalanine, ornithine, citrulline, and homocysteine ​​are all considered amino acids. Other variants or analogs of amino acids are known in the art. Thus, a polypeptide can include synthetic peptidomimetic structures such as peptides. See Simon et al., 1992, Proceedings of the National Academy of Sciences USA, 89, 9367, which is incorporated herein by reference in its entirety. See also Chin et al., 2003, Science 301, 964, and Chin et al., 2003, Chemistry & Biology 10, 511, each of which is incorporated herein by reference in its entirety.

[0102] In some embodiments, the target object evaluated according to some embodiments of the disclosed system and method can also have any number of post-translational modifications.Therefore, the target object can include polymers that are modified by acylation, alkylation, amidation, biotinylation, formylation, γ-carboxylation, glutamylation, glycosylation, glycation, hydroxylation, iodination, isoprenylation, lipoylation, cofactor addition (such as heme, flavin, metal, etc.), nucleoside and their derivative addition, oxidation, reduction, PEGylation, phosphatidylinositol addition, phosphopantetheinylation, phosphorylation, pyroglutamic acid formation, racemization, tRNA-mediated amino acid addition (such as arginylation), sulfation, selenoylation, ISG addition, sumoylation, ubiquitination, chemical modification (such as quetrination and deamidation), and other enzyme (such as protease, phosphotase and kinase) processing. Other types of post-translational modifications are known in the art and are also included.

[0103] In some embodiments, the target entity is an organometallic complex. An organometallic complex is a chemical compound that contains a bond between carbon and a metal. In some cases, organometallic compounds are distinguished by the prefix "organic," e.g., organopalladium compounds.

[0104] In some embodiments, the target substance is a surfactant. A surfactant is a compound that reduces the surface tension of a liquid, the interfacial tension between two liquids, or the interfacial tension between a liquid and a solid. Surfactants can act as detergents, wetting agents, emulsifiers, foaming agents, and dispersants. Surfactants are typically organic compounds that are amphiphilic, meaning that they contain both a hydrophobic group (their tail) and a hydrophilic group (their head). Thus, surfactant molecules contain both water-insoluble (or oil-soluble) and water-soluble components. Surfactant molecules diffuse in water and adsorb to the air-water interface, or, when water is mixed with oil, to the oil-water interface. The insoluble hydrophobic group can extend from the bulk water phase into the air or oil phase, while the water-soluble head group remains in the water phase. This orientation of surfactant molecules at the surface modifies the surface properties of water at the water / air or water / oil interface.

[0105] Examples of ionic surfactants include ionic surfactants such as anionic surfactants, cationic surfactants, or zwitterionic (amphoteric) surfactants. In some embodiments, the targeting entity is a reverse micelle or a liposome.

[0106] In some embodiments, the target entity is a fullerene. A fullerene is any molecule composed entirely of carbon in the form of a hollow sphere, ellipsoid, or tube. Spherical fullerenes are also called buckyballs, as they resemble the balls used in association football. Cylindrical ones are called carbon nanotubes or buckytubes. Fullerenes are structurally similar to graphite, which is composed of stacked graphene sheets of linked hexagonal rings, but they can also contain pentagonal (or sometimes heptagonal) rings.

[0107] In some embodiments, the target object is a polymer, and the spatial coordinates are a set of three-dimensional coordinates {x,...,x} of the crystalline structure of the polymer resolved to a resolution of 2.5 Å or better. N} (208), where N is an integer greater than or equal to 2 (e.g., greater than or equal to 10, greater than or equal to 20, etc.). In some embodiments, the target object is a polymer and the spatial coordinates are a set of three-dimensional coordinates {x1,...,xN} of the crystalline structure of the polymer resolved at a resolution of 3.3 Å or greater (210). In some embodiments, the target object is a polymer and the spatial coordinates are a set of three-dimensional coordinates {x1,...,xN} of the crystalline structure of the polymer resolved (e.g., by X-ray crystallography) at a resolution of 3.3 Å or greater, 3.2 Å or greater, 3.1 Å or greater, 3.0 Å or greater, 2.5 Å or greater, 2.2 Å or greater, 2.0 Å or greater, 1.9 Å or greater, 1.85 Å or greater, 1.80 Å or greater, 1.75 Å or greater, or 1.70 Å or greater. N}.

[0108] In some embodiments, the target entity is a polymer, and the spatial coordinates are an ensemble of 10 or more, 20 or more, 30 or more three-dimensional coordinates of the polymer as determined by nuclear magnetic resonance, the ensemble having a backbone RMSD of 1.0 Å or more, 0.9 Å or more, 0.8 Å or more, 0.7 Å or more, 0.6 Å or more, 0.5 Å or more, 0.4 Å or more, 0.3 Å or more, or 0.2 Å or more. In some embodiments, the spatial coordinates are determined by neutron diffraction or cryo-electron microscopy.

[0109] In some embodiments, the target object comprises two different types of polymers, such as a nucleic acid bound to a polypeptide. In some embodiments, the natural polymer comprises two polypeptides bound to each other. In some embodiments, the natural polymer under study comprises one or more metal ions (e.g., a metalloprotease having one or more zinc atoms). In such cases, the metal ions and / or small organic molecules can be included in the spatial coordinates of the target object.

[0110] In some embodiments, the target entity is a polymer, and the polymer has 10 or more, 20 or more, 30 or more, 50 or more, 100 or more, 100-1000, or less than 500 residues.

[0111] In some embodiments, the spatial coordinates of the target object are determined using modeling methods such as ab initio methods, density functional methods, semi-empirical and empirical methods, molecular mechanics, chemical dynamics, or molecular dynamics.

[0112] In embodiments, the spatial coordinates are represented by Cartesian coordinates of the centers of atoms comprising the target object. In some alternative embodiments, the spatial coordinates of the target object are represented by the electron density of the target object, e.g., as measured by X-ray crystallography. For example, in some embodiments, the spatial coordinates are represented by 2F observed -F calculated Includes electron density maps, F observed is the observed structure factor amplitude of the target object, and Fc is the structure factor amplitude calculated from the calculated atomic coordinates of the target object.

[0113] Thus, spatial coordinates of target entities can be received as input data from a variety of sources, including, but not limited to, structural ensembles generated by solution NMR, co-complexes interpreted from X-ray crystallography, neutron diffraction, or cryo-electron microscopy, sampling from computational simulations, homology modeling or rotamer library sampling, as well as combinations of these techniques.

[0114] In some embodiments, block 210 includes obtaining spatial coordinates of the target object. Block 210 further includes modeling each subject with the target object in each of a plurality of different poses, thereby creating a plurality of voxel maps, each respective voxel map in the plurality of voxel maps including the subject in a respective one of the plurality of different poses.

[0115] In some embodiments, the target object is a polymer having an active site, and each test object is a chemical compound, and modeling each test object with the target object in each of a plurality of different poses includes docking the test object into the active site of the target object. In some embodiments, each test object is docked onto the target object multiple times to form a plurality of poses (e.g., each docking represents a different pose). In some embodiments, the test object is docked onto the target object two, three, four, five or more times, ten or more times, fifty or more times, one hundred or more times, or one thousand or more times. Each such docking represents a different pose of each test object docked onto the target object. In some embodiments, each target object is a polymer having an active site, and the test object is docked into the active site in each of a plurality of different ways, each such way representing a different pose. It is assumed that many of these poses are incorrect, meaning that such poses do not represent the true interactions between each test object and the target object that actually occur. Without intending to be limited to any particular theory, it is hypothesized that intersubject (e.g., intermolecular) interactions observed between incorrect poses will cancel each other out like white noise, whereas intersubject interactions formed by correct poses formed by the test subjects will reinforce each other. In some embodiments, test subjects are docked using either random pose generation techniques or biased pose generation. In some embodiments, test subjects are docked using Markov Chain Monte Carlo sampling. In some embodiments, such sampling allows for full flexibility of test subjects in the docking calculation and a scoring function that is the sum of the interaction energies between the test subject and target subject, as well as the conformational energy of the test subject.See, e.g., Liu and Wang, 1999, "MCDOCK: A Monte Carlo simulation approach to the molecular docking problem," Journal of Computer-Aided Molecular Design 13, 435-451, which is incorporated herein by reference.

[0116] In some embodiments, algorithms such as DOCK (Shoichet, Bodian, and Kuntz, 1992, "Molecular docking using shape descriptors," Journal of Computational Chemistry 13(3), pp. 380-397, and Knegtel, Kuntz, and Oshiro, 1997, "Molecular docking to ensembles of protein structures," Journal of Molecular Biology 266, pp. 424-440, each of which is incorporated herein by reference) are used to find multiple poses for each respective test object relative to each target object. Such algorithms model the target and test objects as rigid bodies. The docked conformations are searched using complementary surfaces to find poses.

[0117] In some embodiments, AutoDOCK (Morris et al., 2009, "AutoDock4 and AutoDockTools4: Automated Docking with Selective Receptor Flexibility," J. Comput. Chem. 30(16), pp. 2785-2791; Sotriffer et al., 2000, "Automated docking of ligands to antibodies: methods and applications," Methods: A Companion to Methods in Enzymology 20, pp. 280-291; and Morris et al., 1998, "Automated Docking Using a Lamarckian Genetic Algorithm and Empirical Binding Free Energy Function," Journal of Computational Chemistry) is used. 19:1639-1662, each of which is incorporated herein by reference) to find multiple poses for each respective test subject relative to each target subject. AutoDOCK uses a kinetic model of the ligand and supports Monte Carlo, simulated annealing, Lamarckian genetic algorithms, and genetic algorithms. Thus, in some embodiments, multiple different poses (for a given test subject-target subject pair) are obtained by Markov chain Monte Carlo sampling, simulated annealing, Lamarckian genetic algorithms, or genetic algorithms using a docking scoring function.

[0118] In some embodiments, an algorithm such as FlexX (Rarey et al., 1996, "A Fast Flexible Docking Method Using an Incremental Construction Algorithm," Journal of Molecular Biology 261, pp. 470-489, which is incorporated herein by reference) is used to find multiple poses for each of the respective subject subsets for each of the target subjects. FlexX uses a greedy algorithm to perform sequential construction of the subject in the active site of the target subject. Thus, in some embodiments, multiple different poses (for a given subject-target pair) are obtained by the greedy algorithm.

[0119] In some embodiments, an algorithm such as GOLD (Jones et al., 1997, "Development and Validation of a Genetic Algorithm for Flexible Docking," Journal of Molecular Biology 267, pp. 727-748, which is incorporated herein by reference) is used to find multiple poses for each of a subset of test subjects for each of the target subjects. GOLD stands for Genetic Optimization for Ligand Docking. GOLD constructs a genetically optimized hydrogen bond network between the test and target subjects.

[0120] In some embodiments, the modeling includes performing molecular dynamics runs of the target object and the test object. During the molecular dynamics run, atoms of the target object and the test object interact for a fixed period of time, allowing for a view of the dynamic evolution of the system. The trajectories of the atoms of the target object and the test object are determined by numerically solving Newton's equations of motion for the system of interacting particles, and the forces between the particles and their potential energies are calculated using interatomic potentials or molecular mechanics force fields. See Alder and Wainwright, 1959, "Studies in Molecular Dynamics. I. General Method," J. Chem. Phys. 31(2):459, and Bibcode, 1959, J. Ch. Ph. 31, 459A, doi:10.1063 / 1.1730376, each of which is incorporated herein by reference. Thus, in this manner, the molecular dynamics run generates trajectories of the target object and the test object over time. The trajectory includes trajectories of atoms of the target object and the test object. In some embodiments, a subset of multiple different poses is obtained by taking snapshots of the trajectory over a period of time. In some embodiments, the poses are obtained from snapshots of several different trajectories, each trajectory including a different molecular dynamics run of the target object interacting with the test object. In some embodiments, prior to the molecular dynamics run, the test object is first docked into the active site of the target object using a docking technique.

[0121] Regardless of the modeling method used, what is achieved for any given subject-target pair is a set of diverse poses of the subject with the target, one or more of which are expected to be sufficiently close to naturally occurring poses to illustrate some of the relevant molecular interactions between the given subject / target pair.

[0122] In some embodiments, an initial pose of the subject at the active site of the target object is generated using any of the techniques described above, and additional poses are generated through application of some combination of rotation, translation, and mirroring operators in any combination of the three X, Y, and Z planes. The subject rotations and translations may be selected randomly (within a range, e.g., plus or minus 5 Å from the origin) or may be generated uniformly at some pre-specified increment (e.g., 5 full degree increments around the circumference). Figure 4 provides a sample illustration of the subject 122 at two different poses (402-1 and 402-2) at the active site of the target object 124.

[0123] After generating each of the poses for each of the target object and / or test subject, in some embodiments, a voxel map is created for each pose, thereby creating multiple voxel maps for a given target object. In some embodiments, each respective voxel map in the multiple voxel maps is created by a method that includes: (i) sampling the test subject at each pose in a plurality of different poses and the target object on a three-dimensional grid basis, thereby forming a corresponding three-dimensional uniform space-filling honeycomb including a corresponding plurality of space-filling (three-dimensional) polyhedral cells; and (ii) for each respective three-dimensional polyhedral cell in the corresponding plurality of three-dimensional cells, filling voxels (discrete sets of regularly spaced polyhedral cells) of the respective voxel map based on attributes (e.g., chemical attributes) of the respective three-dimensional polyhedral cell. Thus, in such an embodiment, if a particular subject has 10 poses relative to the target object, 10 corresponding voxel maps are created, if a particular subject has 100 poses relative to the target object, 100 corresponding voxel maps are created, etc. Examples of space-filling honeycombs include cubic honeycombs with parallelogram cells, hexagonal prism honeycombs with hexagonal prism cells, rhombic dodecahedrons with rhombic dodecahedron cells, prolonged dodecahedrons with prolonged dodecahedron cells, and truncated octahedrons with truncated octahedron cells.

[0124] In some embodiments, the space-filling honeycomb is a cubic honeycomb with cubic cells, and the dimensions of such voxels determine their resolution. For example, a resolution of 1 Å may be selected, meaning that in such an embodiment, each voxel represents a corresponding cube of geometric data having dimensions of 1 Å (e.g., 1 Å x 1 Å x 1 Å in each height, width, and depth of each cell). However, in some embodiments, a finer grid spacing (e.g., 0.1 Å, or even 0.01 Å) or a coarser grid spacing (e.g., 4 Å) is used, which results in an integer number of voxels to cover the input geometric data. In some embodiments, sampling is performed at a resolution between 0.1 Å and 10 Å. By way of example, for a 40 Å input cube, a resolution of 1 Å would result in 40*40*40=64,000 input voxels.

[0125] In some embodiments, each test object is a first compound and the target object is a second compound, and the atomic characteristics resulting from sampling (i) are arranged in a single voxel of each voxel map by filling (ii), with each voxel in the plurality of voxels representing the characteristics of at most one atom. In some embodiments, the atomic characteristics consist of a list of atom types. As an example, for biological data, some embodiments of the disclosed systems and methods are configured to represent the presence of every atom in a given voxel of the voxel map as a different number for that entry; for example, if carbon is present in a voxel, the value 6 is assigned to that voxel because carbon's atomic number is 6. However, such encoding may imply that atoms with similar atomic numbers behave similarly, which may not be particularly useful in some applications. Furthermore, elements may behave more similarly within groups (columns on the periodic table), and therefore such encoding poses additional work for convolutional neural networks to decode.

[0126] In some embodiments, atomic properties are encoded into voxels as binary categorical variables. In such embodiments, atom types are encoded in what is referred to as "one-hot" encoding: every atom type has a separate channel. Thus, in such embodiments, each voxel has multiple channels, and at least a subset of the multiple channels represents an atom type. For example, one channel in each voxel may represent carbon, while another channel in each voxel may represent oxygen. When a given atom type is found in the three-dimensional grid element corresponding to a given voxel, the channel for that atom type in the given voxel is assigned a first value of the binary categorical variable, such as "1," and when the atom type is not found in the three-dimensional grid element corresponding to a given voxel, the channel for that atom type in the given voxel is assigned a second value of the binary categorical variable, such as "0."

[0127] There are over 100 elements, most of which are not encountered in biology. However, even representations of the most common biological elements (e.g., H, C, N, O, F, P, S, Cl, Br, I, Li, Na, Mg, K, Ca, Mn, Fe, Co, Zn) may result in 18 channels per voxel, or 10,483*18=188,694 inputs to a receptor field. Thus, in some embodiments, each respective voxel in a voxel map of multiple voxel maps includes multiple channels, each channel in the multiple channels representing a different attribute that may occur in the three-dimensional space-filling polyhedral cell corresponding to the respective voxel. The number of possible channels for a given voxel is even greater in those embodiments in which additional properties of the atom (e.g., partial charge, presence in a ligand versus protein target, electronegativity, or SYBYL atom type) are additionally presented as separate channels for each voxel, requiring more input channels to distinguish between otherwise equivalent atoms.

[0128] In some embodiments, each voxel has five or more input channels. In some embodiments, each voxel has 15 or more input channels. In some embodiments, each voxel has 20 or more input channels, 25 or more input channels, 30 or more input channels, 50 or more input channels, or 100 or more input channels. In some embodiments, each voxel has five or more input channels selected from the descriptors found in Table 1 below. For example, in some embodiments, each voxel has five or more channels, each channel encoded as a binary categorical variable, where each channel represents a SYBYL atom type selected from Table 1 below. For example, in some embodiments, each respective voxel of the voxel map includes a C.3 (sp3 carbon) atom type channel, meaning that if the grid in the space of a given subject-target object complex represented by the respective voxel contains sp3 carbon, the channel adopts a first value (e.g., "1"), and otherwise a second value (e.g., "0"). [Table 1-1] [Table 1-2]

[0129] In some embodiments, each voxel includes 10 or more input channels, 15 or more input channels, or 20 or more input channels selected from the descriptors found in Table 1 above. In some embodiments, each voxel includes a channel for a halogen.

[0130] In some embodiments, a structural protein-ligand interaction fingerprint (SPLIF) score is generated for each pose of each test object relative to the target object, and this SPLIF score is used as an additional input to the target model or individually encoded into the voxel map. For a description of SPLIF, see Da and Kireev, 2014, J. Chem. Inf. Model. 54, pp. 2555-2561, "Structural Protein-Ligand Interaction Fingerprints (SPLIF) for Structure-Based Virtual Screening: Method and Benchmark Study," which is incorporated herein by reference. SPLIF implicitly encodes all possible interaction types (e.g., π-π, CH-π, etc.) that may occur between the test object interacting fragment and the target object. In the first step, the test object-target object complex (pose) is examined for intermolecular contacts. Two atoms are considered to be in contact if the distance between them is within a specified threshold (e.g., within 4.5 Å). For each such intermolecular atom pair, the respective test atom and target atom are expanded into circular fragments, e.g., fragments containing the atom in question and its contiguous neighbors up to a certain distance. Each type of circular fragment is assigned an identifier. In some embodiments, such identifiers are coded into individual channels of each voxel. In some embodiments, the extended connectivity fingerprint to the first nearest neighbor (ECFP2) defined in Pipeline Pilot software can be used. See Pipeline Pilot, version 8.5, Accelrys Software Inc., 2009, which is incorporated herein by reference. ECFP retains information about all atom / bond types and uses one unique integer identifier to represent one substructure (e.g., a circular fragment). The SPLIF fingerprint encodes all circular fragment identifiers found.In some embodiments, the SPLIF fingerprints act as separate and independent inputs in the target model rather than as encoded individual voxels.

[0131] In some embodiments, rather than or in addition to SPLIF, a structural interaction fingerprint (SIFt) is calculated for each pose of a given subject relative to a target object and provided independently as an input to the target model or encoded into the voxel map. For calculation of SIFt, see Deng et al., 2003, "Structural Interaction Fingerprint (SIFt): A Novel Method for Analyzing Three-Dimensional Protein-Ligand Binding Interactions," J. Med. Chem. 47(2), pp. 337-344, which is incorporated herein by reference.

[0132] In some embodiments, rather than or in addition to SPLIF and SIFT, atom pair-based interaction fragments (APIFs) are calculated for each pose of a given subject relative to a target object and provided independently as input to the target model or individually encoded into voxel maps. For calculation of APIFs, see Perez-Nueno et al., 2009, "APIF: a new interaction fingerprint based on atom pairs and its application to virtual screening," J. Chem. Inf. Model. 49(5), pp. 1245-1260, which is incorporated herein by reference.

[0133] Data representations may be encoded with biological data in a manner that allows for the expression of various structural relationships associated with molecules / proteins, for example. Geometric representations may be implemented in a variety of ways and topographies according to various embodiments. Geometric representations are used for data visualization and analysis. For example, in embodiments, geometric shapes may be represented using voxels laid out on various topographies, such as 2D, 3D Cartesian / Euclidean space, 3D non-Euclidean space, manifolds, etc. For example, FIG. 5 illustrates a sample three-dimensional grid structure 500 including a series of subcontainers, according to an embodiment. Each subcontainer 502 may correspond to a voxel. A coordinate system may be defined for the grid such that each subcontainer has an identifier. In some embodiments of the disclosed systems and methods, the coordinate system is a Cartesian system in 3D space, but in other embodiments of the system, the coordinate system may be any other type of coordinate system, such as an oblate sphere, a cylindrical or spherical coordinate system, a polar coordinate system, or other coordinate systems designed for various manifolds and vector spaces, among others. In some embodiments, voxels may have particular values ​​associated with them, which may be represented, for example, by applying labels and / or determining the positioning of these voxels, among other things.

[0134] In some embodiments, block 210 includes expanding each voxel map in the plurality of voxel maps into a corresponding vector, thereby creating a plurality of vectors, each vector in the plurality of vectors being the same size. In some embodiments, each respective vector in the plurality of vectors is input to a target model. In some embodiments, the target model includes (i) an input layer for sequentially receiving the plurality of vectors, (ii) a plurality of convolutional layers, and (iii) a scorer, wherein the plurality of convolutional layers includes an initial convolutional layer and a final convolutional layer, and each layer in the plurality of convolutional layers is associated with a different set of weights. In such an embodiment, in response to input of each vector in the plurality of vectors, the input layer provides a first plurality of values ​​to an initial convolutional layer as a first function of the values ​​of the respective vectors, and each respective convolutional layer other than the final convolutional layer provides intermediate values ​​to another convolutional layer in the plurality of convolutional layers as a second function of (i) a different set of weights associated with the respective convolutional layer and (ii) each of the input values ​​received by the respective convolutional layer, and the final convolutional layer provides a final value to the scorer as a third function of (i) a different set of weights associated with the final convolutional layer and (ii) the input value received by the final convolutional layer. In this manner, a plurality of scores is obtained from the scorer, each score in the plurality of scores corresponding to the input of a vector in the plurality of vectors to the input layer. The plurality of scores is then used to provide a corresponding target outcome for each subject. In some embodiments, the target outcome is a weighted average of the plurality of scores. In some embodiments, the target outcome is a measure of central tendency of the plurality of scores. Examples of measures of central tendency include the arithmetic mean, weighted mean, mid-range, mid-hinge, three-point mean, winsorized mean, median, or mode of multiple scores.

[0135] In some embodiments, the scorer includes a plurality of fully connected layers and an evaluation layer where a fully connected layer in the plurality of fully connected layers feeds the evaluation layer. In some embodiments, the scorer includes a decision tree, a multiple additive regression tree, a clustering algorithm, principal component analysis, nearest neighbor analysis, linear discriminant analysis, quadratic discriminant analysis, a support vector machine, an evolutionary method, a projection pursuit, and ensembles thereof. In some embodiments, each vector in the plurality of vectors is a one-dimensional vector. In some embodiments, the plurality of distinct poses includes two or more poses, ten or more poses, one hundred or more poses, or one thousand or more poses. In some embodiments, the plurality of distinct poses is obtained using a docking scoring function in one of Markov chain Monte Carlo sampling, simulated annealing, a Lamarckian genetic algorithm, or a genetic algorithm. In some embodiments, the plurality of distinct poses is obtained by sequential search using a greedy algorithm.

[0136] Blocks 212 and 214. In some embodiments, the target model has a higher computational complexity than the predictive model. In some such embodiments, applying the target model to all subjects in the subject dataset is computationally prohibitive. For this reason, the target model is typically applied to a subset of subjects rather than every subject in the subject dataset. In some embodiments, a degree of diversity in the subject subset (e.g., a subject subset including subjects with a range of structural or functional qualities) is desired. In some embodiments, the subset of subjects includes at least 1,000 subjects, at least 5,000 subjects, at least 10,000 subjects, at least 25,000 subjects, at least 50,000 subjects, at least 75,000 subjects, at least 100,000 subjects, at least 250,000 subjects, at least 500,000 subjects, at least 750,000 subjects, at least 1 million subjects, at least 2 million subjects, at least 3 million subjects, at least 4 million subjects, at least 5 million subjects, at least 6 million subjects, at least 7 million subjects, at least 8 million subjects, at least 9 million subjects, or at least 10 million subjects.

[0137] To ensure this, and referring to block 212 of FIG. 2A, in some embodiments, a subset of subjects is selected from the subject dataset on a randomized basis (e.g., the subset of subjects is selected from the subject dataset using any random method known in the art).

[0138] Referring to block 214 of FIG. 2A , in other embodiments, a subset of subjects is selected from a dataset of subjects based on evaluation of one or more features of the subject's feature vector. In some such embodiments, evaluating the features includes selecting subjects from a plurality of subjects based on clustering (e.g., selecting subjects from a plurality of clusters when forming each subset of subjects). The subset of subjects is then selected based at least in part on the redundancy of subjects in individual clusters within the plurality of clusters (e.g., to obtain subsets of subjects representing different types of chemical compounds). For example, consider a case where subjects in a dataset of subjects are clustered into 100 different clusters based on the subject's feature vectors. One approach to selecting a subset of subjects is to select a fixed number of subjects (e.g., 10, 100, 1000, etc.) from each of the different clusters to form the subset of subjects. Within each cluster, subjects can be selected in a random manner. Alternatively, within each cluster, the subject closest to the center of each cluster is selected based on the fact that such subject best represents the characteristics of each cluster of these subjects. In some embodiments, the form of clustering used is unsupervised clustering. The benefit of clustering multiple subjects from a subject dataset is that it provides more accurate training of predictive models. For example, if all or most of the subjects in a subset of subjects are similar chemical compounds (e.g., contain the same chemical group, have similar structures, etc.), there is a risk that the predictive model will be biased or overfitted to that particular type of chemical compound. This may, in some cases, have a negative impact on downstream training (e.g., it may be difficult to efficiently retrain the predictive model to accurately analyze subjects from different types of chemical compounds).

[0139] To illustrate how subject feature vectors are used in clustering, consider the case where a set of 10 common features (the same 10 features) in each feature vector is used for clustering. In some embodiments, each subject in the subject dataset can have values ​​for each of the 10 features. In some embodiments, each subject in the subject dataset has measurements of some of the features, and missing values ​​are filled using interpolation techniques or ignored (underestimated). In some embodiments, each subject in the subject dataset has values ​​of some of the features, and missing values ​​are filled using constraints. The values ​​from the feature vectors of the subjects in the subject dataset define the vectors: X1, X2, X3, X4, X5, X6, X7, X8, X9, X10, X11, X12, X13, X14, X15, X16, X17, X18, X19, X20, X21, X22, X23, X24, X25, X26, X27, X28, X29, X30, X31, X32, X33, X34, X35, X36, X37, X38, X39, X40, X41, X42, X43, X44, X45, X46, X47, X48, X49, X50, X51, X52, X53, X54, X55, X56, X57, X58, X59, X60, X61, X62, X63, X64, X65, X66, X67, X68, X69, X70, X71, X72, X73, X74, X75, X76, X77, X78, X79, X80, X81, X82, X83, X84, X85, X86, X87, X88, X89, X90, X91, X92, X93, X9 10 , where X i is the value of the i-th feature in a particular subject's feature vector. If there are Q subjects in the subject dataset, then the selection of 10 features can define Q vectors. In clustering, members of the subject dataset that exhibit similar measurement patterns across their respective feature vectors tend to cluster together.

[0140] Specific exemplary clustering techniques that may be used include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest neighbor, farthest neighbor, average linkage, centroid, or sum-of-squares algorithms), k-means clustering, fuzzy k-means clustering, Jarvis-Patrick clustering, density-based spatial clustering algorithms, partitional clustering algorithms, supervised clustering algorithms, or ensembles thereof. Such clustering may be on features within each subject's feature vector, or principal components (or other forms of reduced components) derived therefrom. In some embodiments, clustering involves unsupervised clustering, in which no preconceived notions are imposed on what clusters may be formed when the subject dataset is clustered.

[0141] Data clustering is an unsupervised process that requires optimization to be effective; for example, using either too few or too many clusters to describe a data set can result in a loss of information. See, e.g., Jain et al. 1999 "Data Clustering: A review" AMC Computing Surveys 31(3), 264-323, and Berkhin 2002 "Survey of clustering datamining techniques" Tech Report, Accrue Software, San Jose, CA, each of which is incorporated herein by reference. In some embodiments, to improve the clustering process, the plurality of subjects are normalized prior to clustering (e.g., one or more dimensions of each feature vector in the plurality of feature vectors are normalized (e.g., to the average value of each of the corresponding dimensions determined from the plurality of feature vectors)).

[0142] In some embodiments, a centroid-based clustering algorithm is used to cluster multiple subjects. Centroid-based clustering organizes data into non-hierarchical clusters, representing all subjects in terms of a centroid vector (where the vector itself may not be part of the data set). The algorithm then calculates a distance measure between each subject and the centroid vector, and clusters the subjects based on their proximity to one of the centroid vectors. In some embodiments, a Euclidean distance measure, a Manhattan distance measure, or a Minkowski distance measure is used to calculate the distance measure between each subject and the centroid vector. In some embodiments, a k-means, k-medoid, CLARA, or CLARANS clustering algorithm is used to cluster multiple subjects. An example of the k-means algorithm is described in Uppada 2014 "Centroid Based Clustering Algorithms - A Clarion Study" Int J Comp Sci and Inform Technol 5(6), 7309-7313, which is incorporated herein by reference.

[0143] In some embodiments, a density-based clustering algorithm is used to cluster multiple subjects. Density-based spatial clustering algorithms identify clusters as regions (e.g., feature vectors) of a dataset with higher concentrations (e.g., regions with higher concentrations of subjects). In some embodiments, density-based spatial clustering can be performed as described in Ester et al. 1996 "A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise" KDD'96: Proceedings of the Second International Conference on Knowledge Discovery and Data Mining, 226-231, which is incorporated herein by reference. In such embodiments, the algorithm allows for arbitrarily shaped distributions and does not assign outliers (e.g., subjects outside the concentrations of other subjects) to clusters.

[0144] In some embodiments, hierarchical clustering (for example, connectivity-based clustering) algorithm is used to perform clustering of multiple subjects. Generally, hierarchical clustering is used to build a series of clusters, and as will be further described below, it can be agglomerative or divisive (for example, there are agglomerative or divisive subsets of hierarchical clustering methods). For example, Rokach et al., incorporated herein by reference, describe various versions of agglomerative clustering methods ("Clustering Methods" 2005 Data Mining and Knowledge Discovery Handbook, 321-352).

[0145] In some embodiments, hierarchical clustering includes divisive clustering, which first groups subjects into one cluster and then divides the subjects into more clusters (e.g., it is a recursive process) until a certain threshold (e.g., the number of clusters) is reached. Examples of different methods of divisive clustering are described, for example, in Chavent et al. 2007 "DIVCLUS-T: a monothetic divisive hierarchical clustering method" Comp Stats Data Anal 52(2), 687-701, Sharma et al. 2017 "Divisive hierarchical maximum likelihood clustering" BMC Bioinform 18(Suppl 16):546, and Xiong et al. 2011 "DHCC: Divisive hierarchical clustering of categorical data" Data Min Knowl Disc doi 10.1007 / s10618-011-0221-2, each of which is incorporated herein by reference.

[0146] In some embodiments, hierarchical clustering includes agglomerative clustering. Agglomerative clustering generally involves first separating a plurality of subjects into multiple distinct clusters (e.g., in some cases, starting with individual subjects defining clusters), and then successively iteratively merging pairs of clusters. Ward's method is an example of agglomerative clustering that uses the sum of squares to reduce the variance between members of each cluster (e.g., it is a minimum variance agglomerative clustering technique). See Murtagh and Legendre 2014 "Ward's Hierarchical Agglomerative Clustering Method" J. Class 31, 274-295, which is incorporated herein by reference. A drawback of many agglomerative clustering methods is their high computational requirements. In some embodiments, agglomerative clustering algorithms can be combined with k-means clustering algorithms. Non-limiting examples of agglomerative and k-means clustering are described in Karthikeyan et al. 2020 "A comparative study of k-means clustering and agglomerative hierarchical clustering" Int J Emer Trends Eng Res 8(5), 1600-1604, which is incorporated herein by reference. For example, a k-means clustering algorithm divides multiple subjects into a discrete set of k clusters in data space (e.g., initial k partitions). In some embodiments, k-means clustering is applied iteratively to multiple subjects (e.g., k-means clustering is applied multiple times, e.g., serially, to multiple subjects). In some embodiments, the use of a combination of agglomerative and k-means clustering is less computationally demanding than either agglomerative clustering or k-means clustering alone.

[0147] Block 216. Referring to block 216, in some embodiments, the target model is a convolutional neural network.

[0148] In some embodiments (e.g., when at least one target object is a polymer having an active site and the test object is a chemical composition), the test object description provided for each target object is obtained by docking an atomic representation of the test object to an atomic representation of the active site of the polymer. Non-limiting examples of such docking include Liu and Wang, 1999, "MCDOCK: A Monte Carlo simulation approach to the molecular docking problem," Journal of Computer-Aided Molecular Design 13, 435-451; Shoichet et al., 1992, "Molecular docking using shape descriptors," Journal of Computational Chemistry 13(3), 380-397; Knegtel et al., 1997, "Molecular docking to ensembles of protein structures," Journal of Molecular Biology 266, 424-440; Morris et al., 2009, "AutoDock4 and AutoDockTools4: Automated Docking with Selective Receptor Flexibility," J Comput Chem 30(16), 2785-2791; Sotriffer et al., 2000, "Automated docking of ligands to antibodies:methods and applications,”Methods:A Companion to Methods in Enzymology 20,280-291, Morris et al.,1998,“Automated Docking Using a Lamarckian Genetic Algorithm and Empirical Binding Free Energy Function,”Journal of Computational Chemistry 19:1639-1662, and Rarey et al.and W., 1996, "A Fast Flexible Docking Method Using an Incremental Construction Algorithm," Journal of Molecular Biology 261, 470-489, each of which is incorporated herein by reference. This pose description of each subject relative to at least one target object is then applied to the target model. In some such embodiments, the subject is a chemical compound, each target object comprises a polymer having a binding pocket, and presenting the subject description relative to each target object comprises docking modeled atomic coordinates for the chemical compound to atomic coordinates for the binding pocket.

[0149] In some embodiments, each test subject is a chemical compound that is presented to one or more target subjects and presented to a target model using any of the techniques disclosed in U.S. Pat. Nos. 10,546,237, 10,482,355, 10,002,312, and 9,373,059, each of which is incorporated herein by reference.

[0150] In some embodiments, the convolutional neural network includes an input layer, multiple individually weighted convolutional layers, and an output scorer, as described in U.S. Patent No. 10,002,312, entitled "Systems and Methods for Applying a Convolutional Network to Spatial Data," issued June 19, 2018, and incorporated herein in its entirety. For example, in some such embodiments, the convolutional layers of the target model include an initial layer and a final layer. In some embodiments, the final layer may include gating using a threshold function or activation function f, which may be a linear or nonlinear function. The activation function may be, for example, a rectified linear unit (ReLU) activation function, a leaky ReLU activation function, or other functions such as a saturated hyperbolic tangent, identity, binary step, logistic, arctangent, soft sine, parametric rectified linear unit, exponential linear unit, soft plus, bent identity, softExponential, sinusoid, sine, Gaussian, or sigmoid function, or any combination thereof.

[0151] In response to the input, in some embodiments, the input layer provides values ​​to the initial convolutional layer. Each respective convolutional layer other than the final convolutional layer, in some embodiments, provides an intermediate value as a function of the respective convolutional layer weight and the respective convolutional layer input value to another of the convolutional layers. In some embodiments, the final convolutional layer provides a value to the scorer as a function of the final layer weight and the input value. In this manner, the scorer may score each of the feature vectors describing each subject (e.g., the input vectors described in U.S. Pat. No. 10,002,312) and use these scores together to provide a corresponding target outcome (e.g., a classification described in U.S. Pat. No. 10,002,312) for each respective subject. In some embodiments, the scorer provides a respective single score for each feature vector and uses a weighted average of these scores to provide a corresponding target outcome for each respective subject.

[0152] In some embodiments, the total number of layers (including input and output layers) used in a convolutional neural network ranges from about 3 to about 200. In some embodiments, the total number of layers is at least 3, at least 4, at least 5, at least 10, at least 15, or at least 20. In some embodiments, the total number of layers is at most 20, at most 15, at most 10, at most 5, at most 4, or at most 3. One skilled in the art will recognize that the total number of layers used in a convolutional neural network can have any value within this range, for example, 8 layers.

[0153] In some embodiments, the total number of learnable or trainable parameters, such as weighting coefficients, biases, or thresholds, used in a convolutional neural network ranges from about 1 to about 10,000. In some embodiments, the total number of learnable parameters is at least 1, at least 10, at least 100, at least 500, at least 1,000, at least 2,000, at least 3,000, at least 4,000, at least 5,000, at least 6,000, at least 7,000, at least 8,000, at least 9,000, or at least 10,000. Alternatively, the total number of learnable parameters is any number less than 100, any number between 100 and 10,000, or greater than 10,000. In some embodiments, the total number of learnable parameters is at most 10,000, at most 9,000, at most 8,000, at most 7,000, at most 6,000, at most 5,000, at most 4,000, at most 3,000, at most 2,000, at most 1,000, at most 500, at most 100, at most 10, or at most 1. Those skilled in the art will recognize that the total number of learnable parameters used can have any value within this range.

[0154] Because convolutional neural networks require a fixed input size, some embodiments of the disclosed systems and methods utilizing convolutional neural networks for target models crop the geometric data (target object-analyte complex) to fit within an appropriate bounding box. For example, a cube with sides of 25-40 Å may be used. In some embodiments where the target object and / or analyte are docked into the active site of the target object, the center of the active site serves as the center of the cube.

[0155] In some embodiments, a square cube of fixed dimensions centered on the active site of the target object is used to divide the space into a voxel grid, although the disclosed system is not so limited. In some embodiments, any of a variety of shapes are used to divide the space into a voxel grid. In some embodiments, a polyhedron, such as a rectangular prism or a polyhedron shape, is used to divide the space.

[0156] In embodiments, the grid structure may be configured to resemble an arrangement of voxels. For example, each substructure may be associated with a channel for each atom being analyzed. Also, an encoding method may be provided for numerically representing each atom.

[0157] In some embodiments, the voxel map describing the interface between the subject and target objects takes into account the factor of time and may therefore be four-dimensional (X, Y, Z, and time).

[0158] In some embodiments, other implementations may be used instead of voxels, such as pixels, points, polygonal shapes, polyhedra, or any other type of shape that is multi-dimensional (e.g., 3D, 4D, etc. shapes).

[0159] In some embodiments, the geometric data are normalized by selecting the origin of the X, Y, and Z coordinates to be the center of mass of the target object's binding site, as determined by a cavity-flooding algorithm. For representative details of such algorithms, see Ho and Marshall, 1990, "Cavity search: An algorithm for the isolation and display of cavity-like binding regions," Journal of Computer-Aided Molecular Design 4, pp. 337-354, and Hendlich et al., 1997, "Ligsite: automatic and efficient detection of potential small molecule-binding sites in proteins," J. Mol. Graph. Model 15, no. 6, which are incorporated herein by reference. Alternatively, in some embodiments, the origin of the voxel map is centered on the center of mass of the entire co-complex (of the test object bound to the target object, of the target object alone, or of the test object alone). Basis vectors may optionally be chosen to be the principal moments of the entire co-complex, of the target object alone, or of the test object alone. In some embodiments, the target object is a polymer having an active site, and sampling is performed in a three-dimensional grid fashion with the origin at the center of mass of the active site, sampling the subject at each of the plurality of different poses for the subject and the active site, with corresponding three-dimensional uniform honeycombs for sampling representing a portion of the polymer and the subject centered at the center of mass. In some embodiments, the uniform honeycombs are regular cubic honeycombs, and the polymer and subject portions are cubes of predetermined fixed dimensions. The use of cubes of predetermined fixed dimensions ensures that, in such embodiments, relevant portions of the geometric data are used and each voxel map is the same size.In some embodiments, the predetermined fixed dimensions of the cube are NÅ×NÅ×NÅ, where N is an integer or real value between 5 and 100, an integer between 8 and 50, or an integer between 15 and 40. In some embodiments, the uniform honeycomb is a rectangular prism honeycomb and is a portion of a polymer, and the subject is predetermined fixed dimensions of the rectangular prisms QÅ×RÅ×SÅ, where Q is a first integer between 5 and 100, R is a second integer between 5 and 100, and S is a third integer or real value between 5 and 100, and at least one number in the set {Q,R,S} is not equal to another value in the set {Q,R,S}.

[0160] In some embodiments, every voxel has one or more input channels, which may have various values ​​associated with them, which in one implementation may be on / off and configured to encode the type of atom. The atom type may represent the element of the atom, or the atom type may be further refined to distinguish other atomic characteristics. The atoms present may then be encoded in each voxel. Various types of encoding may be utilized using various techniques and / or methodologies. As an exemplary encoding method, the atomic number of the atom may be utilized, resulting in one value per voxel, ranging from 1 for hydrogen to 118 for ununoctium (or any other element).

[0161] However, as discussed above, other encoding methods may be utilized, such as "one-hot encoding," in which every voxel has many parallel input channels, each of which is either on or off and encodes a certain type of atom. The atom type may represent the element of the atom, and the atom type may be further refined to distinguish other atomic characteristics. For example, SYBYL atom types distinguish single-bonded carbon from double-bonded carbon, triple-bonded carbon, or aromatic carbon. For a description of SYBYL atom types, see Clark et al., 1989, "Validation of the General Purpose Tripos Force Field," 1989, J. Comput. Chem. 10, pp. 982-1012, which is incorporated herein by reference.

[0162] In some embodiments, each voxel further includes one or more channels for distinguishing atoms that are part of the target object or cofactors for a part of the test object. For example, in one embodiment, each voxel further includes a first channel for the target object and a second channel for the test object. If the atoms in the portion of the space represented by the voxel are from the target object, the first channel is set to a value such as "1," and is zero otherwise (e.g., because this portion of the space represented by the voxel does not contain atoms from the test object or contains one or more atoms). Furthermore, if the atoms in the portion of the space represented by the voxel are from the test object, the second channel is set to a value such as "1," and is zero otherwise (e.g., because this portion of the space represented by the voxel does not contain atoms from the target object or contains one or more atoms). Similarly, other channels may additionally (or alternatively) specify further information such as partial charge, polarizability, electronegativity, solvent-accessible space, and electron density. For example, in some embodiments, an electron density map of a target object is overlaid with a set of three-dimensional coordinates, and the creation of a voxel map further samples the electron density map. Examples of suitable electron density maps include, but are not limited to, multiple isomorphic replacement maps, single isomorphic replacement with anomalous signal maps, single wavelength anomalous dispersion maps, multi-wavelength anomalous dispersion maps, and 2F observable -F calculated Maps include: McRee, 1993, Practical Protein Crystallography, Academic Press, which is incorporated herein by reference.

[0163] In some embodiments, voxel encoding in accordance with the disclosed systems and methods may include additional optional encoding refinements. The following two are provided as examples:

[0164] In a first encoding refinement, the required memory can be reduced by reducing the set of atoms represented by a voxel (e.g., by reducing the number of channels represented by a voxel) based on the fact that most elements occur rarely in biological systems. Atoms can be mapped to share the same channel in a voxel either by combining rare atoms (which may therefore only rarely affect system performance) or by combining atoms with similar properties (which may therefore minimize inaccuracies from combinations).

[0165] Another encoding refinement is to have a voxel represent an atomic position by partially activating neighboring voxels. This results in the partial activation of neighboring neurons in the subsequent neural network, moving from one-hot encoding to "several-hot" encoding. For example, this would be the case if a voxel has a van der Waals diameter of 3.5 Å and therefore a voxel with a van der Waals diameter of 1 Å is activated. 3 The volume of the grid when placed is 22.4 Å 3Considering a chlorine atom where σ is a chlorine atom, one can illustrate that voxels inside the chlorine atom will be completely filled, while voxels at the edge of the atom will only be partially filled. Therefore, channels representing chlorine in partially filled voxels will be turned on in proportion to the amount of such voxels inside the chlorine atom. For example, if 50% of a voxel's volume is within a chlorine atom, then 50% of the channels in the voxel representing chlorine will be activated. This may result in a "smoothed" and more accurate representation compared to individual one-hot encoding. Thus, in some embodiments, the test object is a first compound and the target object is a second compound, and the atomic features resulting from sampling are spread across a subset of voxels in the respective voxel maps, where the subset of voxels includes 2 or more voxels, 3 or more voxels, 5 or more voxels, 10 or more voxels, or 25 or more voxels. In some embodiments, the atomic properties consist of a list of atom types (e.g., one of the SYBYL atom types).

[0166] Thus, voxelization (rasterization) of the encoded geometric data (docking of target objects onto test objects) is based on various rules applied to the input data.

[0167] 6 and 7 provide diagrams of two subjects 602 encoded on a two-dimensional grid 600 of voxels, according to some embodiments. FIG. 6 provides two subjects superimposed on the two-dimensional grid. FIG. 7 provides one-hot encoding that uses different shading patterns to encode the presence of oxygen, nitrogen, carbon, and empty space, respectively. As noted above, such encoding may be referred to as "one-hot" encoding. FIG. 7 shows the grid 500 of FIG. 6 without the subjects 502. FIG. 8 shows a diagram of the two-dimensional grid of voxels of FIG. 7 with the voxels numbered.

[0168] In some embodiments, feature geometry is represented in forms other than voxels. Figure 9 provides illustrations of various representations in which features (e.g., atom centers) are represented as 0-D points (representation 902), 1-D points (representation 904), 2-D points (representation 906), or 3-D points (representation 908). Initially, the spacing between points may be chosen randomly. However, as the target model is trained, the points may move closer together or further apart. Figure 10 illustrates the range of possible positions for each point.

[0169] In embodiments in which interactions between subject and target objects are encoded as voxel maps, each voxel map is optionally expanded into a corresponding vector, thereby creating a plurality of vectors, each vector in the plurality of vectors being the same size. In some embodiments, each vector in the plurality of vectors is a one-dimensional vector. For example, in some embodiments, a 20 Å cube on each side is centered on the active site of the target object and sampled with a three-dimensional fixed grid spacing of 1 Å to form corresponding voxels of the voxel map that retain basic voxel structural features, such as atom type, in each channel, as well as, optionally, more complex subject-target object descriptors, as discussed above. In some embodiments, the voxels of this three-dimensional voxel map are expanded into one-dimensional floating-point vectors. In some embodiments in which the target model is a convolutional neural network, the vectorized representation of the voxel map is subjected to the convolutional network.

[0170] In some embodiments, a convolutional layer in a plurality of convolutional layers includes a set of filters (also referred to as kernels). Each filter has a fixed three-dimensional size that is convolved (stepped at a predetermined step rate) across the depth, height, and width of the input volume of the convolutional layer, computing dot products (or other functions) between the entries (weights) of the filter and the input, thereby creating a multidimensional activation map for that filter. In some embodiments, the filter step rate is 1 element, 2 elements, 3 elements, 4 elements, 5 elements, 6 elements, 7 elements, 8 elements, 9 elements, 10 elements, or more than 10 elements of the input space. Thus, if a filter is of size 5, 3 In some embodiments, the filter computes the dot product (or other mathematical function) between connected cubes of input space having a depth of 5 elements, a width of 5 elements, and a height of 5 elements, for a total count of 125 input space values ​​per voxel channel.

[0171] The input space to the initial convolutional layer (e.g., output from the input layer) is formed from either a voxel map or a vectorized representation of the voxel map. In some embodiments, the vectorized representation of the voxel map is a one-dimensional vectorized representation of the voxel map that serves as the input space to the initial convolutional layer. Nevertheless, when a filter convolves its input space and the input space is a one-dimensional vectorized representation of the voxel map, the filter still obtains from the one-dimensional vectorized representation those elements that represent a corresponding connected cube of the fixed space within the target subject-test subject complex. In some embodiments, the filter uses standard bookkeeping techniques to select those elements from the one-dimensional vectorized representation that form the corresponding connected cube of the fixed space of the target subject-test subject complex. Thus, in some cases, this entails taking a non-connected subset of the elements of the one-dimensional vectorized representation to obtain the element values ​​of the corresponding connected cube of the fixed space of the target subject-test subject complex.

[0172] In some embodiments, the filter is initialized (e.g., to Gaussian noise) or trained to have 125 corresponding weights (per input channel) and performs a dot product (or some other form of mathematical operation, such as a function of the 125 input space values, to calculate a first single value (or set of values) of the activation layer corresponding to the filter). In some embodiments, the values ​​calculated by the filter are summed, weighted, and / or biased. To calculate additional values ​​of the activation layer corresponding to the filter, the filter is then stepped (convolved) into one of the three dimensions of the input volume by the step rate (stride) associated with the filter, at which point a dot product or some other form of mathematical operation between the filter weights and the 125 input space values ​​(per channel) is performed at the new location of the input volume. This stepping (convolution) is repeated until the filter has sampled the entire input space according to the step rate. In some embodiments, the boundaries of the input space are zero-padded to control the spatial volume of the output space produced by the convolution layer. In typical embodiments, each of the filters in a convolutional layer canvasses the entire three-dimensional input volume in this manner, thereby forming a corresponding activation map. The collection of activation maps from the filters in a convolutional layer collectively forms the three-dimensional output volume of one convolutional layer, thereby serving as the three-dimensional (three spatial dimensions) input for a subsequent convolutional layer. Thus, every entry in the output volume can also be interpreted as the output of a single neuron (or set of neurons) that looks at a small region of the input space to the convolutional layer and shares parameters with neurons in the same activation map. Thus, in some embodiments, a convolutional layer in a plurality of convolutional layers has multiple filters, each filter in the plurality of filters having a stride of N (in the three spatial dimensions) with stride Y. 3 where N is an integer greater than or equal to 2 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or greater than 10) and Y is a positive integer (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or greater than 10).

[0173] Each layer in the plurality of convolutional layers is associated with a different set of weights. More specifically, each layer in the plurality of convolutional layers includes a plurality of filters, and each filter includes a plurality of independent weights. In some embodiments, the convolutional layers are of dimension 5 3 For example, a convolutional layer may have 128 filters, 128×5×5, or 16,000 weights per channel in the voxel map. Thus, if there are five channels in the voxel map, the convolutional layer would have 16,000×5 weights, or 80,000 weights. In some embodiments, some or all such weights (and optionally biases) of every filter in a given convolutional layer may be tied together, e.g., constrained to be the same.

[0174] In response to input of each vector in the plurality of vectors, the input layer provides a first plurality of values ​​to the initial convolutional layer as a first function of the value of the each vector.

[0175] Each respective convolutional layer other than the final convolutional layer supplies intermediate values ​​to another convolutional layer in the plurality of convolutional layers as a second respective function of (i) a different set of weights associated with the respective convolutional layer and (ii) the input values ​​received by the respective convolutional layer. For example, each respective filter in each convolutional layer canvases the input volume to the convolutional layer (in three spatial dimensions) at each respective filter location according to the characteristic three-dimensional stride of the convolutional layer, and takes the dot product (or some other mathematical function) of the filter weights of the respective filter and the values ​​of the input volume (a connected cube that is a subset of the total input space) at each filter location, thereby generating a calculated point (or set of points) on the activation layer corresponding to the respective filter location. The activation layers of the filters in each convolutional layer collectively represent the intermediate values ​​of the respective convolutional layer.

[0176] The final convolutional layer supplies a final value to the scorer as a third function of (i) a distinct set of weights associated with the final convolutional layer and (ii) the input value received by the final convolutional layer. For example, each respective filter in the final convolutional layer canvases the input volume (in three spatial dimensions) at each respective filter location up to the final convolutional layer according to the characteristic three-dimensional stride of the convolutional layer, and takes the dot product (or some other mathematical function) of the filter's filter weights and the value of the input volume at each location, thereby computing a point (or set of points) on the activation layer corresponding to each filter location. The activation layers of the filters in the final convolutional layer collectively represent the final value supplied to the scorer.

[0177] In some embodiments, the convolutional neural network has one or more activation layers. In some embodiments, the activation layer is a layer of neurons that applies a non-saturating activation function f(x) = max(0,x). This increases the nonlinearity of the decision function and the overall network without affecting the receptive field of the convolutional layer. In other embodiments, the activation layer uses other functions to increase nonlinearity, such as the saturated hyperbolic tangent function f(x) = tanh, f(x) = |tanh(x)|, and the sigmoid function f(x) = (1 + e -x ) -1 In some embodiments of the neural network, non-limiting examples of other activation functions found in other activation layers may include, but are not limited to, logistic (or sigmoid), softmax, Gaussian, Boltzmann weighted averaging, absolute value, linear, rectified linear, bounded rectified linear, soft rectified linear, parameterized rectified linear, mean, max, min, some vector norm LP (for p=1, 2, 3, ..., ∞), square, square root, polynomial, inverse quadratic, inverse polynomial, polyharmonic spline, thin plate spline.

[0178] In some embodiments, zero or more of the layers of the target model (in embodiments in which the target model is a convolutional neural network) may consist of pooling layers. As in convolutional layers, a pooling layer is a set of function computations that apply the same function to different spatially localized patches of the input. For a pooling layer, the output is given by a pooling operator, e.g., some vector-norm LP for p=1, 2, 3, ..., ∞, over some voxels. Pooling is typically performed per channel rather than across channels. Pooling divides the input space into a set of three-dimensional boxes and outputs the maximum value for each such subregion. The pooling operation provides a form of translation invariance. The function of a pooling layer is to gradually reduce the spatial size of the representation, reducing the amount of parameters and computation in the network and therefore also controlling overfitting. In some embodiments, pooling layers are inserted between successive convolutional layers in a target model in the form of a convolutional neural network. Such pooling layers operate independently on every depth slice of the input and spatially vary their size. In addition to max pooling, the pooling unit may also perform other functions such as average pooling or even L2 norm pooling.

[0179] In some embodiments, zero or more of the layers of the target model (in embodiments in which the target model is a convolutional neural network) may consist of normalization layers, such as local response normalization or local contrast normalization, which may be applied across channels at the same location or to specific channels across several locations. These normalization layers may promote diversity in the response of several function computations to the same input.

[0180] In some embodiments, the scorer (in embodiments in which the target model is a convolutional neural network) includes multiple fully connected layers and an evaluation layer in which the fully connected layers of the multiple fully connected layers feed into an evaluation layer. Neurons in the fully connected layer have full connections to all activations of the previous layer, as in typical neural networks. Thus, their activations can be calculated by matrix multiplication followed by a bias offset. In some embodiments, each fully connected layer has 512 hidden units, 1024 hidden units, or 2048 hidden units. In some embodiments, the scorer has zero fully connected layers, one fully connected layer, two fully connected layers, three fully connected layers, four fully connected layers, five fully connected layers, six or more fully connected layers, or ten or more fully connected layers.

[0181] In some embodiments, the evaluation layer discriminates between multiple activity classes, hi some embodiments, the evaluation layer includes a logistic regression cost layer spanning two activity classes, three activity classes, four activity classes, five activity classes, or six or more activity classes.

[0182] In some embodiments, the evaluation tier includes a logistic regression cost tier across multiple activity classes, hi some embodiments, the evaluation tier includes a logistic regression cost tier across two activity classes, three activity classes, four activity classes, five activity classes, or six or more activity classes.

[0183] In some embodiments, the evaluation layer distinguishes between two activity classes, the first activity class (first classification) being a test subject's IC for the target subject that exceeds a first binding value. 50 , E.C. 50 , Kd, ​​or KI, and the second activity class (second classification) represents an IC of the test subject against the target subject that is lower than the first binding value. 50 , E.C. 50, Kd, ​​or KI. In some such embodiments, the target result is an indication that the subject has the first activity or the second activity. In some embodiments, the first binding value is 1 nanomolar, 10 nanomolar, 100 nanomolar, 1 micromolar, 10 micromolar, 100 micromolar, or 1 millimolar.

[0184] In some embodiments, the evaluation layer includes a logistic regression cost layer across two activity classes, where the first activity class (first classification) is determined by the IC of the subject against the target subject that exceeds a first binding value. 50 , E.C. 50 , Kd, ​​or KI, and the second activity class (second classification) represents an IC of the test subject against the target subject that is lower than the first binding value. 50 , E.C. 50 , Kd, ​​or KI. In some such embodiments, the target result is an indication that the subject has the first activity or the second activity. In some embodiments, the first binding value is 1 nanomolar, 10 nanomolar, 100 nanomolar, 1 micromolar, 10 micromolar, 100 micromolar, or millimolar.

[0185] In some embodiments, the evaluation layer distinguishes between three activity classes, with the first activity class (first classification) being a test subject's IC20 binding to the target subject that exceeds a first binding value. 50 , E.C. 50 , Kd, ​​or KI, and the second activity class (second classification) represents the IC of the test subject against the target subject between the first binding value and the second binding value. 50 , E.C. 50 , Kd, ​​or KI, and a third activity class (third classification) is one in which the IC of the test subject against the target subject is less than the second binding value. 50 , E.C. 50 , Kd, ​​or KI, and the first binding value is other than the second binding value. In some such embodiments, the target result is an indication that the subject has the first activity, the second activity, or the third activity.

[0186] In some embodiments, the evaluation layer includes a logistic regression cost layer across three activity classes, where the first activity class (first classification) is determined by the IC of the subject against the target subject that exceeds a first binding value. 50 , E.C. 50 , Kd, ​​or KI, and the second activity class (second classification) represents the IC of the test subject against the target subject between the first binding value and the second binding value. 50 , E.C. 50 , Kd, ​​or KI, and a third activity class (third classification) is one in which the IC of the test subject against the target subject is less than the second binding value. 50 , E.C. 50 , Kd, ​​or KI, and the first binding value is other than the second binding value. In some such embodiments, the target result is an indication that the subject has the first activity, the second activity, or the third activity.

[0187] In some embodiments, the scorer (in embodiments in which the target model is a convolutional neural network) comprises a fully connected single-layer or multi-layer perceptron. In some embodiments, the scorer comprises a support vector machine, a random forest, or a nearest neighbor. In some embodiments, the scorer assigns a numerical score indicating the strength (or confidence or probability) of classifying the inputs into various output categories. In some cases, the categories are classified as binders and non-binders, or alternatively, potency levels (e.g., IC<1 molar, <1 millimolar, <100 micromolar, <10 micromolar, <1 micromolar, <100 nanomolar, <10 nanomolar, <1 nanomolar). 50 , E.C. 50 or KI efficacy). In some such embodiments, the target outcome is that the indication is the discrimination of the subject into one of these categories.

[0188] Details for obtaining a target result for a composite target model of a subject and a target object have been described above. As discussed above, in some embodiments, each subject is docked in multiple poses with respect to the target object. Presenting all such poses to the target model at once may require a very large input field (e.g., an input field of size equal to the number of voxels * the number of channels * the number of poses when the target model is a convolutional neural network). In some embodiments, all poses are presented to the target model simultaneously, while in other embodiments, each such pose is processed into a voxel map, vectorized, and serves as a sequential input to the target model (e.g., when the target model is a convolutional neural network). In this manner, multiple scores are obtained from the target model, and each score in the multiple scores corresponds to the input of a vector in the multiple vectors to the input layer of the scorer for the target model. In some embodiments, the scores for each of the poses of a given subject with a given target object are combined together (e.g., as a weighted average of the scores, as a measure of central tendency of the scores, etc.) to generate a final target result for each subject.

[0189] In some embodiments where the scorer outputs of the target model are numerical, the outputs may be combined using any of the activation functions described herein, or any known or to be developed activation functions. Examples include, but are not limited to, the non-saturating activation function f(x)=max(0,x), the saturated hyperbolic tangent function f(x)=tanh, f(x)=|tanh(x)|, the sigmoid function f(x)=(1+e -x ) -1 , logistic (or sigmoid), softmax, Gaussian, Boltzmann weighted averaging, absolute value, linear, rectified linear, bounded rectified linear, soft rectified linear, parameterized rectified linear, mean, max, min, some vector norm LP (for p=1, 2, 3, ..., ∞), square, square root, polynomial, inverse quadratic, inverse polynomial, polyharmonic spline, thin plate spline may be mentioned.

[0190] In some embodiments of the present disclosure, the target model may be configured to utilize a Boltzmann distribution to combine the outputs, since this is consistent with the physical probability of the poses when the outputs are interpreted as representing binding energies. In other embodiments of the present disclosure, the max() function may also provide a reasonable approximation to the Boltzmann, and is computationally efficient.

[0191] In some embodiments where the scorer outputs of the target model are not numerical, the scorer may be configured to combine the outputs using various ensemble voting schemes, which may include, as illustrative, non-limiting examples, majority voting, weighted averaging, Condorcet, Borda counting, among others, to form the corresponding target result.

[0192] In some embodiments, the system may be configured to apply an ensemble of scorers, for example, to generate an index of binding affinity.

[0193] In some embodiments, the subject is a chemical compound, and characterizing the subject (e.g., determining a classification) using the multiple scores (from multiple poses of the subject) includes taking a measure of central tendency of the multiple scores. When the measure of central tendency meets a predetermined threshold or a predetermined threshold range, the subject is deemed to have a first classification. When the measure of central tendency falls short of meeting a predetermined threshold or a predetermined threshold range, the subject is deemed to have a second classification. In some such embodiments, the target result output by the target model for each subject is an indication of one of these classifications.

[0194] In some embodiments, using a plurality of scores to characterize the subject includes taking a weighted average of the plurality of scores (from a plurality of poses of the subject). When the weighted average meets a predetermined threshold or a predetermined threshold range, the subject is considered to have a first classification. When the weighted average does not meet a predetermined threshold or a predetermined threshold range, the subject is considered to have a second classification. In some embodiments, the weighted average is a Boltzmann average of the plurality of scores. In some embodiments, the first classification is determined by the subject's IC for the target object exceeding a first binding value (e.g., 1 nanomolar, 10 nanomolar, 100 nanomolar, 1 micromolar, 10 micromolar, 100 micromolar, or 1 millimolar). 50 , E.C. 50 , Kd, ​​or KI, and the second classification is an IC of the test subject for the target subject that is below the first binding value. 50 , E.C. 50 , Kd, ​​or KI. In some such embodiments, the target result output by the target model for each subject is a representation of one of these categories.

[0195] In some embodiments, using the plurality of scores to provide a target outcome for the subject includes taking a weighted average of the plurality of scores (from the plurality of poses of the subject). When the weighted average satisfies a respective threshold range in the plurality of threshold ranges, the subject is considered to have a respective classification in the plurality of respective classifications that uniquely corresponds to a respective threshold range. In some embodiments, each respective classification in the plurality of classifications represents the subject's IC for the target subject. 50 , E.C. 50 , Kd, ​​or KI range (e.g., 1 micromolar to 10 micromolar, 1 nanomolar to 100 nanomolar).

[0196] In some embodiments, a single pose of each respective subject relative to a given target object is run through the target model, and the respective scores assigned by each target model for each subject based thereon are used to classify the subject.

[0197] In some embodiments, a weighted average of target model scores of one or more poses of the subject for each of a plurality of target objects evaluated by a target model using the techniques disclosed herein is used to provide a target result for the subject. For example, in some embodiments, the plurality of target objects are derived from a molecular dynamics run, and each target object in the plurality of target objects represents the same polymer at a different time step during the molecular dynamics run. Each voxel map of one or more poses of the subject for each of these target objects is evaluated by the target model to obtain a score for each independent pose-target object pair, and a weighted average of these scores, or some other measure of central tendency of these scores, is used to provide a target result for the target object.

[0198] Block 218. Referring to block 218 of FIG. 2A, in some embodiments, at least one target entity is a single entity (e.g., each target entity is a respective single entity). In some embodiments, the single entity is a polymer. In some embodiments, the polymer includes an active site (e.g., the polymer is an enzyme having an active site). In some embodiments, the polymer is an assembly of a protein, a polypeptide, a polynucleic acid, a polyribonucleic acid, a polysaccharide, or any combination thereof. In some embodiments, the single entity is an organometallic complex. In some embodiments, the single entity is a surfactant, a reverse micelle, or a liposome.

[0199] In some embodiments, each test subject in the plurality of test subjects comprises a respective chemical compound that may or may not bind to an active site of at least one target subject with a corresponding affinity (e.g., affinity for forming a chemical bond to at least one target subject).

[0200] In some embodiments, the at least one target object includes at least two target objects, at least three target objects, at least four target objects, at least five target objects, or at least six target objects. In some embodiments, each target object is a respective single object (e.g., a single protein, a single polypeptide, etc.), as described above. In some embodiments, one or more target objects of the at least one target object include multiple objects (e.g., a protein complex and / or an enzyme with multiple subunits, such as a ribosome).

[0201] Block 220. Referring to block 220 of FIG. 2B, the method proceeds by training an initial state predictive model using at least i) a subset of test subjects as independent variables and ii) a corresponding subset of target outcomes as dependent variables, thereby updating the predictive model to an updated trained state. That is, the predictive model is trained to predict what the target outcome (target model score) will be for a given test compound without incurring the computational expense of the target model. Moreover, in some embodiments, the predictive model does not utilize at least one target object. In such embodiments, the predictive model attempts to predict the target model score solely based on information provided for the subject in the test subject dataset (e.g., the chemical structure of the subject), rather than on interactions between the subject and one or more target objects.

[0202] Referring to block 222, in some embodiments, the target model exhibits a first computational complexity when evaluating each subject, and the predictive model exhibits a second computational complexity when evaluating each subject, the second computational complexity being less than the first computational complexity (e.g., the predictive model requires less time and / or less computational effort to provide each predicted result for a subject than the target model requires to provide the corresponding target result for the same subject).

[0203] As used herein, the term "computational complexity" is interchangeable with the term "time complexity" and refers to the time required to obtain results when applying a model to a subject and at least one target subject using a given number of processors, and also refers to the number of processors required to obtain results when applying a model to a subject and at least one target subject within a given time period, where each processor has a given amount of processing power. Thus, as used herein, computational complexity refers to the prediction complexity of a model. However, in some embodiments, the target model exhibits a first training computational complexity, and the prediction model exhibits a second training computational complexity, where the second training computational complexity is less than the first training computational complexity. Table 2 below lists some exemplary prediction models for making predictions and their estimated computational complexities (prediction complexities). [Table 2]

[0204] In Table 2, p is the number of subject features evaluated by the classifier in providing the classifier result, and n trees is the number of trees (for various tree-based methods), and O refers to the Bachmann-Landau notation, which refers to an upper bound on the growth rate of a function. See, e.g., Arora and Barak, 2009, Computational Complexity: A Modern Approach, Cambridge University Press, Cambridge England. In contrast, one estimate of the total time complexity of a convolutional neural network, which is one form of training model, is

number

[0205] Block 224. Referring to block 224 of FIG. 2B, in some embodiments, the predictive model in its initial trained state includes an untrained or partially trained classifier. For example, in some embodiments, the predictive model is partially trained on the subject or on other forms of data, such as assay data, that are separate and distinct from the data provided by the multiple subjects in the subject dataset, e.g., using transfer learning techniques. In one example, the predictive model is partially trained on binding affinity data for a set of compounds, which may or may not be present in the subject dataset, using transfer learning techniques.

[0206] Referring to block 226, in some embodiments, the updated trained state predictive model includes untrained or partially trained classifiers that are distinct from the initial trained state predictive model (e.g., one or more weights of the predictive model have been modified). The ability to retrain or update existing classifiers is particularly useful when the training dataset is subject to changes (e.g., when the training dataset increases the size and / or number of classes).

[0207] In some embodiments, a boosting algorithm is used to update (train) the predictive model. Boosting algorithms are generally described by Dai et al. 2007 "Boosting for transfer learning" in Proc 24th Int Conf on Mach Learn, which is incorporated herein by reference. Boosting algorithms can include reweighting data (e.g., a subset of subjects) previously used to train a predictive model when new data (e.g., an additional subset of subjects) is added to the dataset used to retrain or update the predictive model. See, for example, Freund et al. 1997 "A decision-theoretic generalization of online learning and an application to boosting" J Computer and System Sciences 55(1), 119-139, which is incorporated herein by reference.

[0208] In some embodiments, as discussed above, depending on the type of algorithm used for the initial trained state predictive model (e.g., if the predictive model is not a single decision tree), a transfer learning method is used to update the predictive model to an updated trained state (e.g., with each successive iteration of the method). Transfer learning generally involves transferring knowledge from a first model to a second model (e.g., knowledge either from a first set of tasks or from a first dataset to a second set of tasks or second dataset). Additional reviews of transfer learning methods can be found in Torrey et al. 2009 "Transfer Learning" in the Handbook of Research on Machine Learning Applications, Pan et al. 2009 "A Survey on Transfer Learning" IEEE Transactions on Knowledge and Data Engineering doi:10.1109 / TKDE.2009.191, and Molochanov et al. 2016 "Pruning Convolutional Neural Networks for Resource Efficient Transfer Learning" arXiv:1611.06440v1, each of which is incorporated herein by reference. In some embodiments, variations on random forests can be used with dynamic training datasets. See Ristin et al. 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3654-3661, which is incorporated herein by reference.

[0209] In some embodiments, the predictive model comprises a random forest tree, a random forest including multiple multi-additive decision trees, a neural network, a graph neural network, a dense neural network, principal component analysis, nearest neighbor analysis, linear discriminant analysis, quadratic discriminant analysis, a support vector machine, an evolutionary method, projection pursuit, regression, a naive Bayes algorithm, or an ensemble thereof.

[0210] Random Forest, Decision Tree, and Boosted Tree Algorithms. Decision trees are generally described by Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, 395-396, which is incorporated herein by reference. A random forest is generally defined as a collection of decision trees. Tree-based methods divide feature space into a set of rectangles and fit a model (such as a constant) to each rectangle. In some embodiments, the decision tree includes random forest regression. One specific algorithm that can be used for predictive models is classification and regression trees (CART). Other specific decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, 396-408 and 411-412, which are incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Random forests in general are described in Breiman, 1999, Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is incorporated herein by reference in its entirety.

[0211] Neural networks, graph neural networks, dense neural networks. Various neural networks may be employed as either or both the target model and / or the predictive model, provided that the predictive model has a lower computational complexity than the target model. Neural network algorithms, including convolutional neural network (CNN) algorithms, are disclosed, for example, in Vincent et al., 2010, J Mach Learn Res 11, 3371-3408; Larochelle et al., 2009, J Mach Learn Res 10, 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference. In some embodiments, other variations of neural network algorithms are used for the predictive model, including, but not limited to, graph neural networks (GNNs) and dense neural networks (DNNs). Graph neural networks are useful for data represented in non-Euclidean spaces (e.g., particularly complex datasets). Overviews of GNNs are provided by Wu et al. (2019), “A Comprehensive Survey on Graph Neural Networks,” arVix:1901.00596, and Zhou et al. (2018), “Graph Neural Networks: A Review of Methods and Applications,” arVix:1812.08434. GNNs can be combined with other data analysis methods to enable drug discovery. See, for example, Altre-Tran et al. (2017), “Low Data Drug Discovery with One-Shot Learning,” ACS Cent Sci 3, 283-293.Dense neural networks generally contain many neurons in each layer and are described in Montavon et al. 2018, “Methods for interpreting and understanding deep neural networks,” Digit Signal Process 73, 1-15, and Finnegan et al. 2017, “Maximum entropy methods for extracting the learned features of deep neural networks,” PLoS Comput Biol. 13(10), 1005836, each of which is incorporated herein by reference.

[0212] Principal Component Analysis. Principal component analysis is one of several methods often used for dimensionality reduction of complex data (e.g., to reduce the number of subjects under consideration). An example of using PCA for data clustering is provided, for example, by Yeung and Ruzzo 2001 "Principal component analysis for clustering gene expression data" Bioinformat 17(9), 763-774, which is incorporated herein by reference. Principal components are typically ordered by the extent of variance present (e.g., only the first n components are considered to convey signal instead of noise) and are unrelated (e.g., each component is orthogonal to the other components).

[0213] Nearest Neighbor Analysis. Nearest neighbor analysis is typically performed using Euclidean distance. An example of nearest neighbor analysis is provided by Weinberger et al. 2006 "Distance metric learning for large margin nearest neighbor classification" in NIPS MIT Press 2,3. Nearest neighbor analysis is beneficial in some embodiments because it is effective in settings with large training data sets. See Sonawane 2015 "A Review on Nearest Neighbor Techniques for Large Data" International Journal of Advances Research in Computer and Communication Engineering 4(11), 459-461, which is incorporated herein by reference.

[0214] Linear Discriminant Analysis. Linear discriminant analysis (LDA) is typically performed to characterize a class of subjects or to identify linear combinations of features that characterize distinct classes. Examples of LDA are provided by Ye et al. 2004, "Two-Dimensional Linear Discriminant Analysis," Advances in Neural Information Processing Systems 17, 1569-1576, and Prince et al. 2007, "Probabilistic Linear Discriminant Analysis for Inferences about Identity," 11th International Conference on Computer Vision, 1-8. LDA is advantageous because it can be applied to both large and small sample sizes and can be used in high dimensions. See Kaipatnen 1997, "Utilizing Geometric Anomalies of High Dimension: When Complexity Makes Computation Easier," Computer-Intensive Methods in Control and Signal Processing, 283-294.

[0215] Quadratic Discriminant Analysis. Quadratic discriminant analysis (QDA) is closely related to LDA, but in QDA, a separate covariance matrix is ​​estimated for every class of interest. See Wu et al. 1996, "Comparison of regularized discriminant analysis, linear discriminant analysis, and quadratic discriminant analysis, applied to NIR data," Analytica Chimica Acta 329, 257-265. Examples of QDA are provided by Zhang 1997, "Identification of protein coding regions in the human genome by quadratic discriminant analysis," PNAS 94, 565-568, and Zhang et al. 2003, "Splice site prediction with quadratic discriminant analysis using diversity measure," Nucleic Acids Res 31(21), 6124-6220, each of which is incorporated herein by reference. QDA is advantageous because it provides more effective parameters than LDA, as described in Wu et al. 1996, "Comparison of regularized discriminant analysis, linear discriminant analysis, and quadratic discriminant analysis, applied to NIR data," Analytica Chimica Acta 329, 257-265, which is incorporated herein by reference.

[0216] Support Vector Machine. Non-limiting examples of support vector machine (SVM) algorithms include Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press; Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al. al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in its entirety. When used for classification, SVMs separate a given binary labeled data training set using a hyperplane that is maximally far from the labeled data. When linear separation is not possible, SVMs can operate in conjunction with the technique of "kernels" that automatically achieve a nonlinear mapping to the feature space. The hyperplane found by the SVM in the feature space corresponds to a nonlinear decision boundary in the input space.

[0217] Linear regression. As used herein, linear regression can encompass simple, multivariate, and / or multivariate linear regression analysis. Linear regression uses a linear approach to model the relationship between a dependent variable (also known as a scalar response) and one or more independent variables (also known as explanatory variables), and therefore can be used as a predictive model in the present disclosure. See Altman et al. 2015 "Simple Linear Regression" Nature Methods 12,999-1000, which is incorporated herein by reference. The relationship is predicted using a linear predictor function, the parameters of which are estimated using a linear model. In some embodiments, simple linear regression is used to model the relationship between a dependent variable and a single independent variable. An example of simple linear regression can be found in Altman et al. 2015 "Simple Linear Regression" Nature Methods 12,999-1000, which is incorporated herein by reference.

[0218] In some embodiments, multiple linear regression is used to model the relationship between a dependent variable and multiple independent variables and can therefore be used as a predictive model in the present disclosure. An example of multiple linear regression can be found in Sousa et al. 2007, "Multiple linear regression and artificial neural networks based on principal components to predict ozone concentration," Environ Model & Soft 22(1), 97-103, which is incorporated herein by reference. In some embodiments, multivariate linear regression is used to model the relationship between multiple dependent variables and any number of independent variables. A non-limiting example of multivariate linear regression can be found in Wang et al. 2016, "Discriminative Feature Extraction via Multivariate Linear Regression for SSVEP-Based BCI," IEEE Transactions on Neural Systems and Rehabilitation Engineering 24(5), 532-541, which is incorporated herein by reference.

[0219] Naive Bayes Algorithm. Naive Bayes classifiers (algorithms) are a family of "probabilistic classifiers" based on applying Bayes' theorem with a strong (naive) independence assumption between features. In some embodiments, they are combined with kernel density estimation. See Hastie, Trevor, 2001, The elements of statistical learning: data mining, inference, and prediction, Tibshirani, Robert, Friedman, JH (Jerome H.), New York: Springer, which is incorporated herein by reference.

[0220] In some embodiments, training the predictive model at an initial trained state using at least i) a subset of the subject subjects as independent variables of the predictive model and ii) a corresponding subset of the target outcomes as dependent variables of the predictive model further comprises iii) using at least one target subject as an independent variable of the predictive model to update the predictive model to an updated trained state.

[0221] Blocks 228-230. Referring to block 228 of FIG. 2B, the method proceeds by applying the updated, trained predictive model (e.g., the retrained predictive model) to a complete plurality of test subjects, thereby obtaining a plurality of instances of predicted outcomes. Referring to block 230, in some embodiments, the plurality of instances of predicted outcomes includes a respective predicted outcome for each test subject in the plurality of test subjects. In this manner, a balance is achieved between a high computational burden and corresponding performance improvement of the target model and a low computational burden and corresponding poorer performance of the predictive model. The target model is used to obtain target outcomes for only a subset of the test subjects, thereby forming a training set for training the predictive model. This training set is likely more accurate due to the performance of the more computationally intensive target model and the fact that it utilizes interactions between at least one target subject and the test subject. For example, in some embodiments, the target subject is an enzyme having an active site, and the target model scores interactions between each test subject and the target subject in the subset of test subjects. The training set is then used to train the predictive model. Thus, in a typical embodiment, a predictive model is trained using a training set, the training set including a target model score for each subject of a subset of test subjects, and chemical data is provided for each such subject in the test subject dataset, thereby enabling the predictive model to predict a target model score without using the target subject (e.g., without docking the test subject to the target subject). The predictive model thus trained is then applied to the complete plurality of test subjects to obtain a plurality of predicted result instances. The predicted result instances include the scores that the trained predictive model predicts to be the target model scores for each subject in the complete plurality of target subjects. In this way, the performance of the more computationally demanding target model with which docking occurs is fully utilized to help reduce the number of subjects in the test dataset. Furthermore, a test result for each subject is obtained to fully utilize the efficiency of the predictive model to reduce the number of subjects in the test dataset.

[0222] Blocks 232-234. Referring to block 232 of FIG. 2B , the method proceeds by eliminating a portion of the test subjects from the plurality of test subjects based at least in part on instances of the plurality of predicted outcomes (e.g., according to any of the exclusion criteria described below). In some embodiments, for each respective test subject of a subset of test subjects from the plurality of test subjects, applying a target model to each test subject and at least one target subject to obtain a corresponding target outcome, thereby obtaining a corresponding subset of target outcomes (block 210), training an initial trained state predictive model (block 220), applying the updated trained state predictive model to the plurality of test subjects, thereby obtaining a plurality of instances of predicted outcomes (block 228), and eliminating a portion of the test subjects from the plurality of test subjects based at least in part on instances of the plurality of predicted outcomes (block 232) is an iterative process that is repeated several times (e.g., two, three, more than three, more than ten, more than fifteen, etc.) as subject to evaluation as described in block 236 below. Each time the process is repeated (at each iteration), a portion of the subjects remaining in the plurality of subjects is removed from the plurality of subjects based at least in part on the most recent instance of the plurality of predicted results from block 228.

[0223] Referring to block 234, in some embodiments, the excluding includes: i) clustering the plurality of subjects, thereby assigning each subject in the plurality of subjects to a respective cluster in the plurality of clusters; and ii) excluding a subset of subjects from the plurality of subjects based at least in part on redundancy of subjects in individual clusters in the plurality of clusters (e.g., to ensure a variety of different chemical compounds in the plurality of subjects). In other words, in such embodiments, in each iteration of block 232, the remaining plurality of subjects are clustered. In some embodiments, this clustering is based on feature vectors of the subjects, as described above. In some embodiments, the clustering of block 234 may be performed using any of the clustering methods described in block 214. Whereas in block 214 such clustering was performed to select a subset of subjects for use with the target model, in block 234, clustering is performed to permanently exclude subjects from the plurality of subjects. Consider an example in which the clustering of block 234 clusters the remaining subjects in the plurality of subjects into Q clusters, where Q is a positive integer greater than or equal to 2 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, more than 10, more than 20, more than 30, more than 100, etc.). In some such embodiments, the same number of subjects in each of these clusters are retained in the plurality of subjects, and all other subjects are removed from the plurality of subjects. In this way, the remaining subjects in the plurality of subjects are balanced across all clusters.

[0224] The plurality of prediction results generated in step 232 represent scores that the predictive model predicts that the target model will call for the plurality of test subjects.

[0225] When scoring is performed in a scheme in which lower scores represent compounds with better affinity for one or more target objects, it is interesting to remove those subjects with high scores. Thus, in some alternative embodiments, clustering is not used, and the removing of block 232 includes: i) ranking the plurality of subjects based on instances of the plurality of predicted results; and ii) removing from the plurality of subjects those subjects in the plurality of subjects that do not have a corresponding predicted score that meets a threshold cutoff (e.g., to ensure that the remaining subjects in the plurality of subjects have high predicted scores). In some embodiments, the threshold cutoff is an upper threshold percentage (e.g., the percentage of the plurality of subjects that are highest ranked based on the plurality of predicted results). In some such embodiments, the upper threshold percentage represents subjects in the plurality of subjects whose predicted outcome is in the top 90 percent, top 80 percent, top 75 percent, top 60 percent, top 50 percent, top 40 percent, top 30 percent, top 25 percent, top 20 percent, top 10 percent, or top 5 percent of the plurality of predicted outcomes. In such embodiments, the corresponding lower percentage of subjects are excluded from the plurality of subjects for further consideration (e.g., thereby reducing the number of subjects in the plurality of subjects).

[0226] When scoring is performed in a scheme in which higher scores represent compounds with better affinity for one or more target objects, it is interesting to remove those subjects with low scores. Thus, in some alternative embodiments, clustering is not used, and the removing of block 232 includes: i) ranking the plurality of subjects based on instances of the plurality of predicted results; and ii) removing from the plurality of subjects those subjects in the plurality of subjects that do not have a corresponding predicted score that meets a threshold cutoff (e.g., to ensure that the remaining subjects in the plurality of subjects have low predicted scores). In some such embodiments, the threshold cutoff is a lower threshold percentage (e.g., the percentage of the plurality of subjects that is lowest ranked based on the plurality of predicted results). In some embodiments, the lower threshold percentage represents subjects in the plurality of subjects whose predicted outcome is in the bottom 90 percent, bottom 80 percent, bottom 75 percent, bottom 60 percent, bottom 50 percent, bottom 40 percent, bottom 30 percent, bottom 25 percent, bottom 20 percent, bottom 10 percent, or bottom 5 percent of the plurality of predicted outcomes. In such embodiments, the corresponding upper percentage of subjects are excluded from the plurality of subjects for further consideration (e.g., thereby reducing the number of subjects in the plurality of subjects).

[0227] In some embodiments, each instance of excluding (e.g., in embodiments in which the method iteratively excludes a portion of subjects from the plurality of subjects) excludes between one-tenth and nine-tenths of the subjects in the plurality of subjects at a particular iteration of block 232. In some embodiments, each instance of excluding excludes more than 5 percent, more than 10 percent, more than 15 percent, more than 20 percent, or more than 25 percent of the subjects present in the plurality of subjects at a particular iteration of block 232.

[0228] In some embodiments, each instance of excluding excludes between 5 percent and 30 percent, between 10 percent and 40 percent, between 15 percent and 70 percent, between 20 percent and 50 percent, or between 25 percent and 90 percent of the plurality of subjects in a particular iteration of block 232. In some embodiments, each instance of excluding excludes between one-quarter and three-quarters of the subjects of the plurality of subjects in a particular iteration of block 232. In some embodiments, each instance of excluding excludes between one-quarter and one-half of the subjects of the plurality of subjects in a particular iteration of block 232.

[0229] In some embodiments, each instance of eliminating (block 232) eliminates a predetermined number (or portion) of test subjects from the plurality of test subjects. For example, in some embodiments, each instance of eliminating (block 232) eliminates 5 percent of the test subjects present in the plurality of test subjects at each instance of eliminating. In some embodiments, one or more instances of eliminating eliminate a different number (or portion) of test subjects. For example, initial instances of eliminating (block 232) may eliminate a higher percentage of the test subjects present in the plurality of test subjects during these initial instances of eliminating 232, while subsequent instances of eliminating may eliminate a lower percentage of the test subjects present in the plurality of test subjects during these subsequent instances of eliminating 232. For example, an initial instance may eliminate 10 percent of the plurality of test compounds, while a subsequent instance may eliminate 5 percent of the plurality of test compounds. In another example, initial instances of excluding (block 232) may exclude a lower percentage of the plurality of test subjects present in the plurality of test subjects during those initial instances of excluding, while subsequent instances of excluding may exclude a higher percentage of the plurality of test subjects present in the plurality of test subjects during those subsequent instances of excluding 232. For example, in an initial instance of excluding, 5 percent of the plurality of test compounds may be excluded, while in a subsequent instance of excluding 232, 10 percent of the plurality of test compounds may be excluded.

[0230] Block 236. Referring to block 236 of FIG. 2C, the method proceeds by determining whether one or more predefined reduction criteria are met. If the one or more predefined reduction criteria are not met, the method further includes: (i) applying a target model to each of the additional subsets of subjects in the plurality of subjects to obtain a corresponding target outcome, thereby obtaining an additional subset of target outcomes. The additional subset of subjects is selected at least in part based on the instances of the plurality of predicted outcomes. (ii) updating the subset of subjects by incorporating the additional subset of subjects into the subset of subjects (e.g., the previous subset of subjects). (iii) updating the subset of target outcomes by incorporating the additional subset of target outcomes into the subset of target outcomes. Thus, the subset of target outcomes grows as the method incrementally repeats executing the target model, training the predictive model, and executing the predictive model. After updating (ii) and (iii), the predictive model is modified (iv) by applying the predictive model to at least 1) a subset of the test subjects as independent variables and a corresponding subset of the target outcomes as corresponding dependent variables, thereby providing an updated trained state predictive model. The applying (block 228), eliminating (block 232), and determining (block 236) are repeated until one or more predefined reduction criteria are met.

[0231] In some embodiments, modifying the predictive model (iv) includes either retraining or training a new partially trained predictive model.

[0232] In some embodiments, if one or more predefined reduction criteria are met, the method further includes: i) clustering the plurality of subjects, thereby assigning each subject in the plurality of subjects to a cluster in the plurality of clusters; and ii) eliminating one or more subjects from the plurality of subjects based at least in part on redundancy of the subjects in individual clusters in the plurality of clusters.

[0233] In some embodiments, clustering the plurality of subjects is performed as described with respect to block 212.

[0234] Referring to block 238, in some embodiments, applying (i) further includes forming an additional subset of subjects by selecting one or more subjects from the plurality of subjects (e.g., by selecting subjects from diverse clusters) based on evaluation of one or more features selected from the plurality of feature vectors, as described above.

[0235] In some embodiments, the additional subset of subjects is the same or similar size as the subset of subjects. In some embodiments, the additional subset of subjects is a different size than the subset of subjects. In some embodiments, the additional subset of subjects is separate from the subset of subjects.

[0236] In some embodiments, the additional subset of subjects comprises at least 1,000 subjects, at least 5,000 subjects, at least 10,000 subjects, at least 25,000 subjects, at least 50,000 subjects, at least 75,000 subjects, at least 100,000 subjects, at least 250,000 subjects, at least 500,000 subjects, at least 750,000 subjects, at least 1 million subjects, at least 2 million subjects, at least 3 million subjects, at least 4 million subjects, at least 5 million subjects, at least 6 million subjects, at least 7 million subjects, at least 8 million subjects, at least 9 million subjects, or at least 10 million subjects.

[0237] In some embodiments, modifying the predictive model (iv) comprises retraining the predictive model (e.g., re-running the training process on an updated subset of subjects and potentially changing some parameters or hyperparameters of the predictive model). In some embodiments, modifying the predictive model (iv) comprises training a new predictive model (e.g., replacing a previous predictive model).

[0238] In some embodiments, modifying (iv) further comprises, in addition to using at least 1) a subset of the test subjects as independent variables and 2) a corresponding subset of the target outcomes as corresponding dependent variables, 3) using at least one target object as an independent variable. In other words, in some embodiments, the predictive model actually docks the test subjects to the target objects to generate a predictive outcome trained on the target outcome of the target model, provided that the predictive model with docking remains computationally less demanding than the target model with simultaneous binding.

[0239] Referring to block 240, in some embodiments, satisfying the one or more predefined reduction criteria includes correlating the plurality of predicted outcomes with corresponding target outcomes from a subset of target outcomes. For example, in some embodiments, the one or more predefined reduction criteria are satisfied if the correlation between the plurality of predicted outcomes and the corresponding target outcomes is 0.60 or greater, 0.65 or greater, 0.70 or greater, 0.75 or greater, 0.80 or greater, 0.85 or greater, or 0.90 or greater.

[0240] Referring to block 240, in some embodiments, satisfying one or more predefined reduction criteria includes determining an average difference between the plurality of predicted outcomes and the corresponding target outcomes on an absolute or normalized scale, and the one or more predefined reduction criteria are satisfied if this average difference is less than a threshold amount. In such embodiments, the threshold amount is application dependent.

[0241] In some embodiments, satisfying one or more predefined reduction criteria includes determining that the number of subjects in the plurality of subjects has fallen below a threshold number of subjects. In some embodiments, the one or more predefined reduction criteria require that the plurality of subjects have 30 or fewer subjects, 40 or fewer subjects, 50 or fewer subjects, 60 or fewer subjects, 70 or fewer subjects, 90 or fewer subjects, 100 or fewer subjects, 200 or fewer subjects, 300 or fewer subjects, 400 or fewer subjects, 500 or fewer subjects, 600 or fewer subjects, 700 or fewer subjects, 800 or fewer subjects, 900 or fewer subjects, or 1000 or fewer subjects.

[0242] In some embodiments, the one or more predefined reduction criteria require the plurality of subjects to have between 2 and 30 subjects, between 4 and 40 subjects, between 5 and 50 subjects, between 6 and 60 subjects, between 5 and 70 subjects, between 10 and 90 subjects, between 5 and 100 subjects, between 20 and 200 subjects, between 30 and 300 subjects, between 40 and 400 subjects, between 40 and 500 subjects, between 40 and 600 subjects, or between 50 and 700 subjects.

[0243] In some embodiments, satisfying one or more predefined reduction criteria includes determining that the number of subjects in the plurality of subjects has been reduced by a threshold percentage of the number of subjects in the subject database. In some embodiments, the one or more predefined reduction criteria require reducing the plurality of subjects by at least 10% of the subject database, at least 20% of the subject database, at least 30% of the subject database, at least 40% of the subject database, at least 50% of the subject database, at least 60% of the subject database, at least 70% of the subject database, at least 80% of the subject database, at least 90% of the subject database, at least 95% of the subject database, or at least 99% of the subject database.

[0244] In some embodiments, the one or more predefined reduction criteria is a single reduction criterion. In some embodiments, the one or more predefined reduction criteria is a single reduction criterion, and the single reduction criterion is any one of the reduction criteria described in this disclosure.

[0245] In some embodiments, the one or more predefined reduction criteria is a combination of reduction criteria, hi some embodiments, the combination of reduction criteria is any combination of reduction criteria described in this disclosure.

[0246] Referring to block 242, in some embodiments, if one or more predefined reduction criteria are met, the method further includes applying the predictive model to the plurality of subject subjects and at least one target subject, thereby causing the predictive model to provide a respective score for each subject subject in the plurality of subject subjects (e.g., each score is for each subject subject and target subject). In some such embodiments, each respective score corresponds to an interaction between each subject subject and at least one target subject. In some embodiments, each score is used to characterize at least one target subject. In some embodiments, the score refers to binding affinity (e.g., between one or more target subjects and each subject subject), as described in U.S. Pat. No. 10,002,312, entitled "Systems and Methods for Applying a Convolutional Network to Spatial Data," which is incorporated herein in its entirety. In some embodiments, the interaction between the subject subject and the target subject is influenced by distance, angle, atom type, molecular charge and / or polarization, and surrounding stabilizing or destabilizing environmental factors.

[0247] In some alternative embodiments, if one or more predefined reduction criteria are met, the method further includes applying a target model to the remaining plurality of test subjects and at least one target subject, thereby causing the target model to provide a respective target score for each remaining test subject in the plurality of test subjects (e.g., each target score is for each test subject and target subject in the one or more target subjects). In some such embodiments, each respective target score corresponds to an interaction between a respective test subject and at least one target subject. In some embodiments, each target score is used to characterize at least one target subject. In some embodiments, the target score refers to binding affinity (e.g., between each test subject with one or more target subjects), as described in U.S. Pat. No. 10,002,312, entitled "Systems and Methods for Applying a Convolutional Network to Spatial Data," which is incorporated herein in its entirety. In some embodiments, the interaction between the test subject and the target subject is influenced by distance, angle, atom type, molecular charge and / or polarization, and surrounding stabilizing or destabilizing environmental factors.

[0248] Example 1 - Use Example. The following are sample use cases provided for illustrative purposes only to describe some applications of some embodiments of the present invention. Other applications may be contemplated, and the examples provided below are non-limiting and may be subject to modification, omission, or inclusion of additional elements.

[0249] Each of the following examples illustrates binding affinity predictions, but these examples may be found to differ in whether the prediction is made on a single molecule, a set of iteratively modified molecules, or a series of iteratively modified molecules, whether the prediction is made on a single target or on multiple targets, whether activity against the target is desired or avoided, and whether the amount of interest is absolute activity or relative activity, or whether the molecule or target set is specifically selected (e.g., for molecules, so as to be existing drugs or pesticides, for proteins, so as to have known toxicity or side effects).

[0250] Hit Discovery. Pharmaceutical companies spend millions of dollars screening compounds to discover new promising drug leads. Large compound collections are tested to find the few compounds that have any interaction with the disease target of interest. Unfortunately, wet lab screening suffers from experimental error in addition to the cost and time to perform assay experiments, and assembling large screening collections imposes significant challenges through storage constraints, shelf stability, or chemical costs. Even the largest pharmaceutical companies have only hundreds of thousands to millions of compounds, compared to tens of millions of commercially available molecules and hundreds of millions of simulable molecules.

[0251] A potentially more efficient alternative to physical experimentation is virtual high-throughput screening. Just as physical simulations can help aerospace engineers evaluate possible wing designs before the models are physically tested, computational screening of molecules can focus experimental testing on a small subset of likely molecules. This can reduce screening costs and time, reduce false negatives, improve success rates, and / or cover a broader range of chemical space.

[0252] In this application, a protein target may serve as the target object. A large set of molecules may also be provided in the form of a test subject dataset. For each test subject remaining during application of the disclosed method, a binding affinity to the protein target is predicted. The resulting scores can be used to rank the remaining molecules, with the best-scoring molecules being most likely to bind to the target protein. Optionally, the ranked molecule list may be analyzed for clusters of similar molecules, with larger clusters being used as stronger predictors of molecular binding, and molecules may be selected between clusters to ensure diversity in confirmatory experiments.

[0253] Off-target side effect prediction. Many drugs can be found to have side effects. Often, these side effects result from interactions with biological pathways other than those responsible for the drug's therapeutic effect. These off-target side effects can be unpleasant or dangerous, limiting the patient population in which drug use is safe. Therefore, off-target side effects are an important criterion for evaluating which drug candidates to further develop. While it is important to characterize the interactions of drugs with many alternative biological targets, such tests can be expensive and time-consuming to develop and execute. Computational prediction can make this process more efficient.

[0254] In applying embodiments of the present invention, a panel of biological targets associated with significant biological responses and / or side effects may be constructed. In this case, the system may be configured to predict binding to each protein in the panel by sequentially treating each such protein as a target. Strong activity against a particular target (i.e., as strong as a compound known to activate off-target proteins) may implicate the molecule in side effects due to off-target effects.

[0255] Toxicity prediction. Toxicity prediction is a particularly important special case of off-target side effect prediction. Approximately half of drug candidates in late-stage clinical trials fail due to unacceptable toxicity. As part of the new drug approval process (and before a drug candidate can be tested in humans), the FDA requires toxicity test data for a set of targets including cytochrome P450 liver enzymes (inhibition of which can lead to toxicity from drug-drug interactions) or the hERG channel (binding of which can lead to QT prolongation leading to ventricular arrhythmias and other adverse cardiac effects).

[0256] In toxicity prediction, the system identifies off-target proteins as the primary anti-targets (e.g., CYP450, hERG, or 5-HT 2B The molecules may be configured to bind to specific proteins (e.g., receptors). Binding affinities for drug candidates can then be predicted for these proteins by treating each of these proteins as a target entity (e.g., in separate, independent runs). Optionally, the molecules may be analyzed to predict a set of metabolites (subsequent molecules produced by the body during metabolism / degradation of the original molecule), which may also be analyzed for binding to the anti-target. Problematic molecules may be identified and modified to avoid toxicity or to halt the development of the molecular train to avoid wasting additional resources.

[0257] Pesticide Design. In addition to pharmaceutical uses, the pesticide industry uses binding predictions in the design of new pesticides. For example, one requirement of a pesticide is that it kills a single species of interest without adversely affecting other species. For ecological safety, one may want to kill weevils without killing bumblebees.

[0258] For this application, a user may input into the system a set of one or more target protein structures from different species under consideration. A subset of the proteins may be designated as proteins that should be activated, while the remainder may be designated as proteins for which the molecules should be inactive. As with the previous use case, several sets of molecules (whether from an existing database or newly generated) are considered as test molecules for each target protein, and the system returns molecules with the greatest efficacy against the first group of proteins while avoiding the second group.

[0259] conclusion For components, operations, or structures described herein as a single instance, multiple instances may be provided. Finally, boundaries between various components, operations, and data stores are somewhat arbitrary, and particular operations are illustrated in the context of specific illustrative configurations. Other allocations of functionality are contemplated and may be within the scope of implementations. In general, structures and functionality presented as separate components of illustrative configurations may be implemented as combined structures or components. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements are within the scope of implementations.

[0260] As used herein, the term "if" may be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "when it is determined that" or "when [a stated condition or event] is detected" may be interpreted to mean "upon determining" or "in response to determining" or "upon detecting (a stated condition or event)" or "in response to detecting (a stated condition or event)," depending on the context.

[0261] Terms such as "first," "second," etc. may be used herein to describe various elements, but it should be understood that these elements are not limited by these terms. These terms are used only to distinguish one element from another. For example, a first object may be referred to as a second object, and similarly, a second object may be referred to as a first object, without departing from the scope of the present disclosure. Although a first object and a second object are both objects, they are not the same object.

[0262] The foregoing description has included example systems, methods, techniques, instruction sequences, and computing machine program products embodying example implementations. For purposes of explanation, numerous specific details have been set forth in order to provide an understanding of various implementations of the inventive subject matter. However, it will be apparent to those skilled in the art that implementations of the inventive subject matter may be practiced without these specific details. Generally, well-known instruction instances, protocols, structures, and techniques have not been shown in detail.

[0263] The foregoing description has been set forth with reference to specific embodiments for purposes of explanation. However, the exemplary discussion above is not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The implementations were chosen and described to best explain the principles and their practical applications so that others skilled in the art can best utilize the implementations and various modifications suitable for the particular use envisioned.

Claims

1. 1. A method for reducing the number of subjects among a plurality of subjects in a subject dataset, comprising: A) obtaining said subject dataset in electronic form; B) for each respective subject of a subset of subjects from the plurality of subjects, applying a target model to the respective subject and at least one target object to obtain a corresponding target result, thereby obtaining a corresponding subset of target results; C) training a predictive model at an initial trained state using at least i) the subset of subjects as independent variables and ii) the corresponding subset of target outcomes as dependent variables, thereby updating the predictive model to an updated trained state; D) applying the updated trained predictive model to the plurality of subjects, thereby obtaining a plurality of instances of predicted results; E) excluding a portion of the subjects from the plurality of subjects based at least in part on the instances of the plurality of predicted outcomes; and F) determining whether one or more predefined reduction criteria are met, and if the one or more predefined reduction criteria are not met, the method further comprises: (i) for each respective subject of an additional subset of subjects from the plurality of subjects, applying the target model to the respective subject and the at least one target subject to obtain a corresponding target outcome, thereby obtaining an additional subset of target outcomes, wherein the additional subset of subjects is selected at least in part based on the instances of the plurality of predicted outcomes; (ii) updating the subset of subjects by incorporating an additional subset of subjects into the subset of subjects; (iii) updating the subset of target outcomes by incorporating the additional subset of target outcomes into the subset of target outcomes; (iv) after updating (ii) and (iii), modifying the predictive model by applying the predictive model to at least 1) the subset of subjects as a plurality of independent variables of the predictive model, and 2) a corresponding subset of the target outcomes as a corresponding plurality of dependent variables of the predictive model, thereby providing an updated, trained state of the predictive model; (v) repeating said applying (D), eliminating (E), and determining (F), wherein said plurality of subjects comprises at least 100 million subjects prior to applying said instance of eliminating (E).

2. the target model exhibits a first computational complexity; the predictive model exhibits a second computational complexity; The method of claim 1 , wherein the second computational complexity is less than the first computational complexity.

3. 3. The method of claim 1 or 2, wherein the subject dataset comprises a plurality of feature vectors, each feature vector for a respective subject in the plurality of subjects.

4. 4. The method of any one of claims 1-3, wherein said applying B) further comprises randomly selecting one or more subjects from said plurality of subjects to form said subset of subjects.

5. 4. The method of claim 3, wherein said applying B) further comprises selecting one or more subjects from said plurality of subjects of said subset of subjects based on evaluation of one or more features selected from said plurality of feature vectors.

6. The method of claim 3 , wherein each feature vector in the plurality of feature vectors is a one-dimensional vector.

7. 5. The method of claim 3 or 4, wherein said applying F)(i) further comprises forming an additional subset of said subjects by selecting one or more subjects from said plurality of subjects based on evaluation of one or more features selected from said plurality of feature vectors.

8. 8. The method of claim 1, wherein satisfying the one or more predefined reduction criteria comprises comparing each predicted outcome in the plurality of predicted outcomes to a corresponding target outcome from the subset of target outcomes.

9. 8. The method of any one of claims 1-7, wherein satisfying the one or more predefined reduction criteria comprises determining that the number of subjects in the plurality of subjects has fallen below a threshold number of subjects.

10. The method of any one of claims 1 to 9, wherein the target model is a convolutional neural network.

11. 10. The method of any one of claims 1 to 9, wherein the predictive model comprises a random forest tree, a random forest comprising multiple multi-additive decision trees, a neural network, a graph neural network, a dense neural network, principal component analysis, nearest neighbor analysis, linear discriminant analysis, quadratic discriminant analysis, support vector machine, evolutionary method, projection pursuit, linear regression, a naive Bayes algorithm, a multi-category logistic regression algorithm, or an ensemble thereof.

12. the at least one target object is a single object; The method of any one of claims 1 to 11, wherein the single object is a polymer.

13. The method of claim 12 , wherein the polymer comprises an active site.

14. 14. The method of claim 12 or 13, wherein the polymer is an assembly of proteins, polypeptides, polynucleic acids, polyribonucleic acids, polysaccharides, or any combination thereof.

15. The polymer is a set of three-dimensional coordinates {x 1 ,... ,x N } applied to the target model.

16. The polymer is a set of three-dimensional coordinates {x 1 ,... ,x N } applied to the target model.

17. 13. The method of claim 12, wherein the polymer is applied to the target model based on spatial coordinates that are an ensemble of three-dimensional coordinates of the polymer determined by nuclear magnetic resonance, neutron diffraction, or cryo-electron microscopy.

18. The plurality of subjects, prior to application of the instance of excluding E), is at least 500 million subjects, at least 1 billion subjects, at least 2 billion subjects, at least 3 billion subjects, at least 4 billion subjects, at least 5 billion subjects, at least 6 billion subjects, at least 7 billion subjects, at least 8 billion subjects, at least 9 billion subjects, at least 10 billion subjects, at least 11 billion subjects, 20. The method of any one of claims 1-19, comprising a subject, at least 15 billion subjects, at least 20 billion subjects, at least 30 billion subjects, at least 40 billion subjects, at least 50 billion subjects, at least 60 billion subjects, at least 70 billion subjects, at least 80 billion subjects, at least 90 billion subjects, at least 100 billion subjects, or at least 110 billion subjects.

19. 10. The method of claim 0, wherein the one or more predefined reduction criteria require the plurality of subjects to have 30 or fewer subjects, 40 or fewer subjects, 50 or fewer subjects, 60 or fewer subjects, 70 or fewer subjects, 90 or fewer subjects, 100 or fewer subjects, 200 or fewer subjects, 300 or fewer subjects, 400 or fewer subjects, 500 or fewer subjects, 600 or fewer subjects, 700 or fewer subjects, 800 or fewer subjects, 900 or fewer subjects, or 1000 or fewer subjects.

20. 20. The method of any one of claims 1 to 19, wherein each subject in said plurality of subjects represents a chemical compound.

21. 21. The method of any one of claims 1 to 20, wherein the predictive model in the initial trained state comprises an untrained or partially trained classifier.

22. 22. The method of claim 1, wherein the predictive model of the updated trained state comprises an untrained or partially trained classifier that is distinct from the predictive model of the initial trained state.

23. 23. The method of any one of claims 1-22, wherein the subset of subjects comprises at least 1,000 subjects, at least 5,000 subjects, at least 10,000 subjects, at least 25,000 subjects, at least 50,000 subjects, at least 75,000 subjects, at least 100,000 subjects, at least 250,000 subjects, at least 500,000 subjects, at least 750,000 subjects, at least 1 million subjects, at least 2 million subjects, at least 3 million subjects, at least 4 million subjects, at least 5 million subjects, at least 6 million subjects, at least 7 million subjects, at least 8 million subjects, at least 9 million subjects, or at least 10 million subjects.

24. 24. The method of any one of claims 1-23, wherein the additional subset of subjects comprises at least 1,000 subjects, at least 5,000 subjects, at least 10,000 subjects, at least 25,000 subjects, at least 50,000 subjects, at least 75,000 subjects, at least 100,000 subjects, at least 250,000 subjects, at least 500,000 subjects, at least 750,000 subjects, at least 1 million subjects, at least 2 million subjects, at least 3 million subjects, at least 4 million subjects, at least 5 million subjects, at least 6 million subjects, at least 7 million subjects, at least 8 million subjects, at least 9 million subjects, or at least 10 million subjects.

25. 25. The method of claim 23 or 24, wherein the additional subset of subjects is distinct from the subset of subjects.

26. 2. The method of claim 1, wherein (iv) F) modifying the predictive model comprises retraining the predictive model.

27. 2. The method of claim 1, wherein the training (C) further comprises, in addition to using at least i) a subset of the subject subjects as a plurality of independent variables of the predictive model, and ii) a corresponding subset of the target outcomes as a plurality of dependent variables of the predictive model, iii) using the at least one target subject as an independent variable of the predictive model.

28. 30. The method of claim 1 or 27, wherein the at least one target object comprises at least two target objects, at least three target objects, at least four target objects, at least five target objects, or at least six target objects.

29. 2. The method of claim 1, wherein the instances of the plurality of predicted outcomes include a respective predicted outcome for each subject in the plurality of subjects.

30. 30. The method of any one of claims 1-29, wherein said modifying F)(iv) further comprises 3) using said at least one target subject as an independent variable, in addition to using at least 1) a subset of said subject subjects as an independent variable and 2) a corresponding subset of said target outcomes as a corresponding dependent variable of said predictive model.

31. If the one or more predefined reduction criteria are met, the method further comprises: i) clustering the plurality of subjects, thereby assigning each subject in the plurality of subjects to a cluster in a plurality of clusters; ii) excluding one or more subjects from the plurality of subjects based at least in part on redundancy of subjects in individual clusters in the plurality of clusters.

32. The method comprises: i) clustering the plurality of subjects, thereby assigning each subject in the plurality of subjects to a respective cluster in a plurality of clusters; 31. The method of any one of claims 1 to 30, further comprising selecting the subset of subjects from the plurality of subjects by: ii) selecting the subset of subjects from the plurality of subjects based at least in part on redundancy of subjects in individual clusters in the plurality of clusters.

33. 33. The method of any one of claims 1 to 32, wherein if the one or more predefined reduction criteria are met, the method further comprises applying the predictive model to the plurality of subject subjects and the at least one target subject, thereby causing the predictive model to provide a respective interaction score for each subject in the plurality of subject subjects.

34. 34. The method of claim 33, wherein each respective interaction score corresponds to an interaction between a respective test subject and the at least one target subject.

35. 35. The method of claim 33 or 34, wherein each respective interaction score is used to characterize the at least one target object.

36. The excluding (E) i) clustering the plurality of subjects, thereby assigning each subject in the plurality of subjects to a respective cluster in a plurality of clusters; ii) eliminating a subset of subjects from said plurality of subjects based at least in part on redundancy of subjects in individual clusters in said plurality of clusters.

37. 37. The method of claim 31, 32, or 36, wherein clustering the plurality of subjects is performed using a density-based spatial clustering algorithm, a partitional clustering algorithm, an agglomerative clustering algorithm, a k-means clustering algorithm, a supervised clustering algorithm, or an ensemble thereof.

38. The excluding (E) ranking the plurality of subjects based on the instances of the plurality of predicted outcomes; and removing from said plurality of subjects those subjects in said plurality of subjects that do not have a corresponding predicted outcome that meets a threshold cutoff.

39. 39. The method of claim 38, wherein the threshold cutoff is an upper threshold percentage.

40. 40. The method of claim 39, wherein the upper threshold percentage is the top 90 percent, top 80 percent, top 75 percent, top 60 percent, or top 50 percent of the plurality of predicted results.

41. 41. The method of any one of claims 1-40, wherein each instance of said excluding (E) excludes between one-tenth and nine-tenths of said subjects in said plurality of subjects.

42. 41. The method of any one of claims 1-40, wherein each instance of said excluding (E) excludes between one-quarter and three-quarters of the subjects in said plurality of subjects.

43. wherein the at least one target object is a single target object, and for each respective subject in a subset of subjects from the plurality of subject, applying to the respective subject and target object to obtain a corresponding target result B); i) obtaining spatial coordinates of the target object; ii) modeling the respective subject with the target object in each of a plurality of different poses, thereby creating a plurality of voxel maps, each respective voxel map in the plurality of voxel maps including the subject in a respective pose of the plurality of different poses; iii) expanding each voxel map in the plurality of voxel maps into a corresponding vector, thereby creating a plurality of vectors, each vector in the plurality of vectors being the same size; iv) inputting each respective vector in the plurality of vectors to the target model, the target model including: (a) an input layer for sequentially receiving the plurality of vectors; (b) a plurality of convolutional layers; and (c) a scorer; the plurality of convolutional layers includes an initial convolutional layer and a final convolutional layer; each layer in the plurality of convolutional layers is associated with a different set of weights; In response to input of each vector in the plurality of vectors, the input layer provides a first plurality of values ​​to the initial convolutional layer as a first function of the value of the each vector; each respective convolutional layer other than the final convolutional layer provides an intermediate value to another convolutional layer in the plurality of convolutional layers as a second respective function of (a) a different set of the weights associated with the respective convolutional layer and (b) an input value received by the respective convolutional layer; the final convolutional layer providing a final value to the scorer as a third function of (a) the distinct set of weights associated with the final convolutional layer and (b) an input value received by the final convolutional layer; v) obtaining a corresponding plurality of scores from the scorer, each score in the corresponding plurality of scores corresponding to the input of a vector in the plurality of vectors to the input layer; vi) using said plurality of scores to calculate said corresponding target outcome.

44. 44. The method of claim 43, wherein the scorer comprises a plurality of fully connected layers and an evaluation layer, a fully connected layer in the plurality of fully connected layers feeding the evaluation layer.

45. 44. The method of claim 43, wherein the scorer comprises a decision tree, a multiple additive regression tree, a clustering algorithm, principal component analysis, nearest neighbor analysis, linear discriminant analysis, quadratic discriminant analysis, support vector machine, evolutionary method, projection pursuit, and ensembles thereof.

46. 44. The method of claim 43, wherein each vector in the plurality of vectors is a one-dimensional vector.

47. 44. The method of claim 43, wherein the plurality of different poses comprises 2 or more poses, 10 or more poses, 100 or more poses, or 1000 or more poses.

48. 44. The method of claim 43, wherein the plurality of different poses are obtained using a docking scoring function in one of: Markup Chain Monte Carlo Sampling, Simulated Annealing, a Lamarckian Genetic Algorithm, or a Genetic Algorithm.

49. 44. The method of claim 43, wherein the plurality of different poses are obtained by sequential search using a greedy algorithm.

50. 44. The method of claim 43, wherein said using said plurality of scores to calculate said corresponding target outcome comprises determining a measure of central tendency of said plurality of scores.

51. 44. The method of claim 43, wherein said using said plurality of scores to calculate said corresponding target outcome comprises taking a weighted average of said plurality of scores.

52. Each of the plurality of convolution layers has a plurality of filters, and each filter in the plurality of filters has N 3 44. The method of claim 43, wherein a cubic input space of N is convolved with stride Y, where N is an integer greater than or equal to 2, and Y is a positive integer.

53. 53. The method of claim 52, wherein a different set of the weights associated with each convolutional layer is associated with each filter in the plurality of filters.

54. 44. The method of claim 43, wherein the scorer comprises a plurality of fully connected layers and a logistic regression cost layer, a fully connected layer in the plurality of fully connected layers feeding the logistic regression cost layer.

55. 1. A computer system for reducing a number of subjects among a plurality of subjects in a subject dataset, comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising: A) obtaining said subject dataset in electronic form; B) for each respective subject of a subset of subjects from the plurality of subjects, applying a target model to the respective subject and at least one target object to obtain a corresponding target result, thereby obtaining a corresponding subset of target results; C) training a predictive model at an initial trained state using at least i) the subset of subjects as independent variables and ii) the corresponding subset of target outcomes as dependent variables, thereby updating the predictive model to an updated trained state; D) applying the updated trained predictive model to the plurality of subjects, thereby obtaining a plurality of instances of predicted results; E) excluding a portion of the subjects from the plurality of subjects based at least in part on the instances of the plurality of predicted outcomes; and F) determining whether one or more predefined reduction criteria are met, and if the one or more predefined reduction criteria are not met, the method: (i) for each respective subject of an additional subset of subjects from the plurality of subjects, applying the target model to the respective subject and at least one target subject to obtain a corresponding target outcome, thereby obtaining an additional subset of target outcomes, wherein the additional subset of subjects is selected at least in part based on the instances of the plurality of predicted outcomes; (ii) updating the subset of subjects by incorporating an additional subset of subjects into the subset of subjects; (iii) updating the subset of target outcomes by incorporating the additional subset of target outcomes into the subset of target outcomes; (iv) after updating (ii) and (iii), modifying the predictive model by applying the predictive model to at least 1) the subset of subjects as a plurality of independent variables of the predictive model, and 2) a corresponding subset of the target outcomes as a corresponding plurality of dependent variables of the predictive model, thereby providing an updated, trained state of the predictive model; (v) repeating said applying (D), eliminating (E), and determining (F), wherein said plurality of subjects comprises at least 100 million subjects prior to applying said instance of eliminating (E).

56. 1. A non-transitory computer-readable storage medium and one or more computer programs embodied therein, the one or more computer programs comprising instructions, when executed by a computer system, that cause the computer system to perform a method for reducing a number of subjects among a plurality of subjects in a subject dataset, the method comprising: A) obtaining said subject dataset in electronic form; B) for each respective subject of a subset of subjects from the plurality of subjects, applying a target model to the respective subject and at least one target object to obtain a corresponding target result, thereby obtaining a corresponding subset of target results; C) training a predictive model at an initial trained state using at least i) the subset of subjects as independent variables and ii) the corresponding subset of target outcomes as dependent variables, thereby updating the predictive model to an updated trained state; D) applying the updated trained predictive model to the plurality of subjects, thereby obtaining a plurality of instances of predicted results; E) excluding a portion of the subjects from the plurality of subjects based at least in part on the instances of the plurality of predicted outcomes; and F) determining whether one or more predefined reduction criteria are met, and if the one or more predefined reduction criteria are not met, the method further comprises: (i) for each respective subject of an additional subset of subjects from the plurality of subjects, applying the target model to the respective subject and at least one target subject to obtain a corresponding target outcome, thereby obtaining an additional subset of target outcomes, wherein the additional subset of subjects is selected at least in part based on the instances of the plurality of predicted outcomes; (ii) updating the subset of subjects by incorporating an additional subset of subjects into the subset of subjects; (iii) updating the subset of target outcomes by incorporating the additional subset of target outcomes into the subset of target outcomes; (iv) after updating (ii) and (iii), modifying the predictive model by applying the predictive model to at least 1) the subset of subjects as a plurality of independent variables of the predictive model, and 2) a corresponding subset of the target outcomes as a corresponding plurality of dependent variables of the predictive model, thereby providing an updated, trained state of the predictive model; (v) repeating the applying (D), eliminating (E), and determining (F), wherein the plurality of subjects comprises at least 100 million subjects prior to applying an instance of eliminating (E).

Citation Information

Patent Citations

  • Data set selecting device and experiment designing system

    JP2007304782A

  • Systems and methods for applying convolutional networks to spatial data

    JP2019501433A

  • Method for the synthesis of DNA conjugates by micellar catalysis

    WO2018149863A1

  • Mass spectrometry distinguishable synthetic compounds, libraries, and methods thereof

    WO2019050504A1