Discovery of molecules for affecting biological pathways
Patent Information
- Application Number
- PCT/IB2025/052397
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-06
- Filing Date
- 2025-03-05
- Publication Date
- 2025-10-02
AI Technical Summary
Conventional drug-discovery methods are inefficient, expensive, and lack accuracy in identifying molecules that selectively target disease components, often relying on trial-and-error experimentation and large-scale testing.
Utilizing machine-learning techniques, particularly reinforcement learning, to iteratively refine a list of molecules by applying them to biological samples, assessing interactions, and optimizing properties such as toxicity, specificity, and immunogenicity until a safe and effective molecule is identified.
This approach enhances the efficiency and accuracy of drug discovery by converging on molecules that effectively target biological pathways, reducing the time and cost associated with traditional methods.
Smart Images

Figure IB2025052397_02102025_PF_FP_ABST
Abstract
Description
[0001] DISCOVERY OF MOLECULES FOR AFFECTING BIOLOGICAL PATHWAYS
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] The present application claims priority from U.S. Provisional Patent Application No. 63 / 561,963 to Springer, filed March 06, 2024, 2024, entitled "Discovery of molecules for affecting biological pathways," which is incorporated herein by reference.
[0004] FIELD OF EMBODIMENTS
[0005] Embodiments of the present disclosure are related generally to the field of medicine, and specifically to the discovery of molecules for affecting biological pathways.
[0006] BACKGROUND
[0007] Common drug-discovery and development methods pose many challenges. It is generally rare to discover a molecule, either a naturally occurring or an artificially made molecule, that efficiently and selectively targets a component implicated in a disease.
[0008] Traditional drug-discovery methods heavily rely on trial-and-error experimentation and large-scale testing techniques that involve examining large numbers of potential drug compounds, in order to identify those with the desired properties. However, these methods are typically expensive and time consuming, and often yield results with low accuracy. In addition, they can be limited by the availability of suitable testing techniques and the inability to accurately predict their effect in the body.
[0009] Some current drug-discovery methods utilize virtual screening and / or other computational techniques, to, for example, identify from a large pool of molecules those that bind to a specific cellular target.
[0010] SUMMARY
[0011] Some embodiments of the present disclosure utilize machine-learning techniques to improve the efficiency and accuracy of drug discovery by identifying molecules that are both safe and functional. In accordance with some embodiments of the present disclosure, a method for identifying a pathway-affecting molecule for affecting a biological pathway, is provided. For some embodiments, the molecule may include proteins (such as enzymes, antibodies, receptors, and / or protein complexes), nucleic acids or other complex molecules. The method includes receiving, from a computer model, a list including one or more molecules, and causing the computer model to revise the list until the list includes the desired pathway-affecting molecule. The list is revised by the computer model by repeatedly (a) applying each of the molecules in the list to a biological sample, such that the molecule interacts with the biological sample, (b) identifying one or more effects of the interaction, and (c) inputting, to the computer model, the effects of the interaction, such that the computer model revises the list based on the effects. Based on the effects of the interaction of the pathway-affecting molecule with the biological sample, the pathway-affecting molecule, is identified.
[0012] It is noted that in the context of the present application, in the specification and in the claims, “biological sample” typically refers to a biological sample that includes organic compounds (e.g., proteins, nucleic acids, carbohydrates and / or lipids), such that the one or more molecules interact with organic compounds within the biological sample. For example, the biological sample comprises an in vitro or in vivo biological sample including multiple biological cells and / or extracellular material, and the one or more molecules interact with organic compounds within the biological sample.
[0013] For some embodiments, the computer model generates the initial list of one or more molecules in response to inputting to the computer model an indication of biological pathway to be affected by the molecule. For example, the input specifies a target in the biological pathway, such as a receptor or a protein binding site, to which the molecule is to bind to affect the biological pathway. Alternatively, the computer model produces the initial list of one or more molecules in response to inputting to the computer model an intended effect of the molecule, e.g., if the molecule is to cure a diseased cell.
[0014] The computer model runs algorithms using machine-learning techniques to generate, typically based on the above-described types of input, a list of candidate molecules (e.g., protein and / or nucleic acids) and to iteratively revise the list until the suitable pathway-affecting molecule, is identified. Typically, with each list of candidate molecules that is generated by the computer model, the candidate molecules are obtained (e.g., manufactured), and experimental assays are performed in order to allow interactions of the molecules with components of a biological sample (e.g., with organic compounds within the biological samples). The effects of the interactions of the molecules with the biological sample are inputted to the computer model such that the computer model revises the list of candidate molecules based on the observed effects. This process continues until the computer model converges onto a safe and effective molecule for affecting the biological pathway. There is therefore provided in accordance with some embodiments of the present disclosure, a method for identifying a pathway-affecting molecule for affecting a biological pathway, the method including: receiving, from a computer model, a list including one or more molecules; causing the computer model to revise the list until the list includes the pathway-affecting molecule, by repeatedly: applying each of the molecules in the list to a biological sample, such that the molecule interacts with the biological sample, identifying one or more effects of the interaction, and inputting, to the computer model, the effects of the interaction, such that the computer model revises the list based on the effects; and based on the effects of the interaction of the pathway-affecting molecule with the biological sample, identifying the pathway-affecting molecule.
[0015] In some embodiments, the method further includes, prior to receiving the list from the computer model, inputting the biological pathway to the computer model, such that the computer model initializes the list responsively to the biological pathway.
[0016] In some embodiments, the method further includes, prior to receiving the list from the computer model, inputting an intended effect of the pathway-affecting molecule to the computer model, such that the computer model initializes the list responsively to the intended effect.
[0017] In some embodiments, prior to receiving the list from the computer model, the biological pathway is unknown.
[0018] In some embodiments, identifying the effects of the interaction includes identifying the effects of the interaction by acquiring one or more images of the biological sample, and inputting the effects of the interaction to the computer model includes inputting the effects of the interaction to the computer model by inputting the images to the computer model.
[0019] In some embodiments, the effects of the interaction relate to a toxicity of the molecule with respect to the biological sample.
[0020] In some embodiments, the molecule includes a protein, and the biological sample includes other proteins, and identifying the effects of the interaction includes identifying the effects of the interaction by identifying cross-interactions of the protein with the other proteins.
[0021] In some embodiments, identifying the cross-interactions includes identifying the crossinteractions in a protein chip assay. In some embodiments, the molecule includes a protein, and identifying the effects of the interaction includes identifying the effects of the interaction by identifying breakdown of the protein in the biological sample.
[0022] In some embodiments, the effects of the interaction relate to an immunogenicity of the molecule with respect to the biological sample.
[0023] In some embodiments, at at least one point in time, the list includes one or more binder proteins designed to bind to respective target proteins in the biological sample.
[0024] In some embodiments, each of the binder proteins is tagged with two fluorescent tags, one of which is hidden when the binder protein is bound to the target protein.
[0025] In some embodiments, the target proteins include all proteins in a proteome corresponding to the biological sample, and, at the at least one point in time, the list includes a respective binder protein for each of the target proteins.
[0026] In some embodiments, the biological sample includes organic compounds and wherein identifying the one or more effects of the interaction comprises identifying one or more effects of the interaction of the one or more molecules with the organic compounds.
[0027] There is further provided in accordance with some embodiments of the present disclosure, apparatus for identifying a pathway-affecting molecule for affecting a biological pathway, the apparatus including: a computer processor configured to: generate, from a computer model, a list including one or more molecules; drive the computer model to revise the list until the list includes the pathway-affecting molecule, by repeatedly: driving the computer model to revise the list, by inputting, to the computer model, one or more effects of applying each of the molecules in the list to a biological sample, such that the molecule interacts with the biological sample, the computer model is configured to revise the list based on the effects; and identify the pathway-affecting molecule based on the effects of the interaction of the pathway-affecting molecule with the biological sample.
[0028] In some embodiments, the computer processor is further configured to, prior to generating the list from the computer model, input the biological pathway to the computer model, such that the computer model initializes the list responsively to the biological pathway. In some embodiments, the computer processor is further configured to, prior to generating the list from the computer model, input an intended effect of the pathway-affecting molecule to the computer model, such that the computer model initializes the list responsively to the intended effect.
[0029] In some embodiments, prior to generating the list from the computer model, the biological pathway is unknown.
[0030] In some embodiments, the effects of the applying the one or more molecules relate to a toxicity of the one or more molecules with respect to the biological sample.
[0031] In some embodiments, the molecule includes a protein, wherein the biological sample includes other proteins, and the computer processor is configured to input to the computer model the effect of applying the molecule to the biological sample by inputting cross-interactions of the protein with the other proteins.
[0032] In some embodiments, inputting the cross-interactions includes inputting the crossinteractions as determined in a protein chip assay.
[0033] In some embodiments, the molecule includes a protein, and the computer processor is configured to input to the computer model the effect of applying the molecule to the biological sample by inputting breakdown of the protein in the biological sample.
[0034] In some embodiments, the effects of applying the molecule to the biological sample relate to an immunogenicity of the molecule with respect to the biological sample.
[0035] In some embodiments, at at least one point in time, the list includes one or more binder proteins designed to bind to respective target proteins in the biological sample.
[0036] In some embodiments, the target proteins include all proteins in a proteome corresponding to the biological sample, and at the at least one point in time, the list includes a respective binder protein for each of the target proteins.
[0037] There is further provided in accordance with some embodiments of the present disclosure, a computer software product, for identifying a pathway-affecting molecule for affecting a biological pathway, the computer software product including a non-transitory computer-readable medium in which program instructions are stored, which instructions, when read by a computer cause the computer to perform the steps of: generating, from a computer model, a list including one or more molecules; driving the computer model to revise the list until the list includes the pathway-affecting molecule, by repeatedly: driving the computer model to revise the list, by inputting, to the computer model, one or more effects of applying each of the molecules in the list to a biological sample, such that the molecule interacts with the biological sample, the computer model is configured to revise the list based on the effects; and identifying the pathway-affecting molecule based on the effects of the interaction of the pathway-affecting molecule with the biological sample.
[0038] There is further provided, in accordance with some applications of the present disclosure, a method for imaging a characteristic of cell morphology in a biological sample, the method including: receiving, from a computer model, a list including one or more binder proteins designed to bind to respective one or more target proteins in the biological sample; applying each of the binder proteins in the list to the biological sample, such that each binder protein interacts with the target protein biological sample, identifying one or more effects of each of the interactions, and inputting, to the computer model, the effects of each of the interactions; and based on the effects of the interactions between the binder proteins and the target protein, identifying the characteristic of cell morphology in the biological sample.
[0039] In some embodiments, identifying the characteristic of cell morphology includes identifying a spatial distribution of the one or more target proteins in a biological sample.
[0040] In some embodiments, each of the binder proteins is tagged with a detection tag, and wherein identifying the characteristic of cell morphology in the biological sample includes identifying the cell morphology through the detection tag.
[0041] In some embodiments, each of the binder proteins is tagged with two fluorescent tags, one of which is hidden when the binder protein is bound to the target protein.
[0042] In some embodiments, the binder protein is a fluorescent binder protein.
[0043] In some embodiments, the target proteins include all proteins in a proteome corresponding to the biological sample, and, at the at least one point in time, the list includes a respective binder protein for each of the target proteins.
[0044] In some embodiments, the target proteins include a subset of proteins in a proteome corresponding to a biological pathway in the biological sample, and, at the at least one point in time, the list includes a respective binder protein for each of the target proteins.
[0045] In some embodiments, the binder proteins are configured to bind to the target protein without affecting interactions of the target protein with other proteins in the cell proteome.
[0046] In some embodiments, the binder proteins are configured to bind to the target protein in a manner that affects interactions of the target protein with other proteins in the cell proteome.
[0047] The present disclosure will be more fully understood from the following detailed description of embodiments thereof, taken together with the drawings, in which:
[0048] BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Fig. 1 is a block diagram of a system for identifying a biological pathway-affecting molecule, in accordance with some embodiments of the present disclosure.
[0050] DETAILED DESCRIPTION
[0051] OVERVIEW
[0052] Conventional techniques for molecule discovery are inefficient. For example, some techniques involve iterating through a vast number (e.g., hundreds of thousands) of compounds, none of which was designed to address the specific condition for which treatment is required.
[0053] To address this problem, embodiments of the present disclosure provide an iterative technique that harnesses machine learning techniques, such as reinforcement learning, to efficiently discover molecules that are both safe and functional. The molecules may include, for example, proteins, enzymes, antibodies, protein complexes, nucleic acids (deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)), or other complex molecules. In addition to therapeutic uses, the molecules can be for industrial uses or for food supplementation. One example application of embodiments of the present disclosure is the discovery of a protein for downregulating the activity of cancerous cells, e.g., so as to slow the reproduction rates of these cells or to kill these cells, by binding to another protein in the cancerous cells or via any other biological pathway.
[0054] More specifically, in embodiments of the present disclosure, a computer model, which typically includes a neural network, is provided with an initial input. In some embodiments, the input indicates a biological pathway that the molecule is to affect. For example, in some embodiments, the input specifies a target, such as a receptor or a protein binding site, to which the molecule is to bind as an agonist or antagonist, or a malfunctioning or mutated protein that the molecule is to replace. In other embodiments, the input does not indicate a biological pathway, but rather, merely indicates an intended effect of the molecule. For example, if the molecule is to cure a diseased cell, the input may include descriptions and, optionally, images of healthy and diseased cells, thereby indicating that the molecule is to cause the diseased cells to be more like the healthy cells.
[0055] Based on the input, the computer model initializes a list of one or more molecules. Subsequently, during each iteration, each of the molecules in the list is applied to a biological sample, such as an in vitro or in vivo biological sample comprising multiple biological cells, such that the molecule interacts with the organic compounds within the biological sample. One or more effects of the interaction are identified and passed to the computer model. Based on this input, the model revises the list, and the molecules in the revised list are applied to the biological sample. This process continues until the model converges onto a safe and effective molecule.
[0056] In general, at any stage in the iterative process, the list can include one or more molecules that are candidates for affecting the biological pathway (if the pathway was input by the user or was discovered by the model), and / or one or more molecules designed for probing the biological sample. In some embodiments, the model is configured to decide, automatically, on probing that is to be performed. In other embodiments, a user specifies any probing that is to be performed.
[0057] As an example of probing, the list can include an antibody or antigen for performing an immunoassay. As another example, the list can include one or more binder proteins designed to bind to specific target proteins so as to measure the expression levels and distributions within the cell of the target proteins. (For some embodiments, the binding site on each target protein is selected such that the binder protein does not interfere with the interaction of the target protein with other proteins in the proteome.) As a particular example of the above, the list can include a respective binder protein for each protein in the proteome.
[0058] For example, each binder protein may be tagged with two detection tags, one of which is hidden when the binder protein is bound to the target protein. Examples of suitable tags include FLAG-tags and / or HiBiT tags. Based on an identification of the tags in an image of the cell, the expression levels and distributions of the target proteins may be ascertained. Alternatively, the binder protein itself may be fluorescent when excited by an excitation light source or via the emission of another fluorescent probe as in the Forster resonance energy transfer (FRET) system.
[0059] Each molecule can be drawn from a predefined library or, alternatively, be de-novo. Each molecule can be specified directly or indirectly. For example, a protein can be specified directly as an amino acid sequence, or indirectly as a messenger ribonucleic acid sequence (mRNA) that is to be translated to the protein. Typically, the computer model is trained for a specific target organism, such as a human, and / or for a specific type of biological sample, such as genetic material.
[0060] During each iteration, various methods can be used to generate / manufacture each molecule in the list. For example, a protein can be generated using cell-free protein synthesis, using a yeast or bacteria host, via lipid-nanoparticle delivery of messenger ribonucleic acid to a target cell, or via transfection of a deoxyribonucleic acid plasmid.
[0061] Following the application of each molecule to the biological sample, the effects of the interaction are assessed, e.g., via an assay. Examples of techniques for assessing the effects include an enzyme linked immunosorbent assay, immunophenotyping (e.g., using flow cytometry), imaging (e.g., fluorescence or brightfield cell imaging), cellular ribonucleic or deoxyribonucleic acid sequencing, fluorescence in situ hybridization, toxicity testing, membranal functional testing, immunogenicity testing using immune-system cells (e.g. peripheral blood mononuclear cells), an assessment of the response of a humanized in-vivo host, and cell migration testing. Different assessments can be performed during different iterations. Based on the assessed effects, the computer model selects or designs each molecule for the next iteration.
[0062] Typically, one or more steps in the generation of the molecules, the application of the molecules to the biological sample, and / or the assessments are automated. For example, deoxyribonucleic acid, ribonucleic acid, or proteins may be synthesized automatically, e.g., using the SYNTAX system of DNA Script, Inc. (CA, USA) or the DNA & RNA Synthesizer of Kilobaser Corporation (Graz, Austria). Alternatively or additionally, liquids may be pipetted automatically by automated liquid handlers, and / or the sample may be assessed using an automatic microscope and / or flow cytometer.
[0063] In general, the effects of the interaction can relate to any properties of the molecule and / or of the biological sample. The computer model is configured to optimize these properties by training on existing data and / or learning from the input received during the iterations.
[0064] For example, in some embodiments, the computer model maximizes the effectiveness of the molecules. For example, the computer model may maximize the binding affinity of the molecule with the target site, e.g., by optimizing the electrostatic properties of the molecule.
[0065] Alternatively, or additionally, the computer model maximizes the specificity of the molecule, i.e., the computer model minimizes the number of other biological pathways that are affected by the molecule. For example, if the molecule is to bind to a particular binding site on a particular protein, the model may minimize the number of other binding sites on the protein to which the molecule binds. Alternatively or additionally, the computer model minimizes toxicity of the molecule. For cases in which the molecules are proteins, toxicity can be due to cross-protein interaction, protein breakdown, or protein overload. In some embodiments, the minimization of cross-protein interaction is based on data collected during the training of the model. In other words, during the training, some or all of the relevant proteome is provided as input, and the model learns to design proteins that have little affinity to these other proteins. Alternatively, or additionally, this minimization is based on testing performed during one or more of the iterations. For example, proteins can be tested using a protein chip assay, such as the HuProt™ Human Proteome Microarray, which allows for testing for affinity with thousands of proteome proteins in parallel. In some embodiments, to minimize toxicity due to protein breakdown or overload, the model learns to limit protein production rates, to increase protein stability, and / or to increase the specificity of peptides that result from the breakdown.
[0066] Alternatively, or additionally, the computer model minimizes (or, in some cases, maximizes) the immunogenicity of the molecule. In general, immune cells identify non-self proteins based on local amino-acid sequences and structure. Hence, in some embodiments, the computer model learns, during training, the sequences and structures that are less likely to provoke an immune response. For example, the computer model may include a language model trained on the human proteome and / or on sequences of deoxyribonucleic acid or messenger ribonucleic acid stored in a large genomic database representing a varied population. Alternatively, or additionally, the computer model is trained to prefer proteins having a three-dimensional structure and / or electrochemical properties similar to those of human proteins at a small scale, e.g., at a scale of around 50 angstroms.
[0067] Alternatively, or additionally, the computer model maximizes the solubility, motility, and / or penetrability (e.g., permeability) of the molecule. (The penetrability can be into the cell, into the nucleus of the cell, or outside of the cell, depending on the intended function of the molecule.) In some embodiments, solubility and / or motility are maximized by optimizing the size of the molecule (e.g., minimizing the radius of gyration).
[0068] Alternatively, or additionally, the computer model maximizes protein stability in one or more acidity levels, such as intracellular and / or extracellular acidity levels.
[0069] Alternatively, or additionally, the computer model optimizes the localization of the molecule, e.g., by designing proteins having greater localization within the cell nucleus.
[0070] Alternatively, or additionally, the computer model minimizes protein aggregation, by minimizing the self-binding strength of the proteins. Alternatively, or additionally, the computer model optimizes production attributes, such as by minimizing toxicity for the host cell system that produces the molecule, maximizing ribosomal production rates, maximizing ease of purification, and / or optimizing the expression rate of the molecule.
[0071] Alternatively, or additionally, the computer model optimizes properties of the biological sample such as cellular function (e.g., cell viability, proliferation, and / or protein expression), cellular morphology, and / or cellular appearance. For example, in some embodiments, based on images of diseased cells in the biological sample, the model repeatedly modifies the molecules so as to cause the cells, following their interaction with the molecules, to appear more like healthy cells. In some embodiments, comparisons between images of diseased cells and images of healthy cells are facilitated by image embeddings. Alternatively, or additionally, image embeddings are used to verify that healthy cells are not adversely affected by the molecules.
[0072] In some embodiments, to modify the stability and / or expression rate of a protein, the computer model modifies the non-coding regions of a ribonucleic or deoxyribonucleic acid string. Examples of such regions include the CAP, 3'UTR, 5'UTR, and ployA tail regions.
[0073] As described hereinabove, the list of molecules can include one or more molecules designed for probing components within the biological sample, to obtain information regarding the sample. Typically, to investigate biological cells, e.g., to detect cellular mechanisms of actions, it may be beneficial to probe many proteins of the proteome, or even the entire proteome of the cells of interest. A possible approach in which this could be done is by designing a set of binder proteins such that each binder protein is configured to selectively attach to a specific target protein of the cell proteome in a manner that does not interfere with the target protein’s interaction with other proteins in the cell proteome. The set of binder proteins can be created for the entire cell proteome or for a smaller set of proteins, e.g., to a set of proteins involved in a specific biological pathway in the cell or a set of pathways. Typically, each binder protein comprises one or more detection tags (e.g., fluorescent tags) allowing its spatial distribution to be imaged via microscopy. In such a manner, when using a set of binder proteins for targeting the entire cell proteome, it is possible to visualize the entire proteome of the cells of interest.
[0074] Alternatively, or additionally, another possible approach for probing the entire, or a subset, of a cell proteome comprises designing a set of binder proteins that interact with specific binding sites of particular target proteins in the proteome, such that the existing interactions of the target proteins with other proteins in the proteome are affected. This can be done for all binding sites of all of the proteins in the proteome or for a subset of the binding sites, or for all of the binding sites in a subset of the proteins of the proteome. In such a manner, cells of interest can be perturbed by one of these designed molecules (typically binder proteins, and / or DNA / RNA ). Cells that are affected or unaffected by the molecules can then be imaged e.g. by a high content imaging assay, such as Cell Painting or other imaging protocols, to assess the effect of this perturbation on the morphology of the cell. For some embodiments, each binder protein comprises one or more detection tags (e.g., fluorescent tags) allowing assessment of the effect of the perturbation on spatial distribution and cell morphology using microscopy.
[0075] There are roughly 30,000 proteins in the proteome, such that, a complete mapping of the spatial proteome and the effects of perturbing proteome pathways is possible using the abovedescribed techniques. In such a manner, datasets can be generated which enable characterizing of new cells. For example, a new diseased cell can be evaluated based on its similarity to cells or groups of cells in the perturbed proteome dataset thereby shedding light on a mechanism of action of a disease and identifying therapeutic targets. The above-described techniques may be practiced in combination with additional techniques such as high throughput testing of the perturbed proteome (e.g., RNA sequencing).
[0076] SYSTEM DESCRIPTION
[0077] Reference is now made to Fig. 1, which is a block diagram showing steps implemented by a system 10 for identifying a biological pathway-affecting molecule, in accordance with some embodiments of the present disclosure.
[0078] A computer model runs machine-learning algorithms (step 20) to generate a list of candidate molecules (step 22) for affecting the biological pathway. For some embodiments, the computer model, which typically includes a neural network, is provided with an initial input. In some embodiments, the input indicates a biological pathway that the molecule is to affect. For example, in some embodiments, the input specifies a target, such as a receptor or a protein binding site, to which the molecule is to bind as an agonist or antagonist, or a malfunctioning or mutated protein that the molecule is to replace. In other embodiments, the input does not indicate a biological pathway, but rather, merely indicates an intended effect of the molecule. Typically, in response to the input, the computer model generates the list of candidate molecules (step 22). The candidate molecules may include, for example, proteins (such as enzymes, antibodies, and / or protein complexes), nucleic acids (deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)), or other complex molecules.
[0079] Subsequently, to generating the list of one or more candidate molecules, each of the molecules in the list is produced / manufactured (step 24) and applied to a biological sample to interact with components, e.g., organic compounds such as proteins and / or nucleic acids, within the sample. This is typically accomplished by performing lab assays (step 26), through which the effects of the molecules on components of the biological sample can be determined. Examples of techniques for assessing the effects include an enzyme linked immunosorbent assay, immunophenotyping (e.g., using flow cytometry), imaging (e.g., fluorescence or brightfield cell imaging), cellular ribonucleic or deoxyribonucleic acid sequencing, fluorescence in situ hybridization, toxicity testing, membranal functional testing, immunogenicity testing using immune-system cells (e.g. peripheral blood mononuclear cells), an assessment of the response of a humanized in-vivo host, and cell migration testing. Different assessments can be performed during different iterations. For some applications, the assays include automated biochemical assays and automated handing of the samples, e.g., by using automated liquid handling of the sample.
[0080] One or more effects of the interaction are identified and passed to the computer model. Based on this input, the model revises the list by refinement and retaining of the machine-learning algorithms (step 28), and the molecules in the revised list again applied to the biological sample. This process continues until the model converges onto a safe and effective molecule.
[0081] Applications of the disclosure described herein can take the form of a computer program product accessible from a computer-usable or computer-readable medium (e.g., a non-transitory computer-readable medium) providing program code for use by or in connection with a computer or any instruction execution system. For the purposes of this description, a computer-usable or computer readable medium can be any apparatus that can comprise, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The medium can be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium. Typically, the computer-usable or computer readable medium is a non-transitory computer-usable or computer readable medium.
[0082] Examples of a computer-readable medium include a semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random-access memory (RAM), a read-only memory (ROM), a rigid magnetic disk and an optical disk. Current examples of optical disks include compact disk-read only memory (CD-ROM), compact disk-read / write (CD-R / W) and DVD.
[0083] A data processing system suitable for storing and / or executing program code will include at least one processor coupled directly or indirectly to memory elements through a system bus. The memory elements can include local memory employed during actual execution of the program code, bulk storage, and cache memories which provide temporary storage of at least some program code in order to reduce the number of times code must be retrieved from bulk storage during execution. The system can read the inventive instructions on the program storage devices and follow these instructions to execute the methodology of the embodiments of the disclosure.
[0084] Network adapters may be coupled to the processor to enable the processor to become coupled to other processors or remote printers or storage devices through intervening private or public networks. Modems, cable modem and Ethernet cards are just a few of the currently available types of network adapters.
[0085] Computer program code for carrying out operations of the present disclosure may be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the C programming language or similar programming languages.
[0086] It will be understood that algorithms described herein (and, in particular, the steps that are described as being performed by a computer model) can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general- purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the algorithms described in the present application. These computer program instructions may also be stored in a computer-readable medium (e.g., a non-transitory computer- readable medium) that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instruction means which implement the function / act specified in the flowchart blocks and algorithms. The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the algorithms described in the present application. The computer processor is typically a hardware device programmed with computer program instructions to produce a special purpose computer. For example, when programmed to perform the algorithms described herein, the computer processor typically acts as a special purpose molecule-discovery computer processor. Typically, the operations described herein that are performed by the computer processor transform the physical state of a memory, which is a real physical article, to have a different magnetic polarity, electrical charge, or the like depending on the technology of the memory that is used.
[0087] It will be appreciated by persons skilled in the art that the present disclosure is not limited to what has been particularly shown and described hereinabove. Rather, the scope of the present disclosure includes both combinations and subcombinations of the various features described hereinabove, as well as variations and modifications thereof that are not in the prior art, which would occur to persons skilled in the art upon reading the foregoing description.
Claims
CLAIMS1. A method for identifying a pathway-affecting molecule for affecting a biological pathway, the method comprising: receiving, from a computer model, a list including one or more molecules; causing the computer model to revise the list until the list includes the pathway-affecting molecule, by repeatedly: applying each of the molecules in the list to a biological sample, such that the molecule interacts with the biological sample, identifying one or more effects of the interaction, and inputting, to the computer model, the effects of the interaction, such that the computer model revises the list based on the effects; and based on the effects of the interaction of the pathway-affecting molecule with the biological sample, identifying the pathway-affecting molecule.
2. The method according to claim 1, further comprising, prior to receiving the list from the computer model, inputting the biological pathway to the computer model, such that the computer model initializes the list responsively to the biological pathway.
3. The method according to claim 1, further comprising, prior to receiving the list from the computer model, inputting an intended effect of the pathway-affecting molecule to the computer model, such that the computer model initializes the list responsively to the intended effect.
4. The method according to claim 3, wherein, prior to receiving the list from the computer model, the biological pathway is unknown.
5. The method according to any one of claims 1-3, wherein identifying the effects of the interaction comprises identifying the effects of the interaction by acquiring one or more images of the biological sample, and wherein inputting the effects of the interaction to the computer model comprises inputting the effects of the interaction to the computer model by inputting the images to the computer model.
6. The method according to any one of claims 1-3, wherein the effects of the interaction relate to a toxicity of the molecule with respect to the biological sample.
7. The method according to any one of claims 1-3, wherein the molecule includes a protein, wherein the biological sample includes other proteins, and wherein identifying the effects of the interaction comprises identifying the effects of the interaction by identifying cross-interactions of the protein with the other proteins.
8. The method according to claim 7, wherein identifying the cross-interactions comprises identifying the cross-interactions in a protein chip assay.
9. The method according to any one of claims 1-3, wherein the molecule includes a protein, and wherein identifying the effects of the interaction comprises identifying the effects of the interaction by identifying breakdown of the protein in the biological sample.
10. The method according to any one of claims 1-3, wherein the effects of the interaction relate to an immunogenicity of the molecule with respect to the biological sample.
11. The method according to any one of claims 1-3, wherein, at at least one point in time, the list includes one or more binder proteins designed to bind to respective target proteins in the biological sample.
12. The method according to claim 11, wherein each of the binder proteins is tagged with two fluorescent tags, one of which is hidden when the binder protein is bound to the target protein.
13. The method according to claim 11, wherein the target proteins include all proteins in a proteome corresponding to the biological sample, and wherein, at the at least one point in time, the list includes a respective binder protein for each of the target proteins.
14. The method according to anyone of claims 1-3, wherein the biological sample includes organic compounds and wherein identifying the one or more effects of the interaction comprises identifying one or more effects of the interaction of the one or more molecules with the organic compounds.
15. Apparatus for identifying a pathway-affecting molecule for affecting a biological pathway, the apparatus comprising: a computer processor configured to: generate, from a computer model, a list including one or more molecules; drive the computer model to revise the list until the list includes the pathway-affecting molecule, by repeatedly: driving the computer model to revise the list, by inputting, to the computer model, one or more effects of applying each of the molecules in the list to a biological sample, such that the molecule interacts with the biological sample, wherein the computer model is configured to revise the list based on the effects; and identify the pathway-affecting molecule based on the effects of the interaction of thepathway-affecting molecule with the biological sample.
16. The apparatus according to claim 15, wherein the computer processor is further configured to, prior to generating the list from the computer model, input the biological pathway to the computer model, such that the computer model initializes the list responsively to the biological pathway.
17. The apparatus according to claim 15, wherein the computer processor is further configured to, prior to generating the list from the computer model, input an intended effect of the pathwayaffecting molecule to the computer model, such that the computer model initializes the list responsively to the intended effect.
18. The apparatus according to claim 17, wherein, prior to generating the list from the computer model, the biological pathway is unknown.
19. The apparatus according to any one of claims 15-17, wherein the effects of the applying the one or more molecules relate to a toxicity of the one or more molecules with respect to the biological sample.
20. The apparatus according to any one of claims 15-17, wherein the molecule includes a protein, wherein the biological sample includes other proteins, and the computer processor is configured to input to the computer model the effect of applying the molecule to the biological sample by inputting cross-interactions of the protein with the other proteins.
21. The apparatus according to claim 20, wherein inputting the computer processor is configured to input the cross-interactions as determined in a protein chip assay.
22. The apparatus according to any one of claims 15-17, wherein the molecule includes a protein, and wherein the computer processor is configured to input to the computer model the effect of applying the molecule to the biological sample by inputting breakdown of the protein in the biological sample.
23. The apparatus according to any one of claims 15-17, wherein the effects of applying the molecule to the biological sample relate to an immunogenicity of the molecule with respect to the biological sample.
24. The apparatus according to any one of claims 15-17, wherein, at at least one point in time, the list includes one or more binder proteins designed to bind to respective target proteins in the biological sample.
25. The apparatus according to claim 24, wherein the target proteins include all proteins in aproteome corresponding to the biological sample, and wherein, at the at least one point in time, the list includes a respective binder protein for each of the target proteins.
26. A computer software product, for identifying a pathway-affecting molecule for affecting a biological pathway, the computer software product comprising a non-transitory computer-readable medium in which program instructions are stored, which instructions, when read by a computer cause the computer to perform the steps of: generating, from a computer model, a list including one or more molecules; driving the computer model to revise the list until the list includes the pathway-affecting molecule, by repeatedly: driving the computer model to revise the list, by inputting, to the computer model, one or more effects of applying each of the molecules in the list to a biological sample, such that the molecule interacts with the biological sample, wherein the computer model is configured to revise the list based on the effects; and identifying the pathway-affecting molecule based on the effects of the interaction of the pathway-affecting molecule with the biological sample.
27. A method for imaging a characteristic of cell morphology in a biological sample, the method comprising: receiving, from a computer model, a list including one or more binder proteins designed to bind to respective one or more target proteins in the biological sample; applying each of the binder proteins in the list to the biological sample, such that each binder protein interacts with the target protein biological sample, identifying one or more effects of each of the interactions, and inputting, to the computer model, the effects of each of the interactions; and based on the effects of the interactions between the binder proteins and the target protein, identifying the characteristic of cell morphology in the biological sample.
28. The method according to claim 27, wherein identifying the characteristic of cell morphology comprises identifying a spatial distribution of the one or more target proteins in a biological sample.
29. The method according to claim 27, wherein each of the binder proteins is tagged with a detection tag, and wherein identifying the characteristic of cell morphology in the biological sample comprises identifying the cell morphology through the detection tag.
30. The method according to claim 27, wherein each of the binder proteins is tagged with two fluorescent tags, one of which is hidden when the binder protein is bound to the target protein.
31. The method according to claim 27, wherein the binder protein is a fluorescent binder protein.
32. The method according to claim 27, wherein the target proteins include all proteins in a proteome corresponding to the biological sample, and wherein, at the at least one point in time, the list includes a respective binder protein for each of the target proteins.
33. The method according to claim 27, wherein the target proteins include a subset of proteins in a proteome corresponding to a biological pathway in the biological sample, and wherein, at the at least one point in time, the list includes a respective binder protein for each of the target proteins.
34. The method according to any one of claims 27-33, wherein the binder proteins are configured to bind to the target protein without affecting interactions of the target protein with other proteins in the cell proteome.
35. The method according to any one of claims 27-33, wherein the binder proteins are configured to bind to the target protein in a manner that affects interactions of the target protein with other proteins in the cell proteome.