Programming Cellular Functions Using Combinatorial Genetic Screens

JP2025509461A5Pending Publication Date: 2026-03-19THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
Filing Date
2023-03-16
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Current methods for cell therapy face challenges in systematically identifying combinations of genetic, epigenetic, and pharmacological interventions that confer polygenic therapeutic functions, due to the vast number of possible combinations and the complexity of human cell biology.

Method used

The method involves creating a library of cells that have undergone various perturbation combinations, analyzing these cells at a single-cell level to measure phenotypes and identify applied perturbations, and calculating scores for both observed and theoretical perturbation combinations to determine their likelihood of generating specific phenotypes.

Benefits of technology

This approach allows for the efficient identification of combinations of gene interventions that confer sustained therapeutic capabilities, overcoming the challenge of combinatorial scaling and revealing complex phenotypes that cannot be observed through single genetic screening techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Described herein is a method for identifying combinations of perturbations that result in a cell phenotype. In some embodiments, the method includes generating a library of cells that have been subjected to combinations of perturbations, analyzing a subset of cells at a single-cell level by measuring the phenotype of the cells and identifying which combinations of perturbations were applied to the cells, and calculating a score for the identified combinations of perturbations and a score for a theoretical combination of perturbations based on the results obtained from the analysis, where each score indicates the likelihood that the combination of perturbations will produce a phenotype.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 63 / 321,582, filed March 18, 2022, which application is incorporated herein by reference.

[0002] government support This invention was made with Government support under Contract Nos. HG007735 and HG009436 awarded by the National Institutes of Health. The Government has certain rights in this invention. [Background technology]

[0003] Modern cell therapies often use engineered (e.g., genetically modified) cells to perform unique tasks for patients. Clinical application of these biologics requires complex phenotypes that often cannot be programmed into cells by modulating a single gene pathway. Many biological processes in human cells are robust to individual gene perturbations due to ubiquitous redundancy, and complex phenotypes often require synergistic activation of multiple genes. This inherent complexity of human cell biology poses significant challenges to traditional one-gene functional genomics, which relies on single-gene perturbations. As a result, there is a critical need to systematically identify combinations of genetic, epigenetic, and pharmacological interventions that confer polygenic (involving multiple gene products) therapeutic functions. Summary of the Invention [Problem to be solved by the invention]

[0004] A key challenge in finding combinatorial solutions to cell engineering problems is that the scale of measurements required to make meaningful inferences is intractable. For example, naively phenotyping all combinations of regulators in a population of only 50 potential regulators would require more than 1000 trillion independent measurements. Phenotyping even low-complexity combinations, e.g., combinations of less than five components, would still require millions of experiments. Thus, there is a critical need for scalable approaches to combinatorial cell engineering. [Means for solving the problem]

[0005] A method for identifying the combination of perturbations that leads to cell phenotype is described herein.In some embodiments, the method may include the steps of: making a library of cells that have been subjected to combination of perturbations; analyzing a subset of cells at single-cell level by measuring the phenotype in cells and identifying which combination of perturbations has been applied to cells; and calculating the scores of identified combination of perturbations (i.e., the combination of perturbations that has been identified in cells) and theoretical combination of perturbations (i.e., the combination of perturbations that has not been identified in cells) based on the results obtained from the analysis, each score indicating the likelihood that the combination of perturbations will generate a phenotype.Figure 1 illustrates part of the principle of this method.

[0006] In some embodiments, the method may be iterative in the sense that it may be performed and then repeated one or more times, with each iteration modifying the library of cells according to the calculated score. For example, the iterations may place more emphasis on perturbation combinations that are more likely to generate a phenotype.

[0007] As illustrated in Fig. 1 and described in more detail below, only a limited number of combinations of perturbations are expressed in the cells analyzed.However, the scores of theoretical perturbation combinations (i.e., combinations of perturbations that have not been identified in cells) can be calculated by learning algorithm based on the data obtained from those cells.For example, in some embodiments, all possible combinations of up to n terms (where n is 5, 6, 7, 8, 9 or 10, up to the total number of perturbations), such as binomial, trinomial, tetranomial, etc., can be scored for likelihood of causing phenotype, where such combinations include theoretical perturbation combinations, i.e., combinations that have not been identified in the cells analyzed. The perturbation combinations scored in the latter steps of the method can include i. "observed" perturbation combinations (i.e., combinations of perturbations identified in cells), and ii. "theoretical" perturbation combinations (i.e., combinations of perturbations not identified in cells), where the theoretical combinations can be i. new perturbation combinations not present in any of the analyzed cells, or ii. subcombinations of the perturbation combinations identified in the analyzed cells. The scores of these perturbation combinations (including theoretical combinations) can be generated by statistical analysis of the aggregated data obtained from the cells, in particular by methods using learning algorithms. In any embodiment, the likelihood scores of all potential combinations of perturbations can be calculated.

[0008] This method, which may be referred to below as "Combinatorial Cellular Programming" (CCP), allows one to systematically program phenotypes that require the processing of multiple genetic components into biological cells. Certain principles of the method can be illustrated using the following hypothetical example. It is now known that simultaneous exogenous expression of four transcription factors, termed the Yamanaka factors (Oct3 / 4, Sox2, Klf4, c-Myc), is required to reprogram somatic cells to a pluripotent state. Without prior knowledge, a vast number of different combinations would need to be tested to link these four transcription factors to a reprogrammed phenotype. In this example, even knowing that four transcription factors are required, if one wanted to test all four-term combinations of the 1500 transcription factors encoded by the human genome, it would take approximately 2 × 10 11 One would have to collect data from every possible combination, which is physically impossible. This phenotype sparsity problem, where the correct answer constitutes a small fraction of all possible combinations, can be overcome by increasing the number of perturbation combinations tested in each individual experiment. This solves the sparsity problem by increasing the frequency of "correct" answers, but it makes it more difficult to infer which of the perturbation sets in each observed correct answer are causal, since there will be many perturbations that do not affect the phenotype.

[0009] The premise of the present approach is that many important clinical phenotypes are combinatorially controlled and cannot be reliably accessed using single gene perturbations. However, combinatorial screening introduces a seemingly intractable scaling problem, namely, it is impossible to select the correct combination of genes to be treated, given the large number of possible combinations. Depending on how the method is implemented, the present method leverages recent advances in machine learning, modern genome editing, and high-throughput single-cell phenotyping to solve this combinatorial scaling problem and efficiently identify combinations of genetic interventions that confer sustained therapeutic function. One of the technical insights of the present approach is that the experimentally intractable problem of combinatorial cell engineering can be transformed into a scalable computational problem. This is achieved by constructing a gene perturbation library from which a combinatorial number of phenotypes can be extracted from each single cell.

[0010] Potential regulators are up- or down-regulated at a multiplicity of perturbation (MoP) of more than 1. This facilitates the analysis of many combinations of perturbations in each cell (experimental compaction). The phenotype of each individual cell is then analyzed to provide perturbation-phenotype paired data to an inference (unfolding) engine that identifies causal regulators. The platform of the present invention should identify new classes of polygenic cellular therapeutics not by sequentially regulating individual genes, but by efficient data-driven exploration of high-dimensional combinatorial perturbations. This approach allows for phenotypic screening of trillions of combinatorial perturbations, revealing complex phenotypes that are not observable by any single-gene screening approach. Moreover, these innovations provide significant improvements in the art.

[0011] Also provided are split-pool methods for exposing cells to perturbations. In these embodiments, the steps include partitioning the cells into a plurality of partitions, selecting a subset of perturbations, applying a partial combination of the subset of perturbations to the partitions, optionally applying all perturbations in the subset to at least one partition, and optionally not applying any perturbation in the subset to at least one partition, pooling the cells, and repeating the method one or more times, each iteration using a different subset of perturbations. Details of the method are described in more detail below.

[0012] These and other advantages will become apparent from a consideration of the detailed description that follows. [Brief description of the drawings]

[0013] [Figure 1] 1 is a flow chart illustrating some principles of the methods of the invention. In this example, the likelihood scores calculated are for all possible combinations of perturbations, including theoretical perturbation combinations that were not identified in the analyzed cells. [Diagram 2] 1 is a flow chart illustrating an implementation of the method of the invention, in which the likelihood scores calculated are for all possible combinations of perturbations, including theoretical perturbation combinations that were not identified in the analyzed cells. [Diagram 3] Figure 1 illustrates the workflow of combinatorial genetic screening, showing how causative factors of a phenotype are identified from a population of potential regulators. Cells are subjected to combinatorial perturbations (Domains 1+2) to enrich for a particular phenotype (Domain 3). Positively selected cells are then genotyped to identify the perturbation responsible for the phenotype (Domain 4). Finally, a new structured perturbation library is constructed based on the information acquired during causal inference (Domain 5). [Figure 4]Figure 1 illustrates how introducing high perturbation multiplicity (MOP) per cell makes rare combinatorial solutions of complex phenotypes more frequently observable. Green represents cells exposed to a critical set of perturbations (with concentration n) required to generate a particular phenotype. These cells may also be exposed to other perturbations that may adversely affect the phenotype of interest. However, during many applications of this technique, we are looking for a small number of causal perturbations among a large population of unrelated perturbations. The number of observations per cell (n) that match the required phenotypic complexity is shown for various MOP levels (dark blue). The observation frequency of a given phenotype is also reported for each MOP design (dark green). [Diagram 5] FIG. 1 illustrates an approach for constructing a combinatorial perturbation library using the split-and-pool method. In this method, each perturbation in the perturbation population U is assigned to at least one of Q groups {P1, P2, …, PQ}. These assignments can be random, guided by conventional biological knowledge (e.g., known synergistic or redundant relationships between epigenetic or genetic factors being perturbed), or designed using active learning techniques in the disclosed method. Next, the progenitor cells of the library are split into K wells, (1) the first well receives no perturbation, (2) the second well receives all perturbations in P1, and (3) the remaining K-2 wells receive the perturbation combinations {S1, S2, …, SK-2} extracted from P1 (for each k, SK is a subset of P1). Finally, all cells are pooled into a single well and then split again into K wells, and the same procedure is repeated for groups P2 through PQ. The distribution of perturbations in the set of hypotheses P2 in which the protein represents a perturbation is shown. This procedure generates a perturbation library with complexity (number of different perturbation combinations) KQ. [Figure 6] FIG. 1 illustrates how T cell receptor (TCR) complexes can be displayed on the surface of non-immune cells. [Figure 7]Figure 1 illustrates proof-of-principle probabilistic inference of causal components required for TCR presentation. The posterior probability of TCR presentation was reported for all models with a complexity of 12 or less (194,129,627 models are shown). Models composed of actual TCR components are indicated by black squares. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0014] definition As used herein, the term "perturbation" refers to any type of cellular treatment, including, but not limited to, the introduction of constructs to express or suppress synthetic or endogenous gene products; or exogenous exposure of cells to drugs, antibodies, small molecules, or proteins; or stimulation by physical forces, including electromagnetic, temperature, pH, salt, or other non-molecular insults.

[0015] As used herein, the phrase "combinatorial perturbation" refers to a collection of perturbations that are applied to a cell.

[0016] As used herein, the phrase "perturbation library" refers to a collection of cells that have been exposed to a collection of perturbations, or equivalently, combinatorial perturbations.

[0017] As used herein, the phrase "combinatorial perturbation library" refers to a "perturbation library" in which a subset of constituent cells have two or more perturbations applied to the cells.

[0018] As used herein, the term "perturbation population" refers to the entire set of perturbations that are capable of or associated with a particular cellular phenotype of interest.

[0019] As used herein, the phrase "perturbation multiplicity" (further herein "MOP") refers to the number of perturbations that are applied to a cell.

[0020] As used herein, the phrase "phenotypic complexity" or "phenotypic complexity" refers to the minimum number of perturbations required to generate a given cellular phenotype.

[0021] As used herein, the phrase "causal perturbations" refers to the set of perturbations that are causally related to the generation of a given cellular phenotype.

[0022] As used herein, the phrase "high MOP" refers to MOP greater than phenotypic complexity.

[0023] As used herein, the phrase "low MOP" refers to MOP that is less phenotypic complex.

[0024] As used herein, the phrase "unstructured perturbation library" refers to a perturbation library in which each cell is randomly assigned a set of perturbations.

[0025] As used herein, the phrase "structured perturbation library" refers to a perturbation library in which each cell is assigned a set of non-random perturbations.

[0026] As used herein, the phrase "combinatorial assignment" refers to an assignment of perturbations that are applied to cells in a perturbation library according to the scheme outlined below.

[0027] As used herein, the phrase "active learning" refers to the process of using previously collected data to identify unobserved perturbation combinations that are most informative for analyzing a phenotype.

[0028] As used herein, the phrase "CRISPR machinery" refers to a collection of techniques that utilize the CRISPR nucleoprotein complex to control endogenous gene expression levels in cells, including, but not limited to, CRISPR-Cas9 editing, CRISPR interference, CRISPR activation, CRISPR nucleoprotein direct delivery, and CRISPR-Cas13 editing.

[0029] As used herein, the phrase "single-cell assays" refers to a collection of techniques that allow for ensemble measurements of molecules in individual cells or cellular compartments, including but not limited to single-cell RNA-seq, single-cell ATAC-seq, single-cell CITE-seq, spatial transcriptomics, and spatial metabolomics.

[0030] Before the present invention is further described, it is to be understood that this invention is not limited to the described embodiments, which may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing certain embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.

[0031] When a range of values ​​is provided, it is understood that each value between the upper and lower limit of that range, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these narrower ranges may independently be included in the narrower ranges and are also encompassed within the invention, subject to any specifically excluded limits in the stated range. When the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this invention belongs.Although any method and material similar or equivalent to those described herein can also be used in the practice or testing of this invention, preferred methods and materials are described herein.All publications mentioned herein are incorporated by reference to disclose and describe the method and / or material cited with the publication.

[0033] It should be noted that, as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "a circuit" includes a plurality of such circuits, and a reference to "the nucleic acid" includes reference to one or more nucleic acids and equivalents thereof known to those of skill in the art, and so forth. It should be further noted that the claims may be drafted to exclude any optional element. Thus, this statement is intended to serve as a predicate basis for using exclusive language such as "solely," "only," and the like in conjunction with the recitation of claim elements or the use of "negative" limitations.

[0034] Certain features of the invention that are described in the context of separate embodiments for clarity of description may also be provided in combination in a single embodiment. Conversely, various features of the invention that are described in the context of a single embodiment for brevity of description may also be provided separately or in any suitable subcombination. All combinations of the embodiments of the invention are expressly embraced by the invention and are disclosed herein as if each and every combination were individually and expressly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also expressly embraced by the invention and are disclosed herein as if each and every such subcombination were individually and expressly disclosed herein.

[0035] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publications by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed.

[0036] Detailed Description The present disclosure provides, inter alia, a method for identifying combinations of perturbations that result in a cellular phenotype. Certain principles of the method are illustrated in FIG.

[0037] With reference to FIG. 1, in some embodiments, a method may include creating a library of cells that have been subjected to a combination of perturbations. This library may be referred to herein as a "perturbation library." The total number of perturbations that a cell has been subjected to may range, for example, from 10 to 5,000 or from 20 to 1,000. The average number of perturbations that a cell has been subjected to may range from at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, at least 50 or at least 100, for example, from 5 to 10,000, 5 to 1000 or 5 to 500.

[0038] In the next step of the method illustrated in Figure 1, a subset of cells of the library is analyzed cell by cell. As shown, the cells are analyzed by (i) measuring the phenotype at the single cell level, and (ii) identifying at the single cell level which combination of perturbations was applied to the cell. These same cells are analyzed in this step, which means that the applied perturbations and phenotypic measurements are determined for a single cell. As shown, the cells analyzed have only received a limited number of possible combinations of perturbations (i.e., a relatively small subset of the "universe" of possible perturbations).

[0039] Phenotypes can be measured using any suitable single-cell analysis method, for example, by analyzing DNA, RNA, protein, and / or epigenetic modifications on a single-cell basis. In these embodiments, the term "measured" is intended to mean a quantitative or qualitative assessment. In some embodiments, phenotypes can be measured by performing single-cell "omics" assays. Such assays include RNA-seq (i.e., scRNA-seq), ATAC-seq (assay for transposase-accessible chromatin using sequencing, or csATAC-seq), CITE-seq (cellular indexing of transcriptomes and epitopes by sequencing), scG&T-seq (single-cell genome and transcriptome sequencing), scMT-seq (single-cell methylome and transcriptome sequencing), scM&T-seq (single-cell methylome and transcriptome sequencing), scTrio-seq (single-cell triple-omics sequencing), scCOOL-seq (single-cell chromatin whole-omics landscape sequencing), and DOGMA-seq (for reviews, see Islam et al. (Genome Research. 2011 21:1160-7), Valihrach et al. (International Journal of Molecular Sciences 2018 19:807), Zhu et al. (Nat Methods.2020 17:11-14), and Bode et al. (Front Immunol.2021 12:702636), etc. In some embodiments, the method may be performed by detecting and / or measuring specific markers of the phenotype (e.g., expression of cell surface markers, etc.) by FACS. In some cases, spatial assays may also be used. As will be apparent, the method may involve quantifying how similar a cell is to a cell having a desired phenotype. In some embodiments, the phenotype of the cells may be measured while enriching (e.g., by FACS).Identifying the particular set of perturbations present in the phenotyped cells may involve direct single-cell measurements of the genetic material (e.g., plasmid DNA or mRNA) that mediates the perturbations. Alternatively, this information may be obtained by single-cell sequencing of independent barcodes that code for the particular set of perturbations in the cells. These barcodes may be delivered transiently or persistently by any convenient method.

[0040] In the present methods, the desired phenotype may be characterized to some extent by prior studies. For example, prior studies may have established that cells with a particular phenotype may have a defined gene expression pattern. In these embodiments, "measuring the phenotype" may be relatively simple in some cases and may involve identifying or quantifying the expression of one or more markers of the phenotype. In other cases, "measuring the phenotype" may be more complex and may involve collecting a large number of measurements on a cell (e.g., by determining the transcriptome via RNA-seq) and then ascertaining how similar the measurements (as a whole) are to the same type of measurements from cells with that phenotype. As an example, if the goal is to identify perturbations that convert stem cells into liver cells, RNA-seq may be used to ascertain whether liver cell markers are expressed in the cells and / or how similar the transcriptome is to the transcriptome of liver cells. Methods for cross-comparing single-cell omics data are known and can be readily applied herein if desired (see, e.g., Alam et al., Nat Genet 2021 53:1275), Adbadaal et al. (Genome Biology 2019, Vol. 20:194), Zhao et al. (Proc Natl Acad Sci USA 2021 118:e2100293118), and Li et al. (Front Immunol. 2021 Feb. 24;12:625881), etc.).

[0041] It should be noted that in some embodiments, phenotypes may be measured and perturbations may be identified in the same assay. For example, when RNA-seq (or multi-omics methods including RNA-seq) are used, the same data may be processed to identify perturbations in the cells and measure phenotypes. In these embodiments, perturbations in the cells may be directly related to the data obtained from the cells, which may make the statistical analysis step of the method more accurate. Thus, in some embodiments, phenotypes may be measured and the combination of perturbations applied to the cells in the same analysis is determined.

[0042] As shown in FIG. 1, the next step of the method may involve calculating scores for combinations of theoretical perturbations that were not identified in the analyzed cells, with each score indicating the likelihood that the combination of perturbations will produce a phenotype. For clarity of explanation, a "theoretical" combination of perturbations contains perturbations that are not found together in the same cell (i.e., perturbations that are found only in separate cells) and perturbations that are only present with other perturbations (i.e., as "subcombinations" of the identified perturbations in the cell). As an example, a combination (A,B) is considered to be a theoretical combination of perturbations if A and B are always present in separate cells. Similarly, a combination (A,B) is considered to be a theoretical combination of perturbations if A and B are only found in a cell with another perturbation, e.g., C. These calculations are based on the results obtained from the analysis step and are therefore based on both (i) the measurement of the phenotype in each of the cells and (ii) the perturbations to which those cells were exposed. In any embodiment, as illustrated in Figure 1, this step may be performed by calculating a score indicating the likelihood that a phenotype will be generated by each possible combination of perturbations, i.e., a "universe" of possible combinations of perturbations. The "universe" of possible perturbations includes theoretical combinations of perturbations (i.e., combinations that have not been found in cells). Algorithms for performing these calculations are described below.

[0043] As illustrated by the hypothetical example shown on the left side of FIG. 1, which is used only to illustrate the principle of the method, if there are five perturbations: A, B, C, D and E, and only four perturbation combinations: (A,B,C), (C,D,E), (D,E,A), and (A,C) are identified in a cell, up to 120 possible perturbation combinations can be scored, such as (A,B), (A,C), (A,D), (A,E), (A,B,C), (A,B,D), (A,B,E), (A,B,C,D), (A,B,C,E), (A,B,C,D,E), (B,C), (B,D), etc. As illustrated, the scored combinations contain many theoretical combinations. In reality, the number of theoretical perturbations to be scored may be even larger, and the combinations of perturbations in those cells will be more complex and numerous.

[0044] In some embodiments, the number of perturbation combinations scored in the method can be at least 1M, at least 10M, at least 100M, at least 1B, at least 10B, at least 100B, or at least 1T, depending on the total number of perturbations analyzed at the beginning of the method. The difference between the number of perturbation combinations identified in a cell and the number of perturbation combinations scored can be large. For example, the number of combinations scored in this step can be at least 10 times, at least 100 times, at least 1000 times, at least 10,000 times, at least 100,000 times, or at least 1M times the number of perturbation combinations identified in a cell. In some embodiments, at least all of the combinations up to n terms (where n is, for example, up to 7, 8, 9, 10, or 20), such as 2-term, 3-term, 4-term, 5-term, 6-term combinations, are scored. In some embodiments, the score is calculated using the results obtained from cells that are positive for the phenotype and the results from cells that are negative for the phenotype. Details of the scoring algorithm can be found below. As will be clear, the scores of combinations of perturbations found in cells can be calculated simultaneously. Thus, in any embodiment, the method can involve calculating the scores of combinations of perturbations identified in cells and calculating the scores of theoretical combinations of perturbations not identified in cells, i.e., the "theoretical" combinations described above. In some embodiments, all possible combinations are scored, including combinations found in cells and theoretical combinations not found in cells.

[0045] In these embodiments, the term "score" is intended to refer to a number, letters, words (e.g., "high", "medium", or "low"), or a descriptor (e.g., "+++" or "++") that can indicate the strength of evidence that each potential combination of perturbations causes a phenotype. A value can contain one component (e.g., a single number) or two or more components, depending on how the value is analyzed. In some embodiments, a score can be expressed as or based on a likelihood, probability, or some other number that can be calculated using an algorithm.

[0046] The subset of cells analyzed may include one or more populations of enriched cells. In some embodiments, two separate populations of enriched cells are analyzed: phenotype-positive cells and phenotype-negative cells, and enrichment is performed by any convenient method, such as cell sorting (FACS), enrichment on a support (e.g., bead enrichment), or cell selection assay. For example, in some embodiments, cells may be enriched from the library by expression of one or more cell surface markers associated with a phenotype, and the phenotype may be measured in those cells. In other embodiments, the subset of cells analyzed may include cells randomly sampled from the library. Regardless of the subset of cells that is generated, this step of the method generally involves analyzing at least 1,000, at least 10,000, at least 100,000, at least 1M, or at least 10M cells.

[0047] As shown in FIG. 1, some embodiments may optionally include repeating the method one or more times (e.g., 2 or more, 5 or more, or 10 or more), in which in each iteration, the set of perturbations applied to the cell at the beginning of the method is changed according to the scores calculated in the previous trial. For example, at least one perturbation may be completely eliminated from the next round because it has a low likelihood of causing a phenotype and / or some combinations of perturbations may be preferred. In a pre-determined combination, if a doublet or triplet of perturbations is calculated to have a relatively high likelihood of causing a phenotype (compared to other combinations), it may be in two or more separate sets of perturbations, or in a set with fewer additional perturbations or only those sets. As will be clear, this step may require ranking the scores and / or applying a threshold to the scores to select the "best" combination.

[0048] As shown in FIG. 1, the method results in the identification of the minimum number of combinations of perturbations that can produce a phenotype.

[0049] Perturbation libraries can be created in a variety of different ways. For example, perturbation libraries can contain random combinations of perturbations. In these embodiments, cells can be exposed to perturbations all at once, for example, to be exposed to a random combination of perturbations. In another embodiment, perturbation libraries can be created by dividing cells into multiple partitions (e.g., at least 4, at least 8, at least 16, or at least 20 partitions), introducing various subsets of perturbations into the partitions all at once, and then pooling the cells. In this embodiment, cells in each partition are exposed to a random combination of perturbations that are added to the partition, and then pooled.

[0050] In other embodiments, the perturbation library may contain predetermined (i.e., non-random) combinations of perturbations. These embodiments may be implemented, for example, using the "split-and-pool" method illustrated in Figure 5. This implementation of the method has multiple advantages, as the combinations of perturbations applied to cells can be designed to maximize the efficiency of the discovery process. For example, if a certain combination of perturbations is calculated to be more likely to cause a phenotype, the library can be designed such that that particular combination is present (potentially along with other perturbations) in more cells.

[0051] Figure 2 is a flow chart illustrating an implementation of the method in which the perturbation library contains a combination of pre-determined perturbations (referred to as the "set of perturbations" in this figure). As shown in this figure, the set of perturbations can be designed before being introduced into the cells. The remainder of the method is similar to that described above, except that the set of perturbations applied to the cells is modified by the calculated score. This step can be implemented using a split-and-pool based method described below.

[0052] Similar to the method shown in Figure 1, the total number of perturbations in the ensemble of Figure 2 can range, for example, from 10 to 5,000 or from 20 to 1,000. The average number of perturbations in the ensemble can be at least 5, at least 10, at least 50 or at least 100, for example, from 5 to 10,000, 5 to 1000 or 5 to 500.

[0053] In any embodiment, a collection of perturbations can be applied to cells using a split-and-pool approach. Split-and-pool based methods are generally used for combinatorial chemistry and indexing samples rather than for introducing combinations of perturbations into cells (see, e.g., Kuchina et al. (Science 2021 371:eaba5257), O'Huallachain et al. (Commun. Biol. 2020 3:279), Cao et al. (Science 2017 357:661-667) and Rosenberg (Science 2018 360:176-182)). These publications are incorporated by reference for their description of how split-and-pool methods can be implemented.

[0054] The split-and-pool principle of the present invention is illustrated in Figure 5. The method may include partitioning cells into a plurality of partitions (e.g., at least 4, at least 8, at least 16, at least 24 partitions, at least 48 partitions, or at least 96 partitions), selecting a subset of perturbations, which may contain one or more, two or more, three or more, four or more, or five or more perturbations, applying a partial combination of the subset of perturbations to the partitions, pooling the cells, and then repeating the same steps one or more times (e.g., at least 2, at least 4, at least 10, or at least 20 times), each iteration being performed using a different subset of perturbations. In some embodiments, the subset of perturbations selected in the first round overlaps with at least one of the subsets of perturbations selected in the iterations. In some embodiments, the subset of perturbations selected in the first round may not overlap with any of the subsets of perturbations selected in the iterations. In many embodiments, in each round, up to half, up to 75%, or up to 90% of the partitions are subjected to a subset of perturbations.

[0055] The sets of perturbations used in the methods may overlap in the sense that in any single experiment, one or more perturbations in one set may also be present in another set. The sets may in some cases have the following characteristics: i. at least some of the sets of perturbations include multiple perturbations, ii. at least some of the perturbations are present in more than one set, iii. at least one of the sets contains some, but not all, of the perturbations in another set, and iv. as a whole, the sets do not contain all potential combinations of perturbations.

[0056] As illustrated by the example, if there are 26 perturbations (A-Z), then a first subset of perturbations may contain perturbations A, B, C, and D, and the subset of perturbations applied to the partitions may include (A,B), (B,C), (C), (A,C,D), and optionally (A,B,C,D). In the iterations, i. if the subsets overlap (D overlaps), then the subset of perturbations may include perturbations D, E, F, and G, and the subset of perturbations applied to the partitions may include (D,E), (F), (E,F,G), and optionally (D,E,F,G), or if the subsets do not overlap, then the subset of perturbations may include perturbations E, F, G, and H, and the subset of perturbations applied to the partitions may include (E,F), (G), (E,G,H), and optionally (E,F,G,H).

[0057] In these embodiments, each cell population in the partition may have at least 100, at least 500, at least 1,000, at least 5,000, or at least 10,000 members, and the total number of cells in the pool may be greater than 1M, for example at least 10M.

[0058] As will be apparent, in the method, the perturbations are not chemically related to each other and there are no chemical addition or deprotection steps.

[0059] In some embodiments, the perturbations are nucleic acid constructs, each construct encoding a perturbation. For example, the constructs may encode a protein, an RNA, or any combination thereof. In some embodiments, the perturbations are applied to the cells by introducing nucleic acid constructs into the cells, where each construct encodes a perturbation, and multiple constructs are introduced into the cells randomly or in a predetermined manner. In these embodiments, the constructs may encode proteins (e.g., signaling proteins, transcription factors, enzymes, or protein fragments, etc.), RNAs (e.g., guide RNAs, siRNAs, aptamers, ribozymes, etc.), or any combination thereof (e.g., guide RNAs and RNA-guided proteins such as RNA-guided endonucleases, etc.), where the term "guide RNA" is intended to refer to an RNA that forms a complex with an RNA-guided protein (e.g., proteins such as AGO2, Cas9, Can13, Cas7-11, Cascade, Cpf1, Cas12, including their variants and fusion proteins with additional enzymatic activity) and guides the complexed protein to a specific site or sequence of a nucleic acid (typically a sequence of the nuclear genome). For example, in some embodiments, the nucleic acid construct may encode an open reading frame library ("ORF library") in which the open reading frames may encode entire proteins, protein fragments, variants of wild-type proteins, or proteins from another species, etc. A typical library may contain, for example, 10-5,000 or 20-1,000 constructs.

[0060] In embodiments using guide RNA, the perturbation may result in genetic changes, such as gene knockout. In other embodiments, the RNA-guided protein may be fused with a methylase or demethylase. In these embodiments, the perturbation may result in changes in methylation patterns. In other embodiments, the perturbation may be protein or RNA expression. The construct may be introduced into the cell by any convenient method, such as lipid nanoparticles, viral transduction, transfection or electroporation. The perturbation may introduce a permanent change into the cell via genomic integration (e.g., viral or transposon-based (e.g., PiggyBac) delivery), or a transient effect (e.g., plasmid, dsDNA or RNA electroporation).

[0061] In other embodiments, the perturbation may be a non-nucleic acid molecule, e.g., a drug, an antibody, a small molecule, or a protein, or may be a stimulus, e.g., a physical stimulus, including electromagnetic, temperature, pH, salt, or other non-molecular damage. In these embodiments, the compartmentalized cells may be barcoded in step (b), and the barcode indicates which perturbation was applied to the cell. This barcode may be present, for example, in a construct that is added to the cell simultaneously with the perturbation. In these embodiments, the construct may be non-functional in the sense that it does not actually encode the perturbation. However, it identifies the perturbation that was added simultaneously. Thus, as the cell accumulates perturbations, it should accumulate barcodes that encode those perturbations.

[0062] Any implementation of the method may use a combination of nucleic acid-based and non-nucleic acid-based perturbations.

[0063] Further details can be found in the description below.

[0064] Phenotype The phenotype measured can be a molecular phenotype (e.g., the level or location of a cell surface protein, a nuclear-localized protein (e.g., a transcription factor), or a cytoplasmic protein (e.g., a cytokine)) or a functional phenotype. In the latter case, phagocytosis, tissue- or signal-specific cellular localization can be measured. In some implementations, a library of perturbed cells can be introduced into an organism (e.g., mouse, monkey, or human) and then extracted for molecular, functional, and / or localization phenotyping. In these embodiments, samples from the organism can also be tested. In some embodiments, phenotyping can be performed hierarchically. For example, a sub-library of cells can be selected by high-throughput molecular phenotyping (e.g., surface protein expression) and then used as an input library for lower throughput molecular (e.g., whole-transcriptome single-cell RNA-seq) or functional (in vitro or in vivo) phenotyping.

[0065] Probabilistic Modeling The posterior probability ("P(c)") that a given perturbation combination (c) confers a phenotype of interest can be estimated by a number of statistical methods. In some embodiments, the posterior probability that a given perturbation combination confers a phenotype of interest can be estimated by training and applying an ensemble decision tree statistical model. In other embodiments, the posterior probability that a given perturbation combination confers a phenotype of interest can be estimated by training and applying a random forest statistical model. In yet other embodiments, the posterior probability that a given perturbation combination confers a phenotype of interest can be estimated by training and applying a neural network model. The posterior probability that the set of perturbations includes a full complement of causal control factors can be calculated directly from P(c). See Appendix A, Section 0.1.

[0066] Active Learning The active learning applied in the present method involves the use of information metrics to identify maximally informative perturbation combinations for analyzing phenotypes. For an example of active learning used by the disclosed method, see Appendix A, Section 0.2.

[0067] Combinatorial perturbation library construction In embodiments involving structured perturbation libraries, one approach to creating highly diverse, high-MOP libraries, hereafter referred to as "combinatorial perturbation assignment", "combinatorial assignment", "combinatorial library construction", or "split-and-pool library construction", is illustrated in Figure 5. Briefly, each perturbation in a perturbation population U is divided into Q groups {P1, P2, ..., P Q}. These assignments can be random, guided by conventional biological knowledge (e.g., known synergistic or redundant relationships between epigenetic or genetic factors being perturbed), or designed using active learning as outlined in the disclosed methods. The library progenitors are then split into K wells, (1) the first well receives no perturbation, (2) the second well receives all perturbations in P1, and (3) the remaining K-2 wells receive perturbation combinations {S1, S2, ..., S K-2}(For each k, S k Finally, all the cells are pooled into a single well and then split again into K wells, and groups P2 to P Q Repeat the same procedure until

[0068] This procedure is Q Generate a perturbation library with complexity (number of unique perturbation combinations) equal to

[0069] For some embodiments of this procedure, the number of wells into which cells are split during each round (K in this example) varies depending on the composition of the perturbation group.

[0070] For some embodiments of this procedure, the unperturbed and / or all perturbed wells may be eliminated.

[0071] example: For K=2, in each round of the split-and-pool procedure, all cells are either exposed to all perturbations in the relevant population or to no perturbations.

[0072] Further examples are given for each perturbation subset S k But group P i This is achieved when all elements of K except for the kth element of K are included in the set, where K=n+2, where n is the number of perturbations in each group.

[0073] Generally speaking, a subset S k is designed by the active learning algorithm described in this disclosure.

[0074] Perturbation Library The perturbation library may be an unstructured or structured perturbation library, details of which can be found below.

[0075] With respect to unstructured perturbation libraries, in one embodiment the disclosure provides a method for identifying a set of perturbations that confer a specified phenotype, comprising: applying a set of perturbations randomly selected from a perturbation population to each cell in the perturbation library, wherein the mean number of perturbations per cell (MOP) is greater than the phenotypic complexity; Identifying a particular perturbation applied to each cell for the population of cells that are positive for a phenotype, and identifying a particular perturbation applied to each cell for the population of cells that are negative for a phenotype (which may or may not involve physical separation of the populations of cells that are positive and negative for a given phenotype; Calculating, for all possible perturbation combinations from the data obtained in (2.2), the probability that the set of perturbations includes all elements of the causal control factor (see Appendix A, Section 0.1, Equation 1) and / or the probability that the set of perturbations confers a phenotype of interest and / or a prioritized list of perturbation combinations for subsequent analysis or experimentation. The present invention provides a method comprising:

[0076] With respect to structured perturbation libraries, in another embodiment, the disclosure provides a method for identifying a set of perturbations that confer a specified phenotype, comprising: applying a set of perturbations assigned by combinatorial assignment (see Combinatorial Perturbation Library Construction below), where the mean number of perturbations per cell (MOP) is greater than the phenotypic complexity; Identifying a particular perturbation applied to each cell for the population of cells that are positive for a phenotype, and identifying a particular perturbation applied to each cell for the population of cells that are negative for a phenotype (which may or may not involve physical separation of the populations of cells that are positive and negative for a given phenotype); Calculating, for all possible perturbation combinations from the data obtained in (3.2), the probability that the set of perturbations contains all elements of the causal control factor (see Appendix A, Section 0.1, Equation 1) and / or the probability that the set of perturbations confers a phenotype of interest and / or a prioritized list of perturbation combinations for subsequent analysis or experimentation. The present invention provides a method comprising:

[0077] With respect to structured perturbation libraries with active learning, in one embodiment, the present disclosure provides a method for identifying a set of perturbations that confers a specified phenotype, comprising: applying a set of perturbations to each cell in the perturbation library, where the perturbations are drawn from the perturbation population by either random selection or combinatorial assignment, such that the mean number of perturbations per cell (MOP) is greater than the phenotypic complexity; Identifying a particular perturbation applied to each cell for the population of cells that are positive for a phenotype, and identifying a particular perturbation applied to each cell for the population of cells that are negative for a phenotype (which may or may not involve physical separation of the populations of cells that are positive and negative for a given phenotype); Calculating, for all possible perturbation combinations from the data obtained in (4.2), the probability that the set of perturbations includes all elements of the causal control factors and / or the probability that the set of perturbations will confer a phenotype of interest and / or a prioritized list of perturbation combinations for subsequent analysis or experimentation; Identify a collection of maximally informative unobserved perturbations that can be combinatorially applied to cells as input to 4.1, and continue iteratively as in an active learning algorithm (see Appendix A, section 0.2, Equation 3). The present invention provides a method comprising:

[0078] statistical inference methods An example of a statistical inference method is provided in Appendix A below: D = {(ν0,ω0),…,(ν n ,ω n )} denote the set of data, where v k is the perturbation set ω k is 1 if observed in a cell that is positive for a given phenotype, otherwise it is 0. ν Let (ω) be the probability that the perturbation combination ω gives the phenotype ν∈{0,1}. ν (ω) is trained using the data D. If M is the set of all possible perturbation combinations and M0 is the set of observed combinations, then for each ω∈M, the posterior probability that ω contains all and only causal controls of the phenotype is

[0079]

number

[0080]

number

[0081]

number

[0082] 0.2 Active Learning: Phenotyping Maximally Informative Combinations of Perturbations The information I obtained by analyzing the representation of the perturbed set σ∈M is I[σ]=H ω -H ω|σ (2) where D is the set of data collected at the selected time point and H ω is the entropy of ω, and H ω|σ is the conditional entropy of ω given σ). Intuitively, I[σ] is the expected reduction in entropy of ω (or information gained) that would be realized if the phenotype of σ was known. With Equation 2, we calculate the estimated information gain from phenotyping σ as:

[0083]

number

[0084] cell In any embodiment, the cell may be a mammalian cell.

[0085] Suitable cells include stem cells, progenitor cells, and partially and fully differentiated cells. Suitable cells include neurons, liver cells, kidney cells, immune cells, cardiac cells, skeletal muscle cells, smooth muscle cells, lung cells, and the like.

[0086] Suitable cells include stem cells (e.g., embryonic stem (ES) cells, induced pluripotent stem (iPS) cells; embryonic cells (e.g., oocytes, sperm, oogonia, spermatogonia, etc.); somatic cells, such as fibroblasts, oligodendrocytes, glial cells, hematopoietic cells, neurons, muscle cells, bone cells, hepatic cells, pancreatic cells, etc.

[0087] Suitable cells include human embryonic stem cells, fetal cardiomyocytes, myofibroblasts, mesenchymal stem cells, autologous expanded cardiomyocytes, adipocytes, totipotent cells, pluripotent cells, blood stem cells, myoblasts, adult stem cells, bone marrow cells, mesenchymal cells, embryonic stem cells, parenchymal cells, epithelial cells, endothelial cells, mesothelial cells, fibroblasts, osteoblasts, chondrocytes, exogenous cells, endogenous cells, stem cells, hematopoietic stem cells, bone marrow derived progenitor cells, cardiomyocytes, skeletal cells, fetal cells, undifferentiated cells, multipotent progenitor cells, unipotent progenitor cells, monocytes, cardiac myoblasts, skeletal myoblasts, macrophages, capillary endothelial cells, xenogeneic cells, allogeneic cells, and post-natal stem cells.

[0088] In some cases, the cell is a stem cell. In some cases, the cell is an induced pluripotent stem cell. In some cases, the cell is a mesenchymal stem cell. In some cases, the cell is a hematopoietic stem cell. In some cases, the cell is an adult stem cell.

[0089] Suitable cells include bronchoalveolar stem cells (BASCs), bulge epithelial stem cells (bESCs), corneal epithelial stem cells (CESCs), cardiac stem cells (CSCs), epidermal neural crest stem cells (eNCSCs), embryonic stem cells (ESCs), endothelial progenitor cells (EPCs), hepatic oval cells (HOCs), hematopoietic stem cells (HSCs), keratinocyte stem cells (KSCs), mesenchymal stem cells (MSCs), neuronal stem cells (NSCs), pancreatic stem cells (PSCs), retinal stem cells (RSCs), and skin-derived progenitor cells (SKPs).

[0090] In some cases, the cell is an immune cell. Suitable mammalian immune cells include primary cells and immortalized cell lines. Suitable mammalian cell lines include human cell lines, non-human primate cell lines, rodent (e.g., mouse, rat) cell lines, and the like. In some cases, the cell is not an immortalized cell line, but instead is a cell (e.g., a primary cell) obtained from an individual. For example, in some cases, the cell is an immune cell, immune cell precursor, or immune stem cell obtained from an individual. As an example, the cell is a lymphoid cell, e.g., a lymphocyte, or a precursor thereof, obtained from an individual. As another example, the cell is a cytotoxic cell, or a precursor thereof, obtained from an individual. As another example, the cell is a stem cell or a progenitor cell obtained from an individual.

[0091] As used herein, the term "immune cells" generally includes white blood cells / leukocytes derived from hematopoietic stem cells (HSCs) produced in bone marrow. "Immune cells" include, for example, lymphoid cells, i.e., lymphocytes (T cells, B cells, natural killer (NK) cells), and myeloid-derived cells (neutrophils, eosinophils, basophils, monocytes, macrophages, dendritic cells). "T cells" include all types of immune cells that express CD3, including helper T cells (CD4+ cells), cytotoxic T cells (CD8+ cells), regulatory T cells (Tregs), and gamma-delta T cells. "Cytotoxic cells" include CD8+ T cells, natural killer (NK) cells, and neutrophils, which can mediate cytotoxic responses. "B cells" include mature and immature cells of the B cell lineage, including cells that express CD19, such as, for example, pre-B cells, immature B cells, mature B cells, memory B cells, and plasmablasts. Immune cells also include derived cells of the B cell lineage, such as B cell precursors, e.g., pro-B cells and plasma cells.

[0092] In other embodiments, the cells may be cancer cells, for example malignant cells grown in culture.

[0093] usefulness The method finds use in identifying perturbations that can generate a particular phenotype, such as for identifying perturbations that cause stem cells to differentiate in a particular way (e.g., into any of the cell types listed above), or for identifying perturbations that can make therapeutic cells more effective (e.g., reduce T cell exhaustion).

[0094] In particular, the method can be used to identify perturbations that cause cell differentiation, reprogramming, and / or transdifferentiation. Examples of uses include: (1) differentiating induced pluripotent stem cells (iPSCs) into human cells with therapeutic or regenerative potential (e.g., cytotoxic or anti-inflammatory T cells), (2) regenerating a pool of non-regenerative cells (e.g., neurons) from a proximal regenerative population (e.g., astrocytes, microglia) by transdifferentiation, (3) stabilizing an existing cell type (e.g., cytotoxic T cells resistance to exhaustion or regulatory T cells resistance to inflammation), or (4) identifying perturbations that can construct hybrid cell types from multiple human or non-human cell types that combine therapeutically advantageous properties. These and other uses should be readily apparent.

[0095] All patents, patent applications, provisional applications, and publications mentioned or cited in this specification are incorporated by reference in their entirety, including all figures and tables, to the extent not inconsistent with the explicit teachings of this specification.

[0096] The following are examples illustrating procedures for carrying out the present invention. These examples should not be construed as limiting. Unless otherwise indicated, all percentages are by weight and all solvent mixture proportions are by volume. EXAMPLES

[0097] Example 1 Combinatorial identification of genes that confer T cell receptor (TCR) cell surface presentation to non-immune cells. To illustrate the method of the present invention, the following proof-of-principle experiment was devised, namely, using the method to identify all molecular components of the T cell receptor complex required for cell surface presentation (Figure 6). Six proteins are required to present the TCR complex to non-immune cells, and the challenge is to distinguish these proteins from 24 other unrelated factors. The population of perturbations in this case is a collection of 30 separate genes (6 TCR components and 24 unrelated factors) that can be overexpressed in target cells. A perturbation library consisting of 86 perturbation types with an average MOP of 14 was constructed. TCR positive and negative cells were isolated by flow cytometry and subjected to single-cell RNA sequencing to identify the set of perturbations applied to each cell. All 86 perturbation types were identified by observing an average of 56 cells for each perturbation type. These data were then used to train a binary classifier using an ensemble (bagging) decision tree statistical model. We then used this trained model to estimate the posterior probability of TCR presentation for each combination of unobserved perturbations (1,073,741,824 models were analyzed). These estimates clearly show that the actual composition of the TCR is the most likely predictor of TCR presentation compared to other models of similar complexity (Figure 7).

Claims

1. A split-and-pool method for exposing cells to perturbations, (a) A step of dividing the cells into multiple sections, (b) step of selecting a subset of the perturbation, (c) A step of applying a partial combination of a subset of the perturbation to the division, (d) Optionally, apply all perturbations in the subset to the divisions. (e) Optional step in which no perturbation applies to the classification. (f) The step of pooling cells after (e), and (g) A step which repeats steps (a) to (f) once or more times, where each iteration is performed using a different subset of the perturbation. The split-and-pool method, which includes [this method].

2. The method according to claim 1, wherein the subset of perturbations selected in (b) overlaps with at least one of the subsets of perturbations selected in the iteration of (g).

3. The method according to claim 1, wherein the subset of perturbations selected in (b) does not overlap with any of the subsets of perturbations selected in the iteration of (g).

4. The method according to any one of claims 1 to 3, wherein neither a chemical substance addition step nor a deprotection step is present.

5. The method according to any one of claims 1 to 4, wherein (g) is the repeating of steps (a) to (f) at least twice.

6. The method according to any one of claims 1 to 5, wherein (a) divides the cells into at least four sections.

7. The method according to any one of claims 1 to 6, wherein at least 1 M cells are present.

8. The method according to any one of claims 1 to 7, wherein the perturbation is a nucleic acid construct, and each construct encodes a perturbation.

9. The method according to claim 8, wherein the construct encodes a protein, RNA, or any combination thereof.

10. The method according to any one of claims 1 to 9, wherein the perturbation is a low molecule, and each low molecule is a perturbation.

11. The method according to any one of claims 1 to 10, wherein the separated cells are barcoded in steps (c) to (d), and the barcode indicates which perturbation was applied to the cells.

12. The method according to any one of claims 1 to 11, wherein the cells are mammalian cells.