Method and system for identifying target gene
Patent Information
- Application Number
- JP2025075680
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-21
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-09
AI Technical Summary
Identifying genetic drivers for cellular reprogramming between differentiated states is challenging due to complex gene interactions and the difficulty in distinguishing causal from correlated genes, requiring extensive experimental assays.
High-content, high-throughput CRISPR screening methods with anomaly detection models to quantify transcriptional reprogramming and identify target genes for therapeutic applications, using single-cell RNA-seq profiling, supervised dimensionality reduction, and pooled CRISPR editing experiments.
Enhances the efficiency, accuracy, and throughput of identifying genomic regions that promote cellular reprogramming, enabling the selection of biomarkers and therapeutic targets for disease symptoms.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Patent Application No. 62 / 865,033, filed June 21, 2019, the entirety of which is incorporated herein by reference. [Background technology]
[0002] In therapeutic applications, the ability to convert cells from one differentiated state to another may hold great promise. However, despite the promise of cellular reprogramming, identifying the genetic drivers that can mediate the transition between one cellular state and another remains challenging for many therapeutic applications. Reprogramming phenotypes can be complex and involve many genes that interact with each other in a hierarchical, nonlinear manner. Disentangling which of these genes are causal versus correlated in a process can be a difficult task, and may require extensive and time-consuming experimental assays and animal models for each gene of interest. Summary of the Invention
[0003] The present specification recognizes the need for an improved method for identifying therapeutically targeted genomic regions that can promote the reprogramming of cells from one phenotypic state to another phenotypic state.The methods and systems presented herein can significantly increase the efficiency, accuracy and / or throughput of identifying therapeutically targeted genomic regions that can promote the reprogramming of cells from one phenotypic state to another phenotypic state.
[0004] The present disclosure generally relates to methods and systems for quantifying transcriptional reprogramming of cells from one differentiated state to another. In particular, the technology relates to high-content, high-efficiency, and high-throughput CRISPR (clustered regularly interspaced short palindromic repeats) screening methods to identify relevant target genes that may mediate reprogramming between phenotypically distinct cell states and / or may be selected as effective therapeutic targets. These screenings can utilize abnormality detection models to quantify reprogramming as a measurable phenotype for each CRISPR-targeted gene. The methods and systems of the present disclosure can establish the quantification of reprogramming as a criterion for selecting biomarkers and therapeutic targets related to disease symptoms of interest.
[0005] In one aspect, the present disclosure provides a method for quantifying transcriptional transitions ("reprogramming") between differentiated or phenotypically distinct cell populations. The method may include: (a) single-cell RNA-seq profiling of distinct cell populations; (b) supervised dimensionality reduction of the single-cell RNA-seq profiles into a latent space of topological representations; (c) identifying endogenous genetic drivers ("genes") that mediate transitions between cell populations using a systems biology approach; (d) collating potential genetic drivers using pooled CRISPR editing experiments; and (e) applying anomaly detection methods to quantify the degree of transcriptional reprogramming from one distinct phenotypic state to another distinct phenotypic state for each collated genetic driver.
[0006] In another aspect, the present disclosure provides methods for identifying biomarkers and potential therapeutic target genes for various disease indications. The methods may include (a) identifying appropriate diseases and target cell populations, (b) identifying potential genetic drivers that mediate transitions between the disease and target cell populations, and (c) quantifying reprogramming of each of the genetic drivers. In other embodiments, multiple biomarkers or target genes can be identified by combinatorial inhibition or activation of multiple genes.
[0007] In some embodiments, the cell population is derived from a healthy individual or from a relevant tissue of a patient with a disease corresponding to the symptom of interest. In other embodiments, the cell population is derived from a primary cell line, a human organoid, an animal model, or other suitable model system. In some cases, the disease cell population is characterized by a specific genotype signature, such as a specific mutation in a gene of interest.
[0008] In some embodiments, the target cell population corresponds to a fully differentiated state derived from healthy tissue, a wild-type primary cell line, an organoid, an animal model, or other suitable model system, while in other embodiments, the target cell population corresponds to an intermediate state, such as stem cells, pre-cancerous cells, senescent cells, or progenitor cells associated with disease progression.
[0009] In some embodiments, the CRISPR system is selected from the group consisting of CRISPR (e.g., active Cas9), CRISPRi (e.g., CRISPR interference, catalytically inactive Cas9 fused to a transcriptional repressor peptide including KRAB), CRISPRa (e.g., CRISPR activity, catalytically inactive Cas9 fused to a transcriptional activator peptide including VPR (HIV viral protein R)), RNAi, and shRNA.
[0010] Another aspect described herein is a method for identifying one or more genomic regions that promote reprogramming of a cell from one phenotypic state to another phenotypic state, the method comprising: providing single-cell ribonucleic acid (RNA) sequencing data of a plurality of diseased cells and a plurality of normal cells of a cell type; mapping the single-cell RNA sequencing data of the plurality of diseased cells and the plurality of normal cells to a latent space corresponding to a plurality of phenotypic states of the cell type; identifying the one or more genomic regions that promote reprogramming of a cell type between a first phenotypic state and a second phenotypic state of the plurality of phenotypic states based at least in part on a topology of the latent space, wherein the one or more genomic regions are edited to promote the reprogramming of the cell type between the first phenotypic state and the second phenotypic state; and electronically outputting the one or more genomic regions.
[0011] In one embodiment, the mapping step comprises using a dimensionality reduction algorithm. In one embodiment, the dimensionality reduction algorithm comprises a uniform manifold approximation and projection (UMAP) algorithm. In one embodiment, the UMAP algorithm is a supervised UMAP algorithm. In one embodiment, the supervised UMAP algorithm is trained on single-cell RNA sequencing data of pure cells of the cell type. In one embodiment, the UMAP algorithm is trained using a minimum distance of about 0.025 to 0.25.
[0012] In one embodiment, the identifying step comprises performing a non-linear reconstruction of cell trajectories in the latent space to construct maximum likelihood progression trajectories between the first and second phenotypic states, and using probabilistic inference based on the maximum likelihood progression trajectories to identify the one or more genomic regions that promote reprogramming of the cell type between the first and second phenotypic states.
[0013] In one embodiment, performing the non-linear cell trajectory reconstruction comprises applying a graph inverse embedding algorithm to the latent space, hi one embodiment, the first phenotypic state is cancer and the second phenotypic state is a wild-type state.
[0014] In one embodiment, the method further comprises, prior to said mapping, removing low frequency genomic regions from the single-cell RNA-sequencing data of the plurality of diseased cells and the plurality of normal cells. In one embodiment, the method further comprises, in each of the one or more genomic regions, editing the respective genomic region using a genome editing tool to promote the reprogramming of cells of the cell type between the first phenotypic state and the second phenotypic state. In one embodiment, the genome editing tool is selected from the group consisting of a CRISPR system, a CRISPRi system, a CRISPRa system, an RNAi system, and an shRNA system.
[0015] In one embodiment, the method further comprises measuring the amount of deviation in the latent space of the cell resulting from editing the respective genomic regions using the genome editing means using an anomaly detection algorithm.
[0016] In one embodiment, the anomaly detection algorithm is trained on latent space profiles of multiple cell types. In one embodiment, the multiple cell types include pancreatic ductal cells, pancreatic acinar cells, pancreatic adenocarcinoma, and / or pancreatic adenocarcinoma. In one embodiment, the anomaly detection algorithm includes one or more of density-based methods, subspace-based outlier detection, correlation-based outlier detection, tensor-based outlier detection, support vector machines (SVMs), single-class vector machines, support vector data description, neural networks, Bayesian networks, hidden Markov models (HMMs), cluster analysis-based outlier detection, association rules and frequent itemset deviations, fuzzy logic-based outlier detection, and ensemble methods. In one embodiment, the anomaly detection algorithm is a support vector machine (SVM), a density-based method, a k-nearest neighbor algorithm, a local outlier factor algorithm, or an isolation forest algorithm. In one embodiment, the anomaly detection algorithm is a support vector machine. In one embodiment, the anomaly detection algorithm is an isolation forest algorithm.
[0017] In one embodiment, the method further comprises measuring a distance (e.g., Chebyshev distance, correlation distance, cosine distance, Euclidean distance, signed Euclidean distance, Hamming distance, Jaccard distance, Kullback-Leibler distance, Mahalanobis distance, Manhattan distance, Minkowski distance, or Spearman distance) of the latent space of the cells resulting from editing the respective genomic regions using the genome editing means. In one embodiment, the method further comprises editing each of the one or more genomic regions using the genome editing means to facilitate the reprogramming of each cell of the cell type between the first phenotypic state and the second phenotypic state, measuring the amount of deviation in the latent space of each of the cells resulting from editing the respective genomic region using the genome editing means, and ranking the one or more genes for therapeutic targeting using the measured amount of deviation.
[0018] In one embodiment, the method further includes measuring the shift in the latent space of the cells resulting from editing the respective genomic regions using the genome editing means using a density estimation function (e.g., probability density estimation, rescaled histogram, parametric density estimation function, non-parametric density estimation function (e.g., kernel density function), or data clustering method (e.g., vector quantization)). In one embodiment, the cell type is a pancreatic cell. In one embodiment, the diseased cell is a cancer cell. In one embodiment, the plurality of diseased cells and the plurality of normal cells are selected from the group consisting of a primary cell line, a human organoid, and an animal model.
[0019] In one embodiment, the method further comprises generating the single-cell RNA-seq data for a plurality of diseased cells and a plurality of normal cells of a cell type. In one embodiment, the second phenotypic state is an intermediate state. In one embodiment, the intermediate state is a pre-cancerous state or a low-grade state.
[0020] In one embodiment, the method further comprises identifying one or more therapeutic targets based on the one or more genomic regions for treating a disease associated with the first phenotypic condition.
[0021] In one embodiment, the method further comprises identifying, based at least in part on the topology of the latent space, one or more first genomic regions that promote reprogramming of the cell type between the first phenotypic state and an intermediate phenotypic state of the plurality of phenotypic states, wherein the one or more first genomic regions are edited to promote reprogramming of the cell type between the first phenotypic state and the intermediate phenotypic state; and identifying, based at least in part on the topology of the latent space, one or more second genomic regions that promote phenotypic reprogramming between the intermediate phenotypic state and the second phenotypic state of the plurality of phenotypic states, wherein the one or more second genomic regions are edited to promote reprogramming of the cell type between the intermediate phenotypic state and the second phenotypic state.
[0022] Another aspect described herein is a method for identifying one or more genomic regions that promote reprogramming of a cell from one phenotypic state to another phenotypic state, the method comprising: providing single-cell ribonucleic acid (RNA) sequencing data of a plurality of diseased cells and a plurality of normal cells of a cell type; using a supervised dimensionality reduction algorithm to map the single-cell RNA sequencing data of the plurality of diseased cells and the plurality of normal cells to a latent space corresponding to a plurality of phenotypic states of the cell type; and identifying the one or more genomic regions that promote reprogramming of the cell type between a first phenotypic state and a second phenotypic state of the plurality of phenotypic states based at least on a topology of the latent space. identifying, based on a set of supervised dimensionality reduction algorithms, one or more genomic regions configured to be edited to promote the reprogramming of the cell type between the first phenotypic state and the second phenotypic state; electronically outputting the one or more genomic regions; editing each of the one or more genomic regions using genome editing means to promote the reprogramming of cells of the cell type between the first phenotypic state and the second phenotypic state at each genomic region of the one or more genomic regions; and measuring a deviation in the latent space of the cells resulting from editing each genomic region using the genome editing means using an anomaly detection algorithm. In one embodiment, the supervised dimensionality reduction algorithm is a variational autoencoder.
[0023] Another aspect described herein is a system for identifying one or more genomic regions that promote reprogramming of a cell from one phenotypic state to another phenotypic state, the system including: a database including single-cell ribonucleic acid (RNA) sequencing data of a plurality of diseased cells and a plurality of normal cells of a cell type; and one or more computer processors individually or collectively programmed to: map the single-cell RNA sequencing data of the plurality of diseased cells and the plurality of normal cells to a latent space corresponding to a plurality of phenotypic states of the cell type; identify the one or more genomic regions that promote reprogramming of the cell type between a first phenotypic state and a second phenotypic state of the plurality of phenotypic states based at least in part on a topology of the latent space, wherein the one or more genomic regions are configured to be edited to promote the reprogramming of the cell type between the first phenotypic state and the second phenotypic state; and electronically output the one or more genomic regions.
[0024] In one embodiment, the mapping comprises using a dimensionality reduction algorithm, hi one embodiment, the dimensionality reduction algorithm comprises a uniform manifold approximation and projection (UMAP) algorithm.
[0025] Another aspect described herein is a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements a method for identifying one or more genomic regions that promote reprogramming of a cell from one phenotypic state to another phenotypic state, the method comprising: providing single-cell ribonucleic acid (RNA) sequencing data of a plurality of diseased cells and a plurality of normal cells of a cell type; mapping the single-cell RNA sequencing data of the plurality of diseased cells and the plurality of normal cells to a latent space corresponding to a plurality of phenotypic states of the cell type; identifying the one or more genomic regions that promote reprogramming of a cell type between a first phenotypic state and a second phenotypic state of the plurality of phenotypic states based at least in part on the topology of the latent space, wherein the one or more genomic regions are edited to promote the reprogramming of the cell type between the first phenotypic state and the second phenotypic state; and electronically outputting the one or more genomic regions.
[0026] In one embodiment, the mapping step comprises using a dimensionality reduction algorithm, hi one embodiment, the dimensionality reduction algorithm comprises a uniform manifold approximation and projection (UMAP) algorithm.
[0027] Another aspect described herein is a system for identifying one or more genomic regions that promote reprogramming of a cell from one phenotypic state to another phenotypic state, the system comprising: a database including single-cell ribonucleic acid (RNA) sequencing data for a plurality of diseased cells and a plurality of normal cells of a cell type; and one or more computer processors, the computer processor using a supervised dimensionality reduction algorithm to map the single-cell RNA sequencing data for the plurality of diseased cells and the plurality of normal cells to a latent space corresponding to a plurality of phenotypic states of the cell type; and identifying the one or more genomic regions that promote reprogramming of the cell type between a first phenotypic state and a second phenotypic state of the plurality of phenotypic states based at least on a topology of the latent space. and one or more computer processors individually or collectively programmed to: identify the one or more genomic regions configured to be edited to promote the reprogramming of the cell type between the first phenotypic state and the second phenotypic state; electronically output the one or more genomic regions; edit each of the one or more genomic regions using a genome editing means to promote, at each genomic region, the reprogramming of cells of the cell type between the first phenotypic state and the second phenotypic state; and measure, using an anomaly detection algorithm, a deviation in the latent space of the cells resulting from editing each genomic region using the genome editing means.
[0028] Another aspect described herein is a non-transitory computer readable medium comprising machine-executable code that, when executed by one or more computer processors, implements a method for identifying one or more genomic regions that promote reprogramming of a cell from one phenotypic state to another phenotypic state, the method comprising: providing single-cell ribonucleic acid (RNA) sequencing data of a plurality of diseased cells and a plurality of normal cells of a cell type; using a supervised dimensionality reduction algorithm to map the single-cell RNA sequencing data of the plurality of diseased cells and the plurality of normal cells into a latent space corresponding to a plurality of phenotypic states of the cell type; and identifying the one or more genomic regions that promote reprogramming of the cell type between a first phenotypic state and a second phenotypic state of the plurality of phenotypic states. identifying, based at least on a topology of the latent space, one or more genomic regions configured to be edited to promote the reprogramming of the cell type between the first phenotypic state and the second phenotypic state; electronically outputting the one or more genomic regions; editing each of the one or more genomic regions using a genome editing means to promote the reprogramming of cells of the cell type between the first phenotypic state and the second phenotypic state, at each genomic region of the one or more genomic regions; and measuring, using an anomaly detection algorithm, a deviation in the latent space of the cells resulting from editing each genomic region using the genome editing means.
[0029] Another aspect described herein is a method for identifying one or more genomic regions for therapeutic targeting, the method comprising: providing single-cell ribonucleic acid (RNA) sequencing data for a plurality of diseased cells and a plurality of normal cells of a cell type; mapping the single-cell RNA sequencing data for the plurality of diseased cells and the plurality of normal cells to a latent space; identifying the one or more genomic regions for therapeutic targeting based at least in part on a topology of the latent space; and electronically outputting the one or more genomic regions for identifying therapeutic targets.
[0030] In one embodiment, the mapping step comprises using a dimensionality reduction algorithm, hi one embodiment, the dimensionality reduction algorithm comprises a uniform manifold approximation and projection (UMAP) algorithm.
[0031] Another aspect of the present disclosure provides a non-transitory computer-readable medium containing machine-executable code that, when executed by one or more computer processors, performs the methods described above or elsewhere herein.
[0032] Another aspect of the present disclosure provides a system including one or more computer processors and a computer memory coupled thereto, the computer memory including machine-executable code that, when executed by the one or more computer processors, performs a method described above or elsewhere herein.
[0033]
[0013] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0034] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are incorporated by reference herein to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede and / or take precedence over any such conflicting material. [Brief explanation of the drawings]
[0035] The novel features of the invention are set forth with particularity in the appended claims. The features and advantages of the present invention will be better understood by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figure" and "FIG."). [Figure 1] FIG. 1 shows an example of a flow chart illustrating a method for identifying reprogramming target genes according to an embodiment of the disclosure. [Figure 2] FIG. 2 illustrates a computer system that is programmed or otherwise configured to perform the methods presented herein. [Figure 3A] 3A shows an example of quantifying and identifying novel therapeutic target gene reprogramming in accordance with disclosed embodiments. By leveraging CRISPR (clustered regularly interspaced short palindromic repeats) gene interrogation, intelligent construction of latent space, and anomaly detection, target genes are quantified according to their ability to program disease cell populations toward a desired target phenotypic state. Target states can be derived from healthy tissue or primary cell lines, or represent intermediate states including, but not limited to, senescent cells, stem cells, precancerous cells, or progenitor cells associated with disease progression. [Figure 3B]FIG. 3B illustrates an example of anomaly detection as a method for demarcating a dense manifold that accurately represents the phase space occupied by different cell populations, according to an embodiment of the disclosure. [Figure 3C] Figure 3C shows an example of leveraging anomaly detection to identify top metabolic reprogramming targets for cancer, according to disclosed embodiments. By utilizing multiple anomaly detectors, genes that maximally reprogram cancer cells to a wild-type primary expression profile can be identified based on the trained model's decision function (e.g., distance to a discrete manifold, such as Chebyshev distance, correlation distance, cosine distance, Euclidean distance, signed Euclidean distance, Hamming distance, Jaccard distance, Kullback-Leibler distance, Mahalanobis distance, Manhattan distance, Minkowski distance, or Spearman distance). Here, apoptotic cells were also included to model potential toxic complications that may occur in healthy cells derived from the target of interest (the "toxic" cluster label). [Figure 4A] FIG. 4A shows a comparison of several dimensionality reduction algorithms, including principal component analysis (PCA), t-distributed stochastic neighbor embedding (t-SNE), and uniform manifold approximation and projection (UMAP), applied to a cell-type mixed dataset, according to disclosed embodiments. [Figure 4B] FIG. 4B shows a comparison of the stability of latent spaces constructed by t-SNE and UMAP trained on pancreatic ductal, acinar, and adenocarcinoma cell lines, according to disclosed embodiments. [Figure 4C] FIG. 4C shows an example of the effect of the "minimum distance" parameter of UMAP on quantification of pancreatic cancer reprogramming, according to an embodiment of the disclosure. [Figure 4D] FIG. 4D shows an example of the effect of the dimensionality of the UMAP latent space on the quantification of pancreatic cancer reprogramming, according to an embodiment of the disclosure. [Figure 5A]FIG. 5A shows an example of a two-dimensional projection of pseudo-temporal ordering generated by a candidate selection pipeline featuring a transition from pancreatic acinar cells (dark shading on the right) to pancreatic ductal cells (medium shading in the middle) and then to high-grade cancer cells (KrasG12D, p53- / -, Myc) (light shading on the left) in accordance with an embodiment of the disclosure. [Figure 5B] FIG. 5B shows an example pipeline for generating candidates from causal inference based on a pseudo-temporal principal tree constructed from high-dimensional single-cell RNA-seq data, according to an embodiment of the disclosure. [Figure 6A] FIG. 6A shows a diagram of two anomaly detection algorithms being trained on two half-moon scatter plots with random Gaussian noise, according to an embodiment of the disclosure. [Figure 6B] Figure 6B shows an example heatmap of z-transformed anomaly detection decision functions across 70 single guide RNAs (sgRNAs) according to disclosed embodiments. Targets are ranked according to their average rank across five trained anomaly detector models. Three of the 70 sgRNAs resulted in adjusted p-values consistent with significant reprogramming to the primary cell state ((32), (52), and (38)). [Figure 6C] Figure 6C shows an example of the effect of different anomaly detection algorithms on quantifying reprogramming, according to disclosed embodiments. In both algorithms, three targets ((32), (52), and (38)) resulted in adjusted p-values consistent with significant reprogramming. They shared 8 of the top 10 targets (80%) for each algorithm. [Figure 7A] Figure 7A shows the corresponding cell lines used for reprogramming analysis across different stages of pancreatic cancer progression. Primary pancreatic ductal cells and immortalized acinar cells were used as wild-type cells. Double-mutant pancreatic cancer cells (KrasG12D, p53- / -) were used as low-grade cancer cells. Triple-mutant pancreatic cancer cells (KrasG12D, p53- / -, Myc) were used as high-grade cancer cells. [Figure 7B]FIG. 7B shows an analysis of triple mutant pancreatic cancer cells (KrasG12D, p53- / -, Myc) reprogrammed into wild-type ductal or acinar cells (FIG. 7B). [Figure 7C] Figure 7C shows a heatmap plot of the z-transformed decision function for anomaly detection across 70 single-guide RNAs (Figure 7C), where targets are ordered according to their average rank across the five trained anomaly detector models. [Figure 7D] FIG. 7D shows an analysis of reprogramming triple mutant pancreatic cancer cells (KrasG12D, p53 − / −, Myc) into double mutant pancreatic cancer cells (KrasG12D, p53 − / −) (FIG. 7D). [Figure 7E] Figure 7E shows a heatmap plot of the z-transformed anomaly detection decision functions for 70 single-guide RNAs (Figure 7E), where targets are ordered according to their average rank across the five trained anomaly detector models. [Figure 7F] Figure 7F shows the analysis of triple-mutant pancreatic cancer cells (KrasG12D, p53- / -, Myc) reprogramming into wild-type ductal or acinar cells, with double-mutant pancreatic cancer cells (KrasG12D, p53- / -) as an intermediate cell type. [Figure 7G] Figure 7G shows a heatmap plot of the z-transformed anomaly detection decision function for 70 single-guide RNAs (Figure 7G), where targets are ordered according to their average rank across the five anomaly detector models trained. DETAILED DESCRIPTION OF THE INVENTION
[0036] While various embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be utilized.
[0037] The term " sequencing " as used herein generally refers to the process of generating or identifying the sequence of biomolecules such as nucleic acid molecules.Such sequence can be the sequence of nucleic acid, and can include the sequence of nucleic acid bases.Sequencing method can be massively parallel array sequencing (for example, Illumina sequencing), and can be carried out using template nucleic acid molecules that are fixed on support such as flow cell or beads. Sequencing methods include, but are not limited to, high-throughput sequencing, next-generation sequencing, sequencing-by-synthesis, flow sequencing, massively parallel sequencing, shotgun sequencing, single-molecule sequencing, nanopore sequencing, pyrosequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), Clonal Single Molecule Array (Solexa), and Maxam-Gilbert sequencing.
[0038] The term "subject," as used herein, typically refers to an individual having a biological sample undergoing processing or analysis. A subject may be an animal or a plant. A subject may be a mammal, such as a human, ape, monkey, chimpanzee, dog, cat, horse, pig, rodent (e.g., mouse or rat), reptile, amphibian, or bird. A subject may have or be suspected of having a disease, such as cancer (e.g., breast cancer, colon cancer, brain tumor, leukemia, lung cancer, skin cancer, liver cancer, pancreatic cancer, lymphoma, esophageal cancer, or cervical cancer) or an infectious disease.
[0039] The term "sample," as used herein, generally refers to a biological sample. Examples of biological samples include tissues, cells, nucleic acid molecules, amino acids, polypeptides, proteins, carbohydrates, fats, metabolites, hormones, and viruses. In one example, a biological sample is a nucleic acid sample containing one or more nucleic acid molecules, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA). The nucleic acid molecule can be acellular or acellular nucleic acid molecule, such as cell-free DNA or cell-free RNA. The nucleic acid molecule can be derived from a variety of sources, including human, mammalian, non-human mammalian, ape, monkey, chimpanzee, reptile, amphibian, or avian sources. Furthermore, samples can be extracted from various animal bodily fluids containing cell-free sequences, including, but not limited to, blood, serum, plasma, vitreous, sputum, urine, tears, sweat, saliva, semen, mucosal discharge, mucus, cerebrospinal fluid, amniotic fluid, lymphatic fluid, etc. Cell-free polynucleotides can be fetal in origin (via bodily fluids collected from a pregnant subject) or can be derived from the subject's own tissues.
[0040] The terms "nucleic acid" or "polynucleotide," as used herein, generally refer to a molecule comprising one or more nucleic acid subunits or nucleotides. A nucleic acid may comprise one or more nucleotides selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), or variants thereof. A nucleotide typically comprises one nucleoside and at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more phosphate (PO) groups. A nucleotide may comprise one nucleobase, a pentose sugar (either ribose or deoxyribose), and one or more phosphate groups.
[0041] A ribonucleotide is a nucleotide whose sugar is ribose. A deoxyribonucleotide is a nucleotide whose sugar is deoxyribose. A nucleotide can be a nucleoside monophosphate or a nucleoside polyphosphate. A nucleotide can be a deoxyribonucleoside polyphosphate, such as a deoxyribonucleoside triphosphate (dNTP), which can be selected from deoxyadenosine triphosphate (dATP), deoxycytidine triphosphate (dCTP), deoxyguanosine triphosphate (dGTP), uridine triphosphate (dUTP), and deoxythymidine triphosphate (dTTP), and includes a detectable tag, such as a luminescent tag or a marker (e.g., a fluorophore). A nucleotide can include any subunit that can be incorporated into a growing nucleic acid chain. Such subunits may be specific for one or more complementary A, C, G, T, or U, or may be A, C, G, T, or U, or any other subunit that is complementary to a purine (i.e., A or G, or a variant thereof) or pyrimidine (i.e., C, T, or U, or a variant thereof). In some examples, the nucleic acid is deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a derivative or variant thereof. The nucleic acid may be single-stranded or double-stranded. In some cases, the nucleic acid molecule is circular.
[0042] The terms "nucleic acid molecule," "nucleic acid sequence," "nucleic acid fragment," "oligonucleotide," and "polynucleotide," as used herein, generally refer to polynucleotides that can be of various lengths, such as deoxyribonucleotides or ribonucleotides (RNA), or analogs thereof. Nucleic acid molecules can be at least about 10 bases, 20 bases, 30 bases, 40 bases, 50 bases, 100 bases, 200 bases, 300 bases, 400 bases, 500 bases, 1 kilobase (kb), 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, 50 kb, or longer in length. Oligonucleotides typically consist of a specific sequence of the four nucleotide bases adenine (A), cytosine (C), guanine (G), and thymine (T) (uracil (U) is substituted for thymine (T) when the polynucleotide is RNA). Thus, the term "oligonucleotide sequence" can refer to the alphabetical representation of a polynucleotide molecule, or the term can apply to the polynucleotide molecule itself. This alphabetical representation can be input into a database of a computer having a central processing unit and used in life science informatics applications such as functional genomics and homology searching. Oligonucleotides can contain one or more non-standard nucleotides, nucleotide analogs and / or modified nucleotides.
[0043] The term "nucleotide analogue" as used herein includes diaminopurine and 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylketosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, Queuosine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-D46-isopentenyladenine, uracil-5-oxyacetic acid(v), wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methyl ester, uracil-5-oxyacetic acid(v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, 2,6-diaminopurine, phosphorothrenoate nucleic acid, etc. In some cases, the nucleotide may contain modifications at the phosphate moiety, including modifications at the triphosphate moiety. Further non-limiting examples of modifications include longer phosphate chains (e.g., phosphate chains having 4, 5, 6, 7, 8, 9, 10, or more than 10 phosphate moieties), modifications with thiol moieties (e.g., α-thiotriphosphate and β-thiotriphosphate), or modifications with selenium moieties (e.g., phosphoroselenoate nucleic acids). Nucleic acid molecules can also be modified at the base moiety (e.g., one or more atoms that can normally form hydrogen bonds with a complementary nucleotide and / or one or more atoms that cannot normally form hydrogen bonds with a complementary nucleotide), the sugar moiety, or the phosphate backbone.Nucleic acid molecules may also contain amine-modified groups such as aminoallyl-dUTP (aa-dUTP) and aminohexyl acrylamide-dCTP (aha-dCTP), which allow for the covalent attachment of amine-reactive moieties such as N-hydroxysuccinimide ester (NHS). Substitution of standard DNA or RNA base pairs in the disclosed oligonucleotides may result in higher density (bits per cubic millimeter (mm)), greater safety (e.g., resistance to accidental or intentional synthesis of natural toxins), easier discrimination in light-programmed polymerases, or reduced secondary structure. Nucleotide analogs may be capable of reacting with or binding to detectable moieties for nucleotide detection.
[0044] The term "free nucleotide analog," as used herein, refers to a nucleotide analog that is not normally bound to another nucleotide or nucleotide analog. A free nucleotide analog can be incorporated into a growing nucleic acid chain by a primer extension reaction.
[0045] The term "primer" as used herein generally refers to a polynucleotide complementary to a template nucleic acid. Complementarity, homology, or sequence identity between a primer and a template nucleic acid can be limited. The length of a primer can be between 8 and 50 nucleotide bases. The length of a primer can be 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 37, 40, 42, 45, 47, or 50 or more nucleotide bases.
[0046] Primer can show sequence identity or homology or complementarity with template nucleic acid.The homology or sequence identity or complementarity between primer and template nucleic acid can be based on the length of primer.For example, when the length of primer is about 20 nucleic acids, it can contain 10 or more consecutive nucleic acid bases that are complementary with template nucleic acid.
[0047] The term "primer extension reaction" as used herein generally refers to the primer binding to template nucleic acid strand, and then the primer is extended.It can also include the double-stranded nucleic acid being denatured, the primer strand being bound to one or both of the denatured template nucleic acid strands, and then the primer is extended.Primer extension reaction can be used to incorporate nucleotide or nucleotide analogue into primer in a template-directed manner by using enzyme (polymerase).
[0048] The term "polymerase," as used herein, generally refers to any enzyme capable of catalyzing a polymerization reaction. Examples of polymerases include, but are not limited to, nucleic acid polymerases. Polymerases can be naturally occurring or synthetic. In some cases, polymerases have relatively high processivity. An example of a polymerase is Φ29 polymerase, or a derivative thereof. A polymerase can be a polymerization enzyme. In some cases, a transcriptase or a ligase is used (i.e., an enzyme that catalyzes the formation of a bond). Examples of polymerases include DNA polymerase, RNA polymerase, thermostable polymerase, wild-type polymerase, modified polymerase, E. coli DNA polymerase I, T7 DNA polymerase, bacteriophage T4 DNA polymerase Φ29 (phi29) DNA polymerase, Taq polymerase, Tth polymerase, Tli polymerase, Pfu polymerase, Pwo polymerase, VENT polymerase, DEEPVENT polymerase, EX-Taq polymerase, LA-Taq polymerase, Sso polymerase, Poc polymerase, Pab polymerase, Mth polymerase, ES4 polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tea polymerase, Tih polymerase, Tfi polymerase, Platinum These include Taq polymerase, Tbr polymerase, Tfl polymerase, Pfutubo polymerase, Pyrobest polymerase, Pwo polymerase, KOD polymerase, Bst polymerase, Sac polymerase, Klenow fragment polymerase with 3'-5' exonuclease activity, and modified products and derivatives. In some cases, the polymerase is a single subunit polymerase. The polymerase may have high processivity, i.e., the ability to continuously incorporate nucleotides into a nucleic acid template without releasing the nucleic acid template.In some cases, the polymerase is a polymerase modified to accept dideoxynucleotide triphosphates, such as, for example, Taq polymerase with a 667Y mutation (see, e.g., Tabor et al., PNAS, 1995, 92, 6339-6343, which is incorporated herein by reference in its entirety for all purposes). In some cases, the polymerase is a polymerase with modified nucleotide linkages that may be useful for nucleic acid sequencing, non-limiting examples of which include ThermoSequenas polymerase (GE Life Sciences), AmpliTaq FS (ThermoFisher) polymerase, and Sequencing Pol polymerase (Jena Bioscience). In some cases, the polymerase is engineered to discriminate against dideoxynucleotides, such as, for example, Sequenase DNA polymerase (ThermoFisher).
[0049] The term "support," as used herein, generally refers to a solid support, such as a slide, bead, resin, chip, array, matrix, membrane, nanopore, or gel. A solid support can be, for example, beads on a flat substrate (e.g., glass, plastic, silicon, etc.) or beads in a well of a substrate. The substrate can have surface properties, such as a texture, pattern, microstructured coating, surfactant, or any combination thereof, to hold the beads in a desired location (e.g., in operable communication with a detector). Detectors on bead-based supports can be configured to maintain substantially the same read rate regardless of bead size. The support can be a flow cell or an open substrate. Furthermore, the support can include a biological support, a non-biological support, an organic support, an inorganic support, or any combination thereof. The support can be in optical communication with the detector, in physical contact with the detector, separated by a distance from the detector, or any combination thereof. The support can have multiple independently addressable locations. Nucleic acid molecules can be immobilized on the support at some of the independently addressable locations. Immobilization of each of the plurality of nucleic acid molecules to the support can be assisted by the use of an adapter. The support can optionally be optically coupled to a detector. Immobilization to the support can be assisted by an adapter.
[0050] The term "label," as used herein, generally refers to a moiety that can be attached to a species, such as, for example, a nucleotide analog. In some cases, the label can be a detectable label that generates a detectable signal (or suppresses a signal already generated). In some cases, such a signal can indicate the incorporation of one or more nucleotides or nucleotide analogs. In some cases, the label can be attached to a nucleotide or nucleotide analog, and the nucleotide or nucleotide analog can be used in a primer extension reaction. In some cases, the label can be attached to a nucleotide analog after the primer extension reaction. In some cases, the label can specifically react with a nucleotide or nucleotide analog. Attachment can be covalent (e.g., via ionic interactions, van der Waals forces, etc.) or non-covalent. In some cases, the linkage may be by a linker that may be cleavable, such as photocleavable (e.g., cleavable under ultraviolet light), chemically cleavable (e.g., by reducing agents such as dithiothreitol (DTT), tris(2-carboxyethyl)phosphine (TCEP)), or enzymatically cleavable (e.g., by esterases, lipases, peptidases, or proteases).
[0051] In some cases, the label may be optically active. In some embodiments, the optically active label is an optically active dye (e.g., a fluorescent dye). Non-limiting examples of dyes include SYBR green, SYBR blue, DAPI, propidium iodide, Hoeste, SYBR gold, ethidium bromide, acridine, proflavine, acridine orange, acriflavine, fluorcoumanin, ellipticine, daunomycin, chloroquine, distamycin D, chromomycin, homidium, mithramycin, ruthenium polypyridyl, anthramycin, phenanthridine and acridine, ethidium bromide, propidium iodide, hexidium iodide, dihydroethidium, ethidium homodimer-1 and -2, ethidium monoazide, and ACMA, Hoechst 33258, Hoechst 33342, Hoechst 34580, DAPI, acridine orange, 7-AAD, actinomycin D, LDS751, hydroxystilbamidine, SYTOX Blue, SYTOX Green, SYTOX Orange, POPO-1, POPO-3, YOYO-1, YOYO-3, TOTO-1, TOTO-3, JOJO-1, LOLO-1, BOBO-1, BOBO-3, PO-PRO-1, PO-PRO-3, BO-PRO-1, BO-PRO-3, TO-PRO-1, TO-PRO-3, TO-PRO-5, JO-PRO-1, LO-PRO-1, YO-PRO-1, YO-PRO-3, PicoGreen, OliGreen, RiboGreen, SYBR gold, SYBR Green I, SYBR Green II, SYBR DX, SYTO-40, -41, -42, -43, -44, -45 (blue), SYTO-13, -16, -24, -21, -23, -12, -11, -20, -22, -15, -14, -25 (green), SYTO-81, -80, -82, -83, -84, -85 (orange), SYTO-64, -17, -59, -61, -62, -60, -63 (red), fluorescein, fluorescein isothiocyanate (FITC), tetramethylrhodamine isothiocyanate (TRITC), rhodamine, tetramethylrhodamine, phycoerythrin, Cy-2, Cy-3, Cy-3.5, Cy-5, Cy5.5, Cy-7, Texas Red, Phar-Red, allophycocyanin (APC), Sybr Green I, Sybr Green II, Sybr Gold, CellTracker Green, 7-AAD, ethidium homodimer I, ethidium homodimer II, ethidium homodimer III, ethidium bromide, umbelliferone, eosin, green fluorescent protein, erythrosine, coumarin, methylcoumarin, pyrene, malachite green, stilbene, Lucifer Yellow, Cascade Blue, dichlorotriazinylamine fluorescein, dansyl chloride, fluorescent lanthanide complexes such as those containing europium and terbium, carboxytetrachlorofluorescein, 5- and / or 6-carboxyfluorescein (FAM), VIC, 5-(or 6-) iodoacetamidofluorescein, 5-{[2(and 3)-5-(acetylmercapto)-succinyl]amino}fluorescein (SAMSA-fluorescein), Lissamine rhodamine B sulfonyl chloride, 5 and / or 6 carboxyrhodamine (ROX), 7-amino-methyl-coumarin, 7-amino-4-methylcoumarin-3-acetic acid (AMCA), BODIPY fluorophores, 8-methoxypyrene-1,3,6-trisulfonic acid trisodium salt, 3,6-disulfonate-4-amino-naphthalimide, phycobiliproteins, AlexaFluor 350, 405, 430, 488, 532, 546, 555, 568, 594, 610, 633, 635, 647, 660, 680, 700, 750, and 790 dyes, DyLight These include 350, 405, 488, 550, 594, 633, 650, 680, 755, and 800 dyes, or other fluorophores.
[0052] In some examples, the label may be a nucleic acid intercalator dye. Examples include, but are not limited to, ethidium bromide, YOYO-1, SYBR Green, and EvaGreen. Near-field interactions between an energy donor and an energy acceptor, between an intercalator and an energy donor, or between an intercalator and an energy acceptor can result in the generation of a specific signal or a change in signal amplitude. For example, such interactions can result in quenching (i.e., energy transfer from the donor to the acceptor, resulting in non-radiative energy decay) or Förster resonance energy transfer (FRET) (i.e., energy transfer from the donor to the acceptor, resulting in radiative energy decay). Other examples of labels include electrochemical labels, electrostatic labels, colorimetric labels, and mass tags.
[0053] The term "quencher," as used herein, generally refers to a molecule capable of suppressing a generated signal. The label can be a quencher molecule. For example, a template nucleic acid molecule can be designed to generate a detectable signal. Incorporation of a nucleotide or nucleotide analog containing a quencher can suppress or eliminate the signal, which is then detected. In some cases, as described elsewhere herein, labeling with a quencher can occur after incorporation of a nucleotide or nucleotide analog. Examples of quenchers include Black Hole Quencher dyes such as BH1-0, BHQ-1, BHQ-3, and BHQ-10 (Biosearch Technologies), QSY dye fluorescence quenchers such as QSY7, QSY9, QSY21, and QSY35 (Molecular Probes / Invitrogen), and other quenchers such as Dabcyl and Dabsyl, Cy5Q, Cy7Q, and Dark Cyanine dyes (GE Healthcare). Examples of donor molecules whose signals can be suppressed or quenched with the quenchers include fluorophores such as Cy3B, Cy3, or Cy5, Dy-Quenchers such as DYQ-660 and DYQ-661 (Dyomics), and N-(7-dimethylamino-4-methylcoumarin-3-yl)maleimide (DACM) and ATTO fluorescence quenchers such as ATTO 540Q, 580Q, 612Q, 647N, Atto-633-iodoacetamide, tetramethylrhodamine iodoacetamide, or Atto-488 iodoacetamide (ATTO-TEC GmbH). In some cases, the label may be of a type that is not self-quenching, such as a bimane derivative, e.g., monobromobimane.
[0054] The term "detector," as used herein, generally refers to an instrument capable of detecting a signal, including a signal indicating the presence or absence of an incorporated nucleotide or nucleotide analog. In some cases, a detector may include optical and / or electronic components capable of detecting a signal. The term "detector" may be used in detection methods. Non-limiting examples of detection methods include optical detection, spectroscopic detection, electrostatic detection, electrochemical detection, and the like. Optical detection methods include, but are not limited to, fluorescence spectroscopy and ultraviolet-visible absorbance. Spectroscopic detection methods include, but are not limited to, mass spectrometry, nuclear magnetic resonance (NMR) spectroscopy, and infrared spectroscopy. Electrostatic detection methods include, but are not limited to, gel-based techniques such as gel electrophoresis. Electrochemical detection methods include, but are not limited to, electrochemical detection of amplification products after separation of the amplification products by high-performance liquid chromatography.
[0055] The terms "sequence" or "sequence read," as used herein, typically refer to a series of nucleotide assignments (e.g., by base calling) made during the sequencing process. Such sequences may be estimated sequence reads generated by performing preliminary base calling, after which further base calling analysis or corrections may be performed to generate a final sequence read. Sequences may contain information corresponding to single or individual cells and may be obtained by single-cell sequencing methods (e.g., single-cell RNA sequencing, or scRNA-seq). Single-cell sequencing may be performed to provide higher resolution information about cellular differences and the function of individual cells in the context of their microenvironment. For example, single-cell DNA sequencing may provide information about mutations present in rare cell populations (e.g., those found in cancer cells), while single-cell RNA sequencing may provide information about the expression of individual cells corresponding to the presence and behavior of other cell types.
[0056] The term "single guide RNA" or "sgRNA," as used herein, typically refers to a single RNA molecule that contains both a custom-designed short CRISPR RNA (crRNA) sequence fused with a scaffold of a trans-activating crRNA (tracrRNA) sequence. sgRNAs can be generated synthetically or made in vitro or in vivo from a DNA template.
[0057] In therapeutic applications, the ability to convert cells from one differentiated state to another may hold great promise. However, despite the promise of cellular reprogramming, identifying the genetic drivers that can mediate the transition between one cellular state and another remains challenging for many therapeutic applications. Reprogramming phenotypes can be complex and involve many genes that interact with each other in a hierarchical, nonlinear manner. Disentangling which of these genes are causal versus correlated in a process can be a difficult task, and may require extensive and time-consuming experimental assays and animal models for each gene of interest.
[0058] The present specification recognizes the need for an improved method for identifying therapeutically targeted genomic regions that can promote the reprogramming of cells from one phenotypic state to another phenotypic state.The methods and systems presented herein can significantly increase the efficiency, accuracy and / or throughput of identifying therapeutically targeted genomic regions that can promote the reprogramming of cells from one phenotypic state to another phenotypic state.
[0059] The present disclosure generally relates to methods and systems for quantifying transcriptional reprogramming of cells from one differentiated state to another. In particular, the technology relates to high-content, high-efficiency, and high-throughput CRISPR (clustered regularly interspaced short palindromic repeats) screening methods to identify relevant target genes that may mediate reprogramming between phenotypically distinct cell states and / or may be selected as effective therapeutic targets. These screenings can utilize abnormality detection models to quantify reprogramming as a measurable phenotype for each CRISPR-targeted gene. The methods and systems of the present disclosure can establish the quantification of reprogramming as a criterion for selecting biomarkers and therapeutic targets related to disease symptoms of interest.
[0060] In one aspect, the present disclosure provides a method for quantifying transcriptional transitions ("reprogramming") between differentiated or phenotypically distinct cell populations. The method may include: (a) single-cell RNA-seq profiling of distinct cell populations; (b) supervised dimensionality reduction of the single-cell RNA-seq profiles into a latent space of topological representations; (c) identifying endogenous genetic drivers ("genes") that mediate transitions between cell populations using a systems biology approach; (d) collating potential genetic drivers using pooled CRISPR editing experiments; and (e) applying anomaly detection methods to quantify the degree of transcriptional reprogramming from one distinct phenotypic state to another distinct phenotypic state for each collated genetic driver.
[0061] In another aspect, the present disclosure provides methods for identifying biomarkers and potential therapeutic target genes for various disease indications. The methods may include (a) identifying appropriate diseases and target cell populations, (b) identifying potential genetic drivers that mediate transitions between the disease and target cell populations, and (c) quantifying reprogramming of each of the genetic drivers. In other embodiments, multiple biomarkers or target genes can be identified by combinatorial inhibition or activation of multiple genes.
[0062] In some embodiments, the cell population is derived from a healthy individual or from a relevant tissue of a patient with a disease corresponding to the symptom of interest. In other embodiments, the cell population is derived from a primary cell line, a human organoid, an animal model, or other suitable model system. In some cases, the disease cell population is characterized by a specific genotype signature, such as a specific mutation in a gene of interest.
[0063] In some embodiments, the target cell population corresponds to a fully differentiated state derived from healthy tissue, a wild-type primary cell line, an organoid, an animal model, or other suitable model system, while in other embodiments, the target cell population corresponds to an intermediate state, such as stem cells, pre-cancerous cells, senescent cells, or progenitor cells associated with disease progression.
[0064] In some embodiments, the CRISPR system is selected from the group consisting of CRISPR (e.g., active Cas9), CRISPRi (e.g., CRISPR interference, catalytically inactive Cas9 fused to a transcriptional repressor peptide including KRAB), CRISPRa (e.g., CRISPR activity, catalytically inactive Cas9 fused to a transcriptional activator peptide including VPR (HIV viral protein R)), RNAi, and shRNA.
[0065] FIG. 1 shows an example flowchart illustrating a method (100) for identifying therapeutic targets, such as reprogramming target genes, according to disclosed embodiments. The method may include providing single-cell ribonucleic acid (RNA) sequencing data for a plurality of diseased cells and a plurality of normal cells of a cell type (as in operation (102)). In some embodiments, the method may include generating scRNA-seq data for the plurality of diseased cells and the plurality of normal cells. The method may then include mapping the single-cell RNA sequencing data for the plurality of diseased cells and the plurality of normal cells (e.g., using a dimensionality reduction algorithm such as the uniform manifold approximation and projection (UMAP) algorithm) to a latent space corresponding to a plurality of phenotypic states of the cell type (as in operation (104)). Alternatively, the mapping may be performed using a supervised dimensionality reduction algorithm, such as a variational autoencoder. The method may then include identifying one or more genomic regions that promote reprogramming of a cell type between a first phenotypic state and a second phenotypic state of the plurality of phenotypic states based at least in part on the topology of the latent space (e.g., the one or more genomic regions are configured to be edited to promote reprogramming of a cell type between the first phenotypic state and the second phenotypic state) (as in operation (106)). For example, the first phenotypic state may be a disease state (e.g., cancer) and the second phenotypic state may be a non-disease state (e.g., a wild-type or progenitor state), an early disease state (e.g., a precancerous state, an early stage cancer state, or a progenitor disease state), or an intermediate disease state (e.g., a low-severity or low-grade disease state).As another example, the method may further include identifying, based at least in part on the topology of the latent space, a first genomic region that promotes reprogramming of a cell type between a first phenotypic state and an intermediate phenotypic state, the first genomic region being edited to promote reprogramming of a cell type between the first phenotypic state and the intermediate phenotypic state, and identifying, based at least in part on the topology of the latent space, a second genomic region that promotes phenotypic reprogramming between the intermediate phenotypic state and a second phenotypic state, the second genomic region being edited to promote reprogramming of a cell type between the intermediate phenotypic state and the second phenotypic state. The method may then include electronically outputting the one or more genomic regions (as in act (108)). In some embodiments, the method may include identifying at least one of the genomic regions as a therapeutic target and / or therapeutically targeting at least one of the genomic regions to treat a subject in need thereof (e.g., having a disease state for which therapeutic targeting of the disease state would be an effective treatment). For example, therapeutic targeting may be achieved using small molecule inhibitors, antibody therapy, RNAi, antisense oligonucleotides, or a combination thereof.
[0066] In some embodiments, the UMAP algorithm is a supervised UMAP algorithm, or an unsupervised, supervised UMAP algorithm. For example, a supervised UMAP algorithm can be trained on a dataset containing single-cell RNA sequencing (scRNA-seq) data of pure cells of a cell type. The ... The method may be trained using a minimum distance of about 525, about 0.55, about 0.575, about 0.6, about 0.625, about 0.65, about 0.675, about 0.7, about 0.725, about 0.75, about 0.775, about 0.8, about 0.825, about 0.85, about 0.875, about 0.9, about 0.925, about 0.95, about 0.975, or about 1.0. In some embodiments, the method may further comprise removing low frequency genomic regions from single-cell RNA sequencing (scRNA-seq) data of the plurality of diseased cells and the plurality of normal cells prior to the mapping step.
[0067] Identifying one or more genomic regions that promote reprogramming of a cell type between a first phenotypic state and a second phenotypic state can be performed based on any of a number of suitable analyses of the topology of the latent space. As an example, a nonlinear reconstruction of cell trajectories can be performed in the latent space (e.g., by applying a graph inverse embedding algorithm to the latent space) to construct a maximum likelihood progression trajectory between the first phenotypic state and the second phenotypic state. Probabilistic inference can then be used based on the maximum likelihood progression trajectory to identify one or more genomic regions that promote reprogramming of a cell type between the first phenotypic state and the second phenotypic state. In some embodiments, one or more therapeutic targets for treating a disease associated with the first phenotypic state can be identified based on the identified genomic regions.
[0068] After the genomic regions are identified, genome editing techniques (e.g., CRISPR, CRISPRi, CRISPRa, RNAi, or shRNA) can be used to edit each genomic region to facilitate reprogramming of cells of a cell type between a first phenotypic state and a second phenotypic state. After editing, an anomaly detection algorithm (e.g., using a density estimation function) can be used to measure the deviation in the latent space of cells resulting from editing each genomic region using the genome editing technique. For example, distance measures (e.g., Chebyshev distance, correlation distance, cosine distance, Euclidean distance, signed Euclidean distance, Hamming distance, Jaccard distance, Kullback-Leibler distance, Mahalanobis distance, Manhattan distance, Minkowski distance, Spearman distance, or distance on a Riemannian manifold) can be used to measure the deviation in the latent space. For example, the density estimation function can include probability density estimation, rescaled histogram, parametric density estimation function, nonparametric density estimation function (e.g., kernel density function), or data clustering method (e.g., vector quantization). The anomaly detection algorithm may include an unsupervised machine learning algorithm, a semi-supervised machine learning algorithm, or a supervised machine learning algorithm, and the anomaly detection algorithm may be trained on profiles in a latent space of multiple cell types, such as diseased cell types (e.g., cancer cells such as pancreatic cancer cells) or non-diseased cell types (e.g., pancreatic cells such as pancreatic ductal cells or pancreatic acinar cells).For example, the anomaly detection algorithm may include one or more of density-based methods (k-nearest neighbors, local outlier factors, isolation forests), subspace-based outlier detection, correlation-based outlier detection, tensor-based outlier detection, support vector machines (SVMs), single-class vector machines, support vector data description, neural networks (e.g., replicator neural networks, autoencoders, long-short-term memory (LSTM) neural networks), Bayesian networks, hidden Markov models (HMMs), cluster analysis-based outlier detection, association rules and frequent itemset deviations, fuzzy logic-based outlier detection, and ensemble methods (e.g., using feature bagging, score normalization, and different diversity sources). Diseased and normal cells may include, for example, primary cell lines, human organoids, and animal models. For example, the multiple cell types may include pancreatic ductal cells, pancreatic acinar cells, pancreatic adenocarcinomas, and / or pancreatic adenocarcinomas. After genome editing tools are used to measure the amount of shift in cellular potential space that results from editing each genomic region, one or more genes can be ranked for therapeutic targeting based on the measured amount.
[0069] In another aspect, the present disclosure provides a system for identifying one or more genomic regions that promote the reprogramming of cells from one phenotypic state to another. The system may include a database containing single-cell RNA sequencing data (e.g., of a plurality of diseased cells and a plurality of normal cells of a cell type). The database may be stored locally (e.g., on a local server, computer, or computer media) or remotely (e.g., on a cloud-based server). The system may further include one or more computer processors individually or collectively programmed to perform the methods of the present disclosure. For example, the computer processor may be individually or collectively programmed to do one or more of: map (e.g., using a UMAP algorithm or a supervised dimensionality reduction algorithm) a database including single-cell RNA sequencing (scRNA-seq) data for a plurality of diseased cells and a plurality of normal cells into a latent space corresponding to a plurality of phenotypic states of the cell type; identify, based at least in part on the topology of the latent space, one or more genomic regions that promote reprogramming of the cell type between a first phenotypic state and a second phenotypic state of the plurality of phenotypic states (the one or more genomic regions are configured to be edited to promote reprogramming of the cell type between the first phenotypic state and the second phenotypic state); and / or electronically output the one or more genomic regions.
[0070] Computer Systems The present disclosure provides a computer system programmed to implement the methods of the present disclosure. Figure 2 illustrates a computer system (201) that can, for example, generate or analyze scRNA-seq data, map the scRNA-seq data (e.g., using a dimensionality reduction algorithm such as UMAP) to a latent space corresponding to a plurality of phenotypic states, identify one or more genomic regions that promote reprogramming of a cell type between a first phenotypic state and a second phenotypic state (e.g., using probabilistic inference), train a supervised algorithm (e.g., supervised UMAP) on the scRNA-seq data, perform nonlinear reconstruction of cell trajectories in the latent space, remove low-frequency genomic regions from the scRNA-seq data, and identify one or more genomic regions using genome editing tools to promote reprogramming of cells between the first phenotypic state and the second phenotypic state. measuring the amount of deviation in the cellular latent space resulting from editing the genomic region using the genome editing means using an anomaly detection algorithm; training with profiles of the latent spaces of a plurality of cell types; measuring the distance of deviation in the cellular latent space resulting from editing the genomic region using the genome editing means; ranking genes for therapeutic targeting based on the measured amount of cellular latent space; using a density estimation function to measure the amount of deviation in the cellular latent space resulting from editing the genomic region using the genome editing means; and identifying therapeutic targets for treating a disease associated with a phenotypic condition.
[0071] The computer system (201) may coordinate various aspects of the methods and systems of the present disclosure, including, for example, generating or analyzing scRNA-seq data, mapping the scRNA-seq data (e.g., using a dimensionality reduction algorithm such as UMAP) to a latent space corresponding to multiple phenotypic states, identifying one or more genomic regions that promote reprogramming of a cell type between a first phenotypic state and a second phenotypic state (e.g., using probabilistic inference), training a supervised algorithm (e.g., supervised UMAP) on the scRNA-seq data, performing nonlinear cell trajectory reconstructions in the latent space, removing low frequency genomic regions from the scRNA-seq data, and identifying the genomic regions of cells between the first and second phenotypic states. The method includes editing a genomic region using a genome editing means to promote reprogramming; measuring the amount of deviation in a cellular latent space resulting from editing the genomic region using the genome editing means using an anomaly detection algorithm; training with profiles of the latent spaces of multiple cell types; measuring the distance of the deviation in the cellular latent space resulting from editing the genomic region using the genome editing means; ranking genes for therapeutic targeting based on the measured amount of cellular latent space; using a density estimation function to measure the amount of deviation in the cellular latent space resulting from editing the genomic region using the genome editing means; and identifying therapeutic targets to treat a disease associated with the phenotypic condition.
[0072] The computer system (201) may be a user's or computer system's electronic device, or may be a computer system located remotely relative to the electronic device. The electronic device may be a portable electronic device. The computer system (201) includes a central processing unit (CPU, also referred to herein as a "processor" and "computer processor") (205), which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system (201) also includes memory or memory locations (210) (e.g., random access memory, read-only memory, flash memory), electronic storage (215) (e.g., a hard disk), a communication interface (220) (e.g., a network adapter) for communicating with one or more other systems, and peripheral devices (225), such as cache, other memory, data storage, and / or an electronic display adapter. The memory (210), storage (215), interface (220), and peripheral devices (225) communicate with the CPU (205) via a communication bus (solid lines), such as a motherboard. The storage device (215) may be a data storage device (or data repository) for storing data. The computer system (201) may be operatively connected to a computer network ("network") (230) with the aid of the communication interface (220). The network (230) may be the Internet, an Internet and / or extranet, an intranet and / or extranet in communication with the Internet, or the like. The network (230) may, in some cases, be a telecommunications and / or data network. The network (230) may include one or more computer servers, which may enable distributed computing, such as cloud computing. The network (230), in some cases with the aid of the computer system (201), may implement a peer-to-peer network, which may enable devices connected to the computer system (201) to act as clients or servers.
[0073] The CPU (205) may execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory (210). The instructions may be directed to the CPU (205), which may then program or configure the CPU (205) to implement the methods of the present disclosure. Examples of operations performed by the CPU (205) may include fetch, decode, execute, and writeback.
[0074] The CPU 205 may be part of a circuit, such as an integrated circuit. One or more other components of the system 201 may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0075] The storage device (215) may store files such as drivers, libraries, and saved programs. The storage device (215) may store user data, such as user preferences and user programs. The computer system (201) may optionally include one or more additional data storage devices external to the computer system (201), such as located on a remote server in communication with the computer system (201) via an intranet or the Internet.
[0076] The computer system 201 may communicate with one or more remote computer systems via the network 230. For example, the computer 201 may communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user may access the computer system 201 via the network 230.
[0077] Methods as described herein may be implemented by machine (e.g., computer processor) executable code stored in an electronic storage location of the computer system (201), such as on memory (210) or electronic storage (215). Machine-executable or machine-readable code may be embodied in the form of software. During use, the code may be executed by the processor (205). In some cases, the code may be retrieved from storage (215) and stored in memory (210) for ready access by the processor (205). In some circumstances, electronic storage (215) may be omitted, and machine-executable instructions may be stored in memory (210).
[0078] The code may be pre-compiled and configured for use on a machine having a suitable processor to execute the code, or may be compiled at run-time. The code may be provided in a programming language that may be selected to enable the code to be executed in a pre-compiled or as-compiled manner.
[0079] Aspects of the systems and methods provided herein, such as the computer system (201), may be embodied in programming. Various aspects of the technology may be considered "products" or "articles of manufacture," typically in the form of machine- (or processor-) executable code and / or associated data carried or embodied on a type of machine-readable medium. The machine-executable code may be stored in electronic storage devices, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage"-type media may include any or all of the tangible memory of a computer or processor, such as various semiconductor memories, tape drives, disk drives, etc., or their associated modules, which may provide non-transitory storage for software programming at any time. All or portions of the software may sometimes be communicated via the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, such as from a management server or host computer to an application server computer platform. Thus, other types of media that may bear software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, via wired and optical landline networks, and through various air-links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered media bearing software. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to media that participate in providing instructions to a processor for execution.
[0080] Thus, machine-readable media such as computer-executable code may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices in a computer or the like, such as those that may be used to implement the databases, etc., shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus within a computer system. Carrier wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, other magnetic media, CD-ROMs, DVDs or DVD-ROMs, other optical media, punch cards, paper tape, other physical storage media with patterns of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, other memory chips or cartridges, carrier waves that carry data or instructions, cables or links that carry such carrier waves, or other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0081] The computer system (201) may include or be in communication with an electronic display (235) and may include a user interface (UI) (240) for providing, for example, scRNA-seq data, mapping, or user selection of other algorithms and databases. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0082] The disclosed methods and systems may be implemented by one or more algorithms. The algorithms may be implemented by software when executed by a central processing unit (205). The algorithms may, for example, generate or analyze scRNA-seq data, map the scRNA-seq data (e.g., using a dimensionality reduction algorithm such as UMAP) to a latent space corresponding to multiple phenotypic states, identify genomic regions that promote cell-type reprogramming between a first phenotypic state and a second phenotypic state (e.g., using probabilistic inference), train a supervised algorithm (e.g., supervised UMAP) on the scRNA-seq data, perform nonlinear cell trajectory reconstruction in the latent space, remove low-frequency genomic regions from the scRNA-seq data, and perform genome editing procedures to promote cell reprogramming between the first phenotypic state and the second phenotypic state. editing a genomic region using a stage; measuring the deviation in a cellular latent space resulting from editing the genomic region using a genome editing tool using an anomaly detection algorithm; training with profiles of the latent spaces of a plurality of cell types; measuring the distance of the deviation in the cellular latent space resulting from editing the genomic region using a genome editing tool; ranking genes for therapeutic targeting based on the measured deviation in the cellular latent space; using a density estimation function to measure the deviation in the cellular latent space resulting from editing the genomic region using a genome editing tool; and identifying therapeutic targets to treat a disease associated with the phenotypic condition. [Example]
[0083] Example 1 - Generation and preprocessing of scRNA-seq data Single-cell RNA sequencing (scRNA-seq) data are generated by isolating and culturing multiple types of normal and tumor pancreatic cells from mice, including highly malignant cancer cells (Kras G12D , p53 - / - , Myc), low-grade cancer cells (Kras G12D , p53- / - These cell lines include immortalized normal beta cells, duct cells, and acinar cells. These cell lines are further genetically modified to stably express catalytically inactive Cas9 (dCas9) fused to the transcriptional repressor peptide KRAB, allowing CRISPR interference (CRISPRi) to inactivate genes of interest. For scRNA-seq, each cell type is isolated as a single cell, and then corresponding RNA and cDNA libraries are prepared according to the manufacturer's instructions (10X Genomics). The cDNA libraries are sequenced using Miseq (Illumina) to obtain cell number information, and then sequenced using NextSeq or Hiseq4000 (Illumina) to obtain scRNA-seq data.
[0084] Preprocessing of scRNA-seq count data is performed as follows: The raw, HGNC-aligned UMI count matrix generated by 10X sequencing is preprocessed and scaled before analysis in the downstream analysis pipeline. Low-abundance genes (e.g., average counts <0.1), genes with read counts in <10% of cells, and cells with non-zero read counts in <10% of all genes are removed from the count matrix (e.g., using the SingleCellExperiment, scran, and scatter libraries in R). To adjust for differences in sequencing depth between individual cells, the count matrix is optionally normalized and scaled before carrying forward to the next analysis. Example normalization methods include globally scaling cell-level counts to the median depth across all cells (scalar adjustment) or solving a linear system to obtain specific scaling coefficients for individual cells (e.g., using ComputeSumFactors in the scran library in R). In some cases, batch effects between samples are corrected by a mutual nearest neighbor algorithm (MNN, for example, using the mnnCorrect function in the scran library in R).
[0085] Example 2 - Latent Space Construction The latent space is constructed as follows: A supervised machine learning algorithm is used to map the high-dimensional single-cell count matrix into a 20-100-dimensional latent space. In the case of pancreatic cancer, the reduction algorithm is trained on a set of pure cell types, including pancreatic acinar cells, pancreatic ductal cells, and pancreatic adenocarcinoma cells. To model potential toxic complications that may arise from the target candidates of interest, cells targeted with essential genes (e.g., PCNA or RPA3) are also included during training of the latent space. Supervised learning labels are selected to correspond to each pure cell type.
[0086] Several algorithms are considered for constructing the latent space, including, but not limited to, uniform manifold approximation and projection (UMAP) and variational autoencoder (VAE). In some cases, the Elbow method (e.g., as described in Richards et al., J Shoulder Elbow Surg 8(4): 351-354 (1999), the entire contents of which are incorporated herein by reference) was used to determine the optimal dimensionality of the latent space. For UMAP, the following parameters were used for model training: a minimum distance of 0.025-0.25, a number of neighbors of 75% of the total number of cells, and Euclidean distance as the distance metric.
[0087] Using UMAP manifold learning (Python), nonlinear cell trajectory reconstruction was performed using the Monocle3 package® on a dimensionality-reduced projection of the labeled population data, with an initial pseudo-time ordering of the cell population defining the transition states between normal and disease endpoints. A dimensionality-reduced principal tree defining the progression trajectory of maximum likelihood inference was constructed through the resulting ordering by implementing a graph back-embedding algorithm in the DDRTree and Genie3 packages.
[0088] We then performed probabilistic inference of driver genes through causal gene regulatory network interference using Moran's spatial autocorrelation test with Genie3 and Scribe packages (registered trademark), performed exploratory metrics on the interactions of candidate drivers that influence each other on the inferred trajectories using the Louvain community detection algorithm, and performed statistical tests to increase the robustness of causal inference by implementing the R of Kendall Tau Correlation and Granger Causality Test.
[0089] Target genes identified by the pooled CRISPRi library were quantified for their ability to reprogram cancer cells back to a wild-type-like expression state. Genes were scored using one of several algorithms, including but not limited to anomaly detection, density estimation, and pairwise Euclidean distance against pure cell populations in latent space.
[0090] Anomaly detection was performed as follows. Separate single-class anomaly detectors were trained on the latent expression profiles of different cell types, including, but not limited to, pancreatic ductal cells, pancreatic acinar cells, pancreatic adenocarcinoma, and pancreatic adenocarcinoma with essential genes (e.g., PCNA or RPA3) targeted (by CRISPR / RNAi) as a toxicity model of healthy tissue. Several algorithms were utilized for anomaly detection, including, but not limited to, support vector machines (SVMs), isolation forests, and support vector data description (SVDD). Each trained anomaly detector model was then used to score candidate genes based on the output of a decision function (e.g., distance to a discrete manifold, such as Chebyshev distance, correlation distance, cosine distance, Euclidean distance, signed Euclidean distance, Hamming distance, Jaccard distance, Kullback-Leibler distance, Mahalanobis distance, Manhattan distance, Minkowski distance, or Spearman distance) applied to the latent expression profiles of single cells targeted with the CRISPRi library.
[0091] Density estimation was performed as follows. Separate density estimators were trained on the latent expression profiles of different cell types, including, but not limited to, pancreatic ductal cells, pancreatic acinar cells, pancreatic adenocarcinoma, and pancreatic adenocarcinoma with essential genes (e.g., PCNA or RPA3) targeted (by CRISPR / RNAi) as a toxicity model of healthy tissue. Several algorithms were used to estimate the density function for each cell type, including, but not limited to, ball tree or KD tree-based estimators and neural network-based approaches (e.g., neural autoregressive flow, as described in Huang et al., "Neural autoregressive flows," arXiv:1804.00779, incorporated herein by reference in its entirety). For the tree-based estimators, one of several kernels was used to train the density function, including, but not limited to, 1) Gaussian, 2) top-hat, 3) uniform, and 4) Epanechnikov. Density estimators trained for each pure cell type were then used to score the potential expression profiles of single cells targeted with the CRISPRi library.
[0092] Five-fold cross-validation was performed to train and evaluate the models. For quantification of reprogramming, cell populations for each target gene were repeatedly sampled with replacement (25–100×) to construct bootstrap confidence intervals for each target gene.
[0093] To determine optimal target candidates, the output of the decision function of the anomaly detector or density estimator (e.g., distance to a discrete manifold, such as Chebyshev distance, correlation distance, cosine distance, Euclidean distance, signed Euclidean distance, Hamming distance, Jaccard distance, Kullback-Leibler distance, Mahalanobis distance, Manhattan distance, Minkowski distance, or Spearman distance) was summarized in several ways. These included, but were not limited to, the average decision function across all cells in which the specific target gene of interest was inactivated, the effect size of the decision function for the specific target gene relative to a non-targeting guide RNA or other control population of interest, and the Bonferroni-corrected p-value of a Kolmogorov-Smirnov test of the decision function for the specific target gene relative to a non-targeting guide RNA or other control population of interest. In some cases, the summary metrics were z-transformed across all target genes. Additionally, the summary metrics for each of the anomaly detectors (e.g., primary cells and cancers with negative guides) were aggregated (e.g., average, Stouffer's method, or Fisher's method). Top scoring targets that also meet the p-value threshold of the Kolmogorov-Smirnov test of the decision function against the negative control population are considered as top reprogramming genes and carried forward for further biovalidation.
[0094] Example 3 - Computational pipeline for quantifying transitions between cellular states and identifying therapeutic target genes Figures 3A-3C illustrate a computational framework for identifying potential target genes mediating transcriptional transitions between differentiated or phenotypically distinct cell states after gene matching. Single-cell transcriptomes corresponding to the disease and target population of interest were isolated and sequenced. Representative latent spaces were generated by supervised dimensionality reduction (e.g., UMAP or VAE) in different cell populations (as shown in Figures 3A and 4A-4D, Example 4). Target genes of interest were then identified by pseudo-temporal ordering and single-cell trajectory analysis from the disease state to the target state of interest. Gene candidates (~100) were then matched by transduction with lentiviruses carrying a pooled CRISPR interference (CRISPRi) library targeting the candidates. Genes with the most extensive reprogramming to the target state of interest were carried forward for further biovalidation.
[0095] 3A shows an example of quantifying and identifying novel therapeutic target gene reprogramming according to disclosed embodiments. By leveraging CRISPR gene matching, intelligent construction of latent space, and anomaly detection, target genes are quantified according to their ability to program disease cell populations toward a desired target phenotypic state. Target states can be derived from healthy tissue or primary cell lines, or represent intermediate states, including, but not limited to, senescent cells, stem cells, precancerous cells, or progenitor cells associated with disease progression.
[0096] As shown in Figure 3B-C, anomaly detection (e.g., density-based methods (e.g., k-nearest neighbors, local outlier factors, and isolation forests), subspace-based outlier detection, correlation-based outlier detection, tensor-based outlier detection, support vector machines (SVMs), single-class vector machines, support vector data description, neural networks (e.g., replicator neural networks, autoencoders, and long-short-term memory (LSTM) neural networks), Bayesian networks, hidden Markov models (HMMs), cluster analysis-based outlier detection, association rules and frequent itemset mismatch, fuzzy logic-based outlier detection, and ensemble methods (e.g., feature bagging, score normalization, and using different sources of diversity)) were used to quantitate the degree of transcriptional reprogramming to the target state.
[0097] FIG. 3B illustrates an example of anomaly detection as a method for demarcating a dense manifold that accurately represents the phase space occupied by different cell populations, according to an embodiment of the disclosure.
[0098] Figure 3C depicts an example of this in the context of pancreatic cancer. Briefly, separate anomaly detection models were trained to generate representative manifolds describing the following differentiated cell populations: pancreatic ductal cells (positive control for reprogramming), pancreatic acinar cells (positive control for reprogramming), Kras mutant pancreatic cancer cells expressing a non-targeting guide RNA (negative control for reprogramming), and Kras mutant pancreatic cancer cells expressing an essential gene to be targeted (positive control for toxic target).
[0099] As shown in Figure 3C, the trained aberration detection model was then applied to score single-cell RNA-seq profiles of Kras mutant pancreatic cancer cells targeted with the CRISPRi library. For each target gene of interest, the degree of transcriptional transition was quantified using the aberration detector model's decision function. In practice, the best targets exhibited larger decision functions and effect sizes compared to the negative control and smaller decision functions and effect sizes compared to the positive reprogramming control.
[0100] Furthermore, as shown in Figure 3C, by utilizing multiple anomaly detectors, we can identify genes that maximally reprogram cancer cells to a wild-type primary expression profile based on the trained model's decision function, where we also included apoptotic cells to model potential toxic complications that may occur in healthy cells derived from the target of interest (the "toxic" cluster label).
[0101] Example 4 - UMAP as a supervised algorithm for achieving separability while preserving fine-scale local structure in single-cell RNA-seq data Figures 4A-4D demonstrate the potential of UMAP for quantifying transitions in single-cell RNA-seq data. Figure 4A shows a comparison of several dimensionality reduction algorithms, including principal component analysis (PCA), t-SNE, and uniform manifold approximation and projection (UMAP), applied to a cell-type mixture dataset. As shown in Figure 4A, supervised dimensionality reduction can be applied to cell-type mixture datasets, with cell types serving as supervised labels in model training. Unlike PCA and t-SNE, UMAP achieved excellent separability while preserving the fine-grained local structure of single-cell data.
[0102] Figure 4B shows a comparison of the stability of the latent spaces constructed by t-SNE and UMAP trained on pancreatic ductal, acinar, and adenocarcinoma cell lines. This conceptually illustrates that UMAP achieves greater stability in constructing supervised latent spaces. Under invariant random conditions, the latent space generated from the 20% mixture of pancreatic ductal, acinar, and adenocarcinoma cell lines more closely aligns with the latent space of the entire dataset in the UMAP algorithm.
[0103] Figure 4C shows an example of the effect of the UMAP "minimum distance" parameter on quantifying pancreatic cancer reprogramming, and Figure 4D shows an example of the effect of the dimensionality of the UMAP latent space on quantifying pancreatic cancer reprogramming. Figures 4C-4D show the effect of UMAP hyperparameters on quantifying pancreatic cancer cell reprogramming. The strong correlation of decision scores across a range of conditions indicates that reprogramming quantification is robust across a reasonable range of UMAP hyperparameters.
[0104] Example 5 - Identification of potential therapeutic targets from causal inference based on pseudo-temporal principal trees constructed from high-dimensional single-cell RNA-seq data Figure 5A shows the progression of pancreatic acinar cells (dark shading on the right) to pancreatic ductal cells (medium shading in the middle), and then to high-grade cancer cells (Kras G12D , p53 - / - The black curve represents a 2D projection of the complete pseudo-temporal ordering formed by the candidate selection pipeline, characterized by transitions to the primates (Myc, Myc) (lightly shaded on the left). For demonstration purposes, this is the resulting 2D projection that preserves maximal separation between these cell populations. The black curve represents a 2D projection of the principal trajectory tree learned using the DDRTree algorithm.
[0105] Figure 5B shows an example pipeline for generating candidates from causal inference based on a pseudo-temporal principal tree constructed from high-dimensional single-cell RNA-seq data. The initial target candidate selection pipeline uses graph inverse embedding to learn a clear principal graph from the single-cell data to order cells, thereby robustly and accurately resolving the complex biological processes of these three cell types. Each cell can be viewed as a point in a high-dimensional space, with each dimension describing the expression of a different gene in the genome. Identifying the program of gene expression changes is equivalent to learning the trajectory that the cell follows across that space, which can then be used to study key gene expression changes through transitions for candidate selection.
[0106] In the preprocessing stage of the pipeline, we used UMAP manifold learning to reduce the dimensionality and noise of the scRNA-seq data and remove outliers (often >90%) based on low expression levels. We then used the community detection capabilities of Monocle3 and the Louvain algorithm to label strongly connected components in the resulting low-dimensional representation of the data, which served as the basis for pseudo-temporal ordering using Monocle3 and the DDRTree algorithm. This approach oriented the samples in a semi-supervised manner using only the labels of the "root" endpoints, labeled as cell populations representing the transitional health and disease states of interest. DDRTree then learned a spanning graph connecting the endpoints through the center of the estimated underlying point distribution for each cell population across the pseudo-temporal phase. We used Moran's spatial autocorrelation test to highlight genes driving the transitions, facilitating the selection of "significant" causal genes by assigning scores to genes based on their influence, not just on the endpoints, but on the overall sample point expression marking across the entire transition phase of the interpolation, as inferred by the learned principal tree. This test generated a rank order of genes driving the transitions of interest, and Kendall correlation and Granger causality tests were performed to filter outliers and validate strong candidates.
[0107] Example 6 - Quantification of reprogramming of latent UMAP space is robust across different anomaly detection algorithms Figures 6A-6C illustrate how top reprogramming genes can be identified regardless of the anomaly detection algorithm chosen. Figure 6A shows diagrams of two anomaly detection algorithms trained on two half-moon scatter plots with random Gaussian noise. This allows visualization of the boundaries of the trained decision manifolds for two different anomaly detection models (single-class support vector machine and isolation forest) trained on half-moon scatter data using Gaussian noise. Figure 6B shows a heat map of the z-transformed decision functions of ~70 target sgRNAs across several anomaly detection models relevant to pancreatic cancer. Here, wild-type cells correspond to the positive reprogramming control, while the negative guide column corresponds to the negative reprogramming control. This heat map illustrates how anomaly detectors across multiple cell populations can be used to identify target genes most likely to reprogram from a differentiated or distinct cell state to another differentiated or distinct cell state with the least toxicity. Targets are ranked according to their average rank across the five anomaly detector models trained. Three of the 70 sgRNAs resulted in adjusted p-values that aligned with significant reprogramming to the primary cell state (( 32 ), ( 52 ), and ( 38 )).
[0108] Figure 6C shows several different anomaly detection algorithms, including density-based methods (k-nearest neighbors, local outlier factors, isolation forests), subspace-based outlier detection, correlation-based outlier detection, tensor-based outlier detection, support vector machines (SVMs), single class vector machines, support vector data description, neural networks (e.g., replicator neural networks, autoencoders, long-short-term memory (LSTM) neural networks), Bayesian networks, hidden Markov models (HMMs), cluster analysis-based outlier detection, association rules and frequent itemset deviations, fuzzy logic-based outlier detection, and ensemble methods. We demonstrate that one of the top reprogramming targets in the UMAP latent space can be identified using one of the following methods (e.g., using feature bagging, score normalization, and different diversity sources). For both algorithms, the same three targets ((32), (52), and (58)) exhibited significant reprogramming, separate from the negative control, as measured by the adjusted p-value of the Kolmogorov-Smirnov test. For these three targets, the decision function also met or exceeded the 90th percentile (z-score = 1.645) across all cell population detectors. Eight of the top 10 targets (80%) were shared across both models.
[0109] Example 7 - Quantifying reprogramming and identifying target genes is robust across different cell types during disease progression Using the disclosed methods and systems, we used anomaly detectors across multiple cell populations to identify target genes most likely to reprogram high-grade to low-grade cancer cells. In particular, we developed a method to reprogram triple-mutated pancreatic cancer cells to wild-type ductal or acinar cells.
[0110] Figures 7A-7G illustrate how to identify top target genes for reprogramming high-grade cancer cells into cells at different stages in cancer progression. Figure 7A shows diagrams of pancreatic cancer progression and the corresponding cells used for reprogramming analysis across different stages of cancer development. Primary pancreatic ductal cells and immortalized acinar cells were used as wild-type cells. Pancreatic cancer cells carrying a double mutation (Kras G12D , p53 - / - ) were used as highly malignant cancer cells. G12D , p53 - / - , Myc) were used.
[0111] Figures 7B and 7C show the results of the triple-mutant pancreatic cancer cells (Kras G12D , p53 - / - , Myc) into wild-type ductal or acinar cells (Figure 7B), and a heatmap of the z-transformed decision function for anomaly detection across 70 single-guide RNAs (Figure 7C). Targets are ordered according to their average rank across the five anomaly detector models trained.
[0112] Figure 7C shows a heat map of the z-transformed decision functions of approximately 70 target sgRNAs across several aberration detection models relevant to pancreatic cancer. Here, wild-type cells (ductal or acinar cells) correspond to the positive control for reprogramming, while the negative guide sequence corresponds to the negative control for reprogramming (triple-mutant pancreatic cancer cells harboring negative sgRNAs). This heat map illustrates how anomaly detectors across multiple cell populations can be used to identify target genes most likely to reprogram aggressive cancer cells to wild-type cells.
[0113] Figures 7D and 7E show the results of the triple-mutant pancreatic cancer cells (Kras G12D , p53 - / - , Myc) in double-mutant pancreatic cancer cells (Kras G12D , p53 - / -), and a heatmap of the z-transformed anomaly detection decision functions for 70 single-guide RNAs, with targets ordered by their average rank across the five trained anomaly detector models (Figure 7E).
[0114] Figure 7E shows a heat map of the z-transformed decision functions of approximately 70 target sgRNAs across several aberration detection models relevant to pancreatic cancer. Here, low-grade cancer cells (ductal or acinar cells) correspond to the positive control for reprogramming, while the negative guide sequence corresponds to the negative control for reprogramming (triple-mutant pancreatic cancer cells harboring negative sgRNAs). This heat map illustrates how anomaly detectors across multiple cell populations can be used to identify target genes most likely to reprogram high-grade cancer cells to low-grade cancer cells.
[0115] Figures 7F to 7G show the results of the triple-mutant pancreatic cancer cells (Kras G12D , p53 - / - Figure 7F shows an analysis of the reprogramming of 70 single-guide RNAs (Myc) into wild-type ductal or acinar cells, with double-mutant pancreatic cancer cells as an intermediate cell type, and a heatmap of the z-transformed anomaly detection decision function for 70 single-guide RNAs (Figure 7G), where targets are ordered according to their average rank across the five trained anomaly detector models. That is, the reprogramming includes a first reprogramming (triple-mutant pancreatic cancer cells into double-mutant pancreatic cancer cells) and a second reprogramming (double-mutant pancreatic cancer cells into wild-type ductal or acinar cells).
[0116] Figure 7G shows a heat map of the z-transformed decision functions of approximately 70 target sgRNAs across several anomaly detection models relevant to pancreatic cancer. Here, wild-type cells (ductal or acinar cells) correspond to the positive control for reprogramming, while the negative guide column corresponds to the negative control for reprogramming (triple-mutant pancreatic cancer cells carrying negative sgRNAs). Additionally, to improve accuracy and robustness, low-grade cancer cells (double-mutant pancreatic cancer cells) were also considered in the reprogramming analysis. This heat map illustrates how using anomaly detectors across multiple cell populations and considering intermediate cell types in the reprogramming analysis identifies target genes with the highest likelihood and highest confidence for reprogramming from high-grade cancer cells to wild-type cells. Notably, adding intermediate cell types in the reprogramming analysis yields clearer and more robust heat map data, providing better guidance (and therefore improved accuracy) for determining the desired target cell state for reprogramming.
[0117] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are offered by way of example only. The present invention is not intended to be limited by the specific examples provided within the specification. While the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it will be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the various embodiments of the present invention described herein may be utilized in practicing the invention. It is therefore contemplated that the present invention shall cover any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. A system for identifying one or more genomic regions that, upon editing of a cell type, cause reprogramming of cells of said cell type from one phenotypic state to another phenotypic state, said system comprising: a database containing single-cell ribonucleic acid (RNA) sequence data of a plurality of diseased cells and a plurality of normal cells of the same cell type as the cell; one or more computer processors, said one or more computer processors comprising: (i) mapping the single-cell RNA-seq data of the plurality of diseased cells and the plurality of normal cells into a latent space corresponding to a plurality of phenotypic states of the cell type, wherein the mapping comprises using a dimensionality reduction algorithm; (ii) constructing a maximum likelihood progression trajectory from a first phenotypic state to a second phenotypic state of the plurality of phenotypic states, wherein said constructing comprises performing a non-linear cell trajectory reconstruction on the latent space; (iii) identifying the one or more genomic regions that, upon editing of the cell type, cause reprogramming of the cell type from the first phenotypic state to the second phenotypic state of the plurality of phenotypic states based at least in part on the topology of the latent space, wherein said identifying comprises applying probabilistic inference to the maximum likelihood inferred progression trajectory; (iv) electronically outputting the one or more genomic regions; one or more computer processors, individually or collectively programmed to: Including, the system.
2. The system described in claim 1, wherein the dimensionality reduction algorithm includes a uniform manifold approximation and projection (UMAP) algorithm.
3. The system described in claim 2, wherein the UMAP algorithm is a supervised UMAP algorithm.
4. The system described in claim 3, wherein the supervised UMAP algorithm is trained on single-cell RNA sequencing data of pure cells of the cell type.
5. The system described in claim 3, wherein the UMAP algorithm is trained using a minimum distance of approximately 0.025 to 0.
25.
6. The system described in claim 1, wherein reconstructing the nonlinear cell trajectories includes applying a graph inverse embedding algorithm to the latent space.
7. The system described in claim 1, wherein the first phenotypic state is cancer and the second phenotypic state is a wild-type state.
8. The system of claim 1, wherein the one or more computer processors are further programmed individually or collectively to remove low frequency genomic regions from the single-cell RNA sequencing data of the plurality of diseased cells and the plurality of normal cells prior to said mapping.
9. The system described in claim 1, further comprising a genome editing means configured to edit in each of the one or more genomic regions within the cell so as to cause the reprogramming of the cell of the cell type between the first phenotypic state and the second phenotypic state.
10. The system described in claim 9, wherein the genome editing means is selected from the group consisting of a CRISPR system, a CRISPRi system, a CRISPRa system, an RNAi system, and an shRNA system.
11. The one or more computer processors map single-cell ribonucleic acid (RNA) sequence data of the cells onto the latent space; and Measuring the amount of deviation in the latent space of the cell resulting from editing each of the genome regions using the genome editing means using an anomaly detection algorithm; 10. The system of claim 9, further programmed, individually or collectively, to:
12. The system described in claim 11, wherein the anomaly detection algorithm is trained on profiles of the latent space of multiple cell types.
13. The system of claim 12, wherein the multiple cell types include pancreatic ductal cells, pancreatic acinar cells and / or pancreatic adenocarcinoma.
14. The system of claim 11, wherein the anomaly detection algorithms include one or more of density-based techniques, subspace-based outlier detection, correlation-based outlier detection, tensor-based outlier detection, support vector machines (SVMs), single class vector machines, support vector data description, neural networks, Bayesian networks, hidden Markov models (HMMs), cluster analysis-based outlier detection, association rules and frequent itemset deviations, fuzzy logic-based outlier detection, and ensemble techniques.
15. The system of claim 14, wherein the anomaly detection algorithm is a support vector machine (SVM), a density-based method, a k-nearest neighbor algorithm, a local outlier factor algorithm, or an isolation forest algorithm.
16. The one or more computer processors mapping single-cell ribonucleic acid (RNA) sequence data of the cells onto the latent space; and Measuring the Euclidean distance of the deviation in the latent space of the cell resulting from editing each of the genome regions using the genome editing means; 10. The system of claim 9, further programmed, individually or collectively, to:
17. The method of claim 1, wherein the genome editing means is configured to edit each of the one or more genomic regions in a plurality of cells to cause the reprogramming of each cell of the cell type between the first phenotypic state and the second phenotypic state; the one or more computer processors: mapping single-cell ribonucleic acid (RNA) sequence data of the plurality of cells onto the latent space; Measuring the amount of deviation in the latent space of each of the plurality of cells resulting from editing the respective genome regions using the genome editing means; and using the measured amount of deviation to rank order the one or more genomic regions for therapeutic targeting; 10. The system of claim 9, wherein the system is individually or collectively programmed to:
18. The one or more computer processors mapping single-cell ribonucleic acid (RNA) sequence data of the cells onto the latent space; and Measuring, using a density estimation function, the amount of deviation in the latent space of the cell that occurs as a result of editing each of the genome regions using the genome editing means; 10. The system of claim 9, further programmed, individually or collectively, to:
19. The system of claim 1, wherein the cell type is a pancreatic cell.
20. The system described in claim 1, wherein the diseased cells are cancer cells.
21. The system described in claim 1, wherein the plurality of diseased cells and the plurality of normal cells are selected from the group consisting of primary cell lines, human organoids, and animal models.
22. The system of claim 1, wherein the one or more computer processors are further programmed individually or collectively to generate the single-cell RNA sequencing data for a plurality of diseased cells and a plurality of normal cells of a cell type.
23. The system described in claim 1, wherein the second phenotypic state is an intermediate state.
24. The system described in claim 23, wherein the intermediate state is a precancerous state or a low-grade malignant state.
25. The system described in claim 1, wherein the one or more computer processors are further programmed individually or collectively to identify one or more therapeutic targets based on the one or more genomic regions for treating a disease associated with the first phenotypic condition.
26. The method of claim 26, wherein the one or more computer processors identify, based at least in part on the topology of the latent space, one or more first genomic regions that, upon editing of the cell type, cause reprogramming of the cell type between the first phenotypic state and an intermediate phenotypic state of the plurality of phenotypic states; and wherein, upon editing, the one or more first genomic regions cause reprogramming of the cell type between the first phenotypic state and the intermediate phenotypic state; and identifying, based at least in part on the topology of the latent space, one or more second genomic regions that, upon editing of the cell type, cause reprogramming of the cell type between the intermediate phenotypic state and the second phenotypic state of the plurality of phenotypic states; and wherein, upon editing, the one or more second genomic regions cause reprogramming of the cell type between the intermediate phenotypic state and the second phenotypic state.
10. The system of claim 1, further programmed individually or collectively to: