Method for classifying cell labels
The method for identifying signal cell labels through probabilistic barcoding and threshold-based analysis addresses errors in cell count estimation, enhancing the accuracy of cell analysis by distinguishing between signal and noise cell labels.
Patent Information
- Application Number
- JP2023020218
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-01-12
- Filing Date
- 2023-02-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2037-11-07
AI Technical Summary
Existing methods for cell analysis, such as probabilistic barcoding, can introduce errors leading to an overestimation of cell counts due to amplification and sequencing biases.
A method for identifying signal cell labels involves barcoding cells with probabilistic barcodes, obtaining sequence data, determining the number of distinct molecular labels for each cell label, generating cumulative sum and second derivative plots, and identifying cell labels as signal or noise based on thresholds.
This method accurately distinguishes signal cell labels from noise, reducing errors in cell count estimation and improving the reliability of cell analysis.
Smart Images

Figure 0007706485000001 
Figure 0007706485000002 
Figure 0007706485000003
Abstract
Description
Technical Field
[0001] Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 419,194, filed Nov. 8, 2016, and U.S. Provisional Patent Application No. 62 / 445,546, filed Jan. 12, 2017. The entire contents of each of these related applications are hereby expressly incorporated by reference herein.
[0002] Field The present disclosure generally relates to the field of molecular barcoding, and more particularly to identifying and correcting noisy cell labels.
Background Art
[0003] Methods and techniques such as probabilistic barcoding are useful for cell analysis, particularly for deciphering gene expression profiles to identify the state of cells using, for example, reverse transcription, polymerase chain reaction (PCR) amplification, and next-generation sequencing (NGS). However, these methods and techniques, if uncorrected, can introduce errors that can lead to an overestimation of cell counts.
Summary of the Invention
Means for Solving the Problems
[0004] Disclosed herein is a method for identifying signal cell labels. In some embodiments, the method comprises: (a) using a plurality of barcodes (e.g., probabilistic barcodes) to barcode (e.g., probabilistically barcode) a plurality of targets in a sample of cells to create a plurality of barcoded targets (e.g., probabilistically barcoded targets), wherein each of the plurality of barcodes comprises a cell label and a molecular label, and creating the barcoded targets by barcoding; (b) obtaining sequence data of the plurality of barcoded targets; (c) determining the number of molecular labels having distinct sequences associated with each of the cell labels of the plurality of barcodes; (d) determining the rank of each of the cell labels of the plurality of barcodes based on the number of molecular labels having distinct sequences associated with each of the cell labels; (e) generating a cumulative sum plot based on the number of molecular labels having distinct sequences associated with each of the cell labels identified in (c) and the rank of each of the cell labels determined in (d); (f) generating a second derivative plot of the cumulative sum plot; (g) identifying the minimum of the second derivative plot of the cumulative sum plot, wherein the minimum of the second derivative plot corresponds to a cell label threshold; and (h) identifying each of the cell labels as a signal cell label or a noise cell label based on the number of molecular labels having distinct sequences associated with each of the cell labels identified in (c) and the cell label threshold determined in (g).
[0005] In some embodiments, the method further comprises, in (h), removing sequence information associated with an identified cell label from the sequence data obtained in (b) if the cell label of the plurality of barcodes is identified as a noise cell label. The method can include removing, from the sequence data obtained in (b), sequence information associated with a molecular label having a distinct sequence associated with a target among the plurality of targets if the number of molecular labels having distinct sequences associated with the target among the plurality of targets exceeds a molecular label occurrence threshold.
[0006] In some embodiments, determining the number of molecular labels having distinct sequences associated with each of the cell labels in (c) involves removing from the sequence data the sequence information associated with the non-unique molecular labels associated with each of the cell labels. The cumulative sum plot can be a log-log plot. The log-log plot can be a log 10 -log 10 plot.
[0007] In some embodiments, generating a cumulative sum plot based on the number of molecular labels having distinct sequences associated with each of the cell labels identified in (c) and the rank of each of the cell labels determined in (d) involves determining the cumulative sum for each rank of cell labels, where the cumulative sum for a rank includes the sum of the number of molecular labels having distinct sequences associated with each of the cell labels having lower ranks. Generating a second derivative plot of the cumulative sum plot can include determining that the difference between the cumulative sum for the first rank of cell labels and the cumulative sum for the second rank of cell labels exceeds the difference between the first rank and the second rank. The difference between the first rank and the second rank can be 1.
[0008] In some embodiments, the minimum is a global minimum. Determining the minimum of the second derivative plot includes determining that the minimum of the second derivative plot exceeds a threshold of the minimum number of molecular labels associated with each of the cell labels.
[0009] In some embodiments, the threshold of the minimum number of molecular labels associated with each of the cell labels is a percentile threshold. The threshold of the minimum number of molecular labels associated with each of the cell labels is determined based on the number of cells in the sample of cells.
[0010] In some embodiments, identifying the minimum of the second derivative plot includes determining that the minimum of the second derivative plot is below a threshold of the maximum number of molecular labels associated with each of the cell labels. The threshold of the maximum number of molecular labels associated with each of the cell labels can be a percentile threshold. The threshold of the maximum number of molecular labels associated with each of the cell labels can be determined based on the number of cells in the sample of cells.
[0011] In some embodiments, each of the cell labels is identified as a signal cell label if the number of molecular labels having molecular labels associated with each of the cell labels identified in (c) is greater than the cell label threshold. Each of the cell labels can be identified as a noise cell label if the number of molecular labels having molecular labels associated with each of the cell labels identified in (c) is not greater than the cell label threshold.
[0012] In some embodiments, the method includes (i) for one or more of a plurality of targets, (1) counting the number of molecular labels having distinct sequences associated with the target in the sequence data, and (2) estimating the number of targets based on the number of molecular labels having distinct sequences associated with the target in the sequence data counted in (1).
[0013] Disclosed herein is a method for identifying signal cell labels. In some embodiments, the method comprises: (a) obtaining sequence data of a plurality of barcoded targets (e.g., stochastically barcoded targets), wherein the plurality of barcoded targets are created from a plurality of targets in a sample of cells that have been barcoded (e.g., stochastically barcoded) using a plurality of barcodes (e.g., stochastic barcodes), and each of the plurality of barcodes comprises a cell label and a molecular label; (b) determining the rank of each of the cell labels of the plurality of barcoded targets (or barcodes) based on the number of molecular labels having distinct sequences associated with each of the cell labels of the plurality of barcoded targets (or barcodes); (c) determining a cell label threshold based on the number of molecular labels having distinct sequences associated with each of the cell labels and the rank of each of the cell labels of the plurality of barcoded targets (or barcodes) determined in (b); and identifying each of the cell labels as a signal cell label or a noise cell label based on the number of molecular labels having distinct sequences associated with each of the cell labels and the cell label threshold identified in (c).
[0014] In some embodiments, the method comprises determining the number of molecular labels having distinct sequences associated with each of the cell labels. Determining the number of molecular labels having distinct sequences associated with each of the cell labels can include removing from the sequence data the sequence information associated with non-unique molecular labels associated with each of the cell labels.
[0015] In some embodiments, determining a cell label threshold based on the number of molecular labels having distinct sequences associated with each of the cell labels of the plurality of barcoded targets includes identifying the cell label having the largest change in the cumulative sum of the cell labels having rank n and the cumulative sum of the cell labels having the next rank n + 1, wherein the number of molecular labels having distinct sequences associated with the cell label corresponds to the cell label threshold.
[0016] In some embodiments, determining a cell labeling threshold based on (a) the number of molecular labels having distinct sequences associated with each of the cell labels of a plurality of barcode - tagged targets and (b) the rank of each of the cell labels of the plurality of barcode - tagged targets determined in (b) comprises identifying the cumulative sum of each rank of cell labels, where the cumulative sum of the ranks includes the sum of the number of molecular labels having distinct sequences associated with each of the cell labels having lower ranks, and determining the rank n of the cell label having the largest change in the cumulative sum and the cumulative sum of the next rank n + 1, where the rank n of the cell label having the largest change in the cumulative sum and the cumulative sum of the next rank n + 1 corresponds to the cell labeling threshold.
[0017] In some embodiments, determining a cell labeling threshold based on (a) the number of molecular labels having distinct sequences associated with each of the cell labels of a plurality of barcode - tagged targets and (b) the rank of each of the cell labels of the plurality of barcode - tagged targets determined in (b) comprises generating a cumulative sum plot based on the number of molecular labels having distinct sequences associated with each of the cell labels and the rank of the cell labels determined in (b), generating a second - derivative plot of the cumulative sum plot, and identifying the minimum of the second - derivative plot of the cumulative sum plot, where the minimum of the second - derivative plot corresponds to the cell labeling threshold. Generating a cumulative sum plot based on the number of molecular labels having distinct sequences associated with each of the cell labels and the rank of the cell labels determined in (b) can include identifying the cumulative sum of each rank of cell labels, where the cumulative sum of the ranks includes the sum of the number of molecular labels having distinct sequences associated with each of the cell labels having lower ranks. Generating a second - derivative plot of the cumulative sum plot can include determining that the difference between the cumulative sum of the first rank of cell labels and the cumulative sum of the second rank of cell labels exceeds the difference between the first rank and the second rank.
[0018] In some embodiments, the difference between the first rank and the second rank is 1. In some embodiments, the method includes, in (d), when cell labels of a plurality of barcode - tagged targets are identified as noise cell labels, removing sequence information associated with the identified cell labels from the sequence data obtained in (a). The method can include, when the number of molecular labels having distinct sequences associated with a target among a plurality of targets exceeds a molecular label generation threshold, removing sequence information associated with the molecular labels having distinct sequences associated with the target among the plurality of targets from the sequence data obtained in (a). The cumulative sum plot can be a log - log plot. The log - log plot can be a log 10 -log 10 plot.
[0019] In some embodiments, the minimum is a global minimum. Identifying the minimum of the second - derivative plot can include determining that the minimum of the second - derivative plot exceeds a threshold of the minimum number of molecular labels associated with each of the cell labels. The threshold of the minimum number of molecular labels associated with each of the cell labels can be a percentile threshold. The threshold of the minimum number of molecular labels associated with each of the cell labels can be determined based on the number of cells in the cell sample.
[0020] In some embodiments, identifying the minimum of the second - derivative plot includes determining that the minimum of the second - derivative plot is below a threshold of the maximum number of molecular labels associated with each of the cell labels. The threshold of the maximum number of molecular labels associated with each of the cell labels can be a percentile threshold. The threshold of the maximum number of molecular labels associated with each of the cell labels can be determined based on the number of cells in the cell sample.
[0021] In some embodiments, each of the cell labels is identified as a signal cell label if the number of molecular labels having a molecular label associated with each of the cell labels identified in (c) is greater than the cell label threshold. Each of the cell labels can be identified as a noise cell label if the number of molecular labels having a molecular label associated with each of the cell labels identified in (c) is not greater than the cell label threshold.
[0022] In some embodiments, the method includes, for one or more of a plurality of targets: (1) counting the number of molecular labels having a distinct sequence associated with the target in the sequence data; and (2) estimating the number of targets based on the number of molecular labels having a distinct sequence associated with the target in the sequence data counted in (1).
[0023] Disclosed herein is a method of identifying signal cell labels. In some embodiments, the method includes: (a) obtaining sequence data of a plurality of targets of a cell, wherein each target is associated with the number of molecular labels having a distinct sequence associated with each of the plurality of cell labels; (b) determining a cell label threshold based on the number of molecular labels having a distinct sequence associated with each of the cell labels; and (c) identifying each of the cell labels as a signal cell label or a noise cell label based on the number of molecular labels having a distinct sequence associated with each of the cell labels and the cell label threshold.
[0024] In some embodiments, obtaining array data comprises barcoding multiple targets of cells using multiple barcodes to create multiple barcoded targets, wherein each of the multiple barcodes comprises a cell label and a molecular label of multiple cell labels, barcoding to create barcoded targets, and determining the number of molecular labels having distinct sequences associated with each of the cell labels of the multiple barcodes. In some embodiments, the method comprises, for one or more of the multiple targets, (1) counting the number of molecular labels having distinct sequences associated with the target in the array data, and (2) estimating the number of targets based on the number of molecular labels having distinct sequences associated with the target in the array data counted in (1). The method can comprise removing, from the array data, sequence information associated with an identified cell label if the cell label of the multiple barcodes is identified as a noise cell label. The method can comprise removing, from the array data, sequence information associated with a molecular label having a distinct sequence associated with a target among the multiple targets if the number of molecular labels having distinct sequences associated with the target among the multiple targets exceeds a molecular label occurrence threshold. In some embodiments, determining the number of molecular labels having distinct sequences associated with each of the cell labels in (c) comprises removing, from the array data, sequence information associated with non-unique molecular labels associated with each of the cell labels.
[0025] In some embodiments, determining the cell labeling threshold includes identifying an inflection point of a cumulative sum plot, the cumulative sum plot being based on the number of molecular labels having distinct sequences associated with each of a plurality of cell labels and the rank of each of the cell labels, the inflection point corresponding to the cell labeling threshold. Identifying the inflection point of the cumulative sum plot can include generating the cumulative sum plot based on the number of molecular labels having distinct sequences associated with each of a plurality of cell labels and the rank of each of the cell labels, generating a second derivative plot of the cumulative sum plot, and identifying a minimum of the second derivative plot of the cumulative sum plot, the minimum of the second derivative plot corresponding to the cell labeling threshold. Determining the cell labeling threshold can include determining the rank of each of a plurality of cell labels based on the number of molecular labels having distinct sequences associated with each of the cell labels. The cumulative sum plot can be a log-log plot such as a log10-log10 plot.
[0026] In some embodiments, generating a cumulative sum plot based on the number of molecular labels having distinct sequences associated with each of the cell labels and the rank of each of the cell labels comprises identifying the cumulative sum for each rank of the cell labels, wherein the cumulative sum of a rank comprises identifying the sum of the number of molecular labels having distinct sequences associated with each of the cell labels having lower ranks. Generating a second derivative plot of the cumulative sum plot can comprise determining that the difference between the cumulative sum of the first rank of cell labels and the cumulative sum of the second rank of cell labels exceeds the difference between the first rank and the second rank. The difference between the first rank and the second rank can be 1. The minimum can be a global minimum. Identifying the minimum of the second derivative plot can comprise determining that the minimum of the second derivative plot exceeds a threshold of the minimum number of molecular labels associated with each of the cell labels. The threshold of the minimum number of molecular labels associated with each of the cell labels can be a percentile threshold. The threshold of the minimum number of molecular labels associated with each of the cell labels can be determined based on the number of cells in a plurality of cells.
[0027] In some embodiments, identifying the minimum of the second derivative plot comprises determining that the minimum of the second derivative plot is below a threshold of the maximum number of molecular labels associated with each of the cell labels. The threshold of the maximum number of molecular labels associated with each of the cell labels can be a percentile threshold. The threshold of the maximum number of molecular labels associated with each of the cell labels can be determined based on the number of cells in a plurality of cells.
[0028] In some embodiments, each of the cell labels can be identified as a signal cell label if the number of molecular labels having molecular labels associated with each of the cell labels is greater than a cell label threshold. Each of the cell labels can be identified as a noise cell label if the number of molecular labels having molecular labels associated with each of the cell labels is not greater than the cell label threshold.
[0029] Disclosed herein is a method for identifying signal cell labels. In some embodiments, the method comprises: (a) using a plurality of barcodes (e.g., probabilistic barcodes) to barcode (e.g., probabilistically barcode) a plurality of targets in a sample of cells to create a plurality of barcoded targets (e.g., probabilistically barcoded targets), wherein each of the plurality of barcodes comprises a cell label and a molecular label, barcoded targets created from different cell targets of the plurality of cells have different cell labels, and barcoded targets created from the same cell targets of the plurality of cells have different molecular labels, barcoding to create barcoded targets; (b) obtaining sequence data of the plurality of barcoded targets; (c) identifying a feature vector for each cell label of the plurality of barcodes (or barcoded targets), the feature vector including the number of molecular labels having distinct sequences associated with each cell label; (d) identifying a cluster for each cell label of the plurality of barcodes (or barcoded targets) based on the feature vectors; and (e) identifying each cell label of the plurality of barcodes (or barcoded targets) as a signal cell label or a noise cell label based on the number of cell labels within the cluster and a cluster size threshold.
[0030] In some embodiments, identifying a cluster for each cell label of the plurality of barcoded targets based on the feature vectors comprises clustering each cell label of the plurality of barcoded targets into a cluster based on the distance of the feature vectors to the cluster in the feature vector space. Identifying a cluster for each cell label of the plurality of barcoded targets based on the feature vectors can comprise projecting the feature vectors from the feature vector space to a lower dimensional space and clustering each cell label into a cluster based on the distance of the feature vectors to the cluster in the lower dimensional space.
[0031] In some embodiments, the lower-dimensional space is a two-dimensional space. Projecting the feature vectors from the feature vector space into a lower-dimensional space can include using the t-distributed stochastic neighbor embedding (tSNE) method to project the feature vectors from the feature vector space into a lower-dimensional space. Clustering each cell label into clusters based on the distance of the feature vectors to the clusters in the lower-dimensional space can include using a density-based method to cluster each cell label into clusters based on the distance of the feature vectors to the clusters in the lower-dimensional space. The density-based method can include the density-based spatial clustering of applications with noise (DBSCAN) method.
[0032] In some embodiments, a cell label is identified as a signal cell label if the number of cell labels within the cluster is below a cluster size threshold. A cell label can be identified as a noise cell label if the number of cell labels within the cluster does not fall below the cluster size threshold. The method can include (f) for one or more of a plurality of targets, (1) counting the number of molecular labels having distinct sequences associated with the target in the array data, and (2) estimating the number of targets based on the number of molecular labels having distinct sequences associated with the target in the array data counted in (1).
[0033] In some embodiments, the method includes determining a cluster size threshold based on the number of cell labels of a plurality of barcoded targets. The cluster size threshold can be a percentage of the number of cell labels of a plurality of barcoded targets. In some embodiments, the method includes determining a cluster size threshold based on the number of cell labels of a plurality of barcodes. The cluster size threshold is a percentage of the number of cell labels of a plurality of barcodes. In some embodiments, the method includes determining a cluster threshold size based on the number of molecular labels having distinct sequences associated with each cell label of a plurality of barcodes.
[0034] Disclosed herein is a method for identifying signal cell labels. In some embodiments, the method comprises: (a) obtaining sequence data of a plurality of barcoded targets (e.g., stochastically barcoded targets), wherein the plurality of barcoded targets are created from a plurality of targets in a sample of cells barcoded (e.g., stochastically barcoded) using a plurality of barcodes (e.g., stochastic barcodes), each of the plurality of barcodes includes a cell label and a molecular label, barcoded targets created from targets of different cells of the plurality of cells have different cell labels, and barcoded targets created from targets of the same cell of the plurality of cells have different molecular labels; (b) identifying a feature vector for each cell label of the plurality of barcoded targets, wherein the feature vector includes the number of molecular labels having a distinct sequence associated with each cell label; (c) identifying clusters for each cell label of the plurality of barcoded targets based on the feature vectors; and (d) identifying each cell label of the plurality of barcoded targets as a signal cell label or a noise cell label based on the number of cell labels within the cluster and a cluster size threshold.
[0035] In some embodiments, identifying clusters for each cell label of the plurality of barcoded targets based on the feature vectors includes clustering each cell label of the plurality of barcoded targets into clusters based on the distance of the feature vectors to the clusters in the feature vector space. Identifying clusters for each cell label of the plurality of barcoded targets based on the feature vectors includes projecting the feature vectors from the feature vector space to a lower-dimensional space and clustering each cell label into clusters based on the distance of the feature vectors to the clusters in the lower-dimensional space. The lower-dimensional space can be a two-dimensional space.
[0036] In some embodiments, projecting the feature vectors from the feature vector space to a lower-dimensional space includes using the t-distributed stochastic neighbor embedding (tSNE) method to project the feature vectors from the feature vector space to a lower-dimensional space. Clustering each cell label into clusters based on the distances of the feature vectors to the clusters in the lower-dimensional space can include using a density-based method to cluster each cell label into clusters based on the distances of the feature vectors to the clusters in the lower-dimensional space. The density-based method can include the density-based spatial clustering of applications with noise (DBSCAN) method.
[0037] In some embodiments, a cell label can be identified as a signal cell label if the number of cell labels within a cluster is below a cluster size threshold. A cell label can be identified as a noise cell label if the number of cell labels within a cluster is not below the cluster size threshold.
[0038] In some embodiments, the method includes determining a cluster size threshold based on the number of cell labels of a plurality of barcoded targets. The cluster size threshold can be a percentage of the number of cell labels of the plurality of barcoded targets. In some embodiments, determining a cluster size threshold based on the number of cell labels of a plurality of barcoded targets. The cluster size threshold can be a percentage of the number of cell labels of the plurality of barcodes. In some embodiments, the method includes determining the cluster threshold size based on the number of molecular labels having distinct sequences associated with each cell label of the plurality of barcodes.
[0039] In some embodiments, the method includes, for one or more of a plurality of targets: (1) counting the number of molecular labels having distinct sequences associated with the target in the sequence data; and (2) estimating the number of targets based on the number of molecular labels having distinct sequences associated with the target in the sequence data counted in (1).
[0040] Disclosed herein is a method for identifying signal cell labels. In some embodiments, the method comprises: (a) obtaining sequence data of a plurality of first targets of cells, wherein each of the first targets is associated with the number of molecular labels having distinct sequences associated with each of the plurality of cell labels; (b) identifying each of the cell labels as a signal cell label or a noise cell label based on the number of molecular labels having distinct sequences associated with each cell label and an identification threshold; and (c) re-identifying at least one of the plurality of cell labels identified as a noise cell label in (b) as a signal cell label, or re-identifying at least one of the cell labels identified as a signal cell label in (b) as a noise cell label. Identifying each cell label, re-identifying at least one of the plurality of cell labels as a signal cell label, or re-identifying at least one of the plurality of cell labels as a noise cell label can be based on the same cell label identification method as the present disclosure or a different cell label identification method. The identification threshold can include a cell label threshold, a cluster size threshold, or any combination thereof. The method can include removing one or more of the plurality of cell labels each associated with the number of molecular labels having distinct sequences below a threshold of the number of molecular labels.
[0041] In some embodiments, re-identifying at least one of a plurality of cell labels identified as noise cell labels in (b) as a signal cell label includes identifying a plurality of second targets among a plurality of first targets, each having one or more diversity indications exceeding a diversity threshold, and for each of the plurality of cell labels, re-identifying at least one of the plurality of cell labels identified as noise cell labels in (b) as a signal cell label based on the number of molecular labels having distinct sequences associated with the plurality of second targets and an identification threshold. One or more diversity indications of a second target can include an average, maximum, median, minimum, variance, or any combination thereof of the number of cell labels among the molecular labels having distinct sequences associated with the second target and the plurality of cell labels in the sequence data. One or more diversity indications of a second target can include a diversity indication standard deviation, a normalized variance, or any combination thereof of a subset of the plurality of second targets. The diversity threshold can be less than or equal to the size of a subset of the plurality of second targets.
[0042] In some embodiments, re-identifying at least one of a plurality of cell labels identified as signal cell labels in (b) as a noise cell label includes identifying a plurality of third targets among a plurality of first targets each having a relevance with a cell label identified as a noise cell label in (c) that exceeds a relevance threshold, and re-identifying at least one of the cell labels identified as signal cell labels in (b) as a noise cell label based on the number of molecular labels having distinct sequences associated with the plurality of third targets and an identification threshold for each of the plurality of cell labels. Identifying a plurality of third targets among a plurality of first targets each having a relevance with a cell label identified as a noise cell label in (c) that exceeds a relevance threshold can include, after re-identifying at least one of the cell labels identified as noise cell labels in (b) as signal cell labels, identifying the remaining plurality of cell labels identified as signal cell labels, and identifying the plurality of third targets, where identifying the plurality of third targets includes identifying the plurality of third targets based on the number of molecular labels having distinct sequences associated with the plurality of targets and, for each of the remaining plurality of cell labels, the number of molecular labels having distinct sequences associated with the plurality of targets.
[0043] Disclosed herein is a computer system for identifying signal cell labels. In some embodiments, the system includes a hardware processor and a non-transitory memory storing instructions that, when executed by the hardware processor, cause the processor to perform any of the methods disclosed herein. Disclosed herein is a computer-readable medium for identifying signal cell labels. In some embodiments, the computer-readable medium includes code for performing any of the methods disclosed herein.
Brief Description of the Drawings
[0044]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7
Figure 8A
Figure 8B
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13A
Figure 13B
Figure 14A
Figure 14B
Figure 14C
Figure 14D
Figure 15A
Figure 15B
Figure 16A
Figure 16B
Figure 17A
Figure 17B
Figure 17C
Figure 17D
DETAILED DESCRIPTION OF THE INVENTION
[0045] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, like symbols typically identify like components, unless the context indicates otherwise. For the purposes of the description set forth in the detailed description, the drawings, and the claims, the embodiments are not meant to be limiting. Other embodiments may be utilized and other changes may be made without departing from the spirit or scope of the present disclosure. Aspects of the present disclosure can be arranged, substituted, combined, separated, and designed in a variety of different configurations, as outlined herein and shown in the figures, all of which are explicitly intended herein and are readily understood to be a part of the present disclosure.
[0046] All patents, published patent applications, other publications, and sequences from GenBank, as well as other databases, cited herein are hereby incorporated by reference in their entirety for the relevant art.
[0047] Quantification of a few diffused or targeted molecules, such as messenger ribonucleic acid (mRNA) molecules, is clinically important, for example, to identify genes expressed at various development stages or under various environmental conditions. However, it can be a very difficult problem to determine the absolute number of nucleic acid molecules (e.g., mRNA molecules), especially when the number of molecules is very small. One way to determine the absolute number of molecules in a sample is digital polymerase chain reaction (PCR). Ideally, PCR generates an ideal copy of the molecule in each cycle. However, PCR has drawbacks such that each molecule is replicated with a statistical probability that varies with the PCR cycle and gene sequence, resulting in amplification bias and inaccurate gene expression measurements.
[0048] The number of molecules can be counted using barcodes (e.g., probabilistic barcodes) having unique molecular labels (also called molecular indices (MI)). By using barcodes having unique molecular labels for each cell label, the number of molecules within each cell can be counted. Non-limiting and exemplary assays for barcoding include the Precise™ assay (Cellular Research, Inc. (Palo Alto, CA)), the Resolve™ assay (Cellular Research, Inc. (Palo Alto, CA)), or the Rhapsody™ assay (Cellular Research, Inc. (Palo Alto, CA)). However, these methods and techniques can introduce errors that, if uncorrected, can lead to an overestimation of cell counts.
[0049] The Rhapsody (trademark) assay can hybridize to all poly(A)-mRNAs in a sample by utilizing a non-depleting pool of barcodes (e.g., stochastic barcodes) having a large number, e.g., 6561 to 65536 unique molecular labels, on poly(T) oligonucleotides during the RT step. In addition to the molecular labels, cell labels of the barcodes can be used to identify each one cell within each well of a microplate. The barcode can include a universal PCR priming site. During RT, the target gene molecules react randomly with the barcode. Each target molecule can hybridize to a barcode (e.g., stochastic barcode), resulting in the generation of barcode-tagged complementary ribonucleic acid (cDNA) molecules (e.g., stochastic barcode-tagged cDNA molecules). After labeling, the barcode-tagged cDNA molecules from the microplate wells can be pooled into one tube for PCR amplification and sequencing. The raw sequencing data can be analyzed to generate the number of barcodes with unique molecular labels.
[0050] Methods and systems for identifying signal cell labels are disclosed herein. In some embodiments, the method comprises: (a) barcoding (e.g., stochastically barcoding) a plurality of targets in a sample of cells using a plurality of barcodes (e.g., stochastic barcodes) to create a plurality of barcoded targets (e.g., stochastic barcoded targets), wherein each of the plurality of barcodes comprises a cell label and a molecular label, the barcoding to create the barcoded targets; (b) obtaining sequence data of the plurality of barcoded targets; (c) determining the number of molecular labels having distinct sequences associated with each of the cell labels of the plurality of barcodes; (d) determining the rank of each of the cell labels of the plurality of barcodes based on the number of molecular labels having distinct sequences associated with each of the cell labels; (e) generating a cumulative sum plot based on the number of molecular labels having distinct sequences associated with each of the cell labels identified in (c) and the rank of each of the cell labels determined in (d); (f) generating a second derivative plot of the cumulative sum plot; (g) identifying a minimum of the second derivative plot of the cumulative sum plot, wherein the minimum of the second derivative plot corresponds to a cell label threshold; and (h) identifying the cell label as a signal cell label or a noise cell label based on the number of molecular labels having distinct sequences associated with each of the cell labels identified in (c) and the cell label threshold.
[0051] In some embodiments, the method comprises: (a) obtaining array data of a plurality of barcoded targets (e.g., stochastically barcoded targets), wherein the array data of the plurality of barcoded targets is from a plurality of targets in a sample of cells that have been barcoded (e.g., stochastically barcoded) using a plurality of barcodes (e.g., stochastic barcodes) to create the plurality of barcoded targets (e.g., stochastically barcoded targets), and each of the plurality of barcodes comprises a cell label and a molecular label; (b) determining a rank for each of the cell labels of the plurality of barcodes based on the number of molecular labels having distinct sequences associated with each of the cell labels; (c) identifying a minimum of a second derivative plot of a cumulative sum plot, wherein the cumulative sum plot is based on the number of molecular labels having distinct sequences associated with each of the cell labels and the rank of each cell label determined in (b), and the minimum of the second derivative plot corresponds to a cell label threshold; and (d) classifying a cell label as a signal cell label (associated with a cell) or a noise cell label (discarded as not associated with a cell) based on the number of molecular labels having distinct sequences associated with the cell label and the cell label threshold.
[0052] Disclosed herein is a method for identifying signal cell labels. In some embodiments, the method comprises: (a) barcoding (e.g., stochastically barcoding) a plurality of targets in a sample of cells using a plurality of barcodes (e.g., stochastic barcodes) to create a plurality of barcoded targets (e.g., stochastic barcoded targets), wherein each of the plurality of barcodes comprises a cell label and a molecular label, barcoded targets created from targets of different cells have different cell labels, and barcoded targets created from one of the plurality of cells have different molecular labels; (b) obtaining sequence data of the barcoded targets; (c) identifying a feature vector of the cell label, the feature vector including the number of molecular labels having distinct sequences associated with the cell label; (d) identifying clusters of cell labels based on the feature vector; and (e) identifying the cell label as a signal cell label or a noise cell label based on the number of cell labels within the cluster and a cluster size threshold.
[0053] Disclosed herein is a system for identifying signal cell labels. In some embodiments, the system comprises a hardware processor and a non-transitory memory storing instructions that, when executed by the hardware processor, cause the processor to perform any of the methods disclosed herein. Disclosed herein is a computer-readable medium for identifying signal cell labels. In some embodiments, the computer-readable medium comprises code for performing any of the methods disclosed herein.
[0054] Definitions Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Springs Harbor Press (Cold Springs Harbor, NY 1989). In this disclosure, the following terms are defined as follows.
[0055] As used herein, the term "adapter" can mean a sequence that facilitates the amplification or sequencing of an associated nucleic acid. The associated nucleic acid can include a target nucleic acid. The associated nucleic acid can include one or more of a spatial label, a target label, a sample label, an indexing label, a barcode, a probabilistic barcode, or a molecular label. The adapter can be linear. The adapter can be a pre-adenylated adapter. The adapter can be double-stranded or single-stranded. One or more adapters can be placed at the 5' or 3' end of a nucleic acid. If the adapter contains a known sequence at the 5' or 3' end, the known sequence can be the same or a different sequence. An adapter placed at the 5' or 3' end of a polynucleotide can hybridize to one or more oligonucleotides immobilized on a surface. The adapter can, in some embodiments, include a universal sequence. The universal sequence can be a region of a nucleotide sequence common to two or more nucleic acid molecules. The two or more nucleic acid molecules can have regions of different sequences. Thus, for example, a 5' adapter can include an identical and / or universal nucleic acid sequence, and a 3' adapter can include an identical and / or universal nucleic acid sequence. Universal sequences that can be present in different members of a plurality of nucleic acid molecules can enable the replication or amplification of the plurality of different sequences using a single universal primer complementary to the universal sequence. Similarly, at least one, two (e.g., a pair), or more universal sequences that can be present in different members of a collection of nucleic acid molecules can enable the replication or amplification of the plurality of different sequences using at least one, two (e.g., a pair), or more single universal primers complementary to the universal sequences. Thus, a universal primer includes a sequence that can hybridize to such a universal sequence. A target nucleic acid sequence-bearing molecule can be modified to attach a universal adapter (e.g., a non-target nucleic acid sequence) to one or both ends of different target nucleic acid sequences.One or more universal primers attached to a target nucleic acid can provide a site to which the universal primer hybridizes. The one or more universal primers attached to the target nucleic acid can be the same as or different from each other.
[0056] As used herein, the terms "associated with" or "associated therewith" can mean that two or more species can be identified as being in the same location at some point in time. Association can mean that two or more species are or were in the same container. Association can be informatics-related, in which case, for example, digital information regarding two or more species is stored and can be used in a determination that one or more of the species were arranged in the same location at some point in time. Association can be physical association. In some embodiments, two or more associated species are "tethered," "attached," or "immobilized" to each other or to a common solid or semi-solid surface. Association can refer to covalent or non-covalent means of attaching a label to a solid or semi-solid support such as a bead. Association can be a covalent bond between a target and a label.
[0057] As used herein, the term "complementary" can refer to the ability of two nucleotides to pair precisely. For example, if the nucleotide at a given position in a nucleic acid can hydrogen bond with a nucleotide in another nucleic acid, those two nucleic acids are considered complementary to each other at that position. Complementarity between two single-stranded nucleic acid molecules where only some of the nucleotides bind can be "partial", or can be complete if complete complementarity exists between the single-stranded molecules. A first nucleotide sequence can be said to be "complementary" to a second sequence if the first nucleotide sequence is complementary to the second nucleotide sequence. A first nucleotide sequence can be said to be "reverse-complementary" to a second sequence if the first nucleotide sequence is complementary to a sequence that is the reverse of the second sequence (i.e., the order of the nucleotides is reversed). As used herein, the terms "complement", "complementary", and "reverse-complementary" can be used synonymously. It is understood from the present disclosure that if a molecule can hybridize to another molecule, that molecule can be complementary to the molecule to which it is hybridizing.
[0058] As used herein, the term "digital count" can refer to a method of estimating the number of target molecules in a sample. A digital count can include the step of identifying the number of unique labels associated with a target in a sample. This probabilistic methodology transforms the problem of counting molecules from finding and identifying identical molecules to a series of yes / no digital questions regarding the detection of a predefined set of labels.
[0059] As used herein, the term "label" or "labels" can refer to nucleic acid codes associated with a target in a sample. The label can be, for example, a nucleic acid label. The label can be an amplifiable label, either in whole or in part. The label can be a sequencable label, either in whole or in part. The label can be a portion of a native nucleic acid that is separately distinguishable. The label can be a known sequence. The label can include a junction of nucleic acid sequences, such as a junction between a native sequence and a non-native sequence. As used herein, the term "label" can be used synonymously with the terms "index", "tag", or "label tag". The label can convey information. For example, in various embodiments, a label can be used to identify information about a sample, the source of the sample, cell identification information, and / or to identify a target.
[0060] As used herein, the term "non-depleting reservoir" can refer to a pool of probabilistic barcodes composed of many different labels. A non-depleting reservoir can include a large number of different probabilistic barcodes such that when a pool of targets is associated with the non-depleting reservoir, each target is likely to be associated with a unique probabilistic barcode. The uniqueness of each labeled target molecule can be determined by the statistics of random selection and depends on the copy number of the same target molecule in the collection compared to the diversity of the labels. The size of the resulting set of labeled target molecules can be determined by the probability of the barcoding process, and the number of target molecules present in the original collection or sample can be calculated by analysis of the number of detected probabilistic barcodes. When the ratio of the number of target molecules present to the number of unique probabilistic barcodes is low, the labeled target molecules are highly unique (i.e., the probability that two or more target molecules are labeled with a given label is very low).
[0061] As used herein, the term "nucleic acid" refers to a polynucleotide sequence or a fragment thereof. A nucleic acid can comprise nucleotides. A nucleic acid can be exogenous or endogenous to a cell. A nucleic acid can be present in a cell-free environment. A nucleic acid can be a gene or a fragment thereof. A nucleic acid can be DNA. A nucleic acid can be RNA. A nucleic acid can comprise one or more analogs (e.g., modified backbone, sugar, or nucleobase). Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotide, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine. The terms "nucleic acid," "polynucleotide," "target polynucleotide," and "target nucleic acid" can be used interchangeably.
[0062] The nucleic acid can comprise one or more modifications (e.g., base modifications, backbone modifications) that provide the nucleic acid with new or enhanced features (e.g., improved stability). The nucleic acid can comprise a nucleic acid affinity tag. A nucleotide can be a base-sugar combination. The base portion of the nucleotide can be a heterocyclic base. Two of the most common classes of such heterocyclic bases are purines and pyrimidines. A nucleotide can be a nucleotide that further comprises a phosphate group covalently attached to the sugar portion of the nucleotide. In the case of a nucleotide comprising a pentofuranosyl sugar, the phosphate group can link to the 2’, 3’, or 5’ hydroxyl moiety of the sugar. In forming a nucleic acid, the phosphate groups can covalently bond adjacent nucleotides to one another to form a linear polymeric compound. And each end of this linear polymeric compound can be further joined to form a cyclic compound, although linear compounds are generally suitable. Additionally, the linear compound can have internal nucleotide base complementarity and thus can fold to produce a complete or partial double-stranded compound. Within the nucleic acid, the phosphate groups can generally be regarded as forming the nucleoside backbone of the nucleic acid. The link or backbone can be a 3’-5’ phosphodiester bond.
[0063] Nucleic acids can include a modified backbone and / or modified nucleoside linkages. Modified backbones can include those that retain a phosphorus atom within the backbone and those that do not have a phosphorus atom in the backbone. Suitable modified nucleic acid backbones that contain a phosphorus atom internally include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkyl phosphotriesters, methyl and 3'-alkylene phosphonates, 5'-alkylene phosphonates, and other alkyl phosphonates including chiral phosphonates, phosphinates, phosphoramidites including 3'-aminophosphoramidite and aminoalkyl phosphoramidite, thionophosphoramidites, thionoalkyl phosphonates, thionoalkyl phosphotriesters, selenophosphates and boranophosphates having a normal 3'-5' linkage, 2'-5' linkage analogs, and backbones with inverted polarity where one or more internucleotide linkages are 3'-3', 5'-5' or 2'-2' linkages.
[0064] Nucleic acids can include polynucleotide backbones formed by short chain alkyl or cycloalkyl nucleosides, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic nucleoside internucleoside linkages. These can include those having morpholino linkages (formed in part from the sugar portion of the nucleoside); siloxane backbones; sulfide, sulfoxide, and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; riboacetyl backbones; alkene-containing backbones; sulfamate backbones; methyleneimino and methylenehydrazino backbones; sulfonate and sulfonamide backbones; amide backbones; and others having portions of mixed N, O, S, and CH2 components.
[0065] The nucleic acid can include a nucleic acid mimetic. The term "mimetic" can include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are substituted with non-furanose groups, and substitution of only the furanose ring can be referred to as a sugar surrogate. The heterocyclic base moiety or modified heterocyclic base moiety can be maintained for hybridization with a suitable target nucleic acid. One such nucleic acid can be a peptide nucleic acid (PNA). In PNA, the sugar backbone of the polynucleotide can be substituted with an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotides can be retained and are directly or indirectly attached to the aza nitrogen atom of the amide portion of the backbone. The backbone in a PNA compound can include two or more linked aminoethylglycine units that give the PNA an amide-containing backbone. The heterocyclic base moiety can be directly or indirectly attached to the aza nitrogen atom of the amide portion of the backbone.
[0066] The nucleic acid can include a morpholino backbone structure. For example, the nucleic acid can include a morpholino six-membered ring in place of the ribose ring. In some of these embodiments, phosphorodiamidate or other non-phosphodiester nucleoside linkages can replace the phosphodiester bond.
[0067] Nucleic acids can include linked morpholino units (i.e., morpholino nucleic acids) having heterocyclic bases attached to a morpholino ring. A linking group can link morpholino monomer units in a morpholino nucleic acid. Nonionic morpholino-based oligomeric compounds can have fewer undesirable interactions with cellular proteins. Morpholino-based polynucleotides can be nonionic mimics of nucleic acids. A variety of compounds within the morpholino class can be joined using different linking groups. A further class of polynucleotides can be called cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in nucleic acid molecules can be replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers can be prepared and used in the synthesis of oligomeric compounds using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid strands can enhance the stability of DNA / RNA hybrids. CeNA oligo-adenyls can form complexes with nucleic acid complements having stability similar to that of native complexes. Further modifications can include locked nucleic acids (LNAs) in which a 2'-hydroxyl group is linked to the 4'-carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene linkage and thereby forming a bicyclic sugar moiety. The linkage can be a methylene (-CH2-) group bridging the 2'-oxygen atom and the 4'-carbon atom, where n is 1 or 2. LNAs and LNA analogs can exhibit very high duplex thermal stability (Tm = +3 to +10 °C) with complementary nucleic acids, stability against 3'-exonuclease degradation, and excellent solubility properties.
[0068] Nucleic acids can also contain nucleobase (often simply referred to as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases can include purine bases (e.g., adenine (A) and guanine (G)), as well as pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleobases include 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-amino-adenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C≡C-CH3) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azauracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and other synthetic and natural nucleobases such as 3-deazaguanine and 3-deazaadenine.Modified nucleobases can include tricyclic pyrimidines such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidines (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), etc. of the G-clamp, substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido 5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), pyridoindole cytidine (H-pyrido(3’,2’:4,5)pyrrolo[2,3-d]pyrimidin-2-one), etc. of the G-clamp.
[0069] As used herein, the term "sample" can refer to a composition containing a target. Samples suitable for analysis by the disclosed methods, devices, and systems include cells, tissues, organs, or organisms.
[0070] As used herein, the term "sampling device" or "device" can refer to a device that can collect a portion of a sample and / or place a portion thereof on a substrate. Sample devices can refer to, for example, fluorescence-activated cell sorting (FACS) machines, cell sorter machines, biopsy needles, biopsy devices, tissue sectioning devices, microfluidic devices, blade grids, and / or microtomes.
[0071] As used herein, the term "solid support" can refer to a discrete solid or semi-solid surface to which a plurality of probabilistic barcodes can be attached. The solid support can include any type of solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar structure composed of a plastic, ceramic, metal, or polymeric material (e.g., hydrogel) to which a nucleic acid can be immobilized (e.g., covalently or non-covalently). The solid support can have a spherical shape (e.g., a small sphere), or can include discrete particles having a non-spherical or irregular shape such as a cube, cubic bone, pyramid, cylinder, cone, ellipse, or disk. The plurality of solid supports spaced apart in an array may not include a substrate. The solid support can be used synonymously with the term "bead".
[0072] The solid support can be referred to as a "substrate". The substrate can be a type of solid support. The substrate can refer to a continuous solid or semi-solid surface on which the methods of the present disclosure can be performed. The substrate can refer to, for example, an array, cartridge, chip, device, and slide.
[0073] As used herein, the term "spatial label" can refer to a label that can be associated with a position in space.
[0074] As used herein, the term "probabilistic barcode" can refer to a polynucleotide sequence that includes a label. The probabilistic barcode can be a polynucleotide sequence that can be used for probabilistic barcoding. The probabilistic barcode can be used to quantify a target in a sample. The probabilistic barcode can be used to control for errors that can occur after the label is associated with the target. For example, the probabilistic barcode can be used to evaluate amplification or sequencing errors. The probabilistic barcode associated with a target can be referred to as a probabilistic barcode-target or a probabilistic barcode-tagged target.
[0075] As used herein, the term "gene-specific stochastic barcode" can refer to a polynucleotide sequence that includes a label and a target-binding region that is gene-specific. A stochastic barcode can be a polynucleotide sequence that can be used for stochastic barcoding. A stochastic barcode can be used to quantify a target within a sample. A stochastic barcode can be used to control for errors that can occur after the label has been associated with the target. For example, a stochastic barcode can be used to evaluate amplification or sequencing errors. A stochastic barcode associated with a target can be referred to as a stochastic barcode-target or a stochastic barcode-tagged target.
[0076] As used herein, the term "stochastic barcoding" can refer to the random labeling (e.g., barcoding) of nucleic acids. Stochastic barcoding can utilize a recursive Poisson method to associate a label with a target and to quantify the label associated with the target. As used herein, "stochastic barcoding" can be used synonymously with "gene-specific stochastic barcoding".
[0077] As used herein, the term "target" can refer to a composition with which a stochastic barcode can be associated. Exemplary targets suitable for analysis by the disclosed methods, devices, and systems include DNA, RNA, mRNA, microRNA, tRNA, etc. The target can be single-stranded or double-stranded. In some embodiments, the target can be a protein. In some embodiments, the target is a lipid.
[0078] As used herein, the term "reverse transcriptase" can refer to a group of enzymes having reverse transcription activity (i.e., catalyzing the synthesis of DNA from an RNA template). Generally, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retrons reverse transcriptases, recently discovered reverse transcriptases, group II intron-derived reverse transcriptases, and variants, mutants, or derivatives thereof. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retrons reverse transciptases, and group II intron reverse transcriptases. Examples of group II intron reverse transcriptases include the Lactococcus lactis Ll.LtrB intron reverse transcriptase, the Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or the Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases can include many classes of non-retroviral reverse transcriptases (i.e., particularly retrons, group II introns, and diversity-generating retroelements).
[0079] Disclosed herein are systems and methods for identifying signal cell labels. In some embodiments, the method comprises: (a) stochastically barcoding a plurality of targets in a sample of cells using a plurality of stochastic barcodes to create a plurality of barcoded targets, wherein each of the plurality of stochastic barcodes comprises a cell label and a molecular label, and stochastically barcoding to create barcoded targets; (b) obtaining sequence data of the plurality of barcoded targets; (c) determining the number of molecular labels having distinct sequences associated with each of the cell labels of the plurality of stochastic barcodes; (d) determining the rank of each of the cell labels of the plurality of stochastic barcodes based on the number of molecular labels having distinct sequences associated with each of the cell labels; (e) generating a cumulative sum plot based on the number of molecular labels having distinct sequences associated with each of the cell labels identified in (c) and the rank of each of the cell labels determined in (d); (f) generating a second derivative plot of the cumulative sum plot; (g) identifying a minimum of the second derivative plot of the cumulative sum plot, wherein the minimum of the second derivative plot corresponds to a cell label threshold; and (h) identifying each of the cell labels as a signal cell label or a noise cell label based on the number of molecular labels having distinct sequences associated with each of the cell labels identified in (c) and the cell label threshold determined in (g).
[0080] Barcode Barcoding, such as probabilistic barcoding, has been described, for example, in U.S. Patent Application Publication No. 20150299784, International Publication No. 2015031691, and Fu et al., Proc Natl Acad Sci U.S.A. 2011 May 31;108(22):9026-31 and Fan et al., Science (2015) 347(6222):1258367, the contents of these publications being hereby incorporated by reference in their entirety. In some embodiments, the barcode disclosed herein can be a probabilistic barcode that can be a polynucleotide sequence that can be used to probabilistically label a target (e.g., barcode, tag). A barcode can be called a probabilistic barcode if the ratio of the number of different barcode sequences of the probabilistic barcode to the number of occurrences of any of the targets to be labeled is 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values, or approximately these values or ranges. The target can be, for example, an mRNA species that includes mRNA molecules having the same or substantially the same sequence. A barcode can be called a probabilistic barcode if the ratio of the number of different barcode sequences of the probabilistic barcode to the number of occurrences of any of the targets to be labeled is at least or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1. The barcode sequences of the probabilistic barcode can be called molecular labels.
[0081] Barcodes, such as probabilistic barcodes, can include one or more labels. Exemplary labels can include universal labels, cell labels, barcode sequences (e.g., molecular labels), sample labels, plate labels, spatial labels, and / or pre-spatial labels. Figure 1 shows an exemplary barcode 104 with a spatial label. The barcode 104 can include a 5’ amine that can link the barcode to a solid support 105. The barcode can include a universal label, a dimensional label, a spatial label, a cell label, and / or a molecular label. The order of different labels (including, but not limited to, universal labels, dimensional labels, spatial labels, cell labels, and molecular labels) in the barcode can be various. For example, as shown in Figure 1, the universal label can be the 5’-most label and the molecular label can be the 3’-most label. The spatial label, dimensional label, and cell label can be in any order. In some embodiments, the universal label, spatial label, dimensional label, cell label, and molecular label can be in any order. The barcode can include a target binding region. The target binding region can interact with a target in a sample (e.g., a target nucleic acid, RNA, mRNA, DNA). For example, the target binding region can include an oligo(dT) sequence that can interact with the poly(A) tail of mRNA. In some cases, the labels of the barcode (e.g., universal labels, dimensional labels, spatial labels, cell labels, and barcode sequences) can be separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20, or more nucleotides.
[0082] Identifiers, such as cell identifiers, can include unique sets of nucleic acid sub-sequences of defined length, for example, 7 nucleotides each (equivalent to the number of bits used in some Hamming error correction codes), designed to provide an error correction function. A set of error correction sub-sequences can include 7 nucleotide sequences, and the combination for any pair of sequences in a set can be designed to exhibit a defined "genetic distance" (or number of mismatched bases), for example, a set of error correction sub-sequences can be designed to exhibit a genetic distance of 3 nucleotides. In this case, amplification errors or sequencing errors can be detected or corrected by review (more fully described below) of the error correction sequences in a set of sequence data of a labeled target nucleic acid molecule. In some embodiments, the length of the nucleic acid sub-sequences used to create the error correction code can vary, for example, it can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50, or the number or range between any two of these values of nucleotide lengths, or can be about these values or ranges. In some embodiments, nucleic acid sub-sequences of other lengths can be used to create the error correction code.
[0083] Barcodes can include a target binding region. The target binding region can interact with a target in a sample. The target can be ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing a poly(A) tail respectively, or any combination thereof, or can include these. In some embodiments, a plurality of targets can include deoxyribonucleic acid (DNA).
[0084] In some embodiments, the target binding region can include an oligo(dT) sequence that can interact with the poly(A) tail of the mRNA. One or more of the barcode labels (e.g., universal label, dimensional label, spatial label, cell label, and barcode sequence (e.g., molecular label)) can be separated by a spacer from one or two of the other remaining barcode labels. The spacer can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides. In some embodiments, none of the barcode labels are separated by a spacer.
[0085] Universal label A barcode can include one or more universal labels. In some embodiments, the one or more universal labels can be the same for all barcodes within a given set of barcodes attached to a solid support. In some embodiments, the one or more universal labels can be the same for all barcodes attached to a plurality of beads. In some embodiments, the universal label can include a nucleic acid sequence that is capable of hybridizing to a sequencing primer. The sequencing primer can be used for sequencing the barcode that includes the universal label. The sequencing primer (e.g., a universal sequencing primer) can include a sequencing primer associated with a high-throughput sequencing platform. In some embodiments, the universal label can include a nucleic acid sequence that is capable of hybridizing to a PCR primer. In some embodiments, the universal label can include a nucleic acid sequence that is capable of hybridizing to both a sequencing primer and a PCR primer. The nucleic acid sequence of the universal label that is capable of hybridizing to a sequencing primer or a PCR primer can be referred to as a primer binding site. The universal label can include a sequence that can be used for initiating transcription of the barcode. The universal label can include a sequence that can be used for amplifying the barcode or a region within the barcode. The universal label can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or a number or range between any two of these values or about these lengths of nucleotides in length. For example, the universal label can include at least about 10 nucleotides. The universal label can include at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. In some embodiments, a cleavable linker or a modified nucleotide can be part of the universal label sequence to allow the barcode to be cleaved from the support.
[0086] Dimensional identifier The barcode can include one or more dimensional identifiers. In some embodiments, the dimensional identifier can include a nucleic acid sequence that provides information about the dimension in which the identifier (e.g., probabilistic identifier) occurred. For example, the dimensional identifier can provide information about the time when the target was probabilistically barcoded. The time of barcoding (e.g., probabilistic barcoding) in the sample can be associated with the dimensional identifier. The dimensional identifier can be activated at the time of the identifier. Different dimensional identifiers can be activated at different times. The dimensional identifier provides information about the order in which the target, target group, and / or sample were probabilistically barcoded. For example, a population of cells can be probabilistically barcoded in the G0 phase of the cell cycle. The cells can be pulsed again with a barcode (e.g., probabilistic barcode) in the G1 phase of the cell cycle. The cells can be pulsed again with a barcode in the S phase of the cell cycle, and so on. The barcode in each pulse (e.g., each phase of the cell cycle) can include different dimensional identifiers. In this way, the dimensional identifier provides information about which target was labeled in which phase of the cell cycle. The dimensional identifier can collate many different biological times. Exemplary biological times include, but are not limited to, the cell cycle, transcription (e.g., transcription initiation), and transcript degradation. In another example, a sample (e.g., a cell, a population of cells) can be probabilistically labeled before and / or after treatment with a drug and / or therapy. Changes in the copy number of distinct targets can indicate the response of the sample to the drug and / or therapy.
[0087] The dimensional label can be activatable. An activatable dimensional label can be activated at a specific time point. An activatable label can, for example, be constantly activated (e.g., not turned off). An activatable dimensional label can, for example, be reversibly activated (e.g., the activatable dimensional label can be switched on and off). The dimensional label can, for example, be reversibly activatable at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times. The dimensional label can, for example, be reversibly activatable at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times. In some embodiments, the dimensional label can be activated using fluorescence, light, chemical events (e.g., cleavage, ligation of another molecule, addition of modifications (e.g., pegylation, SUMOylation, acetylation, methylation, deacetylation, demethylation), photochemical events (e.g., photocaging), and introduction of unnatural nucleotides).
[0088] In some embodiments, the dimensional label can be the same for all barcodes (e.g., probabilistic barcodes) attached to a given solid support (e.g., beads), but can also be different for different solid supports (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% of the barcodes on the same solid support can contain the same dimensional label. In some embodiments, at least 60% of the barcodes on the same solid support can contain the same dimensional label. In some embodiments, at least 95% of the barcodes on the same solid support can contain the same dimensional label.
[0089] represented in a plurality of solid supports (e.g., beads) 10 6There can be many unique dimensional label arrays of one or more. The dimensional label can be one, two, three, four, five, ten, fifteen, twenty, twenty-five, thirty, thirty-five, forty, forty-five, fifty nucleotides in length, or a number or range between any two of these values, or can be about these values or ranges. The dimensional label can be at least or at most one, two, three, four, five, ten, fifteen, twenty, twenty-five, thirty, thirty-five, forty, forty-five, fifty, one hundred, two hundred, or three hundred nucleotides in length. The dimensional label can include from about 5 to about 200 nucleotides. The dimensional label can include from about 10 to about 150 nucleotides. The dimensional label can include a length of from about 20 to about 125 nucleotides.
[0090] Spatial label The barcode can include one or more spatial labels. In some embodiments, the spatial label can include a nucleic acid sequence that provides information about the spatial orientation of a target molecule associated with the barcode. Coordinates in the sample can be associated with the spatial label. The coordinates can be fixed coordinates. For example, the coordinates can be fixed with reference to a substrate. The spatial label can refer to a two-dimensional or three-dimensional grid. The coordinates can be fixed with reference to a landmark. The landmark can be identifiable in space. The landmark can be a structure that can be imaged. The landmark can be a biological structure, such as an anatomical landmark. The landmark can be a cellular landmark, such as an organelle. The landmark can be a non-natural landmark, such as a structure having an identifiable identifier such as a color code, barcode, magnetism, fluorescence, radioactivity, or unique size or shape. A physical partition (e.g., a well, container, or droplet) can be associated with the spatial label. In some embodiments, multiple spatial labels are used together to encode one or more positions in space.
[0091] The spatial identifier can be the same for all barcodes attached to a given solid support (e.g., beads), but can also be different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes of the same solid support that contain the same spatial identifier can be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the percentage of barcodes of the same solid support that contain the same spatial identifier can be at least or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, at least 60% of the barcodes of the same solid support can contain the same spatial identifier. In some embodiments, at least 95% of the barcodes of the same solid support can contain the same spatial identifier.
[0092] There can be more than 10 6 unique spatial identifier arrays represented in a plurality of solid supports (e.g., beads). The spatial identifier can be one, two, three, four, five, ten, fifteen, twenty, twenty-five, thirty, thirty-five, forty, forty-five, fifty, or a number or range between any two of these values in length of nucleotides, or approximately these values or ranges. The spatial identifier can be at least or at most one, two, three, four, five, ten, fifteen, twenty, twenty-five, thirty, thirty-five, forty, forty-five, fifty, one hundred, two hundred, or three hundred nucleotides in length. The spatial identifier can contain from about 5 to about 200 nucleotides. The spatial identifier can contain from about 10 to about 150 nucleotides. The spatial identifier can have a length of from about 20 to about 125 nucleotides.
[0093] Cell Label The barcode can include one or more cell labels. In some embodiments, the cell label can include a nucleic acid sequence that provides information for determining which target nucleic acid came from which cell. In some embodiments, the cell label can be the same for all barcodes attached to a given solid support (e.g., beads), but can also be different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support that contain the same cell label can be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values or about these values or ranges. In some embodiments, the percentage of barcodes on the same solid support that contain the same cell label can be at least or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. For example, at least 60% of the barcodes on the same solid support can contain the same cell label. As another example, at least 95% of the barcodes on the same solid support can contain the same cell label.
[0094] There can be 10 6 or more many unique cell label sequences represented in a plurality of solid supports (e.g., beads). The cell label can be of a length of 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or a number or range between any two of these values or about these values or ranges of nucleotide numbers, or can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. For example, the cell label can include from about 5 to about 200 nucleotides. As another example, the cell label can include from about 10 to about 150 nucleotides. The cell label can include a length of from about 20 to about 125 nucleotides.
[0095] Barcode sequence The barcode can include one or more barcode arrays. In some embodiments, the barcode array can include a nucleic acid sequence that provides identification information for a particular type of target nucleic acid species hybridized to the barcode. The barcode array can include a nucleic acid sequence that provides a counter (e.g., providing a rough approximation) of a particular occurrence of a target nucleic acid species hybridized to the barcode (e.g., the target binding region).
[0096] In some embodiments, diverse sets of barcode arrays attach to a given solid support (e.g., beads). In some embodiments, there are 10 2 individuals, 10 3 individuals, 10 4 individuals, 10 5 individuals, 10 6 individuals, 10 7 individuals, 10 8 individuals, 10 9 individuals or any two of these values or a unique molecular labeling sequence within a range, or can be a unique molecular labeling sequence of about these values or ranges. For example, a plurality of barcodes can include about 6561 barcode arrays having distinct sequences. As another example, a plurality of barcodes can include about 65536 barcode arrays having distinct sequences. In some embodiments, there can be at least or at most 10 2 individuals, 10 3 individuals, 10 4 individuals, 10 5 individuals, 10 6 individuals, 10 7 individuals, 10 8 individuals, or 10 9 individuals of unique barcode arrays. The unique molecular labeling sequence can attach to a given solid support (e.g., beads).
[0097] The barcode can be one, two, three, four, five, ten, fifteen, twenty, twenty-five, thirty, thirty-five, forty, forty-five, fifty nucleotides in length, or a number or range between any two of these values, or can be approximately these values or ranges. The barcode can be at least or at most one, two, three, four, five, ten, fifteen, twenty, twenty-five, thirty, thirty-five, forty, forty-five, fifty, one hundred, two hundred, or three hundred nucleotides in length.
[0098] Molecular label The probabilistic barcode can include one or more molecular labels. The molecular label can include a barcode sequence. In some embodiments, the molecular label can include a nucleic acid sequence that provides identification information for a particular type of target nucleic acid species that hybridizes to the probabilistic barcode. The molecular label can include a nucleic acid sequence that provides a counter for a particular occurrence of a target nucleic acid species that hybridizes to the probabilistic barcode (e.g., a target binding region).
[0099] In some embodiments, diverse sets of molecular labels attach to a given solid support (e.g., beads). In some embodiments, there are ten 2 ten 3 ten 4 ten 5 ten 6 ten 7 ten 8 ten 9 unique molecular label sequences, or a number or range, or can be unique molecular label sequences of approximately these values or ranges. For example, multiple probabilistic barcodes can include approximately 6561 molecular labels with distinct sequences. As another example, multiple probabilistic barcodes can include approximately 65536 molecular labels with distinct sequences. In some embodiments, at least or at most ten 2 ten 3 ten 4 ten 5 ten 6 ten 7 ten8 one, or 10 9 There can be one, or 10 unique molecular label sequences. The unique molecular label sequences can be attached to a given solid support (e.g., beads).
[0100] In the case of probabilistic barcoding using multiple probabilistic barcodes, the ratio of the number of different molecular label sequences to the number of occurrences of any target can be 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1 or a ratio between any two of these values or ranges, or can be a unique ratio of approximately these values or ranges. The target can be an mRNA species comprising mRNA transcripts having the same or substantially the same sequence. In some embodiments, the ratio of the number of different molecular label sequences to the number of occurrences of any target can be at least or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1.
[0101] The molecular label can be one, two, three, four, five, ten, fifteen, twenty, twenty-five, thirty, thirty-five, forty, forty-five, fifty, or a number or range between any two of these values in length in terms of the number of nucleotides, or can be approximately these values or ranges. The molecular label can be at least or at most one, two, three, four, five, ten, fifteen, twenty, twenty-five, thirty, thirty-five, forty, forty-five, fifty, one hundred, two hundred, or three hundred nucleotides in length.
[0102] Target binding region A barcode can include one or more target binding regions such as capture probes. In some embodiments, the target binding region can hybridize to a target of interest. In some embodiments, the target binding region can include a nucleic acid sequence that specifically hybridizes to a target (e.g., a target nucleic acid, a target molecule, e.g., a cellular nucleic acid to be analyzed), e.g., a specific gene sequence. In some embodiments, the target binding region can include a nucleic acid sequence that can attach (e.g., hybridize) to a specific location of a specific target nucleic acid. In some embodiments, the target binding region can include a nucleic acid sequence that can specifically hybridize to a restriction enzyme site overhang (e.g., an EcoRI sticky end overhang). Next, the barcode can be ligated to any nucleic acid molecule that includes a sequence complementary to the restriction site overhang.
[0103] In some embodiments, the target binding region can include a non-specific target nucleic acid sequence. A non-specific target nucleic acid sequence can refer to a sequence that can bind to a plurality of target nucleic acids independently of a specific sequence of the target nucleic acid. For example, the target binding region can include a random multimer sequence or an oligo(dT) sequence that hybridizes to a poly(A) tail on an mRNA molecule. A random multimer sequence can be, for example, a random dimer, trimer, tetramer, pentamer, hexamer, heptamer, octamer, nonamer, decamer, or a multimer sequence higher than any of these lengths. In some embodiments, the target binding region is the same for all barcodes attached to a given bead. In some embodiments, the target binding regions of a plurality of barcodes attached to a given bead can include two or more different target binding sequences. The target binding region can be of a length of 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides, or a number or range between any two of these values, or can be approximately these values or ranges. The target binding region can be at most about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length.
[0104] In some embodiments, the target binding region can include oligo(dT) that can hybridize with mRNA containing a polyadenylated end. The target binding region can be gene-specific. For example, the target binding region can be configured to hybridize to a specific region of the target. The target binding region can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a number or range between any two of these values of nucleotide lengths, or can be about these values or ranges. The target binding region can be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The target binding region can be about 5 to about 30 nucleotides in length. When the barcode includes a gene-specific target binding region, herein, the barcode can be referred to as a gene-specific barcode.
[0105] Orientation characteristics The barcode can include one or more orientation characteristics that can be used for barcode orientation (e.g., alignment). The barcode can include an isoelectric focusing moiety. Different barcodes can include different isoelectric focusing points. When these barcodes are introduced into a sample, the sample can undergo isoelectric focusing to orient the barcodes as known. In this way, the orientation characteristics can be used to create a known map of barcodes in the sample. Exemplary orientation characteristics can include electrophoretic mobility (e.g., based on the size of the barcode), isoelectric point, spin, conductivity, and / or self-assembly. For example, a barcode having a self-assembly orientation characteristic can self-assemble into a specific orientation (e.g., a nucleic acid nanostructure).
[0106] Affinity characteristics A barcode can include one or more affinity characteristics. For example, a spatial label can include an affinity characteristic. An affinity characteristic can include a chemical and / or biological moiety that can facilitate the binding of the barcode to another entity (e.g., a cell receptor). For example, an affinity characteristic can include an antibody, e.g., an antibody specific for a particular moiety (e.g., a receptor) on a sample. In some embodiments, the antibody can guide the barcode to a specific cell type or molecule. Targets in and / or in the vicinity of a specific cell type or molecule can be labeled probabilistically. An affinity characteristic, in some embodiments, can provide spatial information in addition to the nucleotide sequence of the spatial label because the antibody can guide the barcode to a specific location. The antibody can be a therapeutic antibody, e.g., a monoclonal or polyclonal antibody. The antibody can be humanized or chimerized. The antibody can be a naked antibody or a fusion antibody.
[0107] The antibody can be an immunoglobulin molecule (e.g., an IgG antibody) that is full-length (i.e., formed by a natural or normal immunoglobulin gene fragment recombination process) or an immunoreactive (i.e., specific binding) portion of an immunoglobulin molecule such as an antibody fragment.
[0108] Antibody fragments can be, for example, parts of antibodies such as F(ab’)2, Fab’, Fab, Fv, sFv, etc. In some embodiments, the antibody fragments can bind to the same antigen recognized by the full-length antibody. Antibody fragments can include isolated fragments consisting of variable regions of antibodies, such as “Fv” fragments consisting of variable regions of heavy and light chains and recombinant single-chain polypeptide molecules where the light and heavy variable regions are connected by a peptide linker (“scFv protein”). Exemplary antibodies can include, but are not limited to, antibodies against cancer cells, antibodies against viruses, antibodies that bind to cell surface receptors (CD8, CD34, CD45), and therapeutic antibodies.
[0109] Universal adapter primer Barcodes can include one or more universal adapter primers. For example, gene-specific barcodes such as gene-specific probabilistic barcodes can include universal adapter primers. A universal adapter primer can refer to a nucleotide sequence that is universal across all barcodes. Universal adapter primers can be used for constructing gene-specific barcodes. A universal adapter primer can be of a length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a number or range between any two of these values of nucleotide counts, or can be approximately these values or ranges. A universal adapter primer can be of a length of at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. A universal adapter primer can be of a length of 5 to 30 nucleotides.
[0110] Linker When the barcode contains two or more types of labels (e.g., two or more cell labels or two or more barcode sequences such as one molecular label), a linker label sequence may be interspersed between the labels. The linker label sequence can be at least about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, or more nucleotides in length. The linker label sequence can be at most about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, or more nucleotides in length. In some cases, the linker label sequence is 12 nucleotides in length. The linker label sequence can be used to facilitate the synthesis of the barcode. The linker label can include an error correction (e.g., Hamming) code.
[0111] Solid support In some embodiments, a solid support can be associated with barcodes such as probabilistic barcodes disclosed in this specification. The solid support can be, for example, synthetic particles. In some embodiments, some or all of the barcode sequences such as molecular labels of probabilistic barcodes (e.g., the first barcode array) of a plurality of barcodes (e.g., the first plurality of barcodes) on the solid support differ by at least one nucleotide. The cell labels of barcodes on the same solid support can be the same. The cell labels of barcodes on different solid supports can differ by at least one nucleotide. For example, the first cell label of the first plurality of barcodes on the first solid support can have the same sequence, and the second cell label of the second plurality of barcodes on the second solid support can have the same sequence. The first cell label of the first plurality of barcodes on the first solid support and the second cell label of the second plurality of barcodes on the second solid support can differ by at least one nucleotide. The cell label can be, for example, about 5 to about 20 nucleotides in length. The barcode sequence can be, for example, about 5 to about 20 nucleotides in length. The synthetic particles can be, for example, beads.
[0112] The beads can be, for example, silica gel beads, glass beads with controlled pores, magnetic beads, Dynabead, Sephadex / Sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The beads can include materials such as polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, latex, sepharose, cellulose, nylon, silicone, or any combination thereof.
[0113] In some embodiments, the beads can be polymer beads, such as deformable beads or gel beads functionalized with a barcode or probabilistic barcode (such as gel beads from 10X Genomics, San Francisco, CA). In some embodiments, the gel beads can comprise a polymeric gel. Gel beads can be generated, for example, by encapsulating one or more polymer precursors in droplets. Exposure of the polymer precursors to a promoter (such as tetramethylethylenediamine (TEMED)) can result in the formation of gel beads.
[0114] In some embodiments, the particles can be degradable. For example, the polymer beads can dissolve, melt, or degrade under desired conditions, for example. The desired conditions can include environmental conditions. The desired conditions can cause the polymer beads to dissolve, melt, or degrade in a controlled manner. Gel beads can dissolve, melt, or degrade due to chemical, physical, biological, thermal, magnetic, electrical, light, or any combination thereof.
[0115] Analytes and / or reagents, such as oligonucleotide barcodes, can be bound / fixed to the inner surface of the gel beads (e.g., the accessible interior via diffusion of the oligonucleotide barcode and / or the materials used to generate the oligonucleotide barcode) and / or the outer surface of the gel beads or any other microcapsules described herein. The binding / fixation can be via any form of chemical bond (e.g., covalent bond, ionic bond) or physical phenomenon (e.g., van der Waals forces, dipole-dipole interactions, etc.). In some embodiments, the binding / fixation of the reagent to the gel beads or any other microcapsules described herein can be reversible, for example, via a labile moiety (e.g., via a chemical crosslinking agent including the chemical crosslinking agents described herein). Upon application of a stimulus, the labile moiety can cleave, releasing the immobilized reagent. In some embodiments, the labile moiety is a disulfide bond. For example, when an oligonucleotide barcode is immobilized on a gel bead via a disulfide bond, exposure of the disulfide bond to a reducing agent can cleave the disulfide bond and release the oligonucleotide barcode from the bead. The labile moiety can be included as part of the gel bead or microcapsule, as part of a chemical linker that links the reagent or analyte to the gel bead or microcapsule, and / or as part of the reagent or analyte. In some embodiments, at least one of the plurality of barcodes can be immobilized on the particle, partially immobilized on the particle, encapsulated within the particle, partially encapsulated within the particle, or any combination thereof.
[0116] In some embodiments, the gel beads can include a wide range of different polymers including, but not limited to, polymers, thermosensitive polymers, photosensitive polymers, magnetic polymers, pH-sensitive polymers, salt-sensitive polymers, chemically sensitive polymers, polyelectrolytes, polysaccharides, peptides, proteins, and / or plastics. The polymers can include, but are not limited to, materials such as poly(N-isopropylacrylamide) (PNIPAAm), poly(styrenesulfonate) (PSS), poly(allylamine) (PAAm), poly(acrylic acid) (PAA), poly(ethyleneimine) (PEI), poly(diallyldimethylammonium chloride) (PDADMAC), poly(pyrrole) (PPy), poly(vinylpyrrolidone) (PVPON), poly(vinylpyridine) (PVP), poly(methacrylic acid) (PMAA), poly(methylmethacrylate) (PMMA), polystyrene (PS), poly(tetrahydrofuran) (PTHF), poly(phthalaldehyde) (PTHF), poly(hexyl viologen) (PHV), poly(L-lysine) (PLL), poly(L-arginine) (PARG), poly(lactic-co-glycolic acid) (PLGA), and the like.
[0117] Many chemical stimuli can be used to trigger the disintegration, dissolution, or decomposition of the beads. Examples of these chemical changes can include, but are not limited to, pH-mediated alterations to the bead wall, disintegration of the bead wall via chemical cleavage of cross-linkages, triggering of depolymerization of the bead wall, and bead wall switching reactions. Bulk changes can also be used to trigger the disintegration of the beads.
[0118] Bulk or physical alterations to the microcapsules through various stimuli also offer many advantages in designing the capsules to release the reagent. The bulk or physical alterations occur on a macroscopic scale, and the rupture of the beads is a result of the mechanical-physical forces induced by the stimulus. These processes can include, but are not limited to, pressure-induced rupture, bead wall melting, or changes in the porosity of the bead wall.
[0119] Biological stimuli can also be used as triggers for the disintegration, dissolution, or decomposition of the beads. Generally, biological stimuli are similar to chemical triggers, but many examples use molecules commonly found in biological systems such as biomolecules or enzymes, peptides, saccharides, fatty acids, nucleic acids, etc. For example, the beads may contain a polymer having a peptide cross-link that is susceptible to cleavage by a specific protease. More specifically, one example may include microcapsules containing a GFLGK peptide cross-link. When a biological trigger such as the protease Cathepsin B is added, the peptide cross-link in the shell wall is cleaved and the contents of the beads are released. In other cases, the protease can be heat-activated. In another example, the beads are provided with a shell wall containing cellulose. The addition of chitosan, a hydrolytic enzyme, functions as a biological trigger for the cleavage of cellulose bonds, the depolymerization of the shell wall, and the release of its contents.
[0120] The application of a thermal stimulus can also be induced to trigger the release of the contents of the beads. Temperature changes can cause various changes in the beads. Thermal changes can melt the beads such that the bead walls disintegrate. In other cases, heat can increase the internal pressure of the internal components of the beads such that the beads disintegrate or explode. In still other cases, heat can convert the beads to a shrinkage and dehydration state. Heat can also act on the heat-sensitive polymer within the bead walls to cause the beads to disintegrate.
[0121] By incorporating magnetic nanoparticles into the bead walls of the microcapsules, the disintegration of the beads can be triggered and the beads can be guided within the array. The devices of the present disclosure can include magnetic beads for any purpose. In one example, the incorporation of Fe3O4 nanoparticles into a polyelectrolyte containing the beads triggers disintegration in the presence of an oscillating magnetic field stimulus.
[0122] Beads can also disintegrate, dissolve, or decompose as a result of electrical stimulation. Similar to the magnetic particles described in the previous section, beads that are susceptible to the effects of electricity can enable both the triggering of bead disintegration and other functions such as alignment in an electric field, conductivity, or redox reactions. In one example, beads containing a material susceptible to the effects of electricity align in an electric field so as to be able to control the release of internal reagents. In another example, an electric field can induce a redox reaction within the bead wall itself, which can increase porosity.
[0123] Light stimulation can also be used for bead disintegration. Many light triggers are possible and can include systems that use various molecules such as nanoparticles and chromophores that can absorb photons in a specific range of wavelengths. For example, a metal oxide coating can be used as a capsule trigger. UV irradiation of a polyelectrolyte capsule coated with SiO2 can cause the bead wall to disintegrate. In yet another example, a light-switchable material such as an azobenzene group can be incorporated into the bead wall. When UV or visible light is applied, these types of chemicals undergo a reversible cis-to-trans isomerization upon absorption of photons. In this manner, the incorporation of a photon switch results in the generation of a bead wall that can disintegrate or become more porous when a light trigger is applied.
[0124] For example, in a non-limiting example of barcoding (e.g., probabilistic barcoding) shown in Figure 2, in block 208, after introducing cells such as a single cell into a plurality of microwells of a microwell array, in block 212, beads can be introduced into the plurality of microwells of the microwell array. Each microwell can contain one bead. The beads can contain a plurality of barcodes. The barcode can contain a 5’ amine region attached to the bead. The beecode can contain a universal label, a barcode sequence (e.g., a molecular label), a target binding region, or any combination thereof.
[0125] The barcodes disclosed herein can be associated with (e.g., attached to) a solid support (e.g., beads). Each barcode associated with a solid support can include a barcode selected from a group of at least 100 or 1000 barcode sequences having unique sequences. In some embodiments, different barcodes associated with a solid support can include barcode sequences of different sequences. In some embodiments, a certain percentage of the barcodes associated with a solid support can include the same cell label. For example, the percentage can be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100% or a number or range between any two of these values, or about these values or ranges. As another example, the percentage can be at least or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, the barcodes associated with a solid support can have the same cell label. Barcodes associated with different solid supports can have different cell labels selected from a group of at least 100 or 1000 cell labels having unique sequences.
[0126] The barcodes disclosed herein can be associated with (e.g., attached to) a solid support (e.g., beads). In some embodiments, probabilistically barcoding multiple targets in a sample can be performed using a solid support comprising a plurality of synthetic particles associated with a plurality of barcodes. In some embodiments, the solid support can comprise a plurality of synthetic particles associated with a plurality of barcodes. The spatial labeling of the plurality of barcodes on different solid supports can differ by at least one nucleotide. The solid support can comprise, for example, a plurality of barcodes in two or three dimensions. The synthetic particles can be beads. The beads can be silica gel beads, glass beads with controlled pores, magnetic beads, Dynabeads, Sephadex / Sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The solid support can comprise a polymer, a matrix, a hydrogel, a needle array device, an antibody, or any combination thereof. In some embodiments, the solid support can float freely. In some embodiments, the solid support can be incorporated into a semi-solid or solid array. It may not be necessary to associate a solid support with the barcode. The barcode can be individual nucleotides. A substrate can be associated with the barcode.
[0127] As used herein, the terms “tethered,” “attached,” or “immobilized” are used synonymously and can refer to covalent or non-covalent means of attaching a barcode to a solid support. Any of a variety of different solid supports can be used to attach a pre-synthesized barcode or as a solid support for in situ solid-phase synthesis of a barcode.
[0128] In some embodiments, the solid support is a bead. The bead can be one or more types of solid, porous, or hollow spheres, balls, bearings, cylinders, or other similar configurations capable of immobilizing (e.g., by covalent or non-covalent bonding) nucleic acids. The beads can be composed of, for example, plastic, ceramic, metal, polymeric materials, or any combination thereof. The beads can be discrete particles that are spheres (e.g., small spheres), or can contain or have non-spherical or irregular shapes such as cubes, cubic bones, pyramids, cylinders, cones, ellipses, or discs. In some embodiments, the beads can be non-spherical.
[0129] The beads can include a variety of materials, including but not limited to paramagnetic materials (e.g., magnesium, molybdenum, lithium, and tantalum), superparamagnetic materials (e.g., ferrite (Fe3O4; magnetite) nanoparticles), ferromagnetic materials (e.g., iron, nickel, cobalt, any alloys thereof, and any rare earth metal compounds), ceramics, plastics, glass, polystyrene, silica, methylstyrene, acrylic polymers, titanium, latex, sepharose, agarose, hydrogels, polymers, cellulose, nylon, or any combination thereof.
[0130] In some embodiments, the beads (e.g., beads to which a label is attached) are hydrogel beads. In some embodiments, the beads contain a hydrogel.
[0131] Some embodiments disclosed herein include one or more particles (e.g., beads). Each particle can include a plurality of oligonucleotides (e.g., barcodes). Each of the plurality of oligonucleotides can include a barcode sequence (e.g., a molecular label), a cell label, and a target binding region (e.g., an oligo(dT) sequence, a gene-specific sequence, a random multimer, or a combination thereof). The cell label sequences of each of the plurality of oligonucleotides can be the same. The cell label sequences of the oligonucleotides on different particles can be different so as to be able to distinguish the oligonucleotides on different particles. The number of different cell label sequences can be different in different embodiments. In some embodiments, the number of cell label sequences is 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 10 7 10 8 10 9 times, or a number or range between any two of these values, or about these values or ranges, or more than 10 9 10 6 10 7 10 8 10 9can be the number. In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more than that number of the plurality of particles contain oligonucleotides having the same cell sequence. In some embodiments, the plurality of particles containing oligonucleotides having the same cell sequence can be at most 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or more than that percentage. In some embodiments, none of the plurality of particles has the same cell labeling sequence.
[0132] The plurality of oligonucleotides on each particle can contain different barcode sequences (e.g., molecular labels). In some embodiments, the number of barcode sequences is 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 、10 7 、10 8 、10 9 in number, or a number or range between any two of these values, or about these values or ranges. In some embodiments, the number of barcode sequences is at least or at most 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 、10 7 、10 8 、or 10 9It can be the number of. For example, at least 100 of the plurality of oligonucleotides contain different barcode sequences. As another example, in one particle, at least 100, 500, 1000, 5000, 10000, 15000, 20000, 50000, or a number or range between any two of these values or more than 50000 of the plurality of oligonucleotides contain different barcode sequences. Some embodiments provide a plurality of particles containing barcodes. In some embodiments, the ratio of the occurrence (or copy or number) of the target to be labeled to different barcode sequences can be at least 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, or a ratio exceeding that. In some embodiments, each of the plurality of oligonucleotides further contains a sample label, a universal label, or both. The particles can be, for example, nanoparticles or microparticles.
[0133] The size of the beads can be various. For example, the diameter of the beads can range from 0.1 μm to 50 μm. In some embodiments, the diameter of the beads can be 0.1 μm, 0.5 μm, 1 μm, 2 μm, 3 μm, 4 μm, 5 μm, 6 μm, 7 μm, 8 μm, 9 μm, 10 μm, 20 μm, 30 μm, 40 μm, 50 μm or a number or range between any two of these values, or approximately these values or ranges.
[0134] The diameter of the beads can be related to the diameter of the wells of the substrate. In some embodiments, the diameter of the beads can be a value longer or shorter than the diameter of the wells by 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or a number or range between any two of these values, about these values or ranges. The diameter of the beads can be related to the diameter of a cell (e.g., one cell captured by the wells of the substrate). In some embodiments, the diameter of the beads can be a value at least or at most 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% longer or shorter than the diameter of the wells. The diameter of the beads can be related to the diameter of a cell (e.g., one cell captured by the wells of the substrate). In some embodiments, the diameter of the beads can be a value longer or shorter than the diameter of the wells by 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, 300%, or a number or range between any two of these values, about these values or ranges. In some embodiments, the diameter of the beads can be a value at least or at most 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, or 300% longer or shorter than the diameter of the wells.
[0135] The beads can be attached to and / or embedded in the substrate. The beads can be attached to and / or embedded in a gel, hydrogel, polymer, and / or matrix. The spatial position of the beads within the substrate (e.g., a gel, matrix, scaffold, or polymer) can be identified using spatial markers present on barcodes on the beads that can function as location addresses.
[0136] Examples of beads include, but are not limited to, streptavidin beads, agarose beads, magnetic beads, Dynabeads®, MACS® microbeads, antibody-conjugated beads (e.g., anti-immunoglobulin microbeads), protein A-conjugated beads, protein G-conjugated beads, protein A / G-conjugated beads, protein L-conjugated beads, oligo(dT)-conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag™ carboxyl-terminal magnetic beads.
[0137] Beads can be associated with (e.g., impregnated with) quantum dots or fluorescent dyes and made to fluoresce in one or more fluorescent optical channels. Beads can be associated with iron oxide or chromium oxide to make them paramagnetic or ferromagnetic. Beads can be made distinguishable. For example, beads can be imaged using a camera. Beads can have a detectable code associated with them. For example, beads can contain a barcode. Beads can change size, for example, due to swelling in an organic or inorganic solution. Beads can be hydrophobic. Beads can be hydrophilic. Beads can be biocompatible.
[0138] Solid supports (e.g., beads) can be visualized. Solid supports can contain visualization tags (e.g., fluorescent dyes). Identifiers (e.g., numbers) can be etched onto solid supports (e.g., beads). The identifiers can be visualized through imaging of the beads.
[0139] The solid support can contain insoluble, semi-soluble, or insoluble substances. When the solid support contains a linker, scaffold, building block, or other reactive moiety attached thereto, it can be called "functionalized", while when there is no such attached reactive moiety, it can be called "non-functionalized". The solid support can be in the form of a microtiter well format; a flow-through format such as in a column; or freely available in a solution such as a display stick.
[0140] The solid support can include a membrane, paper, plastic, coated surface, plane, glass, slide, chip, or any combination thereof. The solid support can take the form of a resin, gel, small spheres, or other geometric configurations. The solid support includes planar supports such as silica chips, microparticles, nanoparticles, plates, arrays, capillaries, glass fiber filters, glass surfaces, metal surfaces (steel, gold, silver, aluminum, silicon, and copper), glass supports, plastic supports, silicon supports, chips, filters, membranes, microtiter plates, slides, multi-well plates or membranes (e.g., formed of polyethylene, polypropylene, polyamide, polyvinylidene fluoride), and / or wafers, combs, pins, or needles (e.g., an array of pins suitable for combinatorial synthesis or analysis), or an array of pins or a wafer (e.g., a silicon wafer), beads within planar nanoliter wells having pins with or without a filter bottom.
[0141] The solid support can contain a polymer matrix (e.g., a gel, hydrogel). The polymer matrix can be penetrable into intracellular spaces (e.g., around organelles). The polymer matrix can be pumpable throughout the circulatory system.
[0142] The solid support can be a biomolecule. For example, the solid support can be a nucleic acid, protein, antibody, histone, cell compartment, lipid, carbohydrate, etc. The solid support that is a biomolecule can be amplified, translated, transcribed, degraded, and / or modified (e.g., pegylated, SUMOylated, acetylated, methylated). The solid support that is a biomolecule can provide spatial information and temporal information in addition to the spatial label attached to the biomolecule. For example, the biomolecule can include a first confirmation when unmodified, but can change to a second confirmation when modified. Different structures can expose the barcodes (e.g., probabilistic barcodes) of the present disclosure to the target. For example, the biomolecule can include barcodes that are inaccessible due to the folding of the biomolecule. When the biomolecule is modified (e.g., acetylated), the biomolecule can change its structure to expose the barcode. The timing of the modification can provide another time dimension to the barcoding method of the present disclosure.
[0143] In some embodiments, the biomolecule comprising the barcode reagent of the present disclosure can be placed within the cytoplasm of a cell. When activated, the biomolecule can move to the cell nucleus, where barcoding can be performed. In this way, the modification of the biomolecule can encode additional spatio-temporal information of the target identified by the barcode.
[0144] Substrate and Microwell Array As used herein, a substrate can refer to a type of solid support. The substrate can include the barcodes and stochastic barcodes of the present disclosure and can refer to a solid support. The substrate can include, for example, a plurality of microwells. For example, the substrate can be a well array including two or more microwells. In some embodiments, a microwell can include a small reaction chamber of a defined volume. In some embodiments, a microwell can incorporate one or more cells. In some embodiments, a microwell can incorporate only one cell. In some embodiments, a microwell can incorporate one or more solid supports. In some embodiments, a microwell can incorporate only one solid support. In some embodiments, a microwell incorporates one cell and one solid support (e.g., beads). A microwell can include the combinatorial barcode reagent of the present disclosure.
[0145] Method of barcoding The present disclosure provides a method for estimating the number of distinct targets at distinct locations in a physical sample (e.g., tissue, organ, tumor, cell). The method can include placing barcodes (e.g., probabilistic barcodes) in the vicinity of the sample, lysing the sample, associating the barcodes with the distinct targets, amplifying the targets, and / or digitally counting the targets. The method can further include analyzing and / or visualizing information obtained from spatial markers on the barcodes. In some embodiments, the method includes visualizing a plurality of targets in the sample. Mapping the plurality of targets to a map of the sample can include generating a two-dimensional map or a three-dimensional map of the sample. The two-dimensional map and the three-dimensional map can be generated before or after barcoding (e.g., probabilistically barcoding) the plurality of targets in the sample. Visualizing the plurality of targets in the sample can include mapping the plurality of targets to a map of the sample. Mapping the plurality of targets to a map of the sample can include generating a two-dimensional map or a three-dimensional map of the sample. The two-dimensional map and the three-dimensional map can be generated before or after barcoding the plurality of targets in the sample. In some embodiments, the two-dimensional map and the three-dimensional map can be generated before or after lysing the sample. Lysing the sample before or after generating the two-dimensional map or the three-dimensional map can include heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof.
[0146] In some embodiments, barcoding the plurality of targets includes hybridizing a plurality of barcodes to the plurality of targets to create barcoded targets (e.g., probabilistically barcoded targets). Barcoding the plurality of targets can include generating an indexed library of barcoded targets. Generating the indexed library of barcoded targets can be performed using a solid support that includes a plurality of barcodes (e.g., probabilistic barcodes).
[0147] Contact between the sample and the barcode The present disclosure provides a method of contacting a sample (e.g., a cell) with a substrate of the present disclosure. For example, a sample containing a thin section of cells, an organ, or a tissue can be contacted with a barcode (e.g., a probabilistic barcode). The cells can be contacted, for example, by gravity flow in which the cells settle and form a monolayer. The sample can be a thin section of tissue. The thin section can be placed on the substrate. The sample can be one-dimensional (e.g., forming a plane). The sample (e.g., a cell) can be spread across the substrate, for example, by growing / culturing the cells on the substrate.
[0148] When the barcode is in the vicinity of the target, the target can hybridize to the barcode. The barcode can be contacted at a non-depletable ratio such that a separate barcode of the present disclosure can be associated with each separate target. To ensure an efficient association between the target and the barcode, the target can be cross-linked to the barcode.
[0149] Cell lysis Following the dispensing of the cells and the barcode, the cells can be lysed to release the target molecules. Cell lysis can be achieved by any of a variety of means, such as chemical or biological means, osmotic shock, or thermal lysis, mechanical lysis, or optical lysis. The cells can be lysed by the addition of a cell lysis buffer containing a detergent (e.g., SDS, lithium dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. To increase the association between the target and the barcode, the diffusion rate of the target molecules can be altered, for example, by lowering the temperature of the lysate and / or increasing the viscosity of the lysate.
[0150] In some embodiments, the sample can be lysed using filter paper. The filter paper can be immersed with its upper part in a lysis buffer. The filter paper can be applied to the sample at a pressure that can facilitate lysis of the sample and hybridization of the sample target to the substrate.
[0151] In some embodiments, lysis can be performed by mechanical lysis, thermal lysis, optical lysis, and / or chemical lysis. Chemical lysis can include the use of digestive enzymes such as Proteinase K, pepsin, and trypsin. Lysis can be performed by adding a lysis buffer to the substrate. The lysis buffer can include Tris-HCl. The lysis buffer can include at least about 0.01M, 0.05M, 0.1M, 0.5M, 1M, or more Tris-HCl. The lysis buffer can include at most about 0.01M, 0.05M, 0.1M, 0.5M, 1M, or more Tris-HCl. The lysis buffer can include about 0.1M Tris-HCl. The pH of the lysis buffer can be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. The pH of the lysis buffer can be at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In some embodiments, the pH of the lysis buffer is about 7.5. The lysis buffer can include a salt (e.g., LiCl). The concentration of the salt in the lysis buffer can be at least about 0.1M, 0.5M, 1M, or more. The concentration of the salt in the lysis buffer can be at most about 0.1M, 0.5M, 1M, or more. In some embodiments, the concentration of the salt in the lysis buffer is about 0.5M. The lysis buffer can include a detergent (e.g., SDS, lithium dodecyl sulfate, Triton X-100, Tween-20, NP-40). The concentration of the detergent in the lysis buffer can be at least about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, or more. The concentration of the detergent in the lysis buffer can be at most about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, or more. In some embodiments, the concentration of the detergent in the lysis buffer can be about 1% lithium dodecyl sulfate.The time used in the dissolution method can depend on the amount of detergent used. In some embodiments, the more detergent used, the shorter the time required for dissolution. The lysis buffer can contain a chelating agent (e.g., EDTA, EGTA). The concentration of the chelating agent in the lysis buffer can be at least about 1 mM, 5 mM, 10 mM, 15 mM, 20 mM, 25 mM, 30 mM, or a concentration exceeding that. The concentration of the chelating agent in the lysis buffer can be at most about 1 mM, 5 mM, 10 mM, 15 mM, 20 mM, 25 mM, 30 mM, or a concentration exceeding that. In some embodiments, the concentration of the chelating agent in the lysis buffer is about 10 mM. The lysis buffer can contain a reducing agent (e.g., β-mercaptoethanol, DTT). The concentration of the reducing agent in the lysis buffer can be at least about 1 mM, 5 mM, 10 mM, 15 mM, 20 mM, or a concentration exceeding that. The concentration of the reducing agent in the lysis buffer can be at most about 1 mM, 5 mM, 10 mM, 15 mM, 20 mM, or a concentration exceeding that. In some embodiments, the concentration of the reducing agent in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer can contain about 0.1 M Tris-HCl, about pH 7.5, about 0.5 M LiCl, about 1% lithium dodecyl sulfate, about 10 mM EDTA and about 5 mM MDTT.
[0152] Lysis can be performed at a temperature of about 4°C, 10°C, 15°C, 20°C, 25°C, 30°C. Lysis can be performed over a time of about 1 minute, 5 minutes, 10 minutes, 15 minutes, 20 minutes, or more. The lysed cells can contain at least about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, or more target nucleic acid molecules. The lysed cells can contain at most about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, or more target nucleic acid molecules.
[0153] Attachment of barcodes to target nucleic acid molecules Following lysis of the cells and release of the nucleic acid molecules therefrom, the nucleic acid molecules can be randomly associated with the barcodes of the solid support in coexistence. The association can include hybridizing the target recognition region of the barcode to the complementary portion of the target nucleic acid molecule (e.g., the oligo(dT) of the barcode can interact with the poly(A) tail of the target). The assay conditions used for hybridization (e.g., buffer pH, ionic strength, temperature, etc.) can be selected to promote the formation of specific stable hybrids. In some embodiments, the nucleic acid molecules released from the lysed cells can be associated with (e.g., hybridized to) a plurality of probes on the substrate. When the probe contains oligo(dT), the mRNA molecule can be hybridized to the probe and reverse transcribed. The oligo(dT) portion of the oligonucleotide can function as a primer for the first-strand synthesis of the cDNA molecule. For example, in the non-limiting example of barcoding shown in Figure 2, in block 216, the mRNA molecule can be hybridized to the barcode on the bead. For example, a single-stranded nucleotide fragment can be hybridized to the target binding region of the barcode.
[0154] Attachment can further include ligation of the target recognition region of the barcode and a portion of the target nucleic acid molecule. For example, the target binding region can include a nucleic acid sequence that allows specific hybridization to a restriction site overhang (e.g., an EcoRI sticky end overhang). The assay procedure can further include treating the target nucleic acid with a restriction enzyme (e.g., EcoRI) to create a restriction site overhang. Next, the barcode can be ligated to any nucleic acid molecule containing a sequence complementary to the restriction site overhang. Two fragments can be joined using a ligase (e.g., T4 DNA ligase).
[0155] For example, in a non-limiting example of barcoding shown in FIG. 2, in block 220, labeled targets from a plurality of cells (or a plurality of samples) (e.g., target barcode molecules) can subsequently be pooled, for example, in a tube. The labeled targets can be pooled, for example, by collecting the barcode and / or beads to which the target barcode molecules are attached.
[0156] Recovery of the solid support-based collection of attached target barcode molecules can be carried out by use of magnetic beads and an externally applied magnetic field. Once the target barcode molecules are pooled, all further processing can proceed within a single reaction vessel. Further processing can include, for example, reverse transcription reactions, amplification reactions, cleavage reactions, separation reactions, and / or nucleic acid extension reactions. The further processing reactions can be carried out within microwells, i.e., without first pooling the labeled target nucleic acid molecules from a plurality of cells.
[0157] Reverse transcription The present disclosure provides a method of creating target barcode conjugates using reverse transcription (e.g., in block 224 of FIG. 2). The target barcode conjugates can include a barcode and all or a partial complementary sequence of the target nucleic acid (i.e., a barcoded cDNA molecule such as a stochastic barcode-attached cDNA molecule). Reverse transcription of the associated RNA molecule can be carried out by addition of a reverse transcription primer together with a reverse transcriptase. The reverse transcription primer can be an oligo(dT) primer, a random hexamer primer, or a target-specific oligonucleotide primer. The oligo(dT) primer can be 12 to 18 nucleotides or about 12 to 18 nucleotides in length and can bind to the endogenous poly(A) tail at the 3' end of mammalian mRNA. The random hexamer primer can bind to mRNA at a variety of complementary sites. The target-specific oligonucleotide primer typically selectively primes the mRNA of interest.
[0158] In some embodiments, reverse transcription of the labeled RNA molecule can be performed by addition of a reverse transcription primer. In some embodiments, the reverse transcription primer is an oligo(dT) primer, a random hexamer primer, or a target-specific oligonucleotide primer. Generally, an oligo(dT) primer is 12 to 18 nucleotides in length and binds to the endogenous poly(A) tail at the 3' end of mammalian mRNA. A random hexamer primer can bind to mRNA at a wide variety of complementary sites. A target-specific oligonucleotide primer typically selectively primes the mRNA of interest.
[0159] Reverse transcription can be performed repeatedly to generate multiple labeled cDNA molecules. The methods disclosed herein can include performing the reverse transcription reaction at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 times. The methods can include performing the reverse transcription reaction at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 times.
[0160] Amplification One or more nucleic acid amplification reactions (e.g., in block 228 of FIG. 2) can be performed to create multiple copies of the labeled target nucleic acid molecule. The amplification can be performed multiplexed, and multiple target nucleic acid sequences are amplified simultaneously. Amplification reactions can be used to add sequencing adapters to the nucleic acid molecules. The amplification reaction can include amplifying at least a portion of the sample label if present. The amplification reaction can include amplifying at least a portion of the cell label and / or barcode sequence (e.g., molecular label). The amplification reaction can include amplifying at least a portion of the sample tag, cell label, spatial label, barcode (e.g., molecular label), target nucleic acid, or combinations thereof. The amplification reaction can include amplifying 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100%, or a range or number between any two of those values of the plurality of nucleic acids. The method can further include performing one or more cDNA synthesis reactions to generate one or more cDNA copies of the target barcode molecule including the sample label, cell label, spatial label, and / or barcode sequence (e.g., molecular label).
[0161] In some embodiments, the amplification can be performed using polymerase chain reaction (PCR). As used herein, PCR can refer to a reaction that amplifies a specific DNA sequence in vitro by simultaneous primer extension of the complementary strands of DNA. As used herein, PCR can include derivatives of the reaction, including but not limited to RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR.
[0162] The amplification of the labeled nucleic acid can also include non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification or circle-to-circle amplification. Other non-PCR-based amplification methods include multi-cycle DNA-dependent RNA polymerase-induced RNA transcription amplification or RNA-dependent DNA synthesis and transcription that amplify DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ) method, use of palindromic probes, strand displacement amplification, oligonucleotide-induced amplification using restriction endonucleases, amplification methods where primers hybridize to nucleic acid sequences and the resulting double-stranded molecules are cleaved prior to the extension reaction and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and ramification extension amplification (RAM). In some embodiments, the amplification does not generate circular transcripts.
[0163] In some embodiments, the methods disclosed herein further comprise performing a polymerase chain reaction on a labeled nucleic acid (e.g., labeled RNA, labeled DNA, labeled cDNA) to generate a labeled amplification product (e.g., a stochastically labeled amplification product). The labeled amplification product can be a double-stranded molecule. The double-stranded molecule can include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule can include a sample label, a spatial label, a cell label, and / or a barcode sequence (e.g., a molecular label). The labeled amplification product can be a single-stranded molecule. The single-stranded molecule can include DNA, RNA, or a combination thereof. The nucleic acids of the present disclosure can include synthetic nucleic acids and modified nucleic acids.
[0164] Amplification can include the use of one or more unnatural nucleotides. The unnatural nucleotides can include photosensitive or triggerable nucleotides. Examples of unnatural nucleotides can include, but are not limited to, peptide nucleic acids (PNAs), morpholinos, and locked nucleic acids (LNAs), as well as glycol nucleic acids (GNAs) and threose nucleic acids (TNAs). The unnatural nucleotides can be added to one or more cycles of the amplification reaction. The addition of the unnatural nucleotides can be used to identify the product as a specific cycle or time point in the amplification reaction.
[0165] Performing one or more amplification reactions can involve the use of one or more primers. The one or more primers can include, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more nucleotides. The one or more primers can include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more nucleotides. The one or more primers can include from 12 to less than 15 nucleotides. The one or more primers can anneal to at least a portion of a plurality of labeled targets (e.g., probabilistically labeled targets). The one or more primers can anneal to the 3' and 5' ends of a plurality of labeled targets. The one or more primers can anneal to an internal region of a plurality of labeled targets. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of a plurality of labeled targets. The one or more primers can include a primer fixation panel. The one or more primers can include at least one or more custom primers. The one or more primers can include at least one or more control primers. The one or more primers can include at least one or more gene-specific primers.
[0166] One or more primers can include universal primers. Universal primers can anneal to a universal primer binding site. One or more custom primers can anneal to a first sample label, a next sample label, a spatial label, a cell label, a barcode sequence (e.g., a molecular label), a target, or any combination thereof. One or more primers can include universal primers and custom primers. Custom primers can be designed to amplify one or more targets. Targets can include a subset of the total nucleic acids in one or more samples. Targets can include a subset of the total labeled targets in one or more samples. One or more primers can include at least 96 or more custom primers. One or more primers can include at least 960 or more custom primers. One or more primers can include at least 9600 or more custom primers. One or more custom primers can anneal to two or more different labeled nucleic acids. Two or more different labeled nucleic acids can correspond to one or more genes.
[0167] In the methods of the present disclosure, any amplification method can be used. For example, in one method, the first PCR can amplify the molecules attached to the beads using gene-specific primers and primers for the universal Illumina sequencing primer 1 sequence. The second PCR can amplify the product of the first PCR using nested gene-specific primers adjacent to the Illumina sequencing primer 2 sequence and primers for the universal Illumina sequencing primer 1 sequence. The third PCR adds P5 and P7 as well as a sample index to make the PCR product an Illumina sequencing library. Sequencing using 150bp×2 sequencing can reveal the cell label and barcode sequence (e.g., molecular label) in read 1, the gene in read 2, and the sample index in the index 1 read.
[0168] In some embodiments, the nucleic acid can be removed from the substrate using chemical cleavage. For example, chemical groups or modified bases present in the nucleic acid can be used to facilitate removal of the nucleic acid from the solid support. For example, an enzyme can be used to remove the nucleic acid from the substrate. For example, the nucleic acid can be removed from the substrate through digestion with a restriction endonuclease. For example, treatment of the nucleic acid containing dUTP or ddUTP with uracil-d-glycosylase (UDG) can be used to remove the nucleic acid from the substrate. For example, the nucleic acid can be removed from the substrate using an enzyme that performs nucleotide removal such as a base excision repair enzyme such as apurinic / apyrimidinic (AP) endonuclease. In some embodiments, the nucleic acid can be removed from the substrate using a photocleavable group and light. In some embodiments, a cleavable linker can be used to remove the nucleic acid from the substrate. For example, the cleavable linker can comprise at least one of biotin / avidin, biotin / streptavidin, biotin / neutravidin, Ig-protein A, a photosensitive linker, an acid- or base-labile linker group, or an aptamer.
[0169] When the probe is gene-specific, the molecule can hybridize to the probe and be reverse transcribed and / or amplified. In some embodiments, the nucleic acid can be amplified after it has been synthesized (e.g., reverse transcribed). Amplification can be performed multiplexed, and multiple target nucleic acid sequences can be amplified simultaneously. Amplification can involve adding sequencing adapters to the nucleic acid.
[0170] In some embodiments, amplification can be performed on a substrate using, for example, bridge amplification. cDNA can be homopolymer tailed to generate compatible ends for bridge amplification using oligo(dT) probes on the substrate. In bridge amplification, a primer complementary to the 3’ end of the template nucleic acid can be the first primer of each pair covalently attached to the solid particles. When a sample containing the template nucleic acid contacts the particles and one thermal cycle is performed, the template molecule can anneal to the first primer, which extends in the forward direction by addition of nucleotides to form a double-stranded molecule consisting of the template molecule and a newly formed DNA strand complementary to the template. In the heating step of the next cycle, the double-stranded molecule can be denatured to release the template molecule from the particles and leave the complementary DNA strand attached to the particles through the first primer. In the annealing step of annealing and the subsequent extension step, the complementary strand can hybridize to a second primer, which is complementary to a segment of the complementary strand at the location removed from the first primer. This hybridization can cause the complementary strand to form a bridge between the first and second primers, where the first primer is covalently immobilized and the second primer is immobilized by hybridization. In the extension step, the second primer extends in the reverse direction by addition of nucleotides to the same reaction mixture, thereby converting the bridge into a double-stranded bridge. Next, the next cycle is initiated and the double-stranded bridge can be denatured to generate two single-stranded nucleic acid molecules, each having one end attached to the particle surface via the first and second primers, respectively, and the other end not attached. In the annealing and extension steps of this second cycle, each strand can hybridize to a further complementary primer on the same particle that was not previously used to form a new single-stranded bridge. The two previously unused primers that are now hybridized extend to convert the two new bridges into double-stranded bridges.
[0171] The amplification reaction can include amplifying at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 100% of a plurality of nucleic acids.
[0172] The amplification of the labeled nucleic acid can include a PCR-based method or a non-PCR-based method. The amplification of the labeled nucleic acid can include exponential amplification of the labeled nucleic acid. The amplification of the labeled nucleic acid can include linear amplification of the labeled nucleic acid. Amplification can be performed by polymerase chain reaction (PCR). PCR can refer to the reaction of in vitro amplification of a specific DNA sequence by simultaneous primer extension of complementary strands of DNA. PCR can include derivatives of the reaction, including but not limited to RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, suppression PCR, semi-PCR suppression, and assembly PCR.
[0173] In some embodiments, the amplification of the labeled nucleic acid includes a non-PCR-based method. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other non-PCR-based amplification methods include multi-cycle DNA-dependent RNA polymerase-induced RNA transcription amplification or RNA-dependent DNA synthesis and transcription that amplify DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ) method, use of palindromic probes, strand displacement amplification, oligonucleotide-induced amplification using restriction endonucleases, amplification methods where primers hybridize to nucleic acid sequences and the resulting double-stranded products are cleaved prior to elongation reactions and amplification, strand displacement amplification using nucleic acid polymerases lacking 5' exonuclease activity, rolling circle amplification, and / or ramification extension amplification (RAM).
[0174] In some embodiments, the methods disclosed herein further comprise performing a nested polymerase chain reaction on the amplified amplification product (e.g., the target). The amplification product can be a double-stranded molecule. The double-stranded molecule can comprise a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule can comprise a sample tag or a molecular identifier label. Alternatively, the amplification product can be a single-stranded molecule. The single-stranded molecule can comprise DNA, RNA, or a combination thereof. The nucleic acids of the present disclosure can comprise synthetic nucleic acids and modified nucleic acids.
[0175] In some embodiments, the method comprises repeatedly amplifying a labeled nucleic acid to generate a plurality of amplification products. The methods disclosed herein can comprise performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Alternatively, the method can comprise performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions.
[0176] Amplification can further comprise adding one or more control nucleic acids to one or more samples comprising a plurality of nucleic acids. Amplification can further comprise adding one or more control nucleic acids to the plurality of nucleic acids. The control nucleic acid can comprise a control label.
[0177] Amplification can include the use of one or more unnatural nucleotides. The unnatural nucleotides can include photosensitive or triggerable nucleotides. Examples of unnatural nucleotides include, without limitation, peptide nucleic acids (PNAs), morpholinos, and locked nucleic acids (LNAs), as well as glycol nucleic acids (GNAs) and threose nucleic acids (TNAs). The unnatural nucleotides can be added to one or more cycles of the amplification reaction. The addition of unnatural nucleotides can be used to identify products as specific cycles or time points in the amplification reaction.
[0178] Performing one or more amplification reactions can include the use of one or more primers. One or more primers can include one or more oligonucleotides. One or more oligonucleotides can include at least about 7 to 9 nucleotides. One or more oligonucleotides can include less than 12 to 15 nucleotides. One or more primers can anneal to at least a portion of a plurality of labeled nucleic acids. One or more primers can anneal to the 3' end and / or 5' end of a plurality of labeled nucleic acids. One or more primers can anneal to an internal region of a plurality of labeled nucleic acids. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of a plurality of labeled nucleic acids. One or more primers can include a primer fixation panel. One or more primers can include at least one or more custom primers. One or more primers can include at least one or more control primers. One or more primers can include at least one or more housekeeping gene primers. One or more primers can include a universal primer. The universal primer can anneal to a universal primer binding site. One or more custom primers can anneal to a first sample tag, next sample tag, molecular identifier label, nucleic acid, or its product. One or more primers can include a universal primer and a custom primer.Custom primers can be designed to amplify one or more target nucleic acids. The target nucleic acids can include a subset of the total nucleic acids in one or more samples. In some embodiments, the primer is a probe attached to the array of the present disclosure.
[0179] In some embodiments, barcoding (e.g., stochastic barcoding) a plurality of targets in a sample further includes generating an indexed library of barcoded fragments. Barcode sequences of different barcodes (e.g., molecular labels of different stochastic barcodes) can be different from each other. Generating an indexed library of barcoded targets (e.g., stochastic barcoded targets) includes generating a plurality of indexed polynucleotides from a plurality of targets in the sample. For example, in the case of an indexed library of barcoded targets including a first indexed target and a second indexed target, the labeled region of the first indexed polynucleotide can be different from the labeled region of the second indexed polynucleotide by at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 nucleotides, or a number or range of numbers between any two of these values. In some embodiments, generating an indexed library of barcoded targets includes contacting a plurality of targets, e.g., mRNA molecules, with a plurality of oligonucleotides including a poly(T) region and a labeled region, and performing first-strand synthesis using reverse transcriptase to generate single-stranded labeled cDNA molecules each including a cDNA region and a labeled region, wherein the plurality of targets includes at least two mRNA molecules of different sequences and the plurality of oligonucleotides includes at least two oligonucleotides of different sequences. Generation of an indexed library of barcoded targets can further include amplifying the single-stranded labeled cDNA molecules to generate double-stranded labeled cDNA molecules, and performing nested PCR on the double-stranded labeled cDNA molecules to generate labeled amplification products. In some embodiments, the method can include generating adapter-labeled amplification products.
[0180] Probabilistic barcoding can use nucleic acid barcodes or tags to label individual nucleic acid (e.g., DNA or RNA) molecules. In some embodiments, it includes adding a DNA barcode or tag to a cDNA molecule when the cDNA molecule is generated from mRNA. Nested PCR can be performed to minimize PCR amplification bias. Adapters can be added, for example, in the case of sequencing using next-generation sequencing (NGS). For example, in block 232 of FIG. 2, the sequencing results can be used to identify cell labels, barcode sequences (e.g., molecular labels), and the sequences of nucleotide fragments of one or more copies of the target.
[0181] FIG. 3 is a schematic diagram showing a non-limiting and exemplary process for generating an indexed library of barcoded targets (e.g., probabilistically barcoded targets), such as mRNA. As shown in step 1, the reverse transcription process can encode a unique barcode sequence (e.g., molecular label), cell label, and universal PCR site to each mRNA. For example, by reverse transcribing RNA molecule 302 and hybridizing a set of barcodes (e.g., probabilistic barcodes) 310 to the poly(A) tail region 308 of RNA molecule 302, a labeled cDNA molecule 304 containing a cDNA region 306 can be generated. Each barcode 310 can include a target binding region, such as a poly(dT) region 312, a barcode sequence or molecular label 314, and a universal PCR region 316.
[0182] In some embodiments, the cell label can comprise from 3 to 20 nucleotides. In some embodiments, the barcode sequence (e.g., molecular label) can comprise from 3 to 20 nucleotides. In some embodiments, each of the plurality of stochastic barcodes further comprises one or more of a universal label and a cell label, the universal label being the same for the plurality of stochastic barcodes on the solid support, and the cell label being the same for the plurality of stochastic barcodes on the solid support. In some embodiments, the universal label can comprise from 3 to 20 nucleotides. In some embodiments, the cell label can comprise from 3 to 20 nucleotides.
[0183] In some embodiments, the labeling region 314 can include a barcode array or molecular label 318 and a cell label 320. In some embodiments, the labeling region 314 can include one or more of a universal label, a dimensional label, and a cell label. The barcode array or molecular label 318 can be one, two, three, four, five, six, seven, eight, nine, ten, twenty, thirty, forty, fifty, sixty, seventy, eighty, ninety, one hundred, or a number or range of numbers between any two of these values in length of nucleotides, can be about these numbers, at least these numbers, or at most these numbers in length. The cell label 320 can be one, two, three, four, five, six, seven, eight, nine, ten, twenty, thirty, forty, fifty, sixty, seventy, eighty, ninety, one hundred, or a number or range of numbers between any two of these values in length of nucleotides, can be about these numbers, at least these numbers, or at most these numbers in length. The universal label can be one, two, three, four, five, six, seven, eight, nine, ten, twenty, thirty, forty, fifty, sixty, seventy, eighty, ninety, one hundred, or a number or range of numbers between any two of these values in length of nucleotides, can be about these numbers, at least these numbers, or at most these numbers in length. The universal label can be the same for a plurality of probabilistic barcodes on the solid support, and the cell label can be the same for a plurality of probabilistic barcodes on the solid support. The dimensional label can be one, two, three, four, five, six, seven, eight, nine, ten, twenty, thirty, forty, fifty, sixty, seventy, eighty, ninety, one hundred, or a number or range of numbers between any two of these values in length of nucleotides, can be about these numbers, at least these numbers, or at most these numbers in length.
[0184] In some embodiments, the labeled region 314 can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or a number or range of numbers between any two of these values in length of nucleotides, and can include different labels such as barcode arrays or molecular labels 318 and cell labels 320 of approximately these numbers, at least these numbers, or at most these numbers. Each label can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range of numbers between any two of these values in length of nucleotides, and can be of approximately these numbers, at least these numbers, or at most these numbers in length. A set of barcodes or probabilistic barcodes 310 can be 10, 20, 40, 50, 70, 80, 90, 10 2 individuals, 10 3 individuals, 10 4 individuals, 10 5 individuals, 10 6 individuals, 10 7 individuals, 10 8 individuals, 10 9 individuals, 10 10 individuals, 10 11 individuals, 10 12 individuals, 10 13 individuals, 10 14 individuals, 10 15 individuals, 10 20 individuals, or a number or range of numbers between any two of these values in length of nucleotides, and can include barcodes or probabilistic barcodes 310 of approximately these numbers, at least these numbers, or at most these numbers. And a set of barcodes or probabilistic barcodes 310 can, for example, each include a unique labeled region 314. The labeled cDNA molecules 304 can be purified to remove excess barcodes or probabilistic barcodes 310. The purification can include Ampure bead purification.
[0185] As shown in Step 2, the products from the reverse transcription process in Step 1 can be pooled in one tube and PCR amplified using a first-generation PCR primer pool and a first-generation universal PCR primer. The pooling is made possible by the unique labeled region 314. In particular, the labeled cDNA-like molecule 304 can be amplified to generate a nested PCR labeled amplification product 322. The amplification can include multiplex PCR amplification. The amplification can include multiplex PCR amplification using 96 multiplex primers in a single reaction volume. In some embodiments, the multiplex PCR amplification can use 10, 20, 40, 50, 70, 80, 90, 10 2 individuals, 10 3 individuals, 10 4 individuals, 10 5 individuals, 10 6 individuals, 10 7 individuals, 10 8 individuals, 10 9 individuals, 10 10 individuals, 10 11 individuals, 10 12 individuals, 10 13 individuals, 10 14 individuals, 10 15 individuals, 10 20 individuals, or a number or range of numbers between any two of these values, approximately these numbers, at least these numbers, or at most these numbers of multiplex primers can be utilized. The amplification can include a first-generation PCR primer pool 324 of custom primers 326A - C targeting a specific gene and a universal primer 328. The custom primer 326 can hybridize to a region within the cDNA portion 306’ of the labeled cDNA molecule 304. The universal primer 328 can hybridize to the universal PCR region 316 of the labeled cDNA molecule 304.
[0186] As shown in step 3 of FIG. 3, the product from the PCR amplification in step 2 can be amplified using a nested PCR primer pool and a second-generation universal PCR primer. Nested PCR can minimize PCR amplification bias. For example, the nested PCR-labeled amplification product 322 can be further amplified by nested PCR. Nested PCR can include multiplex PCR having, in a single reaction volume, a nested PCR primer pool 330 of nested PCR primers 332a-c and a second-generation universal PCR primer 328'. The nested PCR primer pool 328 can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or a number or range between any two of these values, about these numbers, at least these numbers, or at most these numbers of different nested PCR primers 330. The nested PCR primer 332 can include an adapter 334 and can hybridize to a region within the cDNA portion 306'' of the labeled amplification product 322. The universal primer 328' can include an adapter 336 and can hybridize to the universal PCR region 316 of the labeled amplification product 322. Thus, step 3 generates an adapter-labeled amplification product 338. In some embodiments, the nested PCR primer 332 and the second-generation universal PCR primer 328' may not include the adapters 334 and 336. Instead, the adapters 334 and 336 can be ligated to the product of the nested PCR to generate an adapter-labeled amplification product 338.
[0187] As shown in step 4, the PCR product from step 3 can be PCR amplified towards sequencing using library amplification primers. In particular, using adapters 334 and 336, one or more additional assays can be performed on the adapter-labeled amplification product 338. Adapters 334 and 336 can hybridize to primers 340 and 342. One or more of primers 340 and 342 can be PCR amplification primers. One or more of primers 340 and 342 can be sequencing primers. One or more of adapters 334 and 336 can be used for further amplification of the adapter-labeled amplification product 338. One or more of adapters 334 and 336 can be used for sequencing of the adapter-labeled amplification product 338. Primer 342 can include a plate index 344 such that amplification products generated using the same set of barcodes or probabilistic barcodes 310 can be sequenced in one sequencing reaction using next-generation sequencing (NGS).
[0188] Errors in cell label identification Barcoding, such as probabilistic barcoding, e.g., Rhapsody (trademark) assay (Cellular Research, Inc. (Palo Alto, California)), can be bead-based. Molecules or targets such as mRNA from different cells can hybridize to barcodes (e.g., probabilistic barcodes) on different beads. Barcodes on different beads can have different cell labels, and barcodes on the same bead can have a cell label. For example, one cell and one bead can be added to a microwell of a microwell plate before the one bead pairs with the one cell. Thus, the cell label is the same for all oligonucleotides on the bead, but different between different beads, and thus all molecules from one cell can be identified using the same cell label in the sequence data. In some embodiments, the raw sequence data from barcoding (e.g., probabilistic barcoding) can contain a greater number of cell labels than the number of cells input into the experiment. For example, some molecules of 1,000 cells can be barcoded (e.g., probabilistically barcoded), but the raw sequence data can indicate 20,000 to 200,000 cell labels.
[0189] The source of a greater number of cell labels can vary in different embodiments. Without being bound by any particular theory, in some embodiments, cells that do not pair with beads can lyse, and their nucleic acid contents can diffuse and become associated with beads that are not paired with any cells, thereby potentially generating false cell labeling signals. In some embodiments, during the bead manufacturing process, the cell label can have mutations that convert one cell label to another. In this case, molecules from the same cell can appear to be from two different cells (e.g., because the cell label has mutated such that their molecules seem to be from two different beads). Further, substitution and non-substitution errors can occur in the cell label during PCR amplification prior to sequencing. In some embodiments, exonuclease treatment (e.g., in step 216 of FIG. 2) can be inefficient such that single-stranded DNA on the beads can hybridize and form PCR chimeras during the PCR process.
[0190] If not corrected, the excessive number of cell labels in the raw sequence data can lead to an overestimation of the cell count. The methods disclosed herein can separate or distinguish signal cell labels (also referred to as true cell labels) from noise cell labels.
[0191] Identification of cell labels as signal cell labels or noise cell labels based on the second derivative Disclosed herein is a method for identifying signal cell labels. In some embodiments, the method comprises: (a) stochastically barcoding a plurality of targets in a sample of cells using a plurality of stochastic barcodes to create a plurality of stochastically barcoded targets, wherein each of the plurality of stochastic barcodes comprises a cell label and a molecular label, and stochastically barcoding to create a stochastically barcoded target; (b) obtaining sequence data of the plurality of stochastically barcoded targets; (c) determining the number of molecular labels having distinct sequences associated with each of the cell labels of the plurality of stochastic barcodes; (d) determining the rank of each of the cell labels of the plurality of stochastic barcodes based on the number of molecular labels having distinct sequences associated with each of the cell labels; (e) generating a cumulative sum plot based on the number of molecular labels having distinct sequences associated with each of the cell labels identified in (c) and the rank of each of the cell labels determined in (d); (f) generating a second derivative plot of the cumulative sum plot; (g) identifying a minimum of the second derivative plot of the cumulative sum plot, wherein the minimum of the second derivative plot corresponds to a cell label threshold; and (h) identifying each of the cell labels as a signal cell label (associated with a cell) or a noise cell label (not associated with a cell) based on the number of molecular labels having distinct sequences associated with each of the cell labels identified in (c) and the cell label threshold determined in (g).
[0192] The cause of the noise cell label can vary in different embodiments. In some embodiments, the noise cell label can result from one or more PCR or sequencing errors. In some embodiments, the noise cell label can result from RNA molecules released from dead cells. In some embodiments, the noise cell label can result from RNA molecules released from cells not associated with beads that are attached to beads not associated with cells.
[0193] In some embodiments, the method comprises: (a) obtaining array data of a plurality of barcoded targets (e.g., stochastically barcoded targets), wherein the plurality of barcoded targets are created from a plurality of targets in a sample of cells that have been barcoded (e.g., stochastically barcoded) using a plurality of barcodes (e.g., stochastic barcodes), and each of the plurality of barcodes includes a cell label and a molecular label; (b) determining a rank for each of the cell labels of the plurality of barcoded targets (or barcodes) based on the number of molecular labels having distinct sequences associated with each of the cell labels of the plurality of barcoded targets (or barcodes); (c) determining a cell label threshold based on the number of molecular labels having distinct sequences associated with each of the cell labels and the rank for each of the cell labels of the plurality of barcoded targets (or barcodes) determined in (b); and identifying each of the cell labels as a signal cell label or a noise cell label based on the number of molecular labels having distinct sequences associated with each of the cell labels and the cell label threshold determined in (c).
[0194] FIG. 4 is a flowchart illustrating a non-limiting and exemplary method 400 for identifying cells as signal cell labels or noise cell labels. At block 404, method 400 optionally can barcode targets (e.g., stochastically barcoded targets) in cells using barcodes (e.g., stochastic barcodes) to create barcoded targets, as described with reference to FIGS. 2 and 3. Each barcode can include a cell label and a molecular label. Barcoded targets created from targets of different cells of the plurality of cells can have different cell labels. Barcoded targets created from targets of the same cell of the plurality of cells can have different molecular labels.
[0195] In block 408, method 400 can obtain array data of barcoded targets (e.g., stochastically barcoded targets), as described in the section titled Sequencing herein. In block 412, method 400 can optionally determine the number of molecular labels having distinct sequences associated with each cell label of the barcodes (or barcoded targets). Determining the number of molecular labels having distinct sequences associated with each cell label of the barcodes (or barcoded targets) can include (1) counting the number of molecular labels having distinct sequences associated with the targets in the array data, and (2) estimating the number of targets based on the number of molecular labels having distinct sequences associated with the targets in the array data counted in (1). In some embodiments, the array data obtained in block 408 includes the number of molecular labels having distinct sequences associated with each of the cell labels of the barcodes (or barcoded targets).
[0196] In some embodiments, the method can include removing, from the sequence data obtained at block 408, sequence information associated with a molecular label having a distinct sequence associated with a target among a plurality of targets when the number of molecular labels having a distinct sequence associated with the target among the plurality of targets exceeds or falls below a molecular label occurrence threshold. The molecular label occurrence threshold can vary in different embodiments. In some embodiments, the molecular label occurrence threshold can be 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the molecular label occurrence threshold can be at least or at most 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000. In some embodiments, the molecular label occurrence threshold can be 1%, 2%, 3%, 4%, 5%, 6%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the molecular label occurrence threshold can be at least or at most 10%, 20%, 30%, 40%, 50%, 60%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%.
[0197] At block 416, method 400 can determine the rank of each cell label of a barcode (or barcoded target). The rank of each cell label of a barcode (or barcoded target) can be based on the number of molecular labels having a distinct sequence associated with each cell label of a plurality of barcodes (or barcoded targets).
[0198] In block 420, method 400 can determine a cell target threshold associated with each cell label of a plurality of barcodes (or barcoded targets) and the rank of each cell label of the plurality of barcodes (or barcoded targets) determined in block 416. In some embodiments, determining a cell label threshold based on the number of molecular labels having distinct sequences associated with each cell label of a plurality of barcodes (or barcoded targets) includes identifying the cell label having the largest change in the cumulative sum of cell labels having rank n and the cumulative sum of cell labels having the next rank n + 1, where the number of molecular labels having distinct sequences associated with the cell label corresponds to the cell label threshold.
[0199] In some embodiments, determining a cell label threshold based on the number of molecular labels having distinct sequences associated with each cell label of a plurality of barcodes (or barcoded targets) and the rank of each cell label of the plurality of barcodes (or barcoded targets) determined in block 416 includes identifying the cumulative sum of each rank of cell labels, where the cumulative sum of ranks includes the sum of the number of molecular labels having distinct sequences associated with each cell label having a lower rank, and determining the rank n of the cell label having the largest change in the cumulative sum of rank n and the cumulative sum of the next rank n + 1, where the rank n of the cell label having the largest change in the cumulative sum and the cumulative sum of the next rank n + 1 corresponds to the cell label threshold.
[0200] In some embodiments, determining a cell label threshold can include generating a cumulative sum plot based on the number of molecular labels having distinct sequences associated with each cell label and the rank of each cell label determined in 416. Determining a cell label threshold can further include generating a second derivative plot of the cumulative sum plot and identifying the minimum of the second derivative plot of the cumulative sum plot. The minimum of the second derivative plot can correspond to the cell label threshold.
[0201] In some embodiments, generating a cumulative sum plot based on the number of molecular labels having distinct sequences associated with each cell label and the rank of each cell label determined at 416 can include identifying the sum of each rank of cell labels, where the cumulative sum of ranks includes the sum of the number of molecular labels having distinct sequences associated with each cell label having a lower rank. Generating a second derivative plot of the cumulative sum plot can include determining that the difference between the cumulative sum of the first rank of cell labels and the cumulative sum of the second rank of cell labels exceeds the difference between the first rank and the second rank. In some embodiments, the difference between the first rank and the second rank is 1. The cumulative sum plot can be a log-log plot. The log-log plot can be a log10-log10 plot.
[0202] In some embodiments, the minimum is a global minimum. Identifying the minimum of the second derivative plot can include determining that the minimum of the second derivative plot exceeds a threshold of the minimum number of molecular labels associated with each cell label. The threshold of the minimum number of molecular labels associated with each cell label can be a percentile threshold. The threshold of the minimum number of molecular labels associated with each cell label can be determined based on the number of cells in the sample of cells. For example, the threshold of the minimum number of molecular labels associated with each cell label can be a larger value when the number of cells in the sample of cells is larger.
[0203] The threshold for the minimum number of molecular labels associated with each cell label can vary in different embodiments. In some embodiments, the threshold for the minimum number of molecular labels associated with each cell label can be 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the threshold for the minimum number of molecular labels associated with each cell label can be at least or at most 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000. In some embodiments, the threshold for the minimum number of molecular labels associated with each cell label can be 1%, 10%, 20%, 30%, 40%, 45%, 50%, 60%, 80%, 90%, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the molecular label generation threshold can be at least or at most 1%, 10%, 20%, 30%, 40%, 45%, 50%, 60%, 80%, or 90%.
[0204] In some embodiments, identifying the minimum of the second derivative plot includes determining that the minimum of the second derivative plot is below the threshold for the maximum number of molecular labels associated with each cell label. The threshold for the maximum number of molecular labels associated with each cell label can be a percentile threshold. The threshold for the maximum number of molecular labels associated with each cell label can be determined based on the number of cells in the sample of cells. For example, the threshold for the maximum number of molecular labels associated with each cell label can be a larger value when the number of cells in the sample of cells is larger.
[0205] The threshold for the maximum number of molecular labels associated with each cell label can vary in different embodiments. In some embodiments, the threshold for the maximum number of molecular labels associated with each cell label can be 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the threshold for the maximum number of molecular labels associated with each cell label can be at least or at most 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, or 100000. In some embodiments, the threshold for the maximum number of molecular labels associated with each cell label can be 10%, 20%, 30%, 40%, 45%, 50%, 60%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the molecular label generation threshold can be at least or at most 10%, 20%, 30%, 40%, 45%, 50%, 60%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%.
[0206] In block 432, method 400 can identify a cell label as a signal cell label or a noise cell label based on the number of molecular labels having distinct sequences associated with the cell label and a cell label threshold. Each cell label is identified as a signal cell label if the number of molecular labels having distinct sequences associated with each cell label specified in (c) is greater than the cell label threshold. Each cell label can be identified as a noise cell label if the number of molecular labels having distinct sequences associated with each cell label specified in (c) is not greater than the cell label threshold. In some embodiments, the method includes, in 432, removing sequence information associated with the identified cell labels from the sequence data obtained in block 408 if the cell labels of a plurality of barcodes (or barcoded targets) are identified as noise cell labels.
[0207] Identification of Cell Labels Based on Clustering as Signal Cell Labels or Noise Cell Labels Disclosed herein is a method for identifying signal cell labels. In some embodiments, the method comprises: (a) using a plurality of barcodes (e.g., probabilistic barcodes) to barcode (e.g., probabilistically barcode) a plurality of targets in a sample of cells to create a plurality of barcoded targets (e.g., probabilistically barcoded targets), wherein each of the plurality of barcodes includes a cell label and a molecular label, barcoding to create a barcoded target, and barcoded targets created from targets of different cells of the plurality of cells have different cell labels, and barcoded targets created from targets of the same cell of the plurality of cells have different molecular labels; (b) obtaining sequence data of the plurality of barcoded targets; (c) identifying a feature vector for each cell label of the plurality of barcodes (or barcoded targets), wherein the feature vector includes the number of molecular labels having a distinct sequence associated with each cell label; (d) identifying a cluster for each cell label of the plurality of barcodes (or barcoded targets) based on the feature vector; and (e) identifying each cell label of the plurality of barcodes (or barcoded targets) as a signal cell label or a noise cell label based on the number of cell labels within the cluster and a cluster size threshold.
[0208] Disclosed herein is a method for identifying signal cell labels. In some embodiments, the method comprises: (a) obtaining sequence data of a plurality of barcoded targets (e.g., stochastically barcoded targets), wherein the plurality of barcoded targets (e.g., stochastically barcoded targets) are created from a plurality of targets in a sample of cells barcoded (e.g., stochastically barcoded) using a plurality of barcodes (e.g., stochastic barcodes), each of the plurality of barcodes comprises a cell label and a molecular label, barcoded targets created from targets of different cells of the plurality of cells have different cell labels, and barcoded targets created from targets of the same cell of the plurality of cells have different molecular labels; (b) identifying a feature vector for each cell label of the plurality of barcoded targets, wherein the feature vector comprises the number of molecular labels having a distinct sequence associated with each cell label; (c) identifying a cluster for each cell label of the plurality of barcoded targets based on the feature vectors; and (d) identifying each cell label of the plurality of barcoded targets as a signal cell label or a noise cell label based on the number of cell labels within the cluster and a cluster size threshold.
[0209] Figure 5 is a flowchart illustrating another non-limiting, exemplary method for identifying cells as signal cell labels or noise cell labels. At block 504, method 500 can optionally barcode (e.g., stochastically barcode) targets in cells using stochastic barcodes, as described with reference to FIGS. 2 and 3, to create barcoded targets (e.g., stochastically barcoded targets). Each barcode comprises a cell label and a molecular label. Barcoded targets created from targets of different cells of the plurality of cells can have different cell labels. Barcoded targets created from targets of the same cell of the plurality of cells can have different molecular labels.
[0210] In block 508, method 500 can obtain array data of barcoded targets. In block 508, method 500 can optionally determine the number of molecular labels having distinct arrays associated with each cell label of the barcode (or barcoded target). Determining the number of molecular labels having distinct arrays associated with each cell label of the barcode (or barcoded target) can include (1) counting the number of molecular labels having distinct arrays associated with the targets in the array data, and (2) estimating the number of targets based on the number of molecular labels having distinct arrays associated with the targets in the array data counted in (1). In some embodiments, the array data obtained in block 508 includes the number of molecular labels having distinct arrays associated with each of the cell labels of the barcode (or barcoded target).
[0211] In block 512, method 500 can determine a feature vector of the cell label. The feature vector can include the number of molecular labels having distinct arrays associated with the cell label. For example, each element of the feature vector can include the number of molecular labels associated with the cell label. As another example, one element of the feature vector can include the number of molecular labels associated with the cell label, and another element of the feature vector can include the number of another molecular label associated with the cell label.
[0212] In block 516, method 500 can identify clusters of cell labels based on the feature vectors. In some embodiments, identifying the clusters of each cell label of the barcode (or barcoded target) based on the feature vectors includes clustering each cell label of the barcode (or barcoded target) into clusters based on the distance of the feature vectors to the clusters in the feature vector space. Identifying the clusters of each cell label of a plurality of barcoded targets based on the feature vectors includes projecting the feature vectors from the feature vector space into a lower-dimensional space and clustering each cell label into clusters based on the distance of the feature vectors to the clusters in the lower-dimensional space. The lower-dimensional space is a two-dimensional space.
[0213] In some embodiments, projecting the feature vectors from the feature vector space into a lower-dimensional space includes using the t-distributed stochastic neighbor embedding (tSNE) method to project the feature vectors from the feature vector space into a lower-dimensional space. Clustering each cell label into clusters based on the distance of the feature vectors to the clusters in the lower-dimensional space can include using a density-based method to cluster each cell label into clusters based on the distance of the feature vectors to the clusters in the lower-dimensional space. The density-based method can include the density-based spatial clustering of applications with noise (DBSCAN) method.
[0214] In block 520, method 500 can identify a cell label as a signal cell label or a noise cell label based on the number of cells in the cluster and a cluster size threshold. In some embodiments, a cell label can be identified as a signal cell label if the number of cell labels in the cluster is below the cluster size threshold. A cell label can be identified as a noise cell label if the number of cell labels in the cluster is not below the cluster size threshold.
[0215] In some embodiments, the method includes determining a cluster size threshold based on the number of cell labels of a plurality of barcodes (or barcoded targets). The cluster size threshold can be a percentage of the number of cell labels of the plurality of barcoded targets. In some embodiments, determining a cluster size threshold based on the number of cell labels of a plurality of barcodes (or barcoded targets). The cluster size threshold can be a percentage of the number of cell labels of the plurality of barcodes (or barcoded targets). In some embodiments, the method includes determining a cluster threshold size based on the number of molecular labels having distinct sequences associated with each cell label of the plurality of barcodes (or barcoded targets).
[0216] Distinguishing cell labels associated with true cells from noise cells Disclosed herein are embodiments of a method for reliably distinguishing labels (e.g., cell labels) associated with true cells from labels (e.g., cell labels) associated with noise cells. Cell labels associated with true cells are referred to herein as signal cell labels. Noise cells are referred to herein as noise cell labels. The method can, in some embodiments, detect or identify the majority of true cells (or signal cell labels) corresponding to different cell types / clusters. The method can potentially automatically eliminate noise cells that are low-expressing within a particular cell type such as single nuclei and plasma.
[0217] FIG. 6A is a flowchart showing a non-limiting and exemplary method 600a for distinguishing labels associated with true cells from noise cells. Method 600a can be based on one or more cell label identification or classification methods (e.g., method 400 or 500 described with reference to FIGS. 4 or 5). Method 600a can, in some embodiments, improve these cell label identifications. The method can be used for classification in the Rhapsody™ pipeline.
[0218] Method 600a includes a plurality of steps or operations. At block 604, method 600a performs (or runs) a cell labeling identification method (e.g., method 400 or 500 described with reference to FIGS. 4 or 5) to identify a plurality of true cells (or signal cell labels called filtered cells (A) in FIG. 6). For example, the cell labeling identification method can be based on a log10-transformed cumulative read curve. Using the cell labeling identification method, an inflection point where the curve starts to level off can be identified. For example, the main inflection point can be the separation between true cells and noise cells.
[0219] Method 600a can include removing noise cells by, for example, restricting to highly variable (e.g., most variable) genes across most cells (e.g., all cells) and performing the cell labeling identification method. For example, method 600a can include re-performing the cell labeling identification method performed at block 604 for the most variable genes across all cells. Method 600a can include, at block 608, identifying highly variable genes across most cells (e.g., all cells). The cell labeling identification method can be performed, at block 612, on the most variable genes identified at block 608 to identify one or more true cells (or signal cell labels, where called noise cells (B) in FIG. 6). To identify highly variable genes, method 600a can optionally include log-transforming the read count of each gene within each cell (e.g., the number of molecular labels having a distinct sequence associated with each gene of each cell label) to identify gene expression. For example, the read count can be log-transformed using the following equation [1]. log(count + 1) Equation [1]
[0220] In block 608, method 600a can include identifying one or more field measures or indicators of expression for each gene, such as average expression (or maximum, median, or minimum expression) and variability (e.g., variance / average). Method 600a can include assigning each gene (or the expression profile of each gene) to one of a plurality of bins. For example, genes can be assigned to 20 bins based on the average (or maximum, median, or minimum) expression of each gene. The number of bins can vary in different embodiments. In some embodiments, the number of bins can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the number of bins can be at least or at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, or 1000.
[0221] In block 608, method 600a can include identifying one or more measures or indicators of variability of all genes within each bin. For example, the average and standard deviation (STD) of the variability measure of all genes can be identified. Method 600a can include, for example, using Equation [2] to identify the normalized variability measure for each gene. Normalized variance = (variance - average) / standard deviation Equation [2]
[0222] In block 608, method 600a can include applying one or more different cutoff values to the normalized variability to identify genes with high variability in expression values (e.g., having variability exceeding a threshold), even when compared to genes with similar average expression. The number of cutoff values can vary in different embodiments. In some embodiments, the number of cutoff values can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the number of cutoff values can be at least or at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, or 1000.
[0223] In some embodiments, method 600a can determine a cell as a noise cell (or cell label or noise cell label) only if the cell is identified as a noise cell at a threshold number of cutoff values or a threshold percentage of all cutoff values (e.g., a minority, majority, or all cutoff values). In some embodiments, the threshold number of cutoff values can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or a number or range between any two of these values or about these values or ranges. In some embodiments, the threshold number of cutoff values can be at least or at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, or 1000.In some embodiments, the cutoff first threshold percentage can be 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, 100%, or a number or range between any two of these values or about these values or ranges. In some embodiments, the cutoff first threshold percentage can be at least or at most 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 100%. In some embodiments, such noise cell identification can improve the accuracy of the identified noise cells (e.g., reduce the possibility of misidentifying true cells as noise cells).FIG. 7 is a non-limiting and exemplary schematic diagram showing the identification of the most variable genes.
[0224] Referring to FIG. 6A, at block 616, method 600a can include identifying or discriminating true cells (or signal cell labels) that may have been incorrectly determined (e.g., not identified) at block 604 by determining, for example, whether there are any dropout genes. If so, method 600a can include, at block 620, performing or re-performing a cell label discrimination method (e.g., the cell label discrimination method used at block 604 or 612) to identify one or more lost true cells (or lost signal cell labels) that were not identified at block 604. The lost true cells identified at block 620 are referred to as loss cells (D) in FIG. 6A. Identifying dropout genes can include, for each gene, identifying the total read count from all cells and the clean-up cells identified at block 625. Clean-up cells can be identified using Equation [3a] or Equation [3b], where C represents the clean-up cells, A represents the filtered or true cells determined at block 604, and B represents the loss cells identified at block 612. C = set_difference(A, set_difference(A, B)) [3a] C = A - (A - B) Equation [3a]
[0225] Identification of dropout genes can include identifying genes that have a large loss (e.g., the largest loss) in counts from the cleaned-up cells compared to the counts from all cells. For example, genes with the largest loss can be identified by plotting the total counts, finding the optimal fit line, and identifying genes that have a large residual (e.g., the largest residual) of the standard deviation threshold number away from the median of the remaining all genes of at least 1 equal (see FIGS. 8A and 8B). In some embodiments, the median can be used instead of the mean to minimize the influence of outliers. The standard deviation threshold number can be different in different embodiments. In some embodiments, the standard deviation threshold number can be 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, or a number or range between any two of these values or about these values or ranges. In some embodiments, the standard deviation threshold number can be at least or at most 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10.
[0226] In block 624, method 600a can include combining the cells (or cell labels) identified in blocks 620 and 624 to identify a final set of true cells (referred to as filtered cells F in FIG. 6).
[0227] FIG. 6B is a flow diagram showing another non-limiting and exemplary method 600b for distinguishing labels associated with true cells from noise cells. The operations performed in blocks 604-628 in FIG. 6B can be similar to the operations performed in the corresponding blocks of method 600a described with reference to FIG. 6A. Method 600b includes executing an algorithm based on a log10-transformed cumulative read curve, and at block 604, an inflection point where the curve levels off can be found. The major inflection point is the separation between cells and noise. Method 600b can include one or more of the following steps. Starting from all cells, obtain the most variable genes using a gene variability measure z-score cutoff. Focus only on the most variable genes and execute through the current algorithm to infer true cells, and denote this set as B. Identify cells that were detected by other cell labeling identification methods using all genes in the panel but not detected by the algorithm using only the most variant genes, i.e., setdiff(A,B), which are identified as noise cells. In some embodiments, to be more conservative, try multiple variability z-cutoff values and identify a cell as noise only if the cell is classified as noise at some, most, or all of the cutoff values. Remove the noise cells from set A and obtain an updated set of cells using the above formula [3a] or [3b].
[0228] Method 600b can include, at block 608, removing noise cells by restricting to the most variable or highly variable genes across all cells, and at block 612, re-executing the algorithm (e.g., executed at block 604). For example, the method can include one or more of the following steps. Remove true cells. For each gene, calculate the total read counts from all cells and the cells in set C. Find genes that are mostly dropped in set C. Focus on the dropped genes and execute through the method executed at block 604 to remove any potentially lost true cells and denote the cells identified in the individual steps as C.
[0229] Method 600b can include recovering true cells that may have been misdetected or misidentified in block 604 by checking whether there are any dropped genes in block 616. If there are dropped genes, method 600b can be restricted to the dropped genes (also called under-evaluated genes) and, in block 620, can include re-running the algorithm (e.g., the one executed in block 604) to pick up the lost true cells. The final cell list F can be specified using Equation [4]. F = union(C, D) Equation [4]
[0230] In some embodiments, in block 632, the cells from block 628 can be cleaned up or polished by removing cells that do not carry a sufficiently large number of molecules. For example, the minimum threshold of the molecule count can be specified by the following rules. Step (a): Find a large gap (e.g., the largest gap, the second largest gap, the third largest gap, etc.) in the total molecule count of the cells in the lower 1 / 4 and determine the cut-off as the value of the gap. Step (b): Find the cells with a molecule count less than the cut-off determined in step (a) and optionally calculate the proportion of the cells removed due to the low molecule count. Step (c): Without using the adaptive cut-off determined above, use a fixed cut-off, e.g., 20 molecules, under one or both of the following two conditions: Condition (i) the proportion of the cells removed due to the low molecule count is greater than or at least the threshold proportion (e.g., 20%), and / or the gap is less than the threshold number (e.g., 500), and Condition (ii) the largest gap in the total molecule count of all cells is, for example, 1. The cells after cleanup are part of the final set of filtered cells detected by method 600b.
[0231] The fixed cut-off in step (c) can be different in different embodiments. In some embodiments, the cut-off can be 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the cut-off can be at least or at most 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, or 100. The threshold ratio in condition (i) can be different in different embodiments. In some embodiments, the threshold ratio can be 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the threshold ratio can be at least or at most 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, or 50%. The threshold number of gaps in condition (i) can be different in different embodiments.In some embodiments, the threshold number of gaps can be 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, or a number or range between any two of these values or approximately these values or ranges. In some embodiments, the threshold number of gaps can be at least or at most 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, or a number or range between any two of these values or approximately these values or ranges. The maximum gap in condition (ii) can be different in different embodiments. In some embodiments, the maximum gap can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values or approximately these values or ranges. In this embodiment, the maximum gap can be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.
[0232] Disclosed herein are embodiments of a method for identifying signal cell labels. In some embodiments, the method comprises: (a) obtaining sequence data of a plurality of first targets of cells, wherein each of the first targets is associated with a number of molecular labels having distinct sequences associated with each of the plurality of cell labels; (b) applying method 400 or 500, and in block 604 of method 600a or 600b, etc., identifying each of the cell labels as a signal cell label or a noise cell label based on the number of molecular labels having distinct sequences associated with each of the cell labels and an identification threshold; (c) using method 400 or 500, in blocks 608, 612, etc. of method 600a or 600b, re-identifying at least one of the plurality of cell labels identified as noise cell labels in (b) as a signal cell label, or in blocks 616, 620, etc. of method 600a or 600b, re-identifying at least one of the cell labels identified as signal cell labels in (b) as a noise cell label. Identifying each cell label, re-identifying at least one of the plurality of cell labels as a signal cell label, or re-identifying at least one of the plurality of cell labels as a noise cell label can be based on the same cell label identification method or different cell label identification methods (such as method 400 or 500 described with reference to FIG. 4 or FIG. 5) of the present disclosure. The identification threshold can include a cell label threshold, a cluster size threshold, or any combination thereof. The method can include removing one or more of the plurality of cell labels each associated with a number of molecular labels having distinct sequences below a threshold of the number of molecular labels in block 628 of method 600b described with reference to FIG. 6A, etc.
[0233] In some embodiments, in block 608 of method 600a or 600b, etc., re-identifying at least one of the plurality of cell labels identified as noise cell labels in (b) as signal cell labels includes identifying a plurality of second targets among the plurality of first targets, each having one or more diversity indicators exceeding a diversity threshold; and in block 612 of method 600a or 600b, etc., for each of the plurality of cell labels, re-identifying at least one of the plurality of cell labels identified as noise cell labels in (b) as signal cell labels based on the number of molecular labels having distinct sequences associated with the plurality of second targets and an identification threshold. One or more diversity indicators of the second target can include the average, maximum, median, minimum, variance, or any combination thereof of the number of cell labels among the molecular labels having distinct sequences associated with the second target and the plurality of cell labels in the sequence data. One or more diversity indicators of the second target can include the diversity indicator standard deviation, normalized variance, or any combination thereof of a subset of the plurality of second targets. The diversity threshold can be less than or equal to the size of a subset of the plurality of second targets.
[0234] In some embodiments, re-identifying at least one of the plurality of cell labels identified as signal cell labels in (b) as noise cell labels includes, for example, at block 616 of method 600a or 600b, identifying a plurality of third targets among the plurality of first targets each having a relevance with a cell label identified as a noise cell label in (c) that exceeds a relevance threshold, and, at block 620 of method 600a or 600b, for each of the plurality of cell labels, re-identifying at least one of the cell labels identified as signal cell labels in (b) as a noise cell label based on the number of molecular labels having distinct sequences associated with the plurality of third targets and an identification threshold. Identifying a plurality of third targets among the plurality of first targets each having a relevance with a cell label identified as a noise cell label in (c) that exceeds a relevance threshold can include, after re-identifying at least one of the cell labels identified as noise cell labels in (b) as a signal cell label, identifying the remaining plurality of cell labels identified as signal cell labels, and identifying a plurality of third targets, where for each of the plurality of cell labels, identifying the plurality of third targets is based on the number of molecular labels having distinct sequences associated with the plurality of targets, and for each of the remaining plurality of cell labels, identifying the plurality of third targets is based on the number of molecular labels having distinct sequences associated with the plurality of targets.
[0235] Sequencing In some embodiments, estimating the number of differently barcoded targets (e.g., stochastically barcoded targets) can include identifying the sequences of labeled targets, spatial labels, molecular labels, sample labels, cell labels, or any products thereof (e.g., labeled amplification products or labeled cDNA molecules). The amplified targets can be subjected to sequencing. Identifying the sequence of a barcoded target (e.g., a stochastically barcoded target) or any product thereof can include performing a sequencing reaction and identifying the sequence of at least a portion of a sample label, spatial label, cell label, molecular label, labeled target (e.g., a stochastically labeled target), its complement, its reverse complement, or any combination thereof.
[0236] Identifying the sequence of a barcoded target or a stochastically barcoded target (e.g., an amplified nucleic acid, a labeled nucleic acid, a cDNA copy of a labeled nucleic acid, etc.) can be performed using a variety of sequencing methods including, but not limited to, sequencing by hybridization (SBH), sequencing by ligation (SBL), quantitative incremental fluorescence nucleotide addition sequencing (QIFNAS), stepwise ligation and cleavage, fluorescence resonance energy transfer (FRET), molecular beacons, TaqMan reporter probe digestion, pyrosequencing, fluorescence in situ sequencing (FISSEQ), FISSEQ beads, wobble sequencing, multiplex sequencing, polymerized colony (POLONY) sequencing; nanogrid rolling circle sequencing (ROLONY), allele-specific oligoligation assays (e.g., oligoligation assay (OLA), single-template molecule OLA using ligation of linear probes and rolling circle amplification (RCA) readout, ligated padlock probes, or single-template molecule OLA using ligation of circular padlock probes and rolling circle amplification (RCA) readout), and the like.
[0237] In some embodiments, identifying the sequence of a barcoded target (e.g., a stochastically barcoded target) or any product thereof may include paired-end sequencing, nanopore sequencing, high-throughput sequencing, shotgun sequencing, terminator sequencing, multiprimer DNA sequencing, primer walking, Sanger dideoxy sequencing, Maxim-Gilbert sequencing, pyrosequencing, tSMS (true single molecule sequencing), or any combination thereof. Alternatively, the sequence of a barcoded target or any product thereof can be identified by electron microscopy or a chemFET (chemical field effect transistor) array.
[0238] High-throughput sequencing methods such as cycle array sequencing using platforms such as the Roche 454, Illumina Solexa, ABI-SOLiD, ION Torrent, Complete Genomics, Pacific Bioscience, Helicos, or Polonator platforms can be utilized. In some embodiments, the sequencing can include MiSeq sequencing. In some embodiments, the sequencing can include HiSeq sequencing.
[0239] The labeled target (e.g., a probabilistically labeled target) can comprise nucleic acids representing from about 0.01% to about 100% of the genes of the genome of an organism. For example, from about 0.01% to about 100% of the genes of the genome of an organism can be sequenced using a target complementary region comprising a plurality of multimers by capturing genes comprising complementary sequences from a sample. In some embodiments, the barcoded target comprises nucleic acids representing from about 0.01% to about 100% of the transcripts of the transcriptome of an organism. For example, from about 0.501% to about 100% of the transcripts of the transcriptome of an organism can be sequenced using a target complementary region comprising a poly(T) tail by capturing mRNA from a sample.
[0240] Identifying the spatial and molecular label sequences of a plurality of barcodes (e.g., probabilistic barcodes) can include sequencing 0.00001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 99%, 100%, or a number or range between any two of these values of the plurality of barcodes. Identifying the label sequences of a plurality of barcodes, e.g., sample labels, spatial labels, and molecular labels, can include sequencing 1, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 10 3 individuals, 10 4 individuals, 10 5 individuals, 10 6 individuals, 10 7 individuals, 10 8 individuals, 10 9 individuals, 10 10 individuals, 10 11 individuals, 10 12 individuals, 10 13 individuals, 10 14 individuals, 10 15 individuals, 10 16 individuals, 10 17 individuals, 10 18pieces, 10 19 pieces, 10 20 It can include sequencing a number or range between any two of these values. Sequencing some or all of the plurality of barcodes can include generating a sequence having a read length of 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000 nucleotides or bases, or a number or range between any two of these values, approximately these numbers, at most or at least these numbers of nucleotides or bases.
[0241] Sequencing can include sequencing at least or at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more barcoded target nucleotides or base pairs. For example, sequencing can include generating sequencing data having a sequence with a read length of 50, 75, 100, or more nucleotides by performing polymerase chain reaction (PCR) amplification on a plurality of barcoded targets. Sequencing can include sequencing at least or at least about 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more barcoded target nucleotides or base pairs. Sequencing can include sequencing at least or at least about 1500, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or more barcoded target nucleotides or base pairs.
[0242] Sequencing can include at least about 200, 300, 400, 500, 600, 700, 800, 900, 1,000, or more sequencing reads per run. In some embodiments, sequencing can include at least or at least about 1500, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or more sequencing reads per run. Sequencing can include up to about 1,600,000,000 sequencing reads per run. Sequencing can include up to about 200,000,000 reads per run.
[0243] Sample In some embodiments, a plurality of targets can be included in one or more samples. A sample can include one or more cells or nucleic acids from one or more cells. A sample can be a single cell or nucleic acids from a single cell. The one or more cells can be of one or more cell types. At least one of the one or more cell types can be a brain cell, a heart cell, a cancer cell, a circulating tumor cell, an organ cell, an epithelial cell, a metastatic cell, a benign cell, a primary cell, a circulating cell, or any combination thereof.
[0244] The sample used in the method of the present disclosure can contain one or more cells. The sample can refer to one or more cells. In some embodiments, the plurality of cells can contain one or more cell types. At least one of the one or more cell types can be a brain cell, a heart cell, a cancer cell, a circulating tumor cell, an organ cell, an epithelial cell, a metastatic cell, a benign cell, a primary cell, a circulating cell, or any combination thereof. In some embodiments, the cells are cancer cells excised from cancerous tissue, such as breast cancer, lung cancer, colon cancer, prostate cancer, ovarian cancer, pancreatic cancer, brain tumor, melanoma or non-melanoma skin cancer, etc. In some embodiments, the cells are cancer-derived but are collected from body fluids (e.g., circulating tumor cells). Non-limiting examples of cancer can include adenoma, adenocarcinoma, squamous cell carcinoma, basal cell carcinoma, small cell carcinoma, large cell undifferentiated carcinoma, chondrosarcoma, and fibrosarcoma. The sample can contain tissue, cell monolayer, fixed cells, tissue sections, or any combination thereof. The sample can contain a biological sample, a clinical sample, an environmental sample, a biological fluid, tissue, or cells from a subject. The sample can be obtained from a human, a mammal, a dog, a rat, a mouse, a fish, a fly, a worm, a plant, a fungus, a bacterium, a virus, a vertebrate, or an invertebrate.
[0245] In some embodiments, the cells are cells infected with a virus and containing viral oligonucleotides. In some embodiments, the viral infection can be caused by a virus such as a single-stranded (+ strand or "sense") DNA virus (e.g., parvovirus) or a double-stranded RNA virus (e.g., reovirus). In some embodiments, the cells are bacteria. These can include gram-positive or gram-negative bacteria. In some embodiments, the cells are fungi. In some embodiments, the cells are protozoa or other parasites.
[0246] As used herein, the term "cell" can refer to one or more cells. In some embodiments, the cell is a normal cell, for example, a human cell at different stages of development or a human cell from a different organ or tissue type. In some embodiments, the cell is a non-human cell, for example, another type of mammalian cell (e.g., mouse, rat, pig, dog, cow, or horse). In some embodiments, the cell is a cell of another type of animal or plant. In other embodiments, the cell can be any prokaryotic or eukaryotic cell.
[0247] In some embodiments, the cells are sorted prior to associating the cells with the beads. For example, the cells can be sorted by fluorescence-activated cell sorting, magnetic-activated cell sorting, or more generally by flow cytometry. The cells can be filtered by size. In some embodiments, the retainer comprises the cells to be associated with the beads. In some embodiments, the flow-through comprises the cells to be associated with the beads.
[0248] A sample can refer to a plurality of cells. A sample can refer to a monolayer of cells. A sample can refer to a thin section (e.g., a thin section of tissue). A sample can refer to a solid or semi-solid collection of cells that can be arranged in a row on an array.
[0249] Execution environment The present disclosure provides a computer system programmed to implement a method (e.g., method 400, method 500, method 600a, or method 600b, as described in connection with FIGS. 4, 5, 6A, and 6B). FIG. 9 shows a computer system 700 programmed or otherwise configured to implement any method disclosed herein. The computer system 700 can be a user's electronic device or a computer system located remotely from the electronic device. The electronic device can be a mobile electronic device.
[0250] The computer system 700 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 705, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 700 also includes a memory or memory location 710 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 715 (e.g., hard disk), a communication interface 720 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 725 such as a cache, other memories, data storage devices, and / or an electronic display adapter. The memory 710, storage unit 715, interface 720, and peripheral devices 725 communicate with the CPU 705 through a communication bus (solid lines), such as a motherboard. The storage unit 715 can be a data storage unit (or data repository) for storing data. The computer system 700 can be operably coupled to a computer network ("network") 730 using the communication interface 720. The network 730 can be the Internet, the Internet and / or an extranet, or an intranet and / or an extranet that communicates with the Internet. In some cases, the network 730 is a telecommunications network and / or a data network. The network 730 can include one or more computer servers that enable distributed computing such as cloud computing. In some cases, the network 730 can implement a peer-to-peer network using the computer system 700, thereby enabling devices coupled to the computer system 700 to act as clients or servers.
[0251] The CPU 705 can execute a machine-readable instruction sequence, and the instruction sequence can be implemented by a program or software. The instructions can be stored in a memory location such as the memory 710. The instructions can be directed to the CPU 705, and subsequently, the CPU 705 can be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU 705 can include fetch, decode, execute, and write-back. The CPU 705 can be part of a circuit such as an integrated circuit. One or more other components of the system 700 can be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0252] The storage unit 715 can store files such as drivers, libraries, and saved programs. The storage unit 715 can store user data, for example, user preferences and user programs. The computer system 700 can, in some cases, include one or more additional data storage units external to the computer system 700, such as being located on a remote server that communicates with the computer system 700 through an intranet or the Internet.
[0253] The computer system 700 can communicate with one or more remote computer systems through the network 730. For example, the computer system 700 can communicate with a remote computer system of a user (e.g., a microbiologist). Examples of remote computer systems include personal computers (e.g., portable PCs), slates or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab), phones, smartphones (e.g., Apple® iPhone®, Android-compatible devices, Blackberry®), or personal digital assistants. The user can access the computer system 700 via the network 730.
[0254] The computer system 700 includes, or can communicate with, an electronic display 735 that includes a user interface (UI) 740 that provides an output indicative of string co-occurrence or interaction of a plurality of taxa of microorganisms, represented, for example, by strings. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0255] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored in an electronic storage location of the computer system 700, such as, for example, the memory 710 or the electronic storage unit 715. The machine executable or machine readable code can be provided in the form of software. In use, the code can be executed by the processor 705. In some cases, the code can be retrieved from the storage unit 715 and stored in the memory 710 for easy access by the processor 705. In some situations, the electronic storage unit 715 can be eliminated and the machine executable instructions can be stored in the memory 710.
[0256] The code can be pre-compiled and configured to be used in conjunction with a machine having a processor configured to execute the code, or can be compiled during runtime. The code can be provided in a programming language that can be selected to pre-compile the code or to enable the code to be executed as compiled.
[0257] Aspects of the systems and methods provided herein, such as computer system 700, can be implemented in programming. Various aspects of the technology can generally be considered as a "product" or "manufacture" in the form of machine-executable code and / or associated data carried on or implemented within a machine-readable medium or a machine (or processor) executable within a machine-readable medium. The machine-executable code can be stored in an electronic memory unit such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A "storage" type medium can include any tangible memory such as a computer, a processor, or associated modules such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or part of the software can sometimes communicate through the Internet or various other electrical communication networks. Such communication can, for example, enable software to be loaded from one computer or processor to another computer or processor, such as from a management server or host computer to an application server computer platform. Thus, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves such as those used over physical interfaces between local devices, through wired and optical terrestrial networks, and via various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered as media for carrying software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable media" refer to any medium that participates in providing instructions to a processor for execution.
[0258] Accordingly, machine-readable media such as computer-executable code can take many forms including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media includes, for example, optical disks or magnetic disks such as any storage device within any computer that can be used in the implementation of, e.g., a database shown in the drawings. Volatile storage media includes dynamic memory such as main memory of such a computer platform. Physical transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that make up a bus within a computer system. Carrier wave transmission media can take the form of electrical, electromagnetic, acoustic, or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROM, DVD, or DVD-ROM, any other optical media, punch cards, paper tape, any other physical storage media with patterns of holes, RAM, ROM, PROM, EPROM, flash EPROM, any other memory chip or cartridge, carrier waves carrying data or instructions, cables or links carrying such carrier waves, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0259] In some embodiments, some or all of the analysis functions of computer system 700 can be packaged into one software package. In some embodiments, a complete set of data analysis functions can include a suite of software packages. In some embodiments, the data analysis software can be a stand-alone package provided to the user independently of the assay device system. In some embodiments, the software can be web-based and can enable users to share data. In some embodiments, off-the-shelf software can be used to perform all or part of the data analysis. For example, Seven Bridges (https: / / www.sbgenomics.com / ) software can be used to compile a table of copy numbers of one or more genes performed on each cell of an entire collection of cells.
[0260] The methods and systems of the present disclosure can be implemented by one or more algorithms or methods. The methods can be implemented by software executed by central processing unit 705. Exemplary uses of algorithms or methods implemented by software include bioinformatics methods for sequence data processing (e.g., integration, filtering, trimming, clustering), alignment and calling, and processing of string data and optical density data (e.g., most probable number and culturable abundance determination).
[0261] In an exemplary embodiment, computer system 700 can perform data analysis on an array data set generated by performing a single-cell probabilistic barcoding assay. Examples of data analysis functions include, but are not limited to, (i) algorithms for decoding / inverse multiplexing target sequence data provided by sequencing sample labels, cell labels, spatial labels, and molecular labels, and the probabilistic barcode library created in the execution of the assay; (ii) algorithms for identifying the number of reads per gene per cell and the number of unique transcript molecules per gene per cell based on the data and creating a summary table; (iii) statistical analysis of the sequence data, such as predicting confidence intervals, for example, clustering cells based on gene expression data or identifying the number of transcript molecules per gene per cell; (iv) algorithms for identifying rare cell subpopulations using, for example, principal component analysis, hierarchical clustering, k-means clustering, self-organizing maps, neural networks, etc.; (v) a sequence alignment function for aligning gene sequence data with known reference sequences and detecting mutations, polymorphic markers, and splice variants; and (vi) automatic clustering of molecular labels to compensate for amplification errors or sequencing errors. In some embodiments, computer system 700 can output sequencing results in a useful graphical format, such as a heat map showing the copy number of one or more genes occurring in each cell of a collection of cells. In some embodiments, computer system 700 can perform an algorithm for extracting biological meaning from sequencing results by correlating, for example, the copy number of one or more genes occurring in each cell of a collection of cells with cell type, rare cell type, or cells derived from a subject having a particular disease or condition. In some embodiments, computer system 700 can perform an algorithm for comparing populations of cells across different biological samples.
Example
[0262] Some aspects of the foregoing in itself will be considered in more detail in the following examples, which are in no way intended to limit the scope of the present disclosure.
[0263] Example 1 Separation of Signal Cell Labels and Noise Cell Labels - Second Derivative This example describes separating signal cell labels (also called true cell labels) from noise cell labels based on the number of leads (or molecules) associated with the cell labels.
[0264] In some situations, noise cell labels may have fewer associated leads (or molecules) than signal cell labels. For example, noise cell labels can arise because cells not paired with beads lyse, their nucleic acid contents diffuse, and become associated with beads that are not paired with any cells. This type of noise cell label can include a portion of the overall diffusible contents of the cell. Thus, molecules from the same cell can appear as if they are from two different cells (e.g., because the cell labels have mutated such that their molecules seem to be from two different beads).
[0265] As another example, noise cell labels can result from mutations during the bead manufacturing process. Also, noise cell labels can result from insufficient exonuclease treatment (e.g., in step 216 of FIG. 2) during the PCR process, where single-stranded DNA on the beads can hybridize and form PCR chimeras. These two types of noise cell labels can occur randomly and rarely.
[0266] Figure 10 shows a non-limiting, exemplary cumulative sum plot. Relationship between the cumulative read count and the sorted cell label index on a log-log scale. The red line indicates the cut-off between true cell labels and noise cell labels. In Figure 10, when all cell labels were sorted based on the read count, a sharp slope change in the cumulative read (or molecule) count was observed. To find the cut-off between true cell labels and noise cell labels, the second derivative of the log-log plot was calculated. Figure 11 shows a non-limiting second derivative plot of the cumulative sum plot in Figure 10. Relationship between the second derivative of the log10-transformed cumulative read count and the second derivative of the log10-transformed sorted cell label index. The global minimum was inferred as the cut-off between true cell labels and noise cell labels.
[0267] In some embodiments, the inferred cell count may not match the cell count input and the cell count observed by image analysis. Instead, the cut-off determined using Figure 11 may reflect either the separation of high-expression level signal cells and low-expression level signal cells or the separation of different types of noise labels. To accurately infer the cell count in these cases, the limit of the ratio of reads (or molecules) in the signal cell label is set in the range of 45% to 92% based on empirical data. The cell count observed from image analysis can be optionally set as a constraint if this value is available.
[0268] Overall, these data show that identifying true cell labels (also called signal cell labels) and noise cell labels can be achieved by identifying the minimum of the second derivative plot corresponding to the cell label threshold that distinguishes true cell labels from noise cell labels.
[0269] Example 2 Separation of Signal Cell Labels and Noise Cell Labels - Clustering This example describes separating signal cell labels (also called true cell labels) and noise cell labels based on the expression pattern (also called the feature vector).
[0270] In some embodiments, samples of stochastic barcoding experiments can include cell types with a wide range of expression levels. In such experiments, some cell types can have molecule counts that closely resemble noise cell labels. When the number of associated molecules is indistinguishable, clustering techniques can be used to classify noise cell labels and each cell type with a low expression level to separate true cell labels from noise cell labels. The method can be based on the assumption that cell labels within the same cell type have expression patterns that are more similar than cell labels between different cell types, and that noise cell labels also have feature vectors that are more similar to each other than to true cell labels.
[0271] Figure 12 shows a non-limiting tSNE plot of signal or noise cell labels. PBMC cells were stochastically barcoded. The 5,450 cell labels in Figure 12 included 240 true cell labels with low expression levels and 5,210 noise cell labels. In particular, the classification was performed by first projecting the expression vectors into a two-dimensional (2D) space using t-distributed stochastic neighbor embedding (tSNE) and then clustering the 2D coordinates by density-based spatial clustering of applications with noise (DBScan). Using the knowledge that the majority of the 5,450 cell labels were noise cell labels, it was concluded that the dominant cluster was the noise label cluster, and the other three small clusters were concluded to be true cell labels of three different cell types.
[0272] Overall, these data show that discriminating true cell labels and noise cell labels can be achieved by clustering the expression patterns associated with the true cell labels.
[0273] Example 3 Discrimination of True Cell Labels and Noise Cell Labels - Second Derivative This example describes separating true cells (also called signal cell labels or true cell labels) and noise (also called noise cells or noise cell labels) based on the number of leads (or molecules) associated with the cells (or cell labels).
[0274] Example dataset 1 This dataset was processed using the BD™ Breast Cancer Gene Panel (BrCa400) with three separate breast cancer cell lines and donor-isolated PBMCs (peripheral blood mononuclear cells). Method 400b, described with reference to FIG. 4, identified 8,017 cells, of which 186 cells were identified as noise cells by method 600a, described with reference to FIG. 6. Method 600a detected an additional 1,263 cells, which were confirmed to be mainly PBMCs. See FIGS. 13A, 13B, 14A - 14D. FIGS. 13A and 13B show a non-limiting, exemplary plot comparing the cells identified by method 400, shown with reference to FIG. 4, for a sample processed using the BD™ Breast Cancer Gene Panel with three separate breast cancer cell lines and donor-isolated PBMCs (FIG. 13A), and the cells identified by method 600a, shown with reference to FIG. 6A (FIG. 13B). The dots labeled blue in both FIGS. 13A and 13B are the common cells detected by both methods. The dots labeled red in FIG. 13A are the cells identified as noise by method 600a. The dots labeled red in FIG. 13B are the additional true cells identified by method 600a. FIG. 14A is a non-limiting, exemplary plot showing the cells identified by method 600a, and the cells labeled red are the additional cells identified (compared to the cells identified by method 400, shown with reference to FIG. 4). The cells are colored by the expression of PBMCs such as B cells (FIG. 14B), NK cells (FIG. 14C), and T cells (FIG. 14D). FIGS. 14B - 14D show that the additional cells identified by method 600a are indeed true cells.
[0275] Example dataset 2 This dataset was processed using the BD™ Blood Gene Panel (Blood500) along with PBMCs isolated from healthy donors. Method 400b, described with reference to FIG. 4, identified 13,950 cells, of which 1,333 cells were identified as noise cells by method 600a, described with reference to FIG. 6. Method 600a detected an additional 3,842 cells, which were confirmed to be mainly T cells and expressed in important genes such as LAT (linker for activation of T cells) and IL7R (interleukin 7 receptor). See FIGS. 15A, 15B, 16A, 16B, and 17A - 17D. FIGS. 15A and 15B show a non - limiting and exemplary plot comparing the cells identified by method 400, shown with reference to FIG. 4, for a sample processed using the BD™ Blood Gene Panel along with PBMCs isolated from healthy donors (FIG. 15A), and the cells identified by method 600a, shown with reference to FIG. 6A (FIG. 15B). The dots labeled blue in both FIGS. 15A and 15B are the common cells detected by both methods. The dots labeled red in FIG. 15A are the cells identified as noise by method 600a. The dots labeled red in FIG. 15B are the additional true cells identified by method 600a. FIGS. 16A and 16B are non - limiting and exemplary plots showing the cells identified by method 400. In FIG. 16A, the dots labeled red are the cells identified as noise by method 600a. In FIG. 16B, the cells are colored by the expression of mononuclear cell marker genes such as CD14 and S100A6. The "noise" cells identified by the improved algorithm mainly had low expression of mononuclear cells. FIG. 17A is a non - limiting and exemplary plot showing the cells identified by method 600a, and the cells labeled red are the additional identified cells. The cells are colored by the expression of T cells (FIG. 17B), the important gene LAT (FIG. 17C), and IL7R (FIG. 17D).
[0276] Overall, data that different embodiments of methods for identifying signal cell labels or true cell labels can have different performances and can be complementary to each other.
[0277] In at least some of the above-described embodiments, one or more of the elements used in an embodiment can be used interchangeably in another embodiment, unless replacement in another embodiment is not technically feasible. Those skilled in the art will understand that various other omissions, additions, and modifications can be made to the methods and structures described above without departing from the scope of the spirit recited in the claims. All such variations and changes are intended to be within the scope of the spirit defined by the appended claims.
[0278] Regarding the use of almost any plural and / or singular terms herein, those skilled in the art can convert from plural to singular and / or from singular to plural as appropriate for the situation and / or application. Various singular / plural substitutions may be explicitly described herein for the sake of clarity. As used in this specification and the appended claims, the singular forms "a", "an", and "the" include the plural form unless the context clearly indicates otherwise. Any reference to "or" herein is intended to include "and / or" unless otherwise stated.
[0279] Generally, it is understood by those skilled in the art that the terms used in this specification, particularly in the appended claims (e.g., the body of the appended claims), are generally intended to be "open" terms (e.g., the term "including" should be construed as "including but not limited to", the term "having" should be construed as "having at least", the term "include" should be construed as "including but not limited to", etc.). When a specific number is intended in the description of the introduced claim, such description is clearly stated in the claim, and it is further understood by those skilled in the art that when there is no such description, such intention does not exist. For example, for the sake of understanding, the following appended claims may introduce the description of the claim using introductory phrases such as "at least one" and "one or more". However, the use of such phrases should not be construed as implying that the introduction of the claim description by the indefinite article "a" or "an", even if the same claim contains an introductory phrase such as "one or more" or "at least one" and the indefinite article "a" or "an", limits any particular claim containing such introduced claim description to an embodiment containing only one such description (e.g., "a" and / or "an" should be construed as meaning "at least one" or "one or more"), and the same applies when the claim description is introduced using a definite article. In addition, those skilled in the art understand that when a specific number is clearly stated in the description of the introduced claim, such description should be construed as meaning at least the stated number (e.g., a simple description of "two items" without other modifiers means at least two items, or two or more items). Further, when an expression such as "at least one of A, B, and C, etc." is used, generally, such expression is intended to have the meaning that those skilled in the art understand such expression (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, a system having only A, only B, only C, both A and B, both A and C, both B and C, and / or all of A, B, and C, etc.).When expressions similar to "at least one of A, B, and C, etc." are used, generally, such expressions are intended to have the meaning that a person skilled in the art would understand them (for example, "a system having at least one of A, B, and C" includes, without limitation, a system having only A, only B, only C, both A and B, both A and C, both B and C, and / or all of A, B, and C, etc.). Further, it should be understood by a person skilled in the art that substantially any disjunctive and / or disjunctive clause representing two or more alternative terms is intended to include one of the terms, any of the terms, or both terms in any of the description, claims, or drawings. For example, the clause "A or B" is understood to include the possibilities of "A" or "B" or "A and B".
[0280] In addition, when a feature or aspect of the present disclosure is described by a Markush group, a person skilled in the art will recognize that thereby the present disclosure is also described from the perspective of any individual element or subgroup of elements of the Markush group.
[0281] As will be understood by those skilled in the art, for all and every purpose of providing a description, all ranges disclosed herein include any and all possible sub-ranges and combinations of such sub-ranges. It will be readily recognized that any recited range sufficiently describes and enables the same range to be subdivided into, for example, at least half, one-third, one-fourth, one-fifth, one-tenth, etc. As a non-limiting example, each range recited in the specification can be readily divided into a lower one-third, a middle one-third, an upper one-third, etc. Also, as will be understood by those skilled in the art, all language such as "up to", "at least", "greater than", "less than", etc., includes the recited number and refers to ranges that can be further subdivided into sub-ranges as described above. Finally, as will be understood by those skilled in the art, ranges include individual elements. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, etc.
[0282] Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and not of limitation, and the true scope and spirit are indicated by the following claims.
Claims
1. A method for identifying signal cell labels, comprising: (a) obtaining sequence data of a plurality of barcoded targets, wherein the plurality of barcoded targets are created from a plurality of targets in a plurality of cells barcoded using a plurality of barcodes, each of the plurality of barcodes includes a cell label and a molecular label, barcoded targets created from targets of different cells of the plurality of cells have different cell labels, and barcoded targets created from targets of the same cell of the plurality of cells have different molecular labels; (b) identifying a feature vector of each cell label of the plurality of barcoded targets, wherein the feature vector includes the number of molecular labels having distinct sequences associated with each cell label; (c) identifying clusters of each cell label of the plurality of barcoded targets based on the feature vectors, including clustering each cell label of the plurality of barcoded targets into the clusters based on the distance of the feature vectors to the clusters in the feature vector space; (d) identifying each cell label of the plurality of barcoded targets as a signal cell label or a noise cell label based on the number of cell labels in the cluster and a cluster size threshold, wherein the cell label is identified as a signal cell label when the number of cell labels in the cluster is below the cluster size threshold; The method comprising the above steps.
2. Identifying the clusters of each cell label of the plurality of barcoded targets based on the feature vectors comprises: projecting the feature vectors from the feature vector space to a lower-dimensional space; clustering each cell label into the clusters based on the distance of the feature vectors to the clusters in the lower-dimensional space; The method according to Claim 1, comprising the above steps.
3. Clustering each cell label into the clusters based on the distance of the feature vectors to the clusters in the lower-dimensional space comprises using a density-based method to cluster each cell label into the clusters based on the distance of the feature vectors to the clusters in the lower-dimensional space. The method according to Claim 2, comprising the above steps.
4. The method according to any one of claims 1 to 3, wherein the cell label is identified as a noise cell label when the number of cell labels in the cluster does not fall below the cluster size threshold.
5. The method according to any one of claims 1 to 4, comprising determining the cluster size threshold based on the number of cell labels of the plurality of barcoded targets.
6. The method according to any one of claims 1 to 5, comprising determining the cluster size threshold based on the number of molecular labels having distinct sequences associated with each cell label of the plurality of barcodes.
7. (e)For one or more of the plurality of targets, (1)counting the number of molecular labels having distinct sequences associated with the target in the sequence data; (2)estimating the number of the targets based on the number of molecular labels having distinct sequences associated with the target in the sequence data counted in (1). The method according to any one of claims 1 to 6, comprising.
8. A method for identifying signal cell labels, comprising: (a)obtaining sequence data of a plurality of first targets of cells, wherein each first target is associated with the number of molecular labels having distinct sequences associated with each cell label of the plurality of cell labels; (b)identifying each of the cell labels as a signal cell label or a noise cell label based on the number of molecular labels having distinct sequences associated with each of the cell labels and an identification threshold; (c)re-identifying at least one of the plurality of cell labels identified as signal cell labels in (b) as a noise cell label, identifying the variation in the number of molecular labels having distinct sequences associated with each of the first targets among the plurality of first targets; identifying a plurality of second targets from the plurality of first targets, wherein the variation in the number of molecular labels having distinct sequences associated with each of the second targets exceeds a variation threshold; for each of the plurality of cell labels, re-identifying at least one of the plurality of cell labels identified as signal cell labels in (b) as a noise cell label based on the number of molecular labels having distinct sequences associated with the plurality of second targets and the identification threshold. re-identifying at least one of the plurality of cell labels as a noise cell label by; A method comprising.
9. The method according to claim 8, wherein the discrimination threshold includes a cell label threshold, a cluster size threshold, or any combination thereof.
10. Comprising removing one or more cell labels of the plurality of cell labels, The method according to claim 8 or 9, wherein the one or more cell labels are each associated with the number of molecular labels having distinct sequences below a threshold of the number of molecular labels.
11. Identifying the variation includes identifying an average, maximum, median, minimum, variance, or any combination thereof of the number of molecular labels having distinct sequences associated with each of the first targets among the first targets and the number of cell labels among the plurality of cell labels in the sequence data. The method according to any one of claims 8 to 10.
12. Identifying the variation includes identifying a standard deviation, a normalized variance, or any combination thereof. The method according to any one of claims 8 to 11.
13. (b) re-identifying at least one of the plurality of cell labels identified as a noise cell label as a signal cell label, identifying a plurality of third targets among the plurality of first targets, and identifying the plurality of third targets includes after re-identifying at least one of the plurality of cell labels identified as a signal cell label in (b) as a noise cell label, determining a set of the remaining signal cell labels among the plurality of cell labels identified as a signal cell label in (b); identifying the plurality of third targets based on the difference in the number of molecular labels having distinct sequences associated with each of the plurality of first targets for (i) the plurality of cell labels and (ii) a set of the remaining signal cell labels; including identifying; for each of the plurality of cell labels, re-identifying at least one of the cell labels identified as a noise cell label in (b) as a signal cell label based on (1) the number of molecular labels having distinct sequences associated with the plurality of third targets and (2) the discrimination threshold; A method according to any one of claims 8 to 12, comprising re-identifying at least one of the plurality of cell labels as a signal cell label by.
14. A computer system for identifying the number of targets, comprising: a hardware processor; and a non-transitory memory storing instructions, wherein when the instructions are executed by the hardware processor, the processor is caused to execute the method according to any one of claims 1 to 13. A computer system.
15. A computer-readable medium including a program for causing a computer system to execute the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Sequencing process
WO2015177570A1
Devices and systems for molecular barcoding of nucleic acid targets in single cells
WO2016118915A1
Methods and compositions for labeling targets and haplotype phasing
WO2016149418A1