Immune Receptor Barcode Error Correction

By using multiple random barcodes in cell analysis for target barcode and data processing through directional adjacency technology, the correction problems of replacement and non-substituted errors in the prior art are solved, and more accurate molecular counting is achieved.

CN111247589BActive Publication Date: 2025-06-03BECTON DICKINSON & CO
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201880068906.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-09-25
Filing Date
2018-09-24
Publication Date
2025-06-03
Estimated Expiration
2038-09-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively correct substitution errors and non-substitution errors in cell analysis, resulting in inaccurate molecular counting.

Method used

The appearance of the target is estimated by using multiple random barcodes to barcode the target, sequencing data is obtained, and data folding and noise molecular marker removal are carried out by directing clusters of putative sequences and molecular marker sequences of the identified target.

Benefits of technology

Effective correction of substitution errors and non-substituted errors is achieved, and the accuracy of molecular counting is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111247589B_ABST
    Figure CN111247589B_ABST
Patent Text Reader

Abstract

Disclosed herein are methods and systems for determining the presence of a target. In some embodiments, the method includes: folding a putative sequence of the target; folding a molecular marker sequence associated with the putative sequence of the target; and estimating the presence of the target, wherein the estimated presence of the target is related to the presence of the molecular marker sequence associated with the putative sequence of the target in the sequencing data after folding the presence of the putative sequence of the target and folding the presence of the noise molecular marker sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of U.S. Provisional Application No. 62 / 562,978, filed on September 25, 2017. The content of the related application is hereby incorporated by reference in its entirety.

[0003] Reference to sequence listing

[0004] This application is being filed together with a sequence listing in electronic format. The sequence listing is provided as a file entitled Sequence_Listing_BDCRI_035WO.txt, created on September 21, 2018, and is 32 kilobytes in size. The information of the sequence listing in electronic format is hereby incorporated by reference in its entirety. Background of the invention Field of technology

[0006] This disclosure generally relates to the field of molecular barcoding and more particularly to using molecular markers to correct substitution errors and non - substitution errors. Background art

[0007] Methods and techniques such as barcoding (including random barcoding) are useful for cell analysis, particularly for decrypting gene expression profiles using, for example, reverse transcription, polymerase chain reaction (PCR) amplification, and next - generation sequencing (NGS) to determine the state of a cell. However, these methods and techniques may introduce errors such as substitution errors (e.g., substitution errors involving one or more base substitutions) and non - substitution errors (e.g., primer cross - over errors and PCR chimera errors), which, if uncorrected, can lead to overestimated molecular counts. Therefore, there is a need for methods and techniques that can correct various errors to obtain accurate molecular counts. Summary of the invention

[0008] The disclosures herein include methods for determining the presence of a target. In some embodiments, the method includes: (a) barcoding (e.g., randomly barcoding) a plurality of targets using a plurality of barcodes (e.g., random barcodes) to create a plurality of barcoded targets (e.g., randomly barcoded targets), wherein each of the plurality of barcodes includes a cell marker and a molecular marker, wherein the molecular markers of at least two of the plurality of barcodes include different molecular marker sequences, and wherein at least two of the plurality of barcodes include cell markers having the same cell marker sequence; (b) obtaining sequencing data of the barcoded targets; and (c) for at least one of the plurality of targets: (i) identifying a putative sequence of the target in the sequencing data; (ii) counting the occurrences of the molecular marker sequences associated with the putative sequence of the target identified in (i); (iii) identifying clusters of the putative sequences of the target; (iv) collapsing the obtained sequencing data using the clusters of the putative sequences of the target identified in (iii); (v) identifying clusters of the molecular marker sequences associated with the putative sequence of the target; (vi) collapsing the sequencing data using the clusters of the molecular marker sequences identified in (v); (vii) identifying clusters of combined sequences, wherein each combined sequence includes a sequence in the sequence of the target and an associated molecular marker sequence in the molecular marker sequences; (viii) collapsing the sequencing data using the clusters of the combined sequences identified in (vii); (ix) identifying one or more putative sequences of the target corresponding to one or more chimeric sequences of the target, wherein the occurrences of the one or more putative sequences of the target corresponding to one or more chimeric sequences of the target are less than the occurrences of the remaining one or more putative sequences of the target that do not correspond to one or more chimeric sequences of the target; (x) removing from the sequencing data the one or more putative sequences of the target identified in (ix) that correspond to one or more chimeric sequences of the target; and (xi) estimating the presence of the target, wherein after collapsing the sequencing data in (iv), (vi), and (viii) and removing in (x) the one or more putative sequences of the target that correspond to one or more chimeric sequences of the target, the estimated presence of the target is related to the number of molecular marker sequences counted in (ii).

[0009] In some embodiments, the plurality of targets include targets of the entire transcriptome of a cell. The plurality of targets may include genes. The genes may include variable sequences encoding immune receptors, such as variable (V) regions, diversity (D) regions, joining (J) regions, or any combination thereof. The genes may be genes encoding T cell receptors. The putative sequences of the target may differ from each other by at least one nucleotide.

[0010] In some embodiments, clustering the putative sequences of the target to identify the clusters of the putative sequences of the target includes using directed adjacency to identify the clusters of the putative sequences of the target. The putative sequences of the target within the cluster may be within a first predetermined directed adjacency threshold of each other. The first directed adjacency threshold may be the Hamming distance, and the Hamming distance is one. The putative sequences of the target within the cluster may include one or more parent sequences and one or more subsequences of the one or more parent sequences, and wherein the occurrence of the parent sequence is greater than or equal to a first predetermined directed adjacency occurrence threshold. The first predetermined directed adjacency occurrence threshold may be twice the occurrence of a subsequence that is less than one.

[0011] In some embodiments, collapsing the sequencing data obtained in (b) using the clusters of the putative sequences of the target identified in (iii) includes: attributing the occurrence of a subsequence in the one or more subsequences to the parent sequence of the subsequence.

[0012] In some embodiments, clustering the molecular marker sequences associated with the putative sequences of the target to identify the clusters of the molecular marker sequences associated with the putative sequences of the target includes using directed adjacency to identify the clusters of the molecular marker sequences associated with the putative sequences of the target. The molecular marker sequences of the target within the cluster may be within a second predetermined directed adjacency threshold of each other. The second directed adjacency threshold may be the Hamming distance, and the Hamming distance is one. The putative molecular marker sequences of the target within the cluster may include one or more parent molecular marker sequences and one or more sub-molecular marker sequences of the one or more parent molecular marker sequences, and wherein the occurrence of the parent molecular marker sequence is greater than or equal to a second predetermined directed adjacency occurrence threshold. The second predetermined directed adjacency occurrence threshold may be twice the occurrence of a sub-molecular marker sequence that is less than one.

[0013] In some embodiments, collapsing the sequencing data using the clusters of the molecular marker sequences associated with the sequence of the target identified in (v) includes: attributing the occurrence of a sub-molecular marker sequence in the one or more sub-molecular marker sequences to the parent molecular marker of the sub-molecular marker sequence.

[0014] In some embodiments, clustering the combined sequences to identify the clusters of the combined sequences includes using directed adjacency to identify the clusters of the combined sequences. The combined sequences within the cluster may be within a third predetermined directed adjacency threshold of each other. The third directed adjacency threshold may be the Hamming distance, and the Hamming distance is one. The combined sequences within the cluster may include one or more parent combined sequences and one or more sub-combined sequences of the one or more parent combined sequences, and wherein the occurrence of the parent combined sequence is greater than or equal to a third predetermined directed adjacency occurrence threshold. The third predetermined directed adjacency occurrence threshold may be twice the occurrence of a sub-combined sequence that is less than one.

[0015] In some embodiments, collapsing the sequencing data using the clusters of the combined sequences identified in (vii) includes: attributing the occurrences of the sub-combined sequences in the one or more sub-combined sequences to the parent combined sequences of the sub-combined sequences.

[0016] In some embodiments, identifying one or more putative sequences of the target corresponding to one or more chimeric sequences of the target: identifying a putative sequence of the target associated with a molecular marker sequence in the plurality of molecular sequences; identifying a putative sequence among the putative sequences of the target associated with the one molecular marker sequence, the occurrence of the one molecular marker sequence being less than a chimeric occurrence threshold corresponding to a chimeric sequence in the one or more chimeric sequences of the target. The value of the chimeric occurrence threshold can be the occurrence of a putative sequence among the putative sequences of the target associated with a molecular marker sequence, the occurrence being greater than the occurrence of any other sequence among the putative sequences of the target.

[0017] In some embodiments, the method further comprises: adjusting the sequencing data after folding the sequencing data in (iv), (vi), and (viii) and removing one or more putative sequences of the target corresponding to one or more chimeric sequences of the target in (x). Adjusting the sequencing data after folding the sequencing data in (iv), (vi), and (viii) and removing one or more putative sequences of the target corresponding to one or more chimeric sequences of the target in (x) may comprise: after folding the sequencing data in (iv), (vi), and (viii) and removing one or more putative sequences of the target corresponding to one or more chimeric sequences of the target in (x), thresholding the molecular marker sequences associated with the putative sequences of the target to determine signal molecular marker sequences and noise molecular marker sequences associated with the sequences of the target in the sequencing data counted in (b). Thresholding the molecular marker sequences associated with the putative sequences of the target may comprise performing a statistical analysis on the molecular marker sequences of the target. Performing the statistical analysis may comprise: fitting the molecular marker sequences associated with the putative sequences of the target and their occurrences to two negative binomial distributions; using the two negative binomial distributions to determine the occurrence n of the signal molecular marker sequences; and after folding the sequencing data in (iv), (vi), and (viii) and removing one or more putative sequences of the target corresponding to one or more chimeric sequences of the target in (x), removing the noise molecular marker sequences from the sequencing data obtained in (b), wherein the noise molecular marker sequences comprise molecular marker sequences whose occurrences are less than the occurrence of the nth most abundant molecular marker, and wherein the signal molecular marker sequences comprise molecular marker sequences whose occurrences are greater than or equal to the occurrence of the nth most abundant molecular marker. The two negative binomial distributions may comprise a first negative binomial distribution corresponding to the signal molecular marker sequences and a second negative binomial distribution corresponding to the noise molecular marker sequences.

[0018] Disclosed herein is a method for determining the occurrence of a target. In some embodiments, the method comprises: (a) receiving sequencing data of a plurality of targets, wherein the sequencing data comprises putative sequences of the targets in the plurality of targets and the occurrences of molecular marker sequences associated with the sequences of the targets in the sequencing data; (b) folding the putative sequences of the target; (c) folding the molecular marker sequences associated with the putative sequences of the target; and (d) estimating the occurrence of the target, wherein the estimated occurrence of the target is related to the occurrences of the molecular marker sequences associated with the putative sequences of the target in the sequencing data after folding the occurrences of the putative sequences of the target in (b) and folding the occurrences of the noise molecular marker sequences determined in (c).

[0019] In some embodiments, the method includes: identifying the sequence of the target in the sequencing data; and counting the occurrences of molecular marker sequences associated with the sequence of the target in the sequencing data.

[0020] In some embodiments, the method includes: folding clusters of combined sequences, where each combined sequence includes a sequence in the sequence of the target and an associated molecular marker sequence in the molecular marker sequences, and where, after folding the occurrences of the combined sequences, the estimated occurrences of the target are related to the occurrences of the molecular marker sequences associated with the sequence of the target in the sequencing data. Folding the clusters of combined sequences may include: using directed adjacency to fold the clusters of combined sequences. Using directed adjacency to fold the clusters of combined sequences may include: using directed adjacency to identify the clusters of combined sequences; and using the identified clusters of combined sequences to fold the sequencing data. Folding the putative sequence of the target includes: using directed adjacency to fold the putative sequence of the target. Using directed adjacency to fold the putative sequence of the target may include: using directed adjacency to identify the clusters of the putative sequence of the target; and using the identified clusters of the putative sequence of the target to fold the sequencing data. In some embodiments, folding the molecular marker sequences associated with the putative sequence of the target includes: using directed adjacency to fold the molecular marker sequences associated with the putative sequence of the target. Using directed adjacency to fold the molecular marker sequences associated with the putative sequence of the target may include: using directed adjacency to identify the clusters of the molecular marker sequences associated with the putative sequence of the target; and using the identified clusters of the molecular marker sequences associated with the putative sequence of the target to fold the sequencing data.

[0021] In some embodiments, the method includes: identifying one or more putative sequences of the target corresponding to one or more chimeric sequences of the target, wherein the occurrence of the one or more putative sequences of the target corresponding to the one or more chimeric sequences of the target is less than the occurrence of the remaining one or more putative sequences of the target that do not correspond to the one or more chimeric sequences of the target; and removing from the sequencing data the one or more putative sequences of the target corresponding to the one or more chimeric sequences of the target. Identifying the one or more putative sequences of the target corresponding to the one or more chimeric sequences of the target may include: identifying a putative sequence of the target associated with a molecular marker sequence among the plurality of molecular sequences; identifying a putative sequence among the putative sequences of the target associated with the one molecular marker sequence, the occurrence of the one molecular marker sequence being less than a chimeric occurrence threshold corresponding to a chimeric sequence among the one or more chimeric sequences of the target. The value of the chimeric occurrence threshold may be the occurrence of a putative sequence among the putative sequences of the target associated with a molecular marker sequence, the occurrence being greater than the occurrence of any other sequence among the putative sequences of the target.

[0022] In some embodiments, the method includes: determining the sequencing status of the target in the sequencing data; and determining the occurrence of a noise molecular marker sequence associated with a putative sequence of the target in the sequencing data, wherein the estimated occurrence of the target is related to the occurrence of the molecular marker sequence associated with the putative sequence of the target in the sequencing data, and the sequencing data is adjusted according to the occurrence of the noise molecular marker sequence. The sequencing status of the target in the sequencing data may be saturated sequencing, under-sequencing, or over-sequencing.

[0023] In some embodiments, the under-sequencing status may be determined by a target having a depth less than a predetermined under-sequencing threshold, and wherein the depth of the target includes the average, minimum, or maximum depth of the molecular marker sequence associated with the putative sequence of the target in the sequencing data. The under-sequencing threshold may be about four. The under-sequencing threshold may be independent of the number of molecular marker sequences. If the sequencing status of the target in the sequencing data is the under-sequencing status, the number of determined noise molecular marker sequences may be zero.

[0024] In some embodiments, the saturated sequencing state is determined by the number of the molecular marker sequences associated with the putative sequence of the target being greater than a saturation threshold. If the molecular marker sequences among the molecular marker sequences associated with the putative sequence of the target have sequences selected from approximately 6,561 molecular marker sequences, the saturation threshold may be approximately 6,557. If the molecular marker sequences among the molecular marker sequences associated with the putative sequence of the target have sequences selected from approximately 65,536 molecular marker sequences, the predetermined saturation threshold may be approximately 65,532. If the sequencing state of the target in the sequencing data is the saturated sequencing state, the number of determined noise molecular marker sequences may be zero.

[0025] In some embodiments, the over-sequenced state is determined by a target having a depth greater than a predetermined over-sequencing threshold, wherein the depth of the target includes the average, minimum, or maximum depth of the molecular marker sequences associated with the putative sequence of the target in the sequencing data. If the molecular marker sequences among the molecular marker sequences associated with the putative sequence of the target have sequences selected from approximately 6,561 molecular marker sequences, the over-sequencing threshold may be approximately 250. In some embodiments, the method may include: if the sequencing state of the target in the sequencing data is the saturated sequencing state or the over-sequenced state: subsampling the number of molecular marker sequences associated with the sequence of the target in the sequencing data to approximately the predetermined over-sequencing threshold.

[0026] In some embodiments, determining the occurrence of noise molecular marker sequences associated with the putative sequence of the target in the sequencing data includes: if the negative binomial distribution fitting condition is satisfied, fitting a signal negative binomial distribution to the occurrence of the molecular marker sequences associated with the sequence of the target in the sequencing data, wherein the signal negative binomial distribution corresponds to the occurrence of the molecular marker sequences associated with the sequence of the target in the sequencing data (as signal molecular marker sequences); fitting a noise negative binomial distribution to the occurrence of the molecular marker sequences associated with the sequence of the target in the sequencing data, wherein the noise negative binomial distribution corresponds to the occurrence of the molecular marker sequences associated with the sequence of the target in the sequencing data (as noise molecular marker sequences); and using the signal negative binomial distribution and the noise negative binomial distribution to determine the occurrence of the noise molecular marker sequences. In some embodiments, the negative binomial distribution fitting condition may include: the sequencing status of the target in the sequencing data is not the under-sequenced status or the over-sequenced status. Using the signal negative binomial distribution and the noise negative binomial distribution to determine the number of noise molecular marker sequences may include: for each of the molecular marker sequences associated with the putative sequence of the target in the sequencing data: determining the signal probability of the molecular marker sequence in the signal negative binomial distribution; determining the noise probability of the molecular marker sequence in the noise negative binomial distribution; and if the signal probability is less than the noise probability, determining the molecular marker sequence as a noise molecular marker.

[0027] In some embodiments, determining the occurrence of noise molecular marker sequences associated with the sequence of the target in the sequencing data includes: if the sequencing status of the target in the sequencing data is not the under-sequenced status or the over-sequenced status and the occurrence of the molecular marker sequences associated with the sequence of the target in the sequencing data is less than the pseudo-point threshold, adding pseudo-points to the occurrence of the molecular marker sequences associated with the sequence of the target in the sequencing data before determining the occurrence of the noise molecular marker sequences associated with the sequence of the target in the sequencing data. The pseudo-point threshold may be ten. Determining the occurrence of noise molecular marker sequences associated with the sequence of the target in the sequencing data may include: if the sequencing status of the target in the sequencing data is not the under-sequenced status or the over-sequenced status and the occurrence of the molecular marker sequences associated with the sequence of the target in the sequencing data is not less than the pseudo-point threshold, removing non-unique molecular marker sequences when determining the occurrence of the noise molecular marker sequences associated with the sequence of the target in the sequencing data.

[0028] In some embodiments, receiving sequencing data for the plurality of targets includes: barcoding (e.g., randomly barcoding) the plurality of targets using a plurality of barcodes (e.g., random barcodes) to create a plurality of barcoded targets (e.g., randomly barcoded targets), wherein each of the plurality of barcodes includes a cell marker and a molecular marker, wherein the molecular markers of at least two of the plurality of barcodes include different molecular marker sequences, and wherein at least two of the plurality of barcodes include cell markers having the same cell marker sequence; and obtaining sequencing data for the barcoded targets. Barcoding (e.g., randomly barcoding) the plurality of targets in the plurality of cells using the plurality of barcodes to create a plurality of barcoded targets for the cells in the plurality of cells can include: barcoding (e.g., randomly barcoding) the plurality of targets using a plurality of barcodes of particles to create a plurality of barcoded targets, wherein the particles include a subset of the plurality of barcodes, wherein each of the subset of barcodes includes the same cell marker sequence and has at least 100 different molecular marker sequences.

[0029] In some embodiments, the particles are beads. The beads can be selected from the group consisting of: streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A-conjugated beads, protein G-conjugated beads, protein A / G-conjugated beads, protein L-conjugated beads, oligo(dT)-conjugated beads, silica beads, silica-like beads, antibiotin microbeads, anti-fluorescent dye microbeads, and any combination thereof. The particles can include materials selected from the group consisting of: polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic substance, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, agarose gel, cellulose, nylon, silicone, and any combination thereof. The barcode (e.g., random barcode) of the particles includes a molecular marker having at least 1000, 10000, or any combination thereof of different molecular marker sequences.

[0030] In some embodiments, the molecular markers of the barcodes (e.g., random barcodes) include random sequences. The particles may include at least 10,000 barcodes. Using the plurality of barcodes (e.g., random barcodes) to barcode the plurality of targets (e.g., random barcode the targets) to create a plurality of barcoded targets (e.g., randomly barcoded targets) may include: (i) contacting a copy of the target with a target binding region of the barcode; and (ii) reverse transcribing the plurality of targets using the plurality of barcodes to create a plurality of reverse transcribed targets. In some embodiments, the method includes: prior to obtaining sequencing data of the plurality of barcoded targets, amplifying the barcoded targets to generate a plurality of barcoded targets (e.g., amplified randomly barcoded targets). Amplifying the barcoded targets to generate a plurality of randomly barcoded targets may include: amplifying the barcoded targets by polymerase chain reaction (PCR).

[0031] In some embodiments, a computer system for determining the occurrence of a target is disclosed. The computer system may include: a hardware processor; and a non-transitory memory having instructions stored thereon that, when executed by the hardware processor, cause the processor to perform the method as described in any of the above. A computer-readable medium is disclosed. In some embodiments, the computer-readable medium includes executable code for performing the method as described in any of the above. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 Non-limiting exemplary random barcodes are shown.

[0033] Figure 2 A non-limiting exemplary workflow of random barcoding and digital counting is shown.

[0034] Figure 3 is a schematic diagram showing a non-limiting exemplary process for generating an indexed library of randomly barcoded targets from a plurality of targets.

[0035] Figure 4 is a schematic diagram showing a non-limiting exemplary distribution of molecular marker errors, sample labeling errors, and true molecular marker signals.

[0036] Figure 5 is a flowchart showing a non-limiting exemplary embodiment of correcting PCR and sequencing errors based on directed adjacency using molecular markers.

[0037] Figure 6 is a flowchart showing a non-limiting exemplary embodiment of correcting PCR and sequencing errors based on recursive substitution error correction and distribution-based error correction.

[0038] Figure 7 is a schematic diagram showing a non - limiting exemplary embodiment of immune receptor barcode correction based on recursive substitution error correction.

[0039] Figure 8 is a flowchart showing a non - limiting exemplary embodiment of correcting errors in nucleotide sequences and molecular markers (using recursive substitution error correction) and correcting errors in sequencing data attributed to PCR chimeras.

[0040] Figure 9 is a schematic diagram of one possible source of immune receptor chimeras.

[0041] Figure 10 shows a non - limiting exemplary instrument suitable for use in the methods of the present disclosure.

[0042] Figure 11 shows a non - limiting exemplary architecture of a computer system that can be used in conjunction with the embodiments of the present disclosure.

[0043] Figure 12 shows a non - limiting exemplary architecture of a network showing multiple computer systems suitable for use in the methods of the present disclosure.

[0044] Figure 13 shows a non - limiting exemplary architecture of a multiprocessor computer system using a shared virtual address storage space according to the methods of the present disclosure.

[0045] Figure 14 is an exemplary graph of the theoretical calculation of unique molecular markers used as input molecule increments.

[0046] Figure 15 is an exemplary graph of the molecular marker coverage of each molecular marker of a microplate of the highly expressed gene - ATCB, where different distributions are observed between the incorrect molecular markers and the true molecular markers.

[0047] Figure 16 is an exemplary graph showing the fitting of two negative binomial distributions to the molecular marker coverage of each molecular marker of a microplate of the highly expressed gene - ATCB. The fitting of the two negative binomial distributions demonstrates that molecular marker errors with lower molecular marker depth and true molecular markers with higher molecular marker depth can be statistically distinguished. The x - axis is the molecular depth.

[0048] Figure 17 shows molecular marker correction where a pairwise Hamming distance of 1 is over - represented. After molecular marker correction, molecular markers that are one Hamming distance apart are clustered and folded into the same parental molecular marker.

[0049] Figure 18A curve showing the number of corrected molecular markers versus the number of corrected read coverages is presented.

[0050] Figure 19 A schematic diagram showing an example of recursive substitution error correction is presented.

[0051] Figures 20A to 20C Exemplary results showing two negative binomial distribution corrections for PCR and sequencing errors based on CD69 are presented.

[0052] Figures 21A to 21C Exemplary results showing two negative binomial distribution corrections for PCR and sequencing errors based on CD3E are presented.

[0053] Figures 22A to 22J An exemplary non - restrictive validation of a dataset corrected using two negative binomial distributions is presented.

[0054] Figures 23A to 23D Precise TM Exemplary t - stochastic neighbor embedding (t - SNE) visualization of a targeted assay from 96 - well mixed Jurkat and breast cancer (BrCa) single cells (86 genes examined) is presented.

[0055] Figures 24A to 24B A non - restrictive exemplary graph showing differential expression analysis between cell clusters of genes with >0 ML in two selected clusters calculated by DBScan and determined by gene marker levels within each cluster is presented.

[0056] Figures 25A to 25D A non - restrictive exemplary graph showing t - stochastic neighbor embedding (t - SNE) visualization of a BD Precise TM targeted assay from 96 - well mixed Jurkat and breast cancer (T47D) single cells (86 genes examined) is presented.

[0057] Figures 26A to 26B Before any error correction step ( Figure 26A the original ML shown), and after RSEC and DBEC corrections ( Figure 26B the adjusted ML shown), a non - restrictive exemplary heatmap showing differential gene expression in molecular marker counts between different cell clusters identified by Figures 25A to 25D is presented.

[0058] Figures 27A to 27B A table showing a non - restrictive example demonstrating immunoreceptor barcode error correction using recursive substitution error correction is presented.

[0059] Figure 28 A histogram showing non - restrictive exemplary results of immunoreceptor barcode error correction is presented. Detailed implementation mode

[0060] In the following detailed description, reference is made to the accompanying drawings which form a part hereof. In the drawings, like reference numerals generally identify like components unless the context otherwise indicates. The illustrative embodiments described in the specific embodiments, the drawings, and the claims are not meant to be limiting. Other embodiments may be utilized and other changes may be made without departing from the spirit or scope of the subject matter presented herein. It is readily understood that aspects of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are explicitly contemplated herein and form a part of the present disclosure content.

[0061] All patents, published patent applications, other publications, and sequences from GenBank, as well as other databases mentioned herein regarding related technologies are incorporated by reference in their entirety.

[0062] Quantifying small amounts of nucleic acids (e.g., messenger ribonucleic acid (mRNA) molecules) is clinically important for determining genes expressed in cells, for example, at different developmental stages or under different environmental conditions. However, determining the absolute number of nucleic acid molecules (e.g., mRNA molecules) is also very challenging, especially when the number of molecules is very small. One method for determining the absolute number of molecules in a sample is digital polymerase chain reaction (PCR). Ideally, PCR produces identical copies of molecules in each cycle. However, PCR can have the drawback that each molecule replication has a random probability, and this probability varies according to the PCR cycle and gene sequence, which results in amplification bias and inaccurate gene expression measurements. Random barcodes with unique molecular tags (also known as molecular indexes (MI) or universal molecular indexes (UMI)) can be used to count the number of molecules and correct amplification bias. Random barcoding such as Precise TM Assay (Cellular Research, Inc. (Palo Alto, California)) can correct biases induced by PCR and library preparation steps by labeling mRNA with a molecular label (ML) during reverse transcription (RT).

[0063] Precise TM Assay can utilize a non-exhaustive pool of a large number of (e.g., 6561 to 65536) random barcodes, unique molecular tags on poly(T) oligonucleotides, to hybridize with all poly(A)-mRNA in the sample during the RT step. In addition to the molecular label, the sample tag of the random barcode (also known as the sample index (SI)) can be used to identify Precise TMEach well of the plate. The random barcodes may include universal PCR primer sites. During RT, the target gene molecules react randomly with the random barcodes. Each target molecule can hybridize with a random barcode, thereby generating randomly barcoded complementary ribonucleic acid (cDNA) molecules. After labeling, the randomly barcoded cDNA molecules from the wells of the microplate can be pooled into a single tube for PCR amplification and sequencing. The raw sequencing data can be analyzed to yield the number of reads, the number of random barcodes with unique molecular tags, and the number of mRNA molecules based on Poisson correction or a correction method based on two negative binomial distributions.

[0064] In addition to bias correction, molecular tagging can also better understand the statistical quality of the results by revealing the starting number of cDNA molecules present in the observed sequencing reads. For example, a large number of reads may indicate a statistically accurate answer, but if the reads are only from a small number of starting mRNA molecules, the measurement accuracy may be affected.

[0065] Although amplification biases induced by the PCR and library preparation steps can be corrected, for example, by molecular tagging, the quantification of the absolute number of molecules remains challenging due to several other factors. First, the estimation of the number of mRNA molecules can be limited by the overall diversity of the molecular tags. In random barcoding, mRNA molecules can react randomly with the available random barcodes. Thus, each mRNA molecule can hybridize with a random barcode; however, its molecular tag may not necessarily be unique for any given gene. When the number of mRNA molecules is small relative to the number of random barcodes, each mRNA molecule may hybridize with a random barcode with a unique molecular tag, and counting the number of molecules can be equivalent to counting the number of molecular tags.

[0066] As the number of mRNA molecules increases, multiple mRNA molecules become increasingly likely to hybridize with random barcodes with the same molecular tag. Thus, counting using unique molecular tags can underestimate the number of molecules. In some cases, the number of mRNA molecules can be estimated based on Poisson correction or a correction based on two negative binomial distributions of the total number of unique molecular tags observed. However, in the extreme case of observing the entire set of 6561 random barcodes, Poisson correction or correction based on two negative binomial distributions is no longer possible. For example, regardless of 65,000 or 100,000 starting mRNA molecules, at most 6561 saturated random barcodes are expected in any case.

[0067] Second, PCR errors (i.e., errors that occur during PCR amplification) can introduce artificial random barcodes and arbitrarily increase molecular marker counts. Third, PCR amplification bias and inefficient PCR can generate low-copy barcode molecules that are indistinguishable from errors. Fourth, sequencing errors (inaccurate calls of random barcode sequences) can introduce artificial random barcodes and increase molecular marker counts. Additionally, sequencing depth can be important, especially when sequencing is too shallow to detect all of the randomly barcoded mRNAs present in a sample library.

[0068] Substitution errors, primer crossover errors, and PCR chimera errors can occur when performing immune receptor sequencing and analysis. For example, such errors can occur when determining the occurrence or copy number of mRNA molecules encoding an immune receptor, such as a T cell receptor. Immune receptors are highly diverse and closely related genes. Thus, the likelihood of such errors when performing immune receptor sequencing and analysis may be higher when compared to other genes. Such errors typically result in an overestimation of the immunoglobulin repertoire diversity. Methods for mitigating these errors are referred to herein as immune receptor barcode error correction. In some embodiments, immune receptor barcode error correction utilizes recursive substitution error correction to correct substitution errors in molecular markers and nucleotide sequences (e.g., substitution errors in complementarity determining region 3 (CDR3)). For a given sample or cell label, many different CDR3s can be associated with the same molecular marker sequence, resulting in an overestimation of immune receptor diversity. The method can correct PCR chimeras that cross before molecular marking and sample marking, and then identify and remove incorrect molecular markers by distribution-based error correction.

[0069] The disclosures herein include methods for determining the presence of a target. In some embodiments, the method includes: (a) barcoding (e.g., randomly barcoding) a plurality of targets using a plurality of barcodes (e.g., random barcodes) to create a plurality of barcoded targets (e.g., randomly barcoded targets), wherein each of the plurality of barcodes includes a cell marker and a molecular marker, wherein the molecular markers of at least two of the plurality of barcodes include different molecular marker sequences, and wherein at least two of the plurality of barcodes include cell markers having the same cell marker sequence; (b) obtaining sequencing data of the barcoded targets; and (c) for at least one of the plurality of targets: (i) identifying a putative sequence of the target in the sequencing data; (ii) counting the occurrences of the molecular marker sequences associated with the putative sequence of the target identified in (i); (iii) identifying clusters of the putative sequences of the target; (iv) collapsing the obtained sequencing data using the clusters of the putative sequences of the target identified in (iii); (v) identifying clusters of the molecular marker sequences associated with the putative sequence of the target; (vi) collapsing the sequencing data using the clusters of the molecular marker sequences identified in (v); (vii) identifying clusters of combined sequences, wherein each combined sequence includes a sequence in the sequence of the target and an associated molecular marker sequence in the molecular marker sequences; (viii) collapsing the sequencing data using the clusters of the combined sequences identified in (vii); (ix) identifying one or more putative sequences of the target corresponding to one or more chimeric sequences of the target, wherein the occurrences of the one or more putative sequences of the target corresponding to one or more chimeric sequences of the target are less than the occurrences of the remaining one or more putative sequences of the target not corresponding to one or more chimeric sequences of the target; (x) removing from the sequencing data the one or more putative sequences of the target identified in (ix) corresponding to one or more chimeric sequences of the target; and (xi) estimating the presence of the target, wherein the estimated presence of the target is related to the number of molecular marker sequences counted in (ii) after collapsing the sequencing data in (iv), (vi), and (viii) and removing the one or more putative sequences of the target corresponding to one or more chimeric sequences of the target in (x).

[0070] Disclosed herein is a method for determining the presence of a target. In some embodiments, the method includes: (a) receiving sequencing data of a plurality of targets, wherein the sequencing data includes the putative sequence of a target among the plurality of targets and the presence of molecular marker sequences associated with the sequence of the target in the sequencing data; (b) folding the putative sequence of the target; (c) folding the molecular marker sequences associated with the putative sequence of the target; and (d) estimating the presence of the target, wherein after folding the presence of the putative sequence of the target in (b) and folding the presence of the noise molecular marker sequences determined in (c), the estimated presence of the target is related to the presence of the molecular marker sequences associated with the putative sequence of the target in the sequencing data.

[0071] Disclosed is a computer system for determining the presence of a target. Disclosed is a non-transitory computer-readable medium containing executable code that, when executed, causes one or more computing devices to determine the presence of a target.

[0072] Definitions

[0073] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. See, e.g., Singleton et al., Dictionary of Microbiology and Molecular Biology, 2nd ed., J. Wiley & Sons, (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Springs Harbor Press (Cold Springs Harbor, NY 1989). For the purposes of this disclosure, the following terms are defined as follows.

[0074] As used herein, the term "adapter" can refer to a sequence that facilitates the amplification or sequencing of an associated nucleic acid. The associated nucleic acid can include a target nucleic acid. The associated nucleic acid can include one or more of a spatial marker, a target marker, a sample marker, an index marker, or a barcode sequence (e.g., a molecular marker). An adapter can be linear. An adapter can be a pre-adenylated adapter. An adapter can be double-stranded or single-stranded. One or more adapters can be located at the 5' or 3' end of a nucleic acid. When an adapter includes known sequences at the 5' and 3' ends, the known sequences can be the same or different sequences. An adapter located at the 5' and / or 3' end of a polynucleotide is capable of hybridizing to one or more oligonucleotides immobilized on a surface. In some embodiments, an adapter can include a universal sequence. A universal sequence can be a region of a nucleotide sequence that is common to two or more nucleic acid molecules. Two or more nucleic acid molecules can also have regions of different sequences. Thus, for example, a 5' adapter can include the same and / or a universal nucleic acid sequence, and a 3' adapter can include the same and / or a universal sequence. A universal sequence that can be present in different members of a plurality of nucleic acid molecules can allow for the replication or amplification of a plurality of different sequences using a single universal primer that is complementary to the universal sequence. Similarly, at least one, two (e.g., a pair), or more universal sequences that can be present in different members of a collection of nucleic acid molecules can allow for the replication or amplification of a plurality of different sequences using at least one, two (e.g., a pair), or more single universal primers that are complementary to the universal sequences. Thus, a universal primer includes a sequence that can hybridize to such a universal sequence. A molecule carrying a target nucleic acid sequence can be modified to attach a universal adapter (e.g., a non-target nucleic acid sequence) to one or both ends of different target nucleic acid sequences. One or more universal primers attached to a target nucleic acid can provide sites for universal primer hybridization. One or more universal primers attached to a target nucleic acid can be the same or different from each other.

[0075] As used herein, the term "associated" or "associated with" can mean that two or more species can be identified as co-localized at a particular point in time. Association can mean that two or more species are or were within a similar container. Association can be an informatics association. For example, digital information about two or more species can be stored and can be used to determine that one or more of the species are co-localized at a particular point in time. Association can also be a physical association. In some embodiments, two or more associated species are "connected", "attached", or "immobilized" to each other or to a common solid or semi-solid surface. Association can refer to a covalent or non-covalent manner for attaching a marker to a solid or semi-solid support (e.g., a bead). Association can be a covalent bond between a target and a marker. Association can include hybridization between two molecules (e.g., a target molecule and a marker).

[0076] As used herein, the term "complementarity" can refer to the ability of two nucleotides to pair precisely. For example, if the nucleotide at a given position of a nucleic acid can hydrogen bond with the nucleotide of another nucleic acid, the two nucleic acids are considered to be complementary to each other at that position. Complementarity between two single-stranded nucleic acid molecules can be "partial", where only some of the nucleotides bind, or can be complete when there is complete complementarity between the single-stranded molecules. If a first nucleotide sequence is complementary to a second nucleotide sequence, the first nucleotide sequence can be considered a "complement" of the second sequence. If a first nucleotide sequence is complementary to a sequence that is opposite to the second sequence (i.e., the nucleotide order is reversed), the first nucleotide sequence can be considered an "antisense complement" of the second sequence. As used herein, the terms "complement", "complementary", and "antisense complement" can be used interchangeably. It will be understood from this disclosure that if a molecule can hybridize to another molecule, it can be a complement of the hybridizing molecule.

[0077] As used herein, the term "digital counting" can refer to a method for estimating the number of target molecules in a sample. Digital counting can include the step of determining the number of unique labels that have been associated with the target in the sample. This method, which can be random in nature, transforms the problem of counting molecules from one of localizing and identifying identical molecules to a series of yes / no digital questions regarding the detection of a set of predefined labels.

[0078] As used herein, the term(s) "(one or more) 'label'" can refer to a nucleic acid code associated with a target in a sample. A label can be, for example, a nucleic acid label. A label can be a fully or partially amplifiable label. A label can be a fully or partially sequenceable label. A label can be part of a naturally occurring nucleic acid that can be identified as distinct. A label can be a known sequence. A label can include junctions of nucleic acid sequences, such as junctions of natural and non-natural sequences. As used herein, the term "label" can be used interchangeably with the terms "index", "tag", or "label-tag". A label can convey information. For example, in various embodiments, a label can be used to determine the identity of a sample, the source of a sample, the identity of a cell, and / or a target.

[0079] As used herein, the term "non-depleting reservoir" can refer to a pool of barcodes (e.g., random barcodes) consisting of many different labels. A non-depleting reservoir can include a large number of different barcodes such that when the non-depleting reservoir is associated with a target pool, each target may be associated with a unique barcode. The uniqueness of each labeled target molecule can be determined by the statistics of random selection and depends on the copy number of the same target molecule in the set compared to the diverse labels. The size of the resulting set of labeled target molecules can be determined by the random nature of the barcoding process, and then analysis of the number of detected barcodes allows calculation of the number of target molecules present in the original set or sample. When the ratio of the copy number of the target molecules present to the number of unique barcodes is low, the labeled target molecules are highly unique (i.e., the probability of labeling more than one target molecule with a given label is very low).

[0080] As used herein, the term "nucleic acid" refers to a polynucleotide sequence, or a fragment thereof. Nucleic acids can include nucleotides. Nucleic acids can be exogenous or endogenous to a cell. Nucleic acids can be present in a cell-free environment. Nucleic acids can be a gene or a fragment thereof. Nucleic acids can be DNA. Nucleic acids can be RNA. Nucleic acids can include one or more analogs (e.g., altered backbone, sugar, or nucleobase). Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotide, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorinated flavin linked to sugar), thiol-containing nucleotides, biotinylated nucleotides, fluorophore analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, wybutosine, and queuosine. The terms "nucleic acid", "polynucleotide", "target polynucleotide", and "target nucleic acid" can be used interchangeably.

[0081] Nucleic acids can include one or more modifications (e.g., base modifications, backbone modifications) to provide the nucleic acid with new or enhanced characteristics (e.g., improved stability). Nucleic acids can include nucleic acid affinity tags. A nucleoside can be a base-sugar combination. The base portion of a nucleoside can be a heterocyclic base. Two of the most common classes of such heterocyclic bases are purines and pyrimidines. A nucleotide can be a nucleoside further including a phosphate group covalently linked to the sugar portion of the nucleoside. For those nucleosides that include ribose, the phosphate group can be linked to the 2’, 3’, or 5’ hydroxyl moiety of the sugar. In forming a nucleic acid, the phosphate groups can covalently link adjacent nucleosides to one another to form a linear polymeric compound. In turn, the respective ends of this linear polymeric compound can be further joined to form a cyclic compound; however, linear compounds are generally suitable. Additionally, linear compounds can have internal nucleobase complementarity and can thus fold in a manner that gives rise to a fully or partially double-stranded compound. In a nucleic acid, the phosphate groups can generally be referred to as forming the internucleoside backbone of the nucleic acid. The linkage or backbone can be a 3’ to 5’ phosphodiester bond.

[0082] Nucleic acids can include modified backbones and / or modified internucleoside linkages. Nucleic acids can include polynucleotide backbones formed from short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatoms, and alkyl or cycloalkyl internucleoside linkages or one or more short chain heteroatomic or heterocyclic internucleoside linkages. Nucleic acids can include nucleic acid mimics. Nucleic acids can include morpholino backbone structures. Nucleic acids can include linked morpholino units having a heterocyclic base attached to the morpholino ring (i.e., morpholino nucleic acids).

[0083] Nucleic acids can also include nucleobase (commonly referred to as “base”) modifications or substitutions. As used herein, “unmodified” or “natural” nucleobases can include purine bases (e.g., adenine (A) and guanine (G)), and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleobases can include other synthetic as well as natural nucleobases.

[0084] As used herein, the term “sample” can refer to a composition that includes a target. Suitable samples for analysis by the disclosed methods, devices, and systems include cells, tissues, organs, or organisms.

[0085] As used herein, the term “sampling device” or “device” can refer to a device that can take a portion of a sample and / or place the portion on a substrate. Sampling devices can refer to, for example, a fluorescence-activated cell sorter (FACS) machine, a cell sorter, a biopsy needle, a biopsy device, a tissue sectioning device, a microfluidic device, a leaf grid, and / or an ultramicrotome.

[0086] As used herein, the term "solid support" can refer to a discrete solid or semi-solid surface to which multiple barcodes (e.g., random barcodes) can be attached. The solid support can include any type of solid, porous, or hollow sphere, ball, pedestal, cylinder, or other similar configuration, which is made of plastic, ceramic, metal, or polymeric material (e.g., hydrogel), to which nucleic acids can be immobilized (e.g., covalently or non-covalently). The solid support can include discrete particles that can be spherical (e.g., microspheres) or have a non-spherical or irregular shape, such as cubic, rectangular, conical, cylindrical, conical, oval, or disc-shaped. The shape of the beads can be non-spherical. A plurality of solid supports spaced apart in an array may not include a substrate. The solid support can be used interchangeably with the term "bead".

[0087] The solid support can refer to a "substrate". The substrate can be a type of solid support. The substrate can refer to a continuous solid or semi-solid surface on which the methods of the present disclosure can be performed. For example, the substrate can refer to an array, cartridge, chip, device, and slide.

[0088] As used herein, the term "spatial marker" can refer to a marker that can be associated with a position in space.

[0089] As used herein, the term "random barcode" can refer to a polynucleotide sequence that includes the markers of the present disclosure. The random barcode can be a polynucleotide sequence that can be used for random barcoding. The random barcode can be used to quantify targets in a sample. The random barcode can be used to control errors that may occur after a marker is associated with a target. For example, the random barcode can be used to evaluate amplification or sequencing errors. The random barcode associated with a target can be referred to as a random barcode-target or a random barcode-tag-target.

[0090] As used herein, the term "gene-specific random barcode" can refer to a polynucleotide sequence that includes a marker and a gene-specific target-binding region. The random barcode can be a polynucleotide sequence that can be used for random barcoding. The random barcode can be used to quantify targets in a sample. The random barcode can be used to control errors that may occur after a marker is associated with a target. For example, the random barcode can be used to evaluate amplification or sequencing errors. The random barcode associated with a target can be referred to as a random barcode-target or a random barcode-tag-target.

[0091] As used herein, the term "random barcoding" can refer to the random labeling (e.g., barcoding) of nucleic acids. Random barcoding can utilize a recursive Poisson strategy to associate and quantify markers associated with a target. As used herein, the term "random barcoding" can be used interchangeably with "randomly labeling".

[0092] As used herein, the term "target" can refer to a composition that can be associated with a barcode (e.g., a random barcode). Exemplary suitable targets for analysis by the disclosed methods, devices, and systems include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, etc. The target can be single-stranded or double-stranded. In some embodiments, the target can be a protein, peptide, or polypeptide. In some embodiments, the target is a lipid. As used herein, "target" can be used interchangeably with "species".

[0093] As used herein, the term "reverse transcriptase" can refer to a group of enzymes having reverse transcriptase activity (i.e., catalyzing the synthesis of DNA from an RNA template). Generally, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retrons reverse transcriptases, bacterial reverse transcriptases, type II intron-derived reverse transcriptases, and mutants, variants, or derivatives thereof. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retrons reverse transcriptases, and type II intron reverse transcriptases. Examples of type II intron reverse transcriptases include Lactococcus lactis LI.LtrB intron reverse transcriptase, Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases can include many types of non-retroviral reverse transcriptases (i.e., retrons, type II introns, and diversity-generating retroelements, etc.).

[0094] The terms "universal adapter primer", "universal primer adapter", or "universal adapter sequence" are used interchangeably and refer to a nucleotide sequence that can be used to hybridize with a barcode (e.g., a random barcode) to generate a gene-specific barcode. The universal adapter sequence can be, for example, a known sequence that is common to all barcodes used in the methods disclosed herein. For example, when multiple targets are labeled using the methods disclosed herein, each target-specific sequence can be ligated to the same universal adapter sequence. In some embodiments, more than one universal adapter sequence can be used in the methods disclosed herein. For example, when multiple targets are labeled using the methods disclosed herein, at least two target-specific sequences are ligated to different universal adapter sequences. The universal adapter primer and its complement can be included in two oligonucleotides, one of which includes the target-specific sequence and the other of which includes the barcode. For example, the universal adapter sequence can be part of an oligonucleotide that includes the target-specific sequence to generate a nucleotide sequence complementary to the target nucleic acid. A second oligonucleotide that includes the complementary sequence of the barcode and the universal adapter sequence can hybridize with the nucleotide sequence and generate a target-specific barcode (e.g., a target-specific random barcode). In some embodiments, the universal adapter primer has a different sequence from the universal PCR primers used in the methods of the present disclosure.

[0095] Disclosed herein are methods and systems for detecting and / or correcting errors that occur during PCR and / or sequencing. The types of errors can vary and can include, for example, but are not limited to substitution errors (one or more bases) and non-substitution errors. Among substitution errors, the occurrence frequency of a single-base substitution error is much higher than those with a single-base difference. The methods and systems can be used, for example, to provide an accurate count of molecular targets through random barcoding.

[0096] Barcode

[0097] Barcoding, such as random barcoding, has been described, for example, in US 20150299784, WO 2015031691, and Fu et al., Proc Natl Acad Sci U.S.A. [Proceedings of the National Academy of Sciences of the United States of America] May 31, 2011; 108(22):9026 - 31, the contents of these publications are incorporated herein in their entirety. In some embodiments, the barcodes disclosed herein can be random barcodes, which can be polynucleotide sequences that can be used to randomly label targets (e.g., barcodes, tags). If the ratio of the number of different barcode sequences of the random barcode to the number of occurrences of any target to be labeled can be, or about 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values, the barcode can be referred to as a random barcode. The target can be an mRNA species comprising mRNA molecules having the same or nearly the same sequence. If the ratio of the number of different barcode sequences of the random barcode to the number of occurrences of any target to be labeled is at least, or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1, the barcode can be referred to as a random barcode. The barcode sequence of the random barcode can be referred to as a molecular marker.

[0098] Barcodes (e.g., random barcodes) can include one or more markers. Exemplary markers can include universal markers, cell markers, barcode sequences (e.g., molecular markers), sample markers, plate markers, spatial markers, and / or pre - spatial markers. Figure 1 An exemplary barcode 104 with a spatial marker is shown. Barcode 104 can include a 5' amine that can connect the barcode to a solid support 105. The barcode can include a universal marker, a dimensional marker, a spatial marker, a cell marker, and / or a molecular marker. The order of different markers (including but not limited to universal markers, dimensional markers, spatial markers, cell markers, and molecular markers) in the barcode can be changed. For example, as Figure 1As shown, the universal tag can be a 5'-end tag, and the molecular tag can be a 3'-end tag. The spatial tag, the dimensional tag, and the cellular tag can be in any order. In some embodiments, the universal tag, the spatial tag, the dimensional tag, the cellular tag, and the molecular tag are in any order. The barcode can include a target binding region. The target binding region can interact with a target in the sample (e.g., a target nucleic acid, RNA, mRNA, DNA). For example, the target binding region can include an oligo(dT) sequence that can interact with the poly(A) tail of mRNA. In some cases, the tags of the barcode (e.g., the universal tag, the dimensional tag, the spatial tag, the cellular tag, and the barcode sequence) can be separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides.

[0099] The tag (e.g., the cellular tag) can include a set of unique nucleic acid subsequences of defined length, e.g., each seven nucleotides (equivalent to the number of bits used in some Hamming error correction codes), which can be designed to provide error correction capabilities. A set of error correction subsequences including seven-nucleotide sequences can be designed such that any pairwise combination of the sequences in the set exhibits a defined "genetic distance" (or number of mismatched bases), e.g., a set of error correction subsequences can be designed to exhibit a genetic distance of three nucleotides. In such cases, review of the error correction sequences in a sequence data set of the labeled target nucleic acid molecules (described more fully below) can allow detection or correction of amplification or sequencing errors. In some embodiments, the length of the nucleic acid subsequences used to generate the error correction code can vary, e.g., they can be, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50, or a number or range between any two of these values of nucleotide lengths. In some embodiments, nucleic acid subsequences of other lengths can be used to generate the error correction code.

[0100] The barcode (e.g., a random barcode) can include a target binding region. The target binding region can interact with a target in the sample. The target can be, or include, ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNAs each containing a poly(A) tail, or any combination thereof. In some embodiments, the plurality of targets can include deoxyribonucleic acid (DNA).

[0101] In some embodiments, the target binding region can include an oligo(dT) sequence that can interact with the poly(A) tail of the mRNA. One or more labels of the barcode (e.g., universal label, dimensional label, spatial label, cell label, and barcode sequence (e.g., molecular label)) can be separated from the other one or two of the remaining labels of the barcode by a spacer. The spacer can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides. In some embodiments, no labels in the barcode are separated by a spacer.

[0102] Universal Marker

[0103] The barcode can include one or more universal labels. In some embodiments, for all barcodes in a barcode set (attached to a given solid support), one or more universal labels can be the same. In some embodiments, for all barcodes attached to multiple beads, one or more universal labels can be the same. In some embodiments, the universal label can include a nucleic acid sequence capable of hybridizing to a sequencing primer. The sequencing primer can be used to sequence the barcode including the universal label. The sequencing primer (e.g., universal sequencing primer) can include a sequencing primer associated with a high-throughput sequencing platform. In some embodiments, the universal label can include a nucleic acid sequence capable of hybridizing to a PCR primer. In some embodiments, the universal label can include a nucleic acid sequence capable of hybridizing to both a sequencing primer and a PCR primer. The nucleic acid sequence of the universal label capable of hybridizing to a sequencing or PCR primer can be referred to as a primer binding site. The universal label can include a sequence that can be used to initiate transcription of the barcode. The universal label can include a sequence that can be used to extend the barcode or a region within the barcode. The length of the universal label can be, or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides, or a number or range of nucleotides between any two of these values. For example, the universal label can include at least about 10 nucleotides. The length of the universal label can be at least, or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides. In some embodiments, a cleavable linker or modified nucleotide can be part of the universal label sequence to enable the barcode to be cleaved from the support.

[0104] Dimension Marker

[0105] Barcodes can include one or more dimensional tags. In some embodiments, the dimensional tags can include nucleic acid sequences that provide information about the dimension in which the tag (e.g., random tag) occurs. For example, the dimensional tag can provide information about the time at which the target is barcoded. The dimensional tag can be associated with the time of barcoding (e.g., random barcoding) in the sample. The dimensional tag can be activated at the time of tagging. Different dimensional tags can be activated at different times. The dimensional tags provide information about the order in which the targets, target groups, and / or samples are barcoded. For example, a cell population can be barcoded in the G0 phase of the cell cycle. In the G1 phase of the cell cycle, the cells can be pulsed again with a barcode (e.g., random barcode). In the S phase of the cell cycle, the cells can be pulsed again with a barcode, and so on. The barcode at each pulse (e.g., each stage of the cell cycle) can include a different dimensional tag. In this way, the dimensional tags provide information about which targets are tagged at which stage of the cell cycle. Dimensional tags can interrogate many different biological stages. Exemplary biological times can include, but are not limited to, the cell cycle, transcription (e.g., transcription initiation), and transcript degradation. In another example, a sample (e.g., cells, cell population) can be tagged before and / or after treatment with a drug and / or therapy. Changes in the copy number of different targets can indicate the response of the sample to the drug and / or therapy.

[0106] The dimensional tag can be activatable. The activatable dimensional tag can be activated at a specific time point. The activatable tag can be activated, for example, constitutively (e.g., not turned off). The activatable dimensional tag can be activated, for example, reversibly (e.g., the activatable dimensional tag can be turned on and off). The dimensional tag can be reversibly activated, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 times or more. The dimensional tag can be reversibly activated, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 times or more. In some embodiments, the dimensional tag can be activated with fluorescence, light, chemical events (e.g., cleavage, ligation of another molecule, addition of a modification (e.g., PEGylation, sumoylation, acetylation, methylation, deacetylation, demethylation), photochemical events (e.g., photo-locking), and introduction of non-natural nucleotides.

[0107] In some embodiments, the dimensional label can be the same for all barcodes (e.g., random barcodes) attached to a given solid support (e.g., bead), but different for different solid supports (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% of the barcodes on the same solid support can include the same dimensional label. In some embodiments, at least 60% of the barcodes on the same solid support can include the same dimensional label. In some embodiments, at least 95% of the barcodes on the same solid support can include the same dimensional label.

[0108] Multiple solid supports (e.g., beads) can exhibit up to 10 6 or more unique dimensional label sequences. The length of the dimensional label can be, or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides, or a number or range of nucleotides between any two of these values. The length of the dimensional label can be at least, or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides. The dimensional label can include between about 5 and about 200 nucleotides. The dimensional label can include between about 10 and about 150 nucleotides. The dimensional label can include nucleotides with a length between about 20 and about 125.

[0109] Spatial Marker

[0110] A barcode can include one or more spatial labels. In some embodiments, a spatial label can include a nucleic acid sequence that provides information about the spatial orientation of a target molecule associated with the barcode. The spatial label can be associated with coordinates in the sample. The coordinates can be fixed coordinates. For example, substrate-fixed coordinates can be referred to. The spatial label can refer to a two-dimensional or three-dimensional grid. Landmark-fixed coordinates can be referred to. Landmarks are identifiable in space. A landmark can be an imaged structure. A landmark can be a biological structure, such as an anatomical landmark. A landmark can be a cellular landmark, such as an organelle. A landmark can be a non-natural landmark, such as a structure with an identifiable identifier (such as a color code, barcode, magnetism, fluorescence, radioactivity, or unique size or shape). The spatial label can be associated with a physical partition (e.g., a well, container, or droplet). In some embodiments, multiple spatial labels are used together to encode one or more positions in space.

[0111] The spatial label can be the same for all barcodes attached to a given solid support (e.g., bead), but different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support that include the same spatial label can be, or about, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values. In some embodiments, the percentage of barcodes on the same solid support that include the same spatial label can be at least, or at most, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, at least 60% of the barcodes on the same solid support can include the same spatial label. In some embodiments, at least 95% of the barcodes on the same solid support can include the same spatial label.

[0112] Multiple solid supports (e.g., beads) can exhibit up to 10 6 or more distinct spatial label sequences. The length of the spatial label can be, or about, 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides, or a number or range of nucleotides between any two of these values. The length of the spatial label can be at least, or at most, 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides. The spatial label can include between about 5 and about 200 nucleotides. The spatial label can include between about 10 and about 150 nucleotides. The spatial label can include nucleotides with a length between about 20 and about 125.

[0113] Cell Marker

[0114] A barcode (e.g., a random barcode) can include one or more cell markers. In some embodiments, the cell marker can include a nucleic acid sequence that provides information for determining which target nucleic acid is from which cell. In some embodiments, the cell marker is the same for all barcodes attached to a given solid support (e.g., a bead), but different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support that include the same cell marker can be, or about, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values. In some embodiments, the percentage of barcodes on the same solid support that include the same cell marker can be, or about, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. For example, at least 60% of the barcodes on the same solid support can include the same cell marker. As another example, at least 95% of the barcodes on the same solid support can include the same cell marker.

[0115] Multiple solid supports (e.g., beads) can exhibit up to 10 6 or more unique cell marker sequences. The length of the cell marker can be, or about, 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides, or a number or range of nucleotides between any two of these values. The length of the cell marker can be at least, or at most, 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides. For example, the cell marker can include between about 5 and about 200 nucleotides. As another example, the cell marker can include between about 10 and about 150 nucleotides. As yet another example, the cell marker can include nucleotides with a length between about 20 and about 125.

[0116] Barcode Sequence

[0117] A barcode can include one or more barcode sequences. In some embodiments, the barcode sequence can include a nucleic acid sequence that provides identification information for a particular type of target nucleic acid species that hybridizes to the barcode. The barcode sequence can include a nucleic acid sequence that provides a counter (e.g., provides a rough approximation) for a particular occurrence of a target nucleic acid species that hybridizes to the barcode (e.g., the target binding region).

[0118] In some embodiments, a set of different barcode sequences is attached to a given solid support (e.g., a bead). In some embodiments, there can be, or about, 10 2 、10 3, 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 individual, or a unique molecular marker sequence of a number or range between any two of these values. For example, a plurality of barcodes can include approximately 6561 barcode sequences having different sequences. As another example, a plurality of barcodes can include approximately 65536 barcode sequences having different sequences. In some embodiments, there can be at least, or at most 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , or 10 9 individual unique barcode sequences. The unique molecular marker sequences can be attached to a given solid support (e.g., a bead).

[0119] In different embodiments, the length of the barcode can be different. For example, the length of the barcode can be, or approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 individual, or a number or range between any two of these values of nucleotides. As another example, the length of the barcode can be at least, or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides.

[0120] Molecular Marker

[0121] A barcode (e.g., a random barcode) can include one or more molecular markers. The molecular marker can include a barcode sequence. In some embodiments, the molecular marker can include a nucleic acid sequence that provides identification information for a specific type of target nucleic acid species that hybridizes to the barcode. The molecular marker can include a nucleic acid sequence that provides a counter for a specific occurrence of a target nucleic acid species that hybridizes to the barcode (e.g., a target binding region).

[0122] In some embodiments, a set of different molecular markers is attached to a given solid support (e.g., a bead). In some embodiments, there can be, or approximately 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9unique molecular tag sequences that are a number or range of numbers between any two of these values. For example, a plurality of barcodes can include approximately 6,561 molecular tags having different sequences. As another example, a plurality of barcodes can include approximately 65,536 molecular tags having different sequences. In some embodiments, there can be at least, or at most 10 2 10 3 10 4 10 5 10 6 10 7 10 8 or 10 9 unique molecular tag sequences. Barcodes having unique molecular tag sequences can be attached to a given solid support (e.g., beads).

[0123] For random barcoding using a plurality of random barcodes, the ratio of the number of different molecular tag sequences to the number of occurrences of any target can be, or about 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range of numbers between any two of these values. The target can be an mRNA species comprising mRNA molecules having the same or nearly the same sequence. In some embodiments, the ratio of the number of different molecular tag sequences to the number of occurrences of any target is at least, or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1.

[0124] The length of the molecular tag can be, or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides, or a number or range of numbers between any two of these values. The length of the molecular tag can be at least, or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides.

[0125] Target Binding Region

[0126] The barcode may include one or more target binding regions, such as capture probes. In some embodiments, the target binding region may hybridize to a target of interest. In some embodiments, the target binding region may include a nucleic acid sequence that specifically hybridizes (e.g., hybridizes to a specific gene sequence) to a target (e.g., a target nucleic acid, a target molecule, such as a cellular nucleic acid to be analyzed). In some embodiments, the target binding region may include a nucleic acid sequence that can attach (e.g., hybridize) to a specific location of a specific target nucleic acid. In some embodiments, the target binding region may include a nucleic acid sequence that can specifically hybridize to a restriction enzyme site overhang (e.g., an EcoRI sticky end overhang). The barcode can then be ligated to any nucleic acid molecule that includes a sequence complementary to the restriction site overhang.

[0127] In some embodiments, the target binding region may include a non-specific target nucleic acid sequence. A non-specific target nucleic acid sequence may refer to a sequence that can bind to multiple target nucleic acids independent of a specific sequence of the target nucleic acid. For example, the target binding region may include a random polymer sequence or an oligo(dT) sequence that hybridizes to the poly(A) tail on an mRNA molecule. The random polymer sequence may be, for example, a random dimer, trimer, tetramer, pentamer, hexamer, heptamer, octamer, nonamer, decamer, or a higher polymer sequence of any length. In some embodiments, the target binding region is the same for all barcodes attached to a given bead. In some embodiments, for multiple barcodes attached to a given bead, the target binding region may include two or more different target binding sequences. The length of the target binding region may be or about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides, or a number or range of nucleotides between any two of these values. The length of the target binding region may be up to about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 or more nucleotides.

[0128] In some embodiments, the target binding region can include oligo(dT), which can hybridize to mRNA comprising a polyadenylated tail. The target binding region can be gene-specific. For example, the target binding region can be configured to hybridize to a specific region of the target. The length of the target binding region can be or be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides, or a number or range of nucleotides between any two of these values. The length of the target binding region can be at least, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. The length of the target binding region can be about 5 to 30 nucleotides. When the barcode includes a gene-specific target binding region, the barcode can be referred to as a gene-specific barcode.

[0129] Orientation Property

[0130] Random barcodes (e.g., random barcodes) can include one or more orienting features that can be used to orient (e.g., align) the barcodes. The barcodes can include a portion for isoelectric focusing. Different barcodes can include different isoelectric focusing points. When these barcodes are introduced into a sample, the sample can be subjected to isoelectric focusing to facilitate positioning the barcodes in a known manner. In this way, the orienting features can be used to develop a known mapping of the barcodes in the sample. Exemplary orienting features can include electrophoretic mobility (e.g., based on the size of the barcode), isoelectric point, spin, conductivity, and / or self-assembly. For example, a barcode has a self-assembly orienting feature and can self-assemble into a specific orientation (e.g., a nucleic acid nanostructure) when activated.

[0131] Affinity Property

[0132] Barcodes (e.g., random barcodes) can include one or more affinity features. For example, spatial markers can include affinity features. Affinity features can include chemical and / or biological moieties that can facilitate the binding of the barcode to another entity (e.g., a cell receptor). For example, an affinity feature can include an antibody, e.g., an antibody specific for a particular moiety (e.g., a receptor) on a sample. In some embodiments, the antibody can direct the barcode to a specific cell type or molecule. Targets at and / or near a specific cell type or molecule can be labeled (e.g., randomly labeled). In some embodiments, in addition to the nucleotide sequence of the spatial marker, the affinity feature can provide spatial information since the antibody can direct the barcode to a specific location. The antibody can be a therapeutic antibody, e.g., a monoclonal or polyclonal antibody. The antibody can be humanized or chimeric. The antibody can be a naked antibody or a fusion antibody.

[0133] An antibody can be a full-length (i.e., naturally occurring or formed by recombinant processes of normal immunoglobulin gene segments) immunoglobulin molecule (e.g., an IgG antibody) or an immunologically active (i.e., specifically binding) portion of an immunoglobulin molecule (such as an antibody fragment).

[0134] Antibody fragments can be, for example, a portion of an antibody, such as F(ab’)2, Fab’, Fab, Fv, sFv, etc. In some embodiments, an antibody fragment can bind to the same antigen recognized by the full-length antibody. Antibody fragments can include isolated fragments consisting of the variable regions of an antibody, such as an “Fv” fragment consisting of the variable regions of the heavy and light chains and a recombinant single-chain polypeptide molecule (“scFv protein”) in which the light and heavy chain variable regions are linked by a peptide linker. Exemplary antibodies can include, but are not limited to, cancer cell antibodies, viral antibodies, antibodies that bind to cell surface receptors (CD8, CD34, CD45), and therapeutic antibodies.

[0135] Universal Adaptor Primer

[0136] The barcode may include one or more universal adapter primers. For example, gene-specific barcodes (such as gene-specific random barcodes) may include universal adapter primers. A universal adapter primer may refer to a nucleotide sequence that is common to all barcodes. Universal adapter primers can be used to construct gene-specific barcodes. The length of the universal adapter primer can be, or can be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides, or a number or range of nucleotides between any two of these values. The length of the universal adapter primer can be at least, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. The length of the universal adapter primer can be 5 to 30 nucleotides.

[0137] Solid Support

[0138] In some embodiments, the barcodes (such as random barcodes) disclosed herein can be associated with a solid support. For example, the solid support can be synthetic particles. In some embodiments, some or all of the barcode sequences (such as the molecular markers of random barcodes (e.g., the first barcode sequence)) of a plurality of barcodes (e.g., the first plurality of barcodes) on the solid support have at least one nucleotide difference. The cell markers of the barcodes on the same solid support can be the same. The cell markers of the barcodes on different solid supports can have at least one nucleotide difference. For example, the first cell markers of the first plurality of barcodes on the first solid support can have the same sequence, and the second cell markers of the second plurality of barcodes on the second solid support can have the same sequence. The first cell markers of the first plurality of barcodes on the first solid support and the second cell markers of the second plurality of barcodes on the second solid support can have at least one nucleotide difference. The cell marker can be, for example, about 5 to 20 nucleotides long. The barcode sequence can be, for example, about 5 to 20 nucleotides long. The synthetic particles can be, for example, beads.

[0139] The beads can be, for example, silica beads, controlled pore glass beads, magnetic beads, Dynabeads, cross-linked dextran / agarose beads, bead cellulose, polystyrene beads, or any combination thereof. The beads can include materials such as polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic substances, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, agarose gel, cellulose, nylon, silicone, or any combination thereof.

[0140] In some embodiments, the beads can be polymeric microspheres (e.g., deformable beads or gel beads), which are functionalized with barcodes or random barcodes (such as gel beads from 10X Genomics (San Francisco, California)). In some implementations, the gel beads can include polymer-based gels. For example, gel beads can be produced by encapsulating one or more polymer precursors into droplets. After exposing the polymer precursors to a promoter (e.g., tetramethylethylenediamine (TEMED)), gel beads can be produced.

[0141] In some embodiments, the particles can be degradable. For example, the polymeric microspheres can dissolve, melt, or degrade, for example, under desired conditions. The desired conditions can include environmental conditions. The desired conditions can cause the polymeric microspheres to dissolve, melt, or degrade in a controlled manner. Due to chemical stimuli, physical stimuli, biological stimuli, thermal stimuli, magnetic stimuli, electrical stimuli, light stimuli, or any combination thereof, the gel beads can dissolve, melt, or degrade.

[0142] An analyte and / or a reagent (such as an oligonucleotide barcode) can be coupled / fixed, for example, to the inner surface of the gel bead (the accessible interior for diffusion of the oligonucleotide barcode and / or the material for generating the oligonucleotide barcode) and / or the outer surface of the gel bead or any other microcapsule described herein. The coupling / fixing can be via any form of chemical bond (e.g., covalent bond, ionic bond) or physical phenomenon (e.g., van der Waals forces, dipole-dipole interactions, etc.). In some embodiments, the coupling / fixing of the reagent to the gel bead or any other microcapsule described herein can be reversible, for example, via a labile moiety (e.g., via a chemical crosslinker, including the chemical crosslinkers described herein). After applying a stimulus, the labile moiety can be cleaved and release the immobilized reagent. In some embodiments, the labile moiety is a disulfide bond. For example, in the case of immobilizing an oligonucleotide barcode to a gel bead via a disulfide bond, exposing the disulfide bond to a reducing agent can cleave the disulfide bond and release the oligonucleotide barcode from the bead. The labile moiety can be included as part of the gel bead or microcapsule, as part of a chemical linker that connects the reagent or analyte to the gel bead or microcapsule, and / or as part of the reagent or analyte. In some embodiments, at least one barcode of a plurality of barcodes can be fixed on the particle, partially fixed on the particle, encapsulated in the particle, partially encapsulated in the particle, or any combination thereof.

[0143] In some embodiments, the gel beads can include a wide variety of different polymers, including but not limited to: polymers, thermosensitive polymers, photosensitive polymers, magnetic polymers, pH-sensitive polymers, salt-sensitive polymers, chemically sensitive polymers, polyelectrolytes, polysaccharides, peptides, proteins, and / or plastics. The polymers can include but are not limited to the following materials: such as poly(N-isopropylacrylamide) (PNIPAAm), poly(styrenesulfonate) (PSS), poly(allylamine) (PAAm), poly(acrylic acid) (PAA), poly(ethyleneimine) (PEI), poly(diallyldimethylammonium chloride) (PDADMAC), poly(pyrrole) (PPy), poly(vinylpyrrolidone) (PVPON), poly(vinylpyridine) (PVP), poly(methyl methacrylate) (PMAA), poly(methyl methacrylate) (PMMA), polystyrene (PS), poly(tetrahydrofuran) (PTHF), poly(phthalaldehyde) (PTHF), poly(hexyl viologen) (PHV), poly(L-lysine) (PLL), poly(L-arginine) (PARG), poly(lactic-co-glycolic acid) (PLGA).

[0144] Many chemical stimuli can be used to trigger the disruption, dissolution, or degradation of the beads. Examples of such chemical alterations can include but are not limited to pH-mediated changes in the bead wall, decomposition of the bead wall via chemical cleavage of crosslink bonds, triggered depolymerization of the bead wall, and bead wall conversion reactions. Bulk changes can also be used to trigger the disruption of the beads.

[0145] Bulk or physical changes to the microcapsules by various stimuli also offer many advantages in designing capsules for reagent release. The bulk or physical changes occur at the macroscopic scale, where bead rupture is the result of mechanical-physical forces induced by the stimuli. These processes can include but are not limited to pressure-induced rupture, melting of the bead wall, or alteration of the porosity of the bead wall.

[0146] Biological stimuli can also be used to trigger the disruption, dissolution, or degradation of the beads. Generally, biological triggers are similar to chemical triggers, but many examples use biomolecules, or molecules common in living systems, such as enzymes, peptides, sugars, fatty acids, nucleic acids, etc. For example, the beads can include a polymer with peptide crosslinks that are sensitive to cleavage by specific proteases. More specifically, one example can include microcapsules containing GFLGK peptide crosslinks. After adding a biological trigger such as the protease cathepsin B, the peptide crosslinks of the shell pores are cleaved and the contents of the beads are released. In other cases, the protease can be heat-activated. In another example, the beads include a shell wall containing cellulose. The addition of the hydrolase chitosan serves as a biological trigger for cellulose bond cleavage, shell wall depolymerization, and release of the internal contents.

[0147] It is also possible to induce the beads to release their contents after applying a heat stimulus. Changes in temperature can cause various changes in the beads. A change in heat may cause the beads to melt, causing the bead walls to disintegrate. In other cases, heat may increase the internal pressure of the components inside the beads, causing the beads to rupture or explode. In still other cases, heat can cause the beads to become a shrunk and dehydrated state. Heat can also act on the thermosensitive polymers within the bead walls, thereby causing the destruction of the beads.

[0148] Including magnetic nanoparticles within the bead walls of the microcapsules can allow for triggered rupture of the beads as well as guiding the beads into an array. The devices of the present disclosure can include magnetic beads for either purpose. In one example, incorporating Fe 3 O 4 nanoparticles into polyelectrolyte-containing beads triggers rupture in the presence of an oscillating magnetic field stimulus.

[0149] As a result of electrical stimulation, the beads may also be destroyed, dissolved, or degraded. Similar to the magnetic particles described in the previous section, electro-sensitive beads can allow for triggered rupture of the beads as well as other functions such as alignment in an electric field, conductivity, or redox reactions. In one example, beads containing electro-sensitive materials are aligned in an electric field, thereby allowing for control of the release of internal reagents. In other examples, the electric field can cause a redox reaction within the bead walls themselves, which can increase porosity.

[0150] Light stimulation can also be used to disrupt the beads. Many light triggers are possible and can include systems using various molecules (such as nanoparticles and chromophores capable of absorbing photons in a specific wavelength range). For example, a metal oxide coating can be used as a capsule trigger. UV irradiation of polyelectrolyte capsules coated with SiO 2 can cause disintegration of the bead walls. In yet another example, a photo-switchable material (such as an azobenzene group) can be incorporated into the bead walls. After applying UV or visible light, chemicals such as these undergo reversible cis-trans isomerization upon absorption of photons. In this regard, the incorporation of photon switching causes the bead walls to be able to disintegrate or become more porous after applying a photo-trigger.

[0151] For example, in the non-limiting example of barcoding (such as random barcoding) shown in Figure 2 , after introducing cells (such as single cells) into multiple microwells of a microwell array at block 208, beads can be introduced into the multiple microwells of the microwell array at block 212. Each microwell can include one bead. The beads can include multiple barcodes. The barcodes can include an attached 5'-amine region to the bead. The barcodes can include a universal label, a barcode sequence (such as a molecular label), a target binding region, or any combination thereof.

[0152] The barcodes disclosed herein can be associated with (e.g., attached to) a solid support (e.g., a bead). Each barcode associated with a solid support can include a barcode sequence selected from the group consisting of at least 100 or 1000 barcode sequences having unique sequences. In some embodiments, different barcodes associated with a solid support can include barcodes having different sequences. In some embodiments, a percentage of the barcodes associated with a solid support includes the same cell marker. For example, the percentage can be, or can be about, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values. As another example, the percentage can be at least, or at most, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, the barcodes associated with a solid support can have the same cell marker. Barcodes associated with different solid supports can have different cell markers selected from the group consisting of at least 100 or 1000 cell markers having unique sequences.

[0153] The barcodes disclosed herein can be associated with (e.g., attached to) a solid support (e.g., a bead). In some embodiments, a solid support including a plurality of synthetic particles associated with a plurality of barcodes can be used to barcode a plurality of targets in a sample. In some embodiments, the solid support can include a plurality of synthetic particles associated with a plurality of barcodes. The spatial markings of the plurality of barcodes on different solid supports can have a difference of at least one nucleotide. The solid support can, for example, include a plurality of barcodes in two or three dimensions. The synthetic particles can be beads. The beads can be silica beads, controlled pore glass beads, magnetic beads, Dynabeads, cross-linked dextran / agarose beads, beaded cellulose, polystyrene beads, or any combination thereof. The solid support can include a polymer, a matrix, a hydrogel, a needle array device, an antibody, or any combination thereof. In some embodiments, the solid support can be free-floating. In some embodiments, the solid support can be embedded in a semi-solid or solid array. The barcode can be not associated with a solid support. The barcode can be a single nucleotide. The barcode can be associated with a substrate.

[0154] As used herein, the terms "tethered," "attached," and "fixed" can be used interchangeably and can refer to covalent or non-covalent means for attaching a barcode to a solid support. Any of a variety of different solid supports can be used as the solid support for attaching a pre-synthesized barcode or for in situ solid-phase synthesis of a barcode.

[0155] In some embodiments, the solid support is a bead. The bead can include one or more types of solid, porous, or hollow spheres, balls, pedestals, cylinders, or other similar configurations, to which nucleic acids can be immobilized (e.g., covalently or non-covalently). The bead can be composed of, for example, plastic, ceramic, metal, polymeric material, or any combination thereof. The bead can be, or include, spherical (e.g., microspheres) or discrete particles having a non-spherical or irregular shape, such as cubic, rectangular, conical, cylindrical, conical, oval, or disc-shaped. In some embodiments, the shape of the bead can be non-spherical.

[0156] The bead can include a variety of materials, including but not limited to paramagnetic materials (e.g., magnesium, molybdenum, lithium, and tantalum), superparamagnetic materials (e.g., ferrite (Fe 3 O 4 ; magnetite) nanoparticles), ferromagnetic materials (e.g., iron, nickel, cobalt, some of their alloys, and some rare earth metal compounds), ceramics, plastics, glass, polystyrene, silica, methylstyrene, acrylic polymers, titanium, latex, cross-linked agarose, agarose, hydrogel, polymers, cellulose, nylon, or any combination thereof.

[0157] In some embodiments, the bead (e.g., the bead to which the label is attached) is a hydrogel bead. In some embodiments, the hydrogel bead is soluble. In some embodiments, the bead includes a hydrogel.

[0158] Some embodiments disclosed herein include one or more particles (e.g., beads). Each of the particles can include a plurality of oligonucleotides (e.g., barcodes). Each of the plurality of oligonucleotides can include a barcode sequence (e.g., a molecular marker sequence), a cell marker, and a target binding region (e.g., an oligo(dT) sequence, a gene-specific sequence, a random polymer, or a combination thereof). The cell marker sequences of each of the plurality of oligonucleotides can be the same. The cell marker sequences of the oligonucleotides on different particles can be different, such that the oligonucleotides on different particles can be identified. In different implementations, the number of different cell marker sequences can be different. In some embodiments, the number of cell marker sequences can be, can be about, can be at least, or can be at most 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 、10 7 、10 8 、10 9, a number or range between any two of these values, or more. In some embodiments, no more than or no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more of the plurality of particles include oligonucleotides having the same cell sequence.

[0159] The plurality of oligonucleotides on each particle may include different barcode sequences (e.g., molecular markers). In some embodiments, the number of barcode sequences may be, may be about, may be at least, or may be at most 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 、10 7 、10 8 、10 9 、a number or range between any two of these values. As another example, in a single particle, at least 100, 500, 1000, 5000, 10000, 15000, 20000, 50000 of the plurality of oligonucleotides, a number or range between any two of these values, or more include different barcode sequences. Some embodiments provide a plurality of particles including barcodes. In some embodiments, the ratio of the target to be labeled to the occurrence (or copy or number) of different barcode sequences may be at least 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, or higher. In some embodiments, each of the plurality of oligonucleotides further includes a sample label, a universal label, or both. The particles can be, for example, nanoparticles or microparticles.

[0160] Barcoding Method

[0161] The present disclosure provides methods for estimating the number of different targets at different positions in a body sample (e.g., tissue, organ, tumor, cell). The methods can include placing barcodes (e.g., random barcodes) close to the sample, lysing the sample, associating different targets with the barcodes, amplifying the targets and / or digitally counting the targets. The methods can further include analyzing and / or visualizing information obtained from spatial markers on the barcodes. In some embodiments, the methods include visualizing multiple targets in the sample. Mapping the multiple targets onto a map of the sample can include generating a two-dimensional map or a three-dimensional map of the sample. The two-dimensional map and the three-dimensional map can be generated before or after barcoding (e.g., random barcoding) the multiple targets in the sample. Visualizing the multiple targets in the sample includes mapping the multiple targets onto a map of the sample. Mapping the multiple targets onto a map of the sample can include generating a two-dimensional map or a three-dimensional map of the sample. The two-dimensional map and the three-dimensional map can be generated before or after barcoding the multiple targets in the sample. In some embodiments, the two-dimensional map and the three-dimensional map can be generated before or after lysing the sample. Lysing the sample before or after generating the two-dimensional map or the three-dimensional map can include heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof.

[0162] In some embodiments, barcoding the multiple targets includes hybridizing multiple barcodes with the multiple targets to create barcoded targets (e.g., randomly barcoded targets). Barcoding the multiple targets can include generating an indexed library of the barcoded targets. Generating an indexed library of the barcoded targets can be performed using a solid support comprising multiple barcodes (e.g., random barcodes).

[0163] Contact Sample with Barcode

[0164] The present disclosure provides methods for contacting a sample (e.g., a cell) with a substrate of the present disclosure. A sample comprising, for example, a cell, an organ, or a tissue slice can be contacted with a barcode (e.g., a random barcode). For example, the cell can be contacted by gravity flow, wherein the cell can be allowed to sediment and a monolayer of cells can be produced. The sample can be a thin tissue section. The thin section can be placed on the substrate. The sample can be one-dimensional (e.g., forming a plane). The sample (e.g., a cell) can be coated on the substrate, for example, by growing / culturing the cell on the substrate.

[0165] When a barcode is close to a target, the target can hybridize with the barcode. The barcodes can be contacted at a non-exhaustible rate such that each different target can be associated with a different barcode of the present disclosure. To ensure an effective association between the target and the barcode, the target can be crosslinked to the barcode.

[0166] Cell Lysis

[0167] After the distribution of cells and barcodes, the cells can be lysed to release the target molecules. Cell lysis can be accomplished by any of a variety of means, such as by chemical or biochemical means, by osmotic shock, or by thermal lysis, mechanical lysis, or optical lysis. Cells can be lysed by adding a cell lysis buffer that includes a detergent (e.g., SDS, lithium dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. To increase the association of the target and the barcode, the diffusion rate of the target molecule can be altered, for example, by lowering the temperature of the lysate and / or increasing the viscosity of the lysate.

[0168] In some embodiments, lysis can be performed by mechanical lysis, thermal lysis, optical lysis, and / or chemical lysis. The lysed cells can include at least about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 target nucleic acid molecules, or more. The lysed cells can include at most about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 target nucleic acid molecules, or more.

[0169] Attach Barcode to Target Nucleic Acid Molecule

[0170] After cell lysis and release of the nucleic acid molecules, the nucleic acid molecules can be randomly associated with the barcodes of the co-localized solid support. The association can include hybridization of the target recognition region of the barcode with the complementary portion of the target nucleic acid molecule (e.g., the oligo(dT) of the barcode can interact with the poly(A) tail of the target). The assay conditions for hybridization (e.g., buffer pH, ionic strength, temperature, etc.) can be selected to promote the formation of specific stable hybrids. In some embodiments, the nucleic acid molecules released from the lysed cells can be associated with a plurality of probes on a substrate (e.g., hybridized with the probes on the substrate). When the probes include oligo(dT), the mRNA molecules can be hybridized with the probes and reverse transcription can be performed. The oligo(dT) portion of the oligonucleotide can serve as a primer for the first-strand synthesis of cDNA molecules. For example, Figure 2 In the non-limiting example of barcoding illustrated in (at box 216), the mRNA molecules can be hybridized with the barcodes on the beads. For example, single-stranded nucleotide fragments can be hybridized with the target binding regions of the barcodes.

[0171] Attachment can further include ligating the target recognition region of the barcode to a portion of the target nucleic acid molecule. For example, the target binding region can include a nucleic acid sequence that may be capable of specifically hybridizing to a restriction site overhang (e.g., an EcoRI sticky end overhang). The assay procedure can further include treating the target nucleic acid with a restriction enzyme (e.g., EcoRI) to generate a restriction site overhang. The barcode can then be ligated to any nucleic acid molecule that includes a sequence complementary to the restriction site overhang. A ligase (e.g., T4 DNA ligase) can be used to ligate the two fragments.

[0172] For example, in Figure 2 a non-limiting example of barcoding illustrated (at block 220), the labeled targets (e.g., target-barcode molecules) from multiple cells (or multiple samples) can subsequently be pooled, e.g., into a tube. The labeled targets can be pooled, for example, by retrieving the barcodes and / or beads to which the target-barcode molecules are attached.

[0173] Retrieval of the solid support-based collection of attached target-barcode molecules can be achieved by using magnetic beads and an externally applied magnetic field. Once the target-barcode molecules have been pooled, all further processing can be performed in a single reaction vessel. Further processing can include, for example, reverse transcription reactions, amplification reactions, cleavage reactions, dissociation reactions, and / or nucleic acid extension reactions. The further processing reactions can be performed within a micro-well, i.e., without first pooling the labeled target nucleic acid molecules from multiple cells.

[0174] Reverse Transcription

[0175] The present disclosure provides methods of using reverse transcription to generate target-barcode conjugates (at block 224 of Figure 2 ). The target-barcode conjugate can include the barcode and a complementary sequence of all or a portion of the target nucleic acid (i.e., a barcoded cDNA molecule, such as a randomly barcoded cDNA molecule). Reverse transcription of the associated RNA molecule can occur by adding a reverse transcription primer along with a reverse transcriptase. The reverse transcription primer can be an oligo(dT) primer, a random hexamer primer, or a target-specific oligonucleotide primer. The length of the oligo(dT) primer can be, or can be about 12 to 18 nucleotides and binds to the endogenous poly(A) tail at the 3' end of mammalian mRNA. The random hexamer primer can bind to the mRNA at multiple complementary sites. The target-specific oligonucleotide primer typically selectively primes the mRNA of interest.

[0176] In some embodiments, reverse transcription of the labeled RNA molecule can be carried out by adding a reverse transcription primer. In some embodiments, the reverse transcription primer is an oligo(dT) primer, a random hexamer primer, or a target-specific oligonucleotide primer. Generally, the oligo(dT) primer has a length of 12 to 18 nucleotides and binds to the endogenous poly(A) tail at the 3' end of mammalian mRNA. The random hexamer primer can bind to mRNA at multiple complementary sites. The target-specific oligonucleotide primer typically selectively primes the mRNA of interest.

[0177] Reverse transcription can occur repeatedly to generate multiple labeled cDNA molecules. The methods disclosed herein can include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. The methods can include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.

[0178] Amplification

[0179] One or more nucleic acid amplification reactions can be carried out (e.g., in Figure 2 box 228) to generate multiple copies of the labeled target nucleic acid molecule. Amplification can be carried out in a multiplex manner, where multiple target nucleic acid sequences are amplified simultaneously. The amplification reaction can be used to add sequencing adapters to the nucleic acid molecule. The amplification reaction can include amplifying at least a portion of the sample label (if present). The amplification reaction can include amplifying at least a portion of the cell label and / or barcode sequence (e.g., molecular tag). The amplification reaction can include amplifying at least a portion of the sample tag, cell label, spatial tag, barcode sequence (e.g., molecular tag), target nucleic acid, or a combination thereof. The amplification reaction can include amplifying 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100%, or a number or range between any two of these values of multiple nucleic acids. The method can further include performing one or more cDNA synthesis reactions to generate one or more cDNA copies of the target-barcode molecule including the sample label, cell label, spatial label, and / or barcode sequence (e.g., molecular tag).

[0180] In some embodiments, polymerase chain reaction (PCR) can be used for amplification. As used herein, PCR can refer to a reaction for in vitro amplification of a specific DNA sequence by simultaneous primer extension of complementary strands of DNA. As used herein, PCR can include derivative forms of the reaction, including but not limited to RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR.

[0181] Amplification of labeled nucleic acids can include non-PCR-based methods. Examples of non-PCR-based methods include but are not limited to multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification or circle-to-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR), and Qβ replicase (Qβ) methods, use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods that hybridize primers to nucleic acid sequences and cleave the resulting duplexes prior to extension reactions and amplification, strand displacement amplification using nucleic acid polymerases lacking 5'-exonuclease activity, rolling circle amplification, and ramification extension amplification (RAM). In some embodiments, the amplification does not produce circular transcripts.

[0182] In some embodiments, the methods disclosed herein further include performing polymerase chain reaction on a labeled nucleic acid (e.g., labeled RNA, labeled DNA, labeled cDNA) to produce a labeled amplicon (e.g., randomly labeled amplicon). The labeled amplicon can be a double-stranded molecule. The double-stranded molecule can include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule can include a sample label, a spatial label, a cell label, and / or a barcode sequence (e.g., a molecular label). The labeled amplicon can be a single-stranded molecule. The single-stranded molecule can include DNA, RNA, or a combination thereof. The nucleic acids disclosed herein can include synthetic or altered nucleic acids.

[0183] Amplification can include using one or more unnatural nucleotides. Unnatural nucleotides can include photo-labile or triggerable nucleotides. Examples of unnatural nucleotides can include, but are not limited to, peptide nucleic acids (PNA), morpholinos, and locked nucleic acids (LNA), as well as glycol nucleic acids (GNA) and threose nucleic acids (TNA). The unnatural nucleotides can be added to one or more cycles of the amplification reaction. The addition of unnatural nucleotides can also be used to identify products at specific cycles or time points in the amplification reaction.

[0184] One or more primers can include universal primers. Universal primers can anneal to universal primer binding sites. One or more custom primers can anneal to a first sample label, a second sample label, a spatial label, a cell label, a barcode sequence (e.g., a molecular tag), a target, or any combination thereof. One or more primers can include universal primers and custom primers. Custom primers can be designed to amplify one or more targets. The target can include a subset of the total nucleic acids in one or more samples. The target can include a subset of the total labeled targets in one or more samples. One or more primers can include at least 96 or more custom primers. One or more primers can include at least 960 or more custom primers. One or more primers can include at least 9600 or more custom primers. One or more custom primers can anneal to two or more different labeled nucleic acids. The two or more different labeled nucleic acids can correspond to one or more genes.

[0185] Any amplification protocol can be used in the methods of the present disclosure. For example, in one protocol, a first round of PCR can use gene-specific primers and primers targeting the universal Illumina sequencing primer 1 sequence to amplify molecules attached to beads. A second round of PCR can use nested gene-specific primers flanking the Illumina sequencing primer 2 sequence and primers targeting the universal Illumina sequencing primer 1 sequence to amplify the first PCR product. A third round of PCR adds P5 and P7 as well as a sample index to enable the PCR product to enter the Illumina sequencing library. Sequencing using 150bp x 2 sequencing can reveal the cell label and barcode sequence (e.g., a molecular tag) on read 1, the gene on read 2, and the sample index on the index 1 read.

[0186] In some embodiments, nucleic acids can be removed from a substrate using chemical cleavage. For example, chemical groups or modified bases present in the nucleic acid can be used to facilitate its removal from the solid support. For example, enzymes can be used to remove nucleic acids from a substrate. For example, nucleic acids can be removed from a substrate by restriction endonuclease digestion. For example, treatment of nucleic acids containing dUTP or ddUTP with uracil-d-glycosylase (UDG) can remove nucleic acids from a substrate. For example, enzymes for nucleotide excision (e.g., base excision repair enzymes (e.g., apurinic / apyrimidinic (AP) endonucleases)) can be used to remove nucleic acids from a substrate. In some embodiments, photocleavable groups and light can be used to remove nucleic acids from a substrate. In some embodiments, cleavable linkers can be used to remove nucleic acids from a substrate. For example, cleavable linkers can include at least one of the following: biotin / avidin, biotin / streptavidin, biotin / neutravidin, Ig protein A, photo-labile linkers, acid- or base-labile linker groups, or aptamers.

[0187] When the probe is gene-specific, the molecule can be hybridized to the probe and reverse transcription and / or amplification can be performed. In some embodiments, after the nucleic acid has been synthesized (e.g., reverse transcribed), it can be amplified. Amplification can be performed in multiplex, where multiple target nucleic acid sequences are amplified simultaneously. Amplification can add sequencing adapters to the nucleic acid.

[0188] In some embodiments, amplification can be performed on a substrate, for example, by bridge amplification. cDNA can be homopolymer tailed to generate compatible ends for bridge amplification using oligo(dT) probes on the substrate. In bridge amplification, the primer complementary to the 3' end of the template nucleic acid can be the first primer of each pair of primers covalently attached to solid particles. When a sample containing the template nucleic acid contacts the particles and undergoes a single thermal cycle, the template molecule can be annealed to the first primer, and the first primer can be extended forward by adding nucleotides to form a duplex molecule composed of the template molecule and the newly formed DNA strand complementary to the template. In the heating step of the next cycle, the duplex molecule can be denatured, releasing the template molecule from the particles, and attaching the complementary DNA strand to the particles through the first primer. In the annealing phase of the subsequent annealing and extension steps, the complementary strand can hybridize with the second primer, which is complementary to a segment of the complementary strand at the position removed from the first primer. The hybridization can result in the formation of a bridge between the first and second primers that are covalently fixed to the first primer through the complementary strand, and the second primer is formed by hybridization. In the extension phase, by adding nucleotides in the same reaction mixture, the second primer can be extended in the opposite direction, thereby converting the bridge into a double-stranded bridge. Then the next cycle begins, and the double-stranded bridge can be denatured to produce two single-stranded nucleic acid molecules, each with one end attached to the particle surface via the first and second primers, respectively, and the other end of each single-stranded nucleic acid molecule is unattached. In the annealing and extension steps of the second cycle, each strand can hybridize with an additional complementary primer that has not been used previously on the same particle to form a new single-stranded bridge. The two previously unused primers that are now hybridized are extended to convert the two new bridges into double-stranded bridges.

[0189] Amplification of the labeled nucleic acid can include PCR-based methods or non-PCR-based methods. Amplification of the labeled nucleic acid can include indexed amplification of the labeled nucleic acid. Amplification of the labeled nucleic acid can include linear amplification of the labeled nucleic acid. Amplification can be performed by polymerase chain reaction (PCR). PCR can refer to a reaction for in vitro amplification of a specific DNA sequence by simultaneous primer extension of complementary strands of DNA. PCR can cover derivative forms of the reaction, including but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, suppression PCR, semi-suppression PCR, and assembly PCR.

[0190] In some embodiments, the amplification of the labeled nucleic acid comprises a non-PCR-based method. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ), use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods that hybridize primers to nucleic acid sequences and cleave the resulting duplexes prior to extension reactions and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and ramification extension amplification (RAM).

[0191] In some embodiments, the methods disclosed herein further comprise performing a nested polymerase chain reaction on the amplified amplicons (e.g., targets). The amplicons can be double-stranded molecules. The double-stranded molecules can include double-stranded RNA molecules, double-stranded DNA molecules, or RNA molecules hybridized to DNA molecules. One or both strands of the double-stranded molecule can include a sample tag or a molecular identifier label. Alternatively, the amplicon can be a single-stranded molecule. The single-stranded molecule can include DNA, RNA, or a combination thereof. The nucleic acids of the present invention can include synthetic or altered nucleic acids.

[0192] In some embodiments, the method comprises repeatedly amplifying the labeled nucleic acid to produce multiple amplicons. The methods disclosed herein can include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Alternatively, the method includes performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions.

[0193] The amplification can further comprise adding one or more control nucleic acids to one or more samples comprising a plurality of nucleic acids. The amplification can further comprise adding one or more control nucleic acids to the plurality of nucleic acids. The control nucleic acids can include control labels.

[0194] Amplification can include using one or more unnatural nucleotides. Unnatural nucleotides can include photo-labile and / or triggerable nucleotides. Examples of unnatural nucleotides include, but are not limited to, peptide nucleic acids (PNA), morpholinos, and locked nucleic acids (LNA), as well as glycol nucleic acids (GNA) and threose nucleic acids (TNA). The unnatural nucleotides can be added to one or more cycles of the amplification reaction. The addition of unnatural nucleotides can also be used to discriminate products at specific cycles or time points in the amplification reaction.

[0195] Performing one or more amplification reactions can include using one or more primers. One or more primers can include one or more oligonucleotides. One or more oligonucleotides can include at least about 7 to 9 nucleotides. One or more oligonucleotides can include fewer than 12 to 15 nucleotides. One or more primers can anneal to at least a portion of a plurality of labeled nucleic acids. One or more primers can anneal to the 3' end and / or 5' end of a plurality of labeled nucleic acids. One or more primers can anneal to an internal region of a plurality of labeled nucleic acids. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the plurality of labeled nucleic acids. One or more primers can include a set of fixed primers. One or more primers can include at least one or more custom primers. One or more primers can include at least one or more control primers. One or more primers can include at least one or more housekeeping gene primers. One or more primers can include universal primers. Universal primers can anneal to universal primer binding sites. One or more custom primers can anneal to a first sample tag, a second sample tag, a molecular identifier label, a nucleic acid, or their products. One or more primers can include universal primers and custom primers. Custom primers can be designed to amplify one or more target nucleic acids. Target nucleic acids can include a subset of the total nucleic acids in one or more samples. In some embodiments, the primers are probes attached to the arrays disclosed herein.

[0196] In some embodiments, barcoding a plurality of targets in a sample (e.g., random barcoding) further comprises generating an indexed library of barcoded targets (e.g., randomly barcoded targets) or an indexed library of barcoded fragments of the targets. The barcode sequences of different barcodes (e.g., the molecular tags of different random barcodes) can be different from each other. Generating an indexed library of barcoded targets comprises generating a plurality of indexed polynucleotides from a plurality of targets in a sample. For example, for an indexed library of barcoded targets comprising a first indexed target and a second indexed target, the tagged region of the first indexed polynucleotide can have, have about, have at least, or have at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 nucleotides of difference, or a number or range of nucleotides of difference between any two of these values, from the tagged region of the second indexed polynucleotide. In some embodiments, generating an indexed library of barcoded targets comprises contacting a plurality of targets (e.g., mRNA molecules) with a plurality of oligonucleotides comprising a poly(T) region and a tagged region; and performing first-strand synthesis using a reverse transcriptase to produce single-stranded tagged cDNA molecules (each comprising a cDNA region and a tagged region), wherein the plurality of targets comprises at least two mRNA molecules of different sequences, and the plurality of oligonucleotides comprises at least two oligonucleotides of different sequences. Generating an indexed library of barcoded targets can further comprise amplifying the single-stranded tagged cDNA molecules to produce double-stranded tagged cDNA molecules; and performing nested PCR on the double-stranded tagged cDNA molecules to produce tagged amplicons. In some embodiments, the method can comprise producing adapter-tagged amplicons.

[0197] Barcoding (e.g., random barcoding) can comprise using nucleic acid barcodes or tags to label individual nucleic acid (e.g., DNA or RNA) molecules. In some embodiments, it involves adding DNA barcodes or tags to cDNA molecules as they are produced from mRNA. Nested PCR can be performed to minimize PCR amplification bias. Adapters can be added for sequencing using, for example, next-generation sequencing (NGS). For example, at Figure 2 box 232, the sequencing results can be used to determine the sequences of the cellular tags, molecular tags, and nucleotide fragments of one or more copies of the target.

[0198] Figure 3It is a schematic diagram showing a non - limiting exemplary process of generating an indexed library of barcoded targets (e.g., randomly barcoded targets), such as an indexed library of barcoded mRNA or fragments thereof. As shown in step 1, the reverse transcription process can encode each mRNA molecule with unique molecular tags, cell tags, and universal PCR sites. Specifically, by hybridizing (e.g., randomly hybridizing) a set of barcodes (e.g., random barcodes) 310 to the poly(A) tail region 308 of the RNA molecule 302, the RNA molecule 302 can be reverse - transcribed to produce a labeled cDNA molecule 304 (including the cDNA region 306). Each of the barcodes 310 can include a target - binding region, such as a poly(dT) region 312, a tag region 314 (e.g., barcode sequence or molecule), and a universal PCR region 316.

[0199] In some embodiments, the cell tag can include 3 to 20 nucleotides. In some embodiments, the molecular tag can include 3 to 20 nucleotides. In some embodiments, each of the multiple random barcodes further includes one or more of a universal tag and a cell tag, where the universal tag is the same for the multiple random barcodes on the solid support and the cell tag is the same for the multiple random barcodes on the solid support. In some embodiments, the universal tag can include 3 to 20 nucleotides. In some embodiments, the cell tag includes 3 to 20 nucleotides.

[0200] In some embodiments, the tag region 314 may include a barcode sequence or molecular tag 318 and a cell tag 320. In some embodiments, the tag region 314 may include one or more of a universal tag, a dimensional tag, and a cell tag. The length of the barcode sequence or molecular tag 318 may be, may be about, may be at least, or may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides, or a number or range of nucleotides between any of these values. The length of the cell tag 320 may be, may be about, may be at least, or may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides, or a number or range of nucleotides between any of these values. The length of the universal tag may be, may be about, may be at least, or may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides, or a number or range of nucleotides between any of these values. For multiple random barcodes on a solid support, the universal tag may be the same, and for multiple random barcodes on a solid support, the cell tag is the same. The length of the dimensional tag may be, may be about, may be at least, or may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides, or a number or range of nucleotides between any of these values.

[0201] In some embodiments, the tag region 314 may include, include about, include at least, or include at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different tags, such as the barcode sequence or molecular tag 318 and the cell tag 320. The length of each tag may be, may be about, may be at least, or may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides, or a number or range of nucleotides between any of these values. A set of barcodes or random barcodes 310 may contain, contain about, contain at least, or may be at most 10, 20, 40, 50, 70, 80, 90, 10 2 、10 3 、10 4 、10 5 、10 6 、10 7, 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 barcodes or random barcodes 310, or numbers or ranges of barcodes or random barcodes 310 between any of these values. And the groups of barcodes or random barcodes 310 can each contain, for example, a unique tag region 314. The labeled cDNA molecules 304 can be purified to remove excess barcodes or random barcodes 310. Purification can include Ampure bead purification.

[0202] As shown in step 2, the products from the reverse transcription process can be pooled into 1 tube in step 1 and PCR amplified with the first PCR primer pool and the first universal PCR primer. Pooling is possible because of the unique tag region 314. In particular, the labeled cDNA molecules 304 can be amplified to produce nested PCR labeled amplicons 322. Amplification can include multiplex PCR amplification. Amplification can include multiplex PCR amplification with 96 multiplex primers in a single reaction volume. In some embodiments, in a single reaction volume, multiplex PCR amplification can utilize, utilize about, utilize at least, or utilize at most 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 multiplex primers, or numbers or ranges of multiplex primers between any of these values. Amplification can include using a first PCR primer pool 324 that includes custom primers 326A-C targeting specific genes and a universal primer 328. The custom primers 326 can hybridize to regions within the cDNA portion 306' of the labeled cDNA molecules 304. The universal primer 328 can hybridize to the universal PCR region 316 of the labeled cDNA molecules 304.

[0203] As Figure 3Step 3 shows that the product from the PCR amplification in Step 2 can be amplified using a nested PCR primer pool and a second universal PCR primer. Nested PCR can minimize PCR amplification bias. In particular, the nested PCR-labeled amplicon 322 can be further amplified by nested PCR. Nested PCR can include multiplex PCR with a nested PCR primer pool 330 of nested PCR primers 332a-c and a second universal PCR primer 328' in a single reaction volume. The nested PCR primer pool 328 can contain, contain about, contain at least, or contain at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different nested PCR primers 330, or a number or range of different nested PCR primers 330 between any of these values. The nested PCR primers 332 can contain adapters 334 and hybridize to a region within the cDNA portion 306" of the labeled amplicon 322. The universal primer 328' can contain an adapter 336 and hybridize to the universal PCR region 316 of the labeled amplicon 322. Thus, Step 3 produces adapter-labeled amplicon 338. In some embodiments, the nested PCR primers 332 and the second universal PCR primer 328' may not contain adapters 334 and 336. Instead, adapters 334 and 336 can be ligated to the product of the nested PCR to produce adapter-labeled amplicon 338.

[0204] As shown in Step 4, the PCR product from Step 3 can be PCR amplified using library amplification primers for sequencing. In particular, adapters 334 and 336 can be used to perform one or more additional assays on the adapter-labeled amplicon 338. Adapters 334 and 336 can hybridize to primers 340 and 342. One or more of the primers 340 and 342 can be PCR amplification primers. One or more of the primers 340 and 342 can be sequencing primers. One or more of the adapters 334 and 336 can be used for further amplification of the adapter-labeled amplicon 338. One or more of the adapters 334 and 336 can be used for sequencing the adapter-labeled amplicon 338. Primer 342 can contain a plate index 344 such that amplicons generated using the same set of barcodes or random barcodes 310 can be sequenced in a single next-generation sequencing (NGS) run.

[0205] Sequencing Data Error

[0206] The methods disclosed herein can be used to identify and / or correct sequencing data errors, such as those that occur in methods for counting one or more target nucleic acids. In some embodiments, the errors can include, or be, deletions of one or more nucleotides, substitutions of one or more nucleotides, additions of one or more nucleotides, or any combination thereof. The errors can be present on a molecular label (ML), a sample label (SL), or other labels on a barcode (e.g., a random barcode). In some embodiments, the sequencing data errors can include, or be, PCR-introduced errors, sequencing-introduced errors, reverse transcription (RT) primer contamination errors, or any combination thereof. PCR-introduced errors can include, or be, the result of PCR amplification errors, PCR amplification bias, PCR amplification insufficiency, or any combination thereof. Sequencing-introduced errors can include, or be, the result of inaccurate base calling, insufficient sequencing, or any combination thereof. RT primer contamination errors can be errors caused by reverse transcription primers entering the PCR.

[0207] As used herein, the term "coverage" or "sequencing depth" can refer to the reads of barcoded targets having a particular ML and a particular SL in sequencing data. For example, a barcoded target can be sequenced multiple times. Thus, a barcoded target having a particular ML and SL can be observed multiple times. As another example, a cell can contain multiple copies of a target (e.g., multiple copies of an mRNA molecule of a gene). Multiple copies of the target can be barcoded. After PCR amplification (e.g., Figure 2 in box 228), there can be multiple copies of the barcoded target having a particular ML and SL. During sequencing, some or all of the multiple copies of the barcoded target having a particular ML and SL can be sequenced. The number of reads of the barcoded target having the same ML and SL observed in the sequencing data can be referred to as "coverage" or "sequencing depth".

[0208] In some embodiments, sequencing data errors can be identified and / or corrected. For example, copies of a target from a cell can be barcoded with different MLs and the same SL. The barcoded target having an ML can have multiple reads in the sequencing data. The barcoded targets having different MLs can have only a small number of reads (e.g., one read). Compared to the latter barcoded targets, the former barcoded targets are more likely to have a true ML (or a true ML or a signal ML). The latter barcoded targets can include an incorrect ML (or a false ML or a noise ML). This can be because it can be expected that the two MLs have similar coverage or sequencing depth. The latter barcoded targets having only a small number of reads can be artifacts or errors generated during sequencing or PCR.

[0209] As another example, barcodes (e.g., random barcodes) entering PCR can result in RT primer contamination errors. In some embodiments, after reverse transcribing mRNA molecules into cDNA molecules (e.g., at Figure 2 box 224 of Figure 2 ), barcodes not incorporated into the cDNA molecules can be removed by purification, e.g., with Ampure beads. The removal method (e.g., Ampure bead purification) may not completely remove barcodes that were not extended by reverse transcription to incorporate into the barcoded cDNA molecules (e.g., randomly barcoded cDNA molecules). For example, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, or a range between any two of these values of barcodes that were not extended by reverse transcription to incorporate into the barcoded cDNA molecules cannot be removed by Ampure bead purification. These unremoved barcodes can result in sequencing data errors during cDNA molecule amplification (e.g., at

[0210] box 228 of Figure 4 ). Barcodes between samples can be highly similar. For example, the sample label of the barcodes can be the same for the samples. Thus, PCR crossover can occur because these unremoved barcodes can hybridize with other nucleic acid molecules from the same sample (e.g., the SL region of barcoded mRNA molecules (such as randomly barcoded mRNA molecules)) during PCR and can result in sequencing data errors called SL errors. Figure 4 As illustrated in

[0211] Orientation-Adjacent-Corrected PCR and Sequencing Error

[0212] Disclosed herein are methods for correcting PCR or sequencing errors. In some embodiments, the method comprises: (a) receiving sequencing data of barcoded targets (e.g., randomly barcoded targets). The barcoded targets can be obtained by barcoding (e.g., randomly barcoding) a plurality of targets with a plurality of barcodes (e.g., random barcodes) to create a plurality of barcoded targets (e.g., randomly barcoded targets), wherein each of the plurality of barcodes comprises a molecular tag. In some embodiments, the method comprises: (b) for one or more of the plurality of targets: (i) counting the number of molecular tags having different sequences associated with the target in the sequencing data; (ii) identifying clusters of the molecular tags of the target using directed adjacency; (iii) folding the sequencing data received in (b) using the clusters of the molecular tags of the target identified in (ii); and (iv) estimating the number of targets, wherein after folding the sequencing data in (ii), the estimated number of targets is related to the number of molecular tags having different sequences associated with the target in the sequencing data counted in (i). The plurality of targets can comprise targets of the entire transcriptome of a cell. In some embodiments, the method further comprises: (c) barcoding (e.g., randomly barcoding) the plurality of targets with the plurality of barcodes to create the plurality of barcoded targets; and (d) sequencing the barcoded targets to generate the sequencing data of the received randomly barcoded targets.

[0213] Figure 5

[0214] ​At block 508, for one or more of the plurality of targets: the number of molecular markers having different sequences associated with the targets in the sequencing data can be counted. At block 512, clusters of molecular markers of the targets can be identified using directed adjacency. The molecular markers of the targets in a cluster can be within a predetermined directed adjacency threshold of each other. The directed adjacency threshold can vary. In some embodiments, the Hamming distance of the predetermined directed adjacency threshold can be, be about, at least, or at most one or two.

[0215] In some embodiments, the molecular markers of the targets within a cluster can include one or more parental molecular markers and sub-molecular markers of the one or more parental molecular markers. The occurrence of a parental molecular marker can be greater than or equal to a predetermined directed adjacency occurrence threshold. In some embodiments, the predetermined directed adjacency occurrence threshold can be, be about, at least, or at most twice the occurrence of a sub-molecular marker less than one. In some embodiments, the predetermined directed adjacency occurrence threshold can be, or be about 1.5 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times the occurrence of a sub-molecular marker, or a number or range between any two of these values. In some embodiments, the predetermined directed adjacency occurrence threshold can be at least or at most 1.5 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, or 10 times the occurrence of a sub-molecular marker.

[0216] In block 520, the sequencing data is folded using the clusters of molecular markers of the targets. Folding the sequencing data can include attributing the occurrences of sub-molecular markers to parental molecular markers. At block 532, the number of targets can be estimated to produce an output after folding the sequencing data. The method ends at 500 at block 536.

[0217] In some embodiments, the method further includes: determining the sequencing depth of a target. If the sequencing depth of the target is higher than a predetermined sequencing depth threshold, estimating the number of targets includes adjusting the sequencing data counted in (i). The predetermined sequencing depth threshold can be between 15 and 20. Adjusting the sequencing data counted in (i) includes: defining a threshold of molecular markers of the target to determine true and false molecular markers associated with the targets in the sequencing data obtained in (b). Defining a threshold of molecular markers of the target includes performing a statistical analysis of the molecular markers of the target. Performing the statistical analysis includes: fitting the distribution of the molecular markers of the target and their occurrences to two distributions, such as two negative binomial distributions; using the two negative binomial distributions to determine the number n of true molecular markers; and removing false molecular markers from the sequencing data obtained in (b), where the false molecular markers include molecular markers having their occurrences lower than the occurrence of the nth most abundant molecular marker, and where the true molecular markers include molecular markers having their occurrences greater than or equal to the occurrence of the nth most abundant molecular marker.

[0218] Correct PCR and Sequencing Errors Based on Orientation Adjacency and Distribution-Based Error Correction

[0219] Disclosed herein are methods for correcting PCR or sequencing errors. The methods can be used to determine the number of targets. In some embodiments, the methods include: (a) receiving sequencing data of barcoded targets (e.g., randomly barcoded targets). The barcoded targets can be obtained by barcoding (e.g., randomly barcoding) a plurality of targets with a plurality of barcodes (e.g., random barcodes) to create a plurality of barcoded targets (e.g., randomly barcoded targets), wherein each of the plurality of barcodes includes a molecular tag. In some embodiments, the methods include (b) for one or more of the plurality of targets: (i) counting the number of molecular tags having different sequences associated with the target in the sequencing data; (ii) determining the number of noise molecular tags having different sequences associated with the target in the sequencing data; and (iii) estimating the number of targets, wherein the estimated number of targets is related to the number of molecular tags having different sequences associated with the target counted in (i), and the number of molecular tags is adjusted according to the number of noise molecular tags determined in (ii). In some embodiments, the methods include determining the sequencing status of the targets in the sequencing data. In some embodiments, the methods further include: (c) barcoding (e.g., randomly barcoding) a plurality of targets with a plurality of barcodes to create a plurality of barcoded targets; and (d) sequencing the barcoded targets to generate the sequencing data of the received barcoded targets.

[0220] Figure 6 is a flowchart showing a non-limiting exemplary embodiment 600 of correcting PCR and sequencing errors based on recursive substitution error correction and distribution-based error correction. After receiving the sequencing data of a plurality of barcoded targets (e.g., randomly barcoded targets), method 600 begins at block 604. In some embodiments, method 600 further includes barcoding (e.g., randomly barcoding) a plurality of targets with a plurality of barcodes (e.g., random barcodes) to create a plurality of barcoded targets, wherein each of the plurality of barcodes includes a molecular tag. In some embodiments, method 600 further includes sequencing the plurality of barcoded targets to obtain the sequencing data.

[0221] At block 608, for one or more of the plurality of targets: the number of molecular markers having different sequences associated with the target in the sequencing data can be counted. At decision block 612, it can be determined whether the sequencing data has a saturated sequencing state. For example, if a target has a number of molecular markers with different sequences greater than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or a number or range between any two of these, the target is considered to have a saturated sequencing state. As another example, if a target has a number of molecular markers with different sequences greater than 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or a number or range between any two of these of the molecular markers with different sequences of barcodes (e.g., random barcodes), the target can be considered to have a saturated sequencing state.

[0222] In some embodiments, a saturated sequencing state can be determined by a target having a number of molecular markers with different sequences greater than a predetermined saturation threshold, and the predetermined saturation threshold can be different in different implementations. For example, the predetermined saturation threshold can be, or can be about 1000, 2000, 3000, 4000, 5000, 6000, 6557, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 65532, 70000, 80000, 90000, 100000, or a number or range between any two of these values. As another example, the predetermined saturation threshold can be at least, or at most 1000, 2000, 3000, 4000, 5000, 6000, 6557, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 65532, 70000, 80000, 90000, or 100000.

[0223] In some embodiments, the saturated sequencing state can depend on the number of molecular markers with different sequences of barcodes (e.g., random barcodes). For example, if a barcode includes about 6561 molecular markers with different sequences, the predetermined saturation threshold can be about 6557. As another example, if a barcode (e.g., random barcode) includes about 65536 molecular markers with different sequences, the predetermined saturation threshold can be about 65532. In some embodiments, the saturated sequencing state may not depend on the number of molecular markers with different sequences of barcodes.

[0224] If the sequencing data do not have a saturated sequencing state at decision block 612, method 600 proceeds to block 616, where the molecular marker counts can be adjusted based on directed adjacencies. In some embodiments, adjusting the molecular marker counts based on directed adjacencies can refer to Figure 5 the description. For example, adjusting the molecular marker counts based on directed adjacencies can include using directed adjacencies to identify clusters of molecular markers of a target; collapsing the sequencing data using the identified clusters of molecular markers of the target; and estimating the number of targets, wherein after collapsing the sequencing data, the estimated number of targets is associated with the number of molecular markers that have different sequences associated with the targets in the counted sequencing data.

[0225] At block 620, the sequencing state of the targets in the sequencing data can be determined. The sequencing state of the targets in the sequencing data can include, or be, under-sequenced. At decision block 624, it can be determined whether the sequencing state of the targets in the sequencing data is an under-sequenced state. For example, if the depth of a target (e.g., average, minimum, or maximum depth) is less than, or less than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these, then the target can be considered to have an under-sequenced state. As another example, if the depth of a target is less than at least, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, then the target can be considered to have an under-sequenced state.

[0226] In some embodiments, an under-sequenced state can be determined by targets having a depth (e.g., average, minimum, or maximum depth) less than a predetermined under-sequencing threshold. In different implementations, the under-sequencing threshold can be different. For example, the under-sequencing threshold can be, or be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values. As another example, the under-sequencing threshold can be at least, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.

[0227] In some embodiments, the under-sequenced state can depend on the number of molecular tags with different sequences of a barcode (e.g., a random barcode). For example, if the barcode includes, or about 1000, 2000, 3000, 4000, 5000, 6000, 6561, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 65532, 70000, 80000, 90000, 100000 molecular tags with different sequences, or a number or range between any two of these values, the under-sequencing threshold can be 10 (or another threshold number). As another example, if the barcode includes at least, or at most 1000, 2000, 3000, 4000, 5000, 6000, 6561, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 65532, 70000, 80000, 90000, or 100000, the under-sequencing threshold can be 10 (or another threshold number). In some embodiments, the under-sequenced state may not depend on the number of molecular tags with different sequences of a barcode (e.g., a random barcode).

[0228] At decision block 624, if the sequencing state of the target in the sequencing data is not the under-sequenced state, method 600 can proceed to block 628 to filter the molecular tag count. Filtering the molecular tag count can include determining, at decision block 632, the number of molecular tags with different sequences associated with the target in the sequencing data that are less than a pseudopoint threshold. In different implementations, the pseudopoint threshold can be different. For example, if the barcode (e.g., a random barcode) includes about 6561 molecular tags with different sequences, the pseudopoint threshold can be, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values. As another example, if the barcode (e.g., a random barcode) includes about 6561 molecular tags with different sequences, the pseudopoint sequencing threshold can be at least, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.

[0229] At decision block 632, if the number of molecular markers with different sequences associated with the target in the sequencing data is less than the pseudopoint threshold, method 600 may optionally proceed to block 636, where pseudopoints may be added to the number of molecular markers with different sequences associated with the target in the sequencing data before determining the number of noisy molecular markers with different sequences associated with the target in the sequencing data. In different implementations, the pseudopoints may have different molecular marker counts. For example, the molecular marker count of a pseudopoint may be, or may be about 0.0001, 0.001, 0.01, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values. As another example, the molecular marker count of a pseudopoint may be at least, or at most 0.0001, 0.001, 0.01, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.

[0230] At decision block 632, if the number of molecular markers with different sequences associated with the target in the sequencing data is not less than the pseudopoint threshold, non-unique molecular markers may be removed at block 640. At block 644, non-unique molecular markers may be removed by determining the number of noisy molecular markers with different sequences associated with the target in the sequencing data. Non-unique molecular markers may include molecular markers with different sequences associated with the target in the sequencing data that are greater than a predetermined recycled molecular marker threshold. In different implementations, the recycled molecular marker threshold may be different. For example, if a barcode (e.g., a random barcode) includes about 6561 molecular markers with different sequences, the recycled molecular marker threshold may be, or may be about 100, 200, 300, 400, 500, 600, 650, 700, 900, 1000, 2000, or a number or range between any two of these values. As another example, if a barcode (e.g., a random barcode) includes about 6561 molecular markers with different sequences, the recycled molecular marker threshold may be at least, or at most 100, 200, 300, 400, 500, 600, 650, 700, 900, 1000, or 2000.

[0231] In some embodiments, removing non-unique molecular markers includes: determining a theoretical number of non-unique molecular markers for the number of molecular markers with different sequences associated with the target in the sequencing data. Removing non-unique molecular markers may include removing a molecular marker that occurs more than the nth most abundant molecular marker with different sequences associated with the target in the sequencing data. The number n may be the theoretical number of non-unique molecular markers.

[0232] At block 644, a distribution-based error correction method can be used to adjust the molecular marker counts. The distribution-based error correction method can include determining the number of noise molecular markers having different sequences associated with a target in the sequencing data. Determining the number of noise molecular markers can include fitting two negative binomial distributions to the number of molecular markers having different sequences associated with a target in the sequencing data. For example, determining the number of noise molecular markers can include fitting a signal negative binomial distribution (one of the two negative binomial distributions) to the number of molecular markers having different sequences associated with a counted target in the sequencing data, where the signal negative binomial distribution corresponds to the number of molecular markers having different sequences associated with a counted target in the sequencing data that are signal molecular markers. Determining the number of noise molecular markers can include fitting a noise negative binomial distribution (the other of the two negative binomial distributions) to the number of molecular markers having different sequences associated with a counted target in the sequencing data, where the noise negative binomial distribution corresponds to the number of molecular markers having different sequences associated with a counted target in the sequencing data that are noise molecular markers. Determining the number of noise molecular markers can include using the fitted signal negative binomial distribution and the fitted noise negative binomial distribution to determine the number of noise molecular markers.

[0233] In some embodiments, using the fitted signal negative binomial distribution and the fitted noise negative binomial distribution to determine the number of noise molecular markers includes, for each of the different sequences associated with a target in the sequencing data: determining the signal probability of the different sequence in the signal negative binomial distribution. And the noise probability of the different sequence in the noise negative binomial distribution can be determined. Additionally, if the signal probability is less than the noise probability, the different sequence can be determined to be a noise molecular marker. In some embodiments, if fewer than two peaks are found (since two peaks are required to determine the signal negative binomial distribution and the noise negative binomial distribution), adjusting the molecular marker counts at block 644 can include removing singletons (e.g., single base substitutions).

[0234] At block 648, the number of targets can be estimated after neighborhood-based and distribution-based error correction to produce an output. At decision block 612, if the sequencing status of the target in the sequencing data is a saturated sequencing status, method 600 can proceed to block 648 to produce an output without adjusting the molecular markers based on directed adjacency and distribution-based error correction. For example, the determined number of noise molecular markers can be zero.

[0235] At decision block 624, if the sequencing status of the target in the sequencing data is an under-sequenced status, method 600 can proceed to block 648 to produce an output without adjusting the molecular markers based on distribution-based error correction. For example, the determined number of noise molecular markers can be zero. Method 600 ends at block 652.

[0236] Immune Receptor Barcode Error Correction

[0237] When performing sequencing and analysis such as immune receptor sequencing and analysis, substitution errors, primer cross errors, and PCR chimera errors may occur. For example, errors may occur when determining the occurrence or copy number of RNA molecules encoding immune receptors such as T cell receptors. Immune receptors include highly diverse and closely related genes. Thus, when compared to other genes, the likelihood of errors in sequencing data when performing immune receptor sequencing and analysis is higher. Such errors typically result in an overestimation of the immune repertoire diversity. Methods for mitigating these errors are referred to herein as immune receptor barcode error correction. In some embodiments, immune receptor barcode error correction utilizes recursive substitution error correction to correct substitution errors in molecular tags and nucleotide sequences (e.g., substitution errors in complementarity determining region 3 (CDR3)). For a given sample tag or cell tag, many different CDR3s can be associated with the same molecular tag sequence, resulting in an overestimation of immune receptor diversity. The correction methods disclosed herein can correct PCR chimeras that cross prior to molecular tagging and sample tagging, and then identify and remove incorrect molecular tags.

[0238] The disclosures herein include methods for determining the occurrence of a target. In some embodiments, the method includes: (a) barcoding (e.g., randomly barcoding) a plurality of targets using a plurality of barcodes (e.g., random barcodes) to create a plurality of barcoded targets (e.g., randomly barcoded targets), wherein each of the plurality of barcodes includes a cell marker and a molecular marker, wherein the molecular markers of at least two of the plurality of barcodes include different molecular marker sequences, and wherein at least two of the plurality of barcodes include cell markers having the same cell marker sequence; (b) obtaining sequencing data of the barcoded targets; and (c) for at least one of the plurality of targets: (i) identifying a putative sequence of the target in the sequencing data; (ii) counting the occurrences of the molecular marker sequences associated with the putative sequence of the target identified in (i); (iii) identifying clusters of the putative sequences of the target; (iv) collapsing the obtained sequencing data using the clusters of the putative sequences of the target identified in (iii); (v) identifying clusters of the molecular marker sequences associated with the putative sequence of the target; (vi) collapsing the sequencing data using the clusters of the molecular marker sequences identified in (v); (vii) identifying clusters of combined sequences, wherein each combined sequence includes a sequence in the sequence of the target and an associated molecular marker sequence in the molecular marker sequences; (viii) collapsing the sequencing data using the clusters of the combined sequences identified in (vii); (ix) identifying one or more putative sequences of the target corresponding to one or more chimeric sequences of the target, wherein the occurrences of the one or more putative sequences of the target corresponding to one or more chimeric sequences of the target are less than the occurrences of the remaining one or more putative sequences of the target that do not correspond to one or more chimeric sequences of the target; (x) removing from the sequencing data the one or more putative sequences of the target identified in (ix) that correspond to one or more chimeric sequences of the target; and (xi) estimating the occurrence of the target, wherein the estimated occurrence of the target is related to the number of molecular marker sequences counted in (ii) after collapsing the sequencing data in (iv), (vi), and (viii) and removing in (x) the one or more putative sequences of the target that correspond to one or more chimeric sequences of the target.

[0239] Disclosed herein are methods for determining the presence of a target. In some embodiments, the method includes: (a) receiving sequencing data of a plurality of targets, wherein the sequencing data includes the putative sequence of a target among the plurality of targets and the presence of molecular marker sequences associated with the sequence of the target in the sequencing data; (b) folding the putative sequence of the target; (c) folding the molecular marker sequences associated with the putative sequence of the target; and (d) estimating the presence of the target, wherein the estimated presence of the target is related to the presence of the molecular marker sequences associated with the putative sequence of the target in the sequencing data after folding the presence of the putative sequence of the target in (b) and folding the presence of the noise molecular marker sequences determined in (c).

[0240] Complementary determining regions (CDRs) are parts of the variable chains of immunoglobulins (antibodies) produced by B cells and T cell receptors produced by T cells. Three CDRs (CDR1, CDR2, and CDR3) are arranged discontinuously on the amino acid sequence of the variable domain of an antigen receptor or immune receptor. Since an antigen receptor typically consists of two variable domains (on two different polypeptide chains: heavy and light chains), there are a total of six CDRs for each pair of heavy and light chains that can contact an antigen together. Since most sequence variations associated with immunoglobulins and T cell receptors are found in the CDRs, it is not easy to distinguish between variations in the nucleotide sequences encoding the CDRs caused by one or more errors during sequencing and variations present in the nucleotide sequences encoding the CDRs. Within the variable domain, CDR1 and CDR2 are found in the variable (V) region of the polypeptide chain, while CDR3 includes a part of the V region and all of the diversity (D) regions and joining (J) regions of the heavy chain. The light chain contains the V region and the J region, but not the D region. CDR3 is the most variable of the CDRs.

[0241] Figure 7FIG. 0 is a schematic diagram showing non-limiting exemplary embodiments of immune receptor barcode correction based on recursive substitution error correction (RSEC, also referred to herein as directed adjacency). Barcodes (e.g., random barcodes) including cell markers and molecular markers can be used to determine the sequence of mRNA molecules encoding immune receptors, or generally for determining a target of interest. CDR3 includes a portion of the V region and the entirety of the D and J regions. As shown, substitution errors during sample preparation and sequencing can occur in the D and J regions (indicated by "*"). Although not shown, substitution errors can also occur in the V region. Additionally, sequencing errors can occur in the molecular markers (ML, indicated by "*"). Recursive substitution error correction (RSEC) can be used to correct such errors. For example, RSEC can be used to first adjust the count of CDR3 sequences. Subsequently, RSEC can be used to adjust the count of molecular markers. Optionally, by treating each CDR3 sequence and associated molecular marker in the sequencing data as a single sequence, RSEC can be used to simultaneously adjust the counts of CDR3 sequences and molecular markers. In some embodiments, RSEC can be used to first adjust the count of molecular markers.

[0242] Figure 8 FIG. 4 is a flowchart of non-limiting exemplary embodiment 800 of correcting errors in nucleotide sequences and molecular markers (based on recursive substitution error correction) and correcting errors in sequencing data attributable to one or more PCR chimeras. After receiving sequencing data for a plurality of barcoded targets (e.g., randomly barcoded targets), method 800 begins at block 804. In some embodiments, method 800 further includes barcoding (e.g., randomly barcoding) a plurality of targets using a plurality of barcodes (e.g., random barcodes) to create a plurality of barcoded targets, wherein each of the plurality of barcodes includes a molecular marker and / or a cell marker. In some embodiments, method 800 further includes sequencing the plurality of barcoded targets to obtain sequencing data. The plurality of targets can include targets of cells such as: the entire transcriptome, genes (e.g., genes encoding immune receptors such as T cell receptors), variable sequences (e.g., variable (V), diversity (D), joining (J) regions encoding immune receptors), or any combination thereof.

[0243] At block 808, method 800 may include, for one or more of the plurality of targets: determining (e.g., counting) the number of molecular markers having different sequences associated with the targets in the sequencing data. In some embodiments, counting the number of molecular markers having different sequences associated with the targets in the sequencing data may include: identifying a putative sequence of the target in the sequencing data (e.g., an immune receptor sequence such as a CDR3 sequence); and counting the occurrences of the molecular marker sequences associated with the identified putative sequence of the target in the sequencing data.

[0244] The putative sequences of the targets (e.g., CDR3) may differ from each other by one or more nucleotides. The sequences are putative in the sense that only one sequence is the true or correct sequence (e.g., there is only one correct CDR3 sequence per cell). The putative sequences of the targets may differ from each other by at least one nucleotide.

[0245] At block 812, method 800 may include adjusting the count of the putative nucleotide sequence of the target of interest based on recursive substitution error correction (also referred to herein as directed adjacency). In some embodiments, adjusting the nucleotide sequence count based on RSEC may be similar to adjusting the molecular marker count based on the directed adjacency Figure 5 described. For example, adjusting the nucleotide sequence count based on directed adjacency may include: identifying a putative sequence of the target in the sequencing data; counting the occurrences of the molecular marker sequences associated with the identified putative sequence of the target in the sequencing data; identifying clusters of the putative sequences of the target; and folding the obtained sequencing data using the identified clusters of the putative sequences of the target. Identifying clusters of the putative sequences of the target may include using RSEC to identify clusters of the putative sequences of the target. Folding the obtained sequencing data using the identified clusters of the putative sequences of the target may include: attributing the occurrences of subsequences in the one or more subsequences to the parental sequences of the subsequences.

[0246] In some embodiments, the putative sequences of the targets within a cluster can be within a first predetermined orientation adjacency threshold of each other. In different embodiments, the first predetermined orientation adjacency threshold can be different. In some embodiments, the first orientation adjacency threshold can be a Hamming distance, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or a number or range between any two of these values. In some embodiments, the first orientation adjacency threshold can be a Hamming distance, at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. The putative sequences of the targets within a cluster can include one or more parental sequences and one or more subsequences of the one or more parental sequences. The number of occurrences of the parental sequence can be greater than or equal to a first predetermined orientation adjacency occurrence threshold. In different embodiments, the first predetermined orientation adjacency occurrence threshold can be different. In some embodiments, the first predetermined orientation adjacency occurrence threshold can be twice the number of occurrences of a subsequence that is less than one. In some embodiments, the first predetermined orientation adjacency occurrence threshold can be, or about 1.5 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times the number of occurrences of a subsequence, or a number or range between any two of these values. In some embodiments, the first predetermined orientation adjacency occurrence threshold can be at least or at most 1.5 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, or 10 times the number of occurrences of a subsequence.

[0247] At block 816, method 800 can include adjusting the count of molecular markers based on RSEC. In some embodiments, adjusting the molecular marker count based on orientation adjacency can refer to Figure 5Description. For example, adjusting molecular marker counts based on directed adjacency can include: identifying clusters of the molecular marker sequences associated with the putative sequence of the target; and collapsing the sequencing data using the identified clusters of the molecular marker sequences. Identifying clusters of the molecular marker sequences associated with the putative sequence of the target can include using directed adjacency to identify clusters of the molecular marker sequences associated with the putative sequence of the target. The molecular marker sequences of the target within a cluster can be within a second predetermined directed adjacency threshold of each other. In different embodiments, the second predetermined directed adjacency threshold can be different. In some embodiments, the second directed adjacency threshold can be a Hamming distance, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or a number or range between any two of these values. In some embodiments, the second directed adjacency threshold can be a Hamming distance, at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. The putative molecular marker sequences of the target within a cluster can include one or more parent molecular marker sequences and one or more child molecular marker sequences of the one or more parent molecular marker sequences. The occurrence of a parent molecular marker sequence can be greater than or equal to a second predetermined directed adjacency occurrence threshold. In different embodiments, the second predetermined directed adjacency occurrence threshold can be different. In some embodiments, the second predetermined directed adjacency occurrence threshold is twice the occurrence of a child molecular marker sequence that is less than one. In some embodiments, the second predetermined directed adjacency occurrence threshold can be, or about, 1.5 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times the occurrence of a subsequence, or a number or range between any two of these values. In some embodiments, the second predetermined directed adjacency occurrence threshold can be at least or at most 1.5 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, or 10 times the occurrence of a subsequence. Collapsing the sequencing data using the identified clusters of the molecular marker sequences associated with the sequence of the target can include: attributing the occurrence of a child molecular marker sequence among the one or more child molecular marker sequences to the parent molecular marker sequence of the child molecular marker sequence.

[0248] At block 820, method 800 may optionally include adjusting the counts of nucleotide sequences and molecular markers simultaneously based on RSEC. Adjusting the counts of nucleotide sequences and molecular markers simultaneously based on directed adjacency may include identifying clusters of combined sequences, where each combined sequence includes a sequence from the sequence of the target and an associated molecular marker sequence from the molecular marker sequence; and collapsing the sequencing data using the identified clusters of combined sequences. Identifying the clusters of combined sequences may include using directed adjacency to identify the clusters of combined sequences. The combined sequences within a cluster may be within a third predetermined directed adjacency threshold of each other. In some embodiments, the third directed adjacency threshold may be a Hamming distance, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or a number or range between any two of these values. In some embodiments, the third directed adjacency threshold may be a Hamming distance, at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. The combined sequences within the cluster may include one or more parent combined sequences and one or more sub-combined sequences of the one or more parent combined sequences, and where the occurrence of the parent combined sequence is greater than or equal to a third predetermined directed adjacency occurrence threshold. In different embodiments, the third predetermined directed adjacency occurrence threshold may be different. In some embodiments, the third predetermined directed adjacency occurrence threshold is twice the occurrence of a sub-molecular marker sequence that is less than one. In some embodiments, the third predetermined directed adjacency occurrence threshold may be, or about 1.5 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times the occurrence of a subsequence, or a number or range between any two of these values. In some embodiments, the third predetermined directed adjacency occurrence threshold may be at least or at most 1.5 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, or 10 times the occurrence of a subsequence. Collapsing the sequencing data using the identified clusters of combined sequences in (vii) may include: attributing the occurrence of sub-combined sequences in the one or more sub-combined sequences to the parent combined sequences of the sub-combined sequences.

[0249] At block 824, the sequencing status of a target in the sequencing data can optionally be determined. The sequencing status of a target in the sequencing data can include, or be, under-sequenced. At decision block 828, it can optionally be determined whether the sequencing status of a target in the sequencing data is an under-sequenced state. For example, if the depth of the target (e.g., average, minimum, or maximum depth) is less than, or less than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these, the target can be considered to have an under-sequenced state. As another example, if the depth of the target is less than at least, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, the target can be considered to have an under-sequenced state.

[0250] In some embodiments, an under-sequenced state can be determined by a target having a depth (e.g., average, minimum, or maximum depth) less than a predetermined under-sequencing threshold. In different implementations, the under-sequencing threshold can be different. For example, the under-sequencing threshold can be, or be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values. As another example, the under-sequencing threshold can be at least, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.

[0251] In some embodiments, the under-sequenced state can depend on the number of molecular tags having different sequences of a barcode (e.g., a random barcode). For example, if the barcode includes, or about 1000, 2000, 3000, 4000, 5000, 6000, 6561, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 65532, 70000, 80000, 90000, 100000 molecular tags having different sequences, or a number or range between any two of these values of molecular tags having different sequences, the under-sequencing threshold can be 10 (or another threshold number). As another example, if the barcode includes at least, or at most 1000, 2000, 3000, 4000, 5000, 6000, 6561, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 65532, 70000, 80000, 90000, or 100000, the under-sequencing threshold can be 10 (or another threshold number). In some embodiments, the under-sequenced state can not depend on the number of molecular tags having different sequences of a barcode (e.g., a random barcode).

[0252] At decision block 828, if the sequencing status of the target in the sequencing data is not the under-sequenced state, method 800 can proceed to decision block 832 to determine if any singletons (e.g., single base substitutions) remain in the nucleotide sequence and / or molecular tags after the adjustments at blocks 812, 816, and 820. If at least one singleton remains, method 800 can proceed to block 836, where the molecular tag counts corresponding to chimeras can be removed.

[0253] At block 836, method 800 can include removing the molecular tag counts corresponding to chimeras. Figure 9 Is a schematic of a possible source of immune receptor chimeras (or target chimeras). As shown, many different CDR3 sequences (or putative sequences of the target) may have the same ML (and sample or cell tag) or be associated with it, which can lead to an overestimation of TCR diversity. Two or more true CDR3 sequences can cross during PCR. For example, after barcoding (e.g., after Figure 2 of block 224), two true CDR3 sequences (at Figure 9The CDR3-1 and CDR3-2) labeled therein can be associated with two different molecular markers (labeled ML-1 and ML-2). Two (or more) true CDR3 sequences can be two different CDR3 sequences from two different cells, which are not caused by PCR crossover. Immunoreceptor chimeras can form during PCR (e.g., at Figure 2 box 228 in

[0254] such that multiple different CDR3 sequences (e.g., two different CDR3 sequences) can have the same ML. For example, ML-1 can be associated with CDR3-1 and a chimera of CDR3-1 and CDR3-2. As another example, ML-2 can be associated with CDR3-2 and a chimera of CDR3-1 and CDR3-2. It may be advantageous to remove the molecular marker count corresponding to the chimera. Figure 8, removing the molecular marker counts corresponding to the chimeras can include identifying one or more putative sequences of the target corresponding to one or more chimeric sequences of the target, wherein the occurrence of the one or more putative sequences of the target corresponding to one or more chimeric sequences of the target is less than the occurrence of the remaining one or more putative sequences of the target that do not correspond to one or more chimeric sequences of the target; and removing from the sequencing data the identified one or more putative sequences of the target corresponding to one or more chimeric sequences of the target. Identifying one or more putative sequences of the target corresponding to one or more chimeric sequences of the target can include: identifying a putative sequence of the target associated with a molecular marker sequence among the plurality of molecular sequences; and identifying a putative sequence among the putative sequences of the target associated with the one molecular marker sequence, wherein the occurrence of the molecular marker sequence is less than a chimeric occurrence threshold corresponding to a chimeric sequence among one or more chimeric sequences of the target. In different embodiments, the chimeric occurrence threshold can be different. In some embodiments, the value of the chimeric occurrence threshold can be the occurrence of a putative sequence among the putative sequences of the target associated with a molecular marker sequence, wherein the occurrence is greater than the occurrence of any other sequence among the putative sequences of the target. In some embodiments, the chimeric occurrence threshold can be the occurrence of a putative sequence among the putative sequences of the target associated with a molecular marker sequence, wherein the occurrence is greater than the occurrence of any other sequence among the putative sequences of the target adjusted by an offset (e.g., subtraction). In different embodiments, the offset can be different. In some embodiments, the offset can be, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values. In some embodiments, the offset can be at least, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.

[0255] After folding the sequencing data at blocks 812, 816, and 820, the estimated occurrence of the target after block 824 is related to the number of molecular marker sequences counted at block 808. After folding the sequencing data at blocks 812, 816, and 820 and removing from the sequencing data at block 836 one or more putative sequences of the target corresponding to one or more chimeric sequences of the target, the estimated occurrence of the target after block 836 is related to the number of molecular marker sequences counted at block 808.

[0256] At block 840, dummy points may optionally be added to the number of molecular markers having different sequences associated with the target in the sequencing data before determining the number of noise molecular markers having different sequences associated with the target in the sequencing data. For example, if the number of molecular markers having different sequences associated with the target in the sequencing data is less than a dummy point threshold, method 800 may include adding dummy points to the number of molecular markers having different sequences associated with the target in the sequencing data before determining the number of noise molecular markers having different sequences associated with the target in the sequencing data. In different implementations, the dummy point threshold may be different. For example, if a barcode (e.g., a random barcode) includes approximately 6561 molecular markers having different sequences, the dummy point threshold may be, or may be approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values. As another example, if a barcode (e.g., a random barcode) includes approximately 6561 molecular markers having different sequences, the dummy point sequencing threshold may be at least, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100. In different embodiments, the added dummy points may have different molecular marker counts. For example, the molecular marker count of the dummy points may be, or may be approximately 0.0001, 0.001, 0.01, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values. As another example, the molecular marker count of the dummy points may be at least, or at most 0.0001, 0.001, 0.01, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.

[0257] At block 844, a distribution-based error correction method can be used to adjust the molecular marker counts. Performing the distribution-based error correction method can include determining the number of noise molecular markers having different sequences associated with a target in the sequencing data. Determining the number of noise molecular markers can include fitting two distributions, such as two negative binomial distributions, to the number of molecular markers having different sequences associated with the target in the sequencing data. For example, determining the number of noise molecular markers can include fitting a signal negative binomial distribution (one of the two negative binomial distributions) to the number of molecular markers having different sequences associated with the target in the counted sequencing data, where the signal negative binomial distribution corresponds to the number of molecular markers having different sequences associated with the target in the counted sequencing data that are signal molecular markers. Determining the number of noise molecular markers can include fitting a noise negative binomial distribution (the other of the two negative binomial distributions) to the number of molecular markers having different sequences associated with the target in the counted sequencing data, where the noise negative binomial distribution corresponds to the number of molecular markers having different sequences associated with the target in the counted sequencing data that are noise molecular markers. Determining the number of noise molecular markers can include using the fitted signal negative binomial distribution and the fitted noise negative binomial distribution to determine the number of noise molecular markers.

[0258] In some embodiments, using the fitted signal negative binomial distribution and the fitted noise negative binomial distribution to determine the number of noise molecular markers includes, for each of the different sequences associated with the target in the sequencing data: determining the signal probability of the different sequence in the signal negative binomial distribution. And the noise probability of the different sequence in the noise negative binomial distribution can be determined. Additionally, if the signal probability is less than the noise probability, the different sequence can be determined to be a noise molecular marker. In some embodiments, if fewer than two peaks are found (since two peaks are needed to determine the signal negative binomial distribution and the noise negative binomial distribution), adjusting the molecular marker counts at block 644 can include removing singletons (e.g., single base substitutions).

[0259] After folding the sequencing data at blocks 812, 816, and 820 and removing one or more putative sequences of the target corresponding to one or more chimeric sequences of the target at block 836, performing distribution-based error correction can include: after folding the sequencing data at blocks 812, 816, and 820 and removing one or more putative sequences of the target corresponding to one or more chimeric sequences of the target at block 836, thresholding the molecular marker sequences associated with the putative sequences of the target to determine signal molecular marker sequences and noise molecular marker sequences associated with the sequences of the target in the sequencing data counted at block 808. Thresholding the molecular marker sequences associated with the putative sequences of the target can include performing a statistical analysis on the molecular marker sequences of the target. Performing the statistical analysis can include: fitting the molecular marker sequences associated with the putative sequences of the target and their occurrences to two negative binomial distributions; using the two negative binomial distributions to determine the occurrence n of the signal molecular marker sequences; and after folding the sequencing data in (iv), (vi), and (viii) and removing one or more putative sequences of the target corresponding to one or more chimeric sequences of the target in (x), removing the noise molecular marker sequences from the sequencing data obtained in (b), where the noise molecular marker sequences include molecular marker sequences whose occurrences are less than the occurrence of the nth most abundant molecular marker, and where the signal molecular marker sequences include molecular marker sequences whose occurrences are greater than or equal to the occurrence of the nth most abundant molecular marker. The two negative binomial distributions can include a first negative binomial distribution corresponding to the signal molecular marker sequences and a second negative binomial distribution corresponding to the noise molecular marker sequences.

[0260] At block 848, the number of targets can be estimated after neighbor-based and distribution-based error correction to produce an output. At decision block 828, if the sequencing status of the target in the sequencing data is an under-sequenced status, method 800 can proceed to block 848 to produce an output without adjusting the molecular markers based on distribution-based error correction. For example, the number of noise molecular markers can be zero. At decision block 832, if there are no singletons in the adjusted sequencing data, method 800 can proceed to block 848 to produce an output without adjusting the molecular markers based on distribution-based error correction. Method 800 ends at block 852.

[0261] Sequencing

[0262] In some embodiments, estimating the number of different labeled or barcoded targets (e.g., randomly barcoded targets) can include determining the sequences of labeled targets, spatial labels, molecular labels, sample labels, cell labels, or any products thereof (e.g., labeled amplicons, or labeled cDNA molecules). The amplified targets can be subjected to sequencing. Determining the sequence of a barcoded target (e.g., a randomly barcoded target) or any product thereof can include performing a sequencing reaction to determine the sequence of at least a portion of a sample label, spatial label, cell label, molecular label, the sequence of at least a portion of the barcoded target, its complement, its reverse complement, or any combination thereof.

[0263] The sequence of a barcoded or randomly barcoded target (e.g., amplified nucleic acid, labeled nucleic acid, cDNA copy of a labeled nucleic acid, etc.) can be determined using a variety of sequencing methods, including but not limited to sequencing by hybridization (SBH), sequencing by ligation (SBL), quantitative incremental fluorescent nucleotide addition sequencing (QIFNAS), fragment ligation and cleavage, fluorescence resonance energy transfer (FRET), molecular beacons, TaqMan reporter probe digestion, pyrosequencing, fluorescence in situ sequencing (FISSEQ), FISSEQ beads, wobble sequencing, multiplex sequencing, polymerized colony (POLONY) sequencing; nanogrid rolling circle sequencing (ROLONY), allele-specific oligo ligation assay (e.g., oligonucleotide ligation assay (OLA), linear probe and rolling circle amplification (RCA) readout using ligation, single template molecule OLA using a locked probe and ligation, or single template molecule OLA using a circular locked probe and rolling circle amplification (RCA) readout), etc.

[0264] In some embodiments, determining the sequence of a barcoded target (e.g., a randomly barcoded target) or any of its products includes paired-end sequencing, nanopore sequencing, high-throughput sequencing, shotgun sequencing, dye-terminator sequencing, multiplex primer DNA sequencing, primer walking, Sanger dideoxy sequencing, Maxim-Gilbert sequencing, pyrosequencing, true single molecule sequencing, or any combination thereof. Alternatively, the sequence of the barcoded target or any of its products can be determined by electron microscopy analysis or a chemFET array.

[0265] High-throughput sequencing methods can also be used, such as cyclic array sequencing using platforms such as the Roche 454, Illumina Solexa, ABI-SOLiD, ION Torrent, Complete Genomics, Pacific Bioscience, Helicos, or Polonator platforms. In some embodiments, the sequencing can include MiSeq sequencing. In some embodiments, the sequencing can include HiSeq sequencing.

[0266] The barcoded target (e.g., a randomly barcoded target) can include nucleic acids representing from about 0.01% to about 100% of the genes of an organism's genome. For example, a target complementary region including multiple polymers can be used to sequence from about 0.01% to about 100% of the genes of an organism's genome by capturing genes containing complementary sequences from the sample. In some embodiments, the barcoded target includes nucleic acids representing from about 0.01% to about 100% of the transcripts of an organism's transcriptome. For example, a target complementary region including a poly(T) tail can be used to sequence from about 0.501% to about 100% of the transcripts of an organism's transcriptome by capturing mRNA from the sample.

[0267] Determining the sequence of the spatial and molecular tags of multiple barcodes (e.g., random barcodes) can include sequencing 0.00001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 99%, 100%, or a number or range between any two of these values of the multiple barcodes. Determining the sequence of the tags (e.g., sample tags, spatial tags, and molecular tags) of multiple barcodes can include sequencing 1, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 10 3 、104 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 16 , 10 17 , 10 18 , 10 19 , 10 20 or sequencing a number or range between any two of these values. Sequencing some or all of the plurality of barcodes can include generating a sequence having, having about, having at least, or having at most 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or range between any two of these values of nucleotides or bases in the read length.

[0268] Sequencing can include sequencing at least or at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more nucleotides or base pairs of a barcoded target (e.g., a randomly barcoded target). For example, sequencing can include generating sequencing data by polymerase chain reaction (PCR) amplification of a plurality of barcoded targets, where the sequence has a read length of 50, 75, or 100 or more nucleotides. Sequencing can include sequencing at least or at least about 200, 300, 400, 500, 600, 700, 800, 900, 1,000 or more nucleotides or base pairs of a barcoded target. Sequencing can include sequencing at least or at least about 1500, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 or more nucleotides or base pairs of a barcoded target.

[0269] Sequencing can include at least about 200, 300, 400, 500, 600, 700, 800, 900, 1,000 or more sequencing reads / runs. In some embodiments, sequencing includes sequencing at least or at least about 1500, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 or more sequencing reads per run. Sequencing can include less than or equal to about 1,600,000,000 sequencing reads / runs. Sequencing can include less than or equal to about 200,000,000 reads / runs.

[0270] Sample

[0271] In some embodiments, multiple targets can be included in one or more samples. The sample can include one or more cells, or nucleic acids from one or more cells. The sample can be a single cell or nucleic acids from a single cell. One or more cells can be one or more cell types. At least one of the one or more cell types can be a brain cell, heart cell, cancer cell, circulating tumor cell, organ cell, epithelial cell, metastatic cell, benign cell, primary cell, circulating cell, or any combination thereof.

[0272] Samples for use in the methods of the present disclosure can include one or more cells. The sample can refer to one or more cells. In some embodiments, multiple cells can include one or more cell types. At least one of the one or more cell types can be a brain cell, heart cell, cancer cell, circulating tumor cell, organ cell, epithelial cell, metastatic cell, benign cell, primary cell, circulating cell, or any combination thereof. In some embodiments, the cell is a cancer cell resected from a cancerous tissue, such as breast cancer, lung cancer, colon cancer, prostate cancer, ovarian cancer, pancreatic cancer, brain cancer, melanoma, and non-melanoma skin cancer, etc. In some embodiments, the cell is derived from cancer, but collected from a body fluid (e.g., circulating tumor cell). Non-limiting examples of cancers can include adenoma, adenocarcinoma, squamous cell carcinoma, basal cell carcinoma, small cell carcinoma, large cell undifferentiated carcinoma, chondrosarcoma, and fibrosarcoma. The sample can include tissue, monolayer cells, fixed cells, tissue sections, or any combination thereof. The sample can include a biological sample, clinical sample, environmental sample, biological fluid, tissue or cells from a subject. The sample can be obtained from a human, mammal, dog, rat, mouse, fish, fly, worm, plant, fungus, bacterium, virus, vertebrate, or invertebrate.

[0273] In some embodiments, the cell is a cell that has been infected with a virus and contains viral oligonucleotides. In some embodiments, the viral infection can be caused by a virus such as a single-stranded (+ strand or "sense") DNA virus (e.g., parvovirus), or a double-stranded RNA virus (e.g., a respiratory enterovirus). In some embodiments, the cell is a bacterium. These can include Gram-positive bacteria or Gram-negative bacteria. In some embodiments, the cell is a fungus. In some embodiments, the cell is a protozoan or other parasite.

[0274] As used herein, the term "cell" can refer to one or more cells. In some embodiments, the cell is a normal cell, e.g., a human cell at different developmental stages, or a human cell from different organ or tissue types. In some embodiments, the cell is a non-human cell, e.g., other types of mammalian cells (e.g., mouse, rat, pig, dog, cow, or horse). In some embodiments, the cell is other types of animal or plant cells. In other embodiments, the cell can be any prokaryotic or eukaryotic cell.

[0275] Data Analysis and Display Software

[0276] Data Analysis and Visualization of Target Spatial Resolution

[0277] The present disclosure provides methods for estimating the number and location of targets using spatially barcoded (e.g., randomly barcoded) and digital counting. Data obtained from the disclosed methods can be visualized on a map. The information generated using the methods described herein can be used to construct a map of the number and location of targets from a sample. The map can be used to locate the physical position of the targets. The map can be used to identify the locations of multiple targets. The multiple targets can be the same target species, or the multiple targets can be multiple different targets. For example, a map of the brain can be constructed to show the digital counts and locations of multiple targets.

[0278] A map can be generated from data from a single sample. A map can be constructed using data from multiple samples, thereby generating a combined map. A map can be constructed using data from dozens, hundreds, and / or thousands of samples. A map constructed from multiple samples can show the distribution of digital counts of targets associated with regions common to the multiple samples. For example, replicate assays can be shown on the same map. At least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, or more replicates (e.g., overlays) can be shown on the same map. At most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, or more replicates (e.g., overlays) can be shown on the same map. The spatial distribution and target number can be represented by various statistics.

[0279] Combining data from multiple samples can increase the positional resolution of the combined map. The orientation of multiple samples can be recorded by common landmarks, where individual position measurements across samples are at least partially discontinuous. A particular example is slicing a sample with an ultramicrotome along one axis and then slicing a second sample along a different path. The combined data set will give three-dimensional spatial positions associated with digital counts of the target. Reusing the above method will allow for a high-resolution three-dimensional map of digital count statistics.

[0280] In some embodiments of the instrument system, the system will include a computer-readable medium that includes code for providing data analysis for sequence data sets generated by performing single-cell, barcoded assays (e.g., random barcoded assays). Examples of data analysis functions that can be provided by data analysis software include, but are not limited to, (i) algorithms for decoding / demultiplexing sample labels, cell labels, spatial labels, and molecular labels, and target sequence data provided by sequencing a barcode library (e.g., a random barcode library) generated during a running assay, (ii) algorithms for determining the number of reads per gene per cell and the number of unique transcript molecules per gene per cell based on the data and creating a summary table, (iii) statistical analysis of the sequence data, such as for clustering cells by gene expression data or for predicting confidence intervals for determining the number of transcript molecules per gene per cell, etc., (iv) algorithms for identifying rare cell subsets, such as using principal component analysis, hierarchical clustering, k-means clustering, self-organizing maps, neural networks, etc., (v) sequence alignment capabilities for aligning gene sequence data with known reference sequences and detecting mutations, polymorphic markers, and splice variants, and (vi) automatic clustering of molecular markers to compensate for amplification or sequencing errors. In some embodiments, commercially available software can be used to perform all or part of the data analysis. For example, Seven Bridges (https: / / www.sbgenomics.com / ) software can be used to compile a table of the copy number of one or more genes present in each cell across an entire cell population. In some embodiments, the data analysis software can include options for outputting the sequencing results in a useful graphical format, such as a heat map indicating the copy number of one or more genes present in each cell of a cell population. In some embodiments, the data analysis software can also include algorithms for extracting biological meaning from the sequencing results, such as by associating the copy number of one or more genes present in each cell of a cell population with a type of cell, a rare cell, or a cell derived from a subject having a particular disease or disorder. In some embodiments, the data analysis software can also include algorithms for comparing cell populations across different biological samples.

[0281] In some embodiments, all data analysis functions can be packaged in a single software package. In some embodiments, the complete set of data analysis capabilities can include a suite of software packages. In some embodiments, the data analysis software can be a stand-alone package that is independent of the assay instrument system and is available to the user. In some embodiments, the software can be web-based and can allow users to share data.

[0282] In some embodiments, all data analysis functions can be packaged in a single software package. In some embodiments, the complete set of data analysis capabilities can include a suite of software packages. In some embodiments, the data analysis software can be a stand-alone package that is independent of the assay instrument system and is available to the user. In some embodiments, the software can be web-based and can allow users to share data.

[0283] System Processor and Network

[0284] Generally, a computer or processor used in a method suitable for the presently disclosed instrument system (as Figure 10 illustrated) can be further understood as a logic device that can read instructions from medium 1011 or network port 1005 and that can optionally be connected to a server 1009 having a fixed medium 1012. The system 1000 (as Figure 10 shown) can include a CPU 1001, a disk drive 1003, an optional input device (such as a keyboard 1015 or a mouse 1016), and an optional monitor 1007. Data communication can be achieved by transmitting from a specified communication medium to a server at a local or remote location. The communication medium can include any means for sending or receiving data. For example, the communication medium can be a network connection, a wireless connection, or an Internet connection. Such a connection can provide communication via the World Wide Web. It is contemplated that data related to the present disclosure can be transmitted through such a network or connection for reception or review by a recipient 1022, as Figure 10 illustrated.

[0285] Figure 11 Illustrated is an exemplary embodiment of a first instance architecture of a computer system 1100 that can be used in conjunction with the exemplary embodiments of the present disclosure. As Figure 11 depicted, the exemplary computer system can include a processor 1102 for processing instructions. Non-limiting examples of processors include: Intel XeonTM processors, AMD Opteron TM processors, Samsung 32-bit RISC ARM 1176JZ(F)-S v1.0 TMA processor, such as the ARM Cortex-A8 Samsung S5PC100TM processor, the ARM Cortex-A8 Apple A4TM processor, the Marvell PXA930TM processor, or a functionally equivalent processor. Multiple execution threads can be used for parallel processing. In some embodiments, a multiprocessor or a processor with multiple cores can also be used, whether in a single computer system, in a cluster, or distributed across systems through a network, the network including multiple computers, mobile phones, or personal data assistant devices.

[0286] As Figure 11 illustrated, the cache 1104 can be connected to or incorporated into the processor 1102 to provide high-speed memory for instructions or data that have recently been or are frequently used by the processor 1102. The processor 1102 is connected to the north bridge 1106 through the processor bus 1108. The north bridge 1106 is connected to the random access memory (RAM) 1110 through the memory bus 1112 and manages the access of the processor 1102 to the RAM 1110. The north bridge 1106 is also connected to the south bridge 1114 through the chipset bus 1116. The south bridge 1114 is in turn connected to the peripheral bus 1118. The peripheral bus can be, for example, PCI, PCI-X, PCI Express, or other peripheral buses. The north bridge and the south bridge are generally referred to as the processor chipset and manage the data transfer between the processor, the RAM, and the peripheral components on the peripheral bus 1118. In some alternative architectures, the functionality of the north bridge can be incorporated into the processor instead of using a separate north bridge chip.

[0287] In some embodiments, the system 1100 can include an accelerator card 1122 attached to the peripheral bus 1118. The accelerator can include a field programmable gate array (FPGA) or other hardware for accelerating certain processing. For example, the accelerator can be used for adaptive data reconstruction or for evaluating algebraic expressions used in extended set processing.

[0288] Software and data are stored in the external memory 1124 and can be loaded into the RAM 1110 or the cache 1104 for use by the processor. The system 1100 includes an operating system for managing system resources; non-limiting examples of the operating system include: Linux, Windows TM , MACOS TM , BlackBerry OS TM , iOS TM , and other functionally equivalent operating systems, as well as application software running on top of the operating system for managing data storage and optimization according to the exemplary embodiments of the present invention.

[0289] In this example, system 1100 also includes network interface cards (NICs) 1120 and 1121 connected to a peripheral bus for providing a network interface to external memories such as network attached storage (NAS) and other computer systems that can be used for distributed parallel processing.

[0290] Figure 12 An exemplary diagram showing network 1200 is presented, the network having multiple computer systems 1202a and 1202b, multiple cell phones and personal data assistants 1202c, and network attached storage (NAS) 1204a and 1204b suitable for the methods of this disclosure. In an example embodiment, systems 1212a, 1212b, and 1212c can manage data storage and optimize data access for data stored in network attached storage (NAS) 1214a and 1214b. Mathematical models can be used with the data and evaluated using distributed parallel processing across computer systems 1212a and 1212b, and cell phones and personal data assistant systems 1212c. Computer systems 1212a and 1212b and cell phones and personal data assistant systems 1212c can also provide parallel processing for adaptive data reconstruction of data stored in network attached storage (NAS) 1214a and 1214b. Figure 12 Only one example is illustrated, and a wide variety of other computer architectures and systems can also be used in conjunction with different embodiments of the present invention. For example, blade servers can be used to provide parallel processing. Processor blades can be connected through a backplane to provide parallel processing. Memory can also be attached to the backplane through a separate network interface or as network attached storage (NAS).

[0291] In some example embodiments, processors can maintain separate storage spaces and transfer data through network interfaces, backplanes, or other connectors for parallel processing by other processors. In other embodiments, some or all processors can use a shared virtual address storage space.

[0292] Figure 13FIG. 1300 shows an exemplary block diagram of a multi-processor computer system 1300 in accordance with an example embodiment, the system using a shared virtual address storage space. The system includes a plurality of processors 1302a-f that can access a shared storage subsystem 1304. The system incorporates a plurality of programmable hardware storage algorithm processors (MAPs) 1306a-f in the storage subsystem 1304. Each of the MAPs 1306a-f can include a memory 1308a-f and one or more field programmable gate arrays (FPGAs) 1310a-f. The MAPs provide configurable functional units and specific algorithms or portions of algorithms can be provided to the FPGAs 1310a-f for processing in close cooperation with the corresponding processors. For example, the MAPs can be used to evaluate algebraic expressions relative to a data model and for performing adaptive data reconstruction in the example embodiment. In this example, each MAP is globally accessible to all processors for these purposes. In one configuration, each MAP can use direct memory access (DMA) to access the associated memory 1308a-f, allowing it to perform tasks independently of and asynchronously with the corresponding microprocessors 1302a-f. In this configuration, a MAP can give results directly to another MAP for pipelining operations and parallel execution of algorithms.

[0293] The above computer architectures and systems are merely examples, and a wide variety of other computer, cellular phone, and personal data assistant architectures and systems can be used in conjunction with the example embodiments, including systems using any combination of general processors, coprocessors, FPGAs, and other programmable logic devices, system-on-chips (SOCs), application specific integrated circuits (ASICs), and other processing and logic elements. In some embodiments, all or part of the computer system can be implemented in software or hardware. Any kind of data storage medium can be used in conjunction with the example embodiments, including random access memory, hard disk drives, flash memory, tape drives, disk arrays, network attached storage (NAS), and other local or distributed data storage devices and systems.

[0294] In the example embodiments, software modules executing on any of the above or other computer architectures and systems can be used to implement the computer subsystems of the present disclosure. In other embodiments, the functionality of the system can be implemented, in part or in whole, in: firmware, programmable logic devices (such as field programmable gate arrays (FPGAs)), system-on-chips (SOLs), application specific integrated circuits (ASICs), or other processing and logic elements. For example, a set processor and optimizer can be implemented by hardware acceleration using a hardware accelerator card (such as an accelerator card).

[0295] Examples

[0296] Some aspects of the embodiments discussed above are further disclosed in detail in the following examples, which are not intended to limit the scope of the present disclosure in any way.

[0297] Example 1

[0298] ML Coverage of Each ML in Plate for High-Expression Gene - ACTB

[0299] This example demonstrates that the different distributions of ML errors derived during sequencing or PCR generally have a different distribution from true ML.

[0300] In addition to absolute gene expression counts and PCR bias correction, ML can also provide a better understanding of the statistical quality of library preparation procedures and sequencing data. When looking at the number of reads presenting the same gene ML (referred to as ML coverage), sequencing error base calls or PCR errors generated during library preparation can be detected. For example, compared to the gene ML of a given SL represented by only a single read, the gene ML of a given SL represented by multiple reads may be a more accurate measurement. In the presence of high ML coverage barcodes in the same library, low ML coverage barcodes are typically artifacts or errors generated during the sequencing run or PCR step during library preparation. ML errors derived during sequencing or PCR generally have a different distribution from true ML. Figure 15 Is an exemplary graph showing the molecular marker coverage of each molecular marker of a microplate for the highly expressed gene - ATCB, where different distributions are observed between the false molecular markers and the true molecular markers. Figure 16 Is an exemplary graph showing the fitting of two negative binomial distributions to the molecular marker coverage of each molecular marker of a microplate for the highly expressed gene - ATCB. The fitting of the two negative binomial distributions proves that molecular marker errors with lower molecular marker depth and true molecular markers with higher molecular marker depth can be statistically distinguished. The x - axis is the molecular depth.

[0301] In summary, these data demonstrate that ML errors derived during sequencing or PCR generally have a different distribution from true ML.

[0302] Example 2

[0303] Correct Molecular Markers Due to PCR or Sequencing Substitution Errors

[0304] This example demonstrates a method for correcting molecular markers due to PCR and sequencing substitution errors, which can be applied to whole - transcriptome assays without the assumption of uniform coverage and without high sequencing coverage of a fully sequenced state.

[0305] Duplicate removal is performed on the first mapped coordinate and unique molecular identifier (UMI) of each read, and reads are assumed to be identical if they have the same starting coordinate, UML, and strand. After duplicate removal, the UML with the highest count for each cluster is retained (Table 13).

[0306] Based on per-gene bias-corrected molecular labels (ML). For each gene, clusters of ML are identified using directed adjacency. If the MLs are within 1 Hamming distance and the parental ML count ≥ 2*(child MI count) - 1, the directed adjacency method clusters the MLs. All MLs in the same cluster are considered to be derived from the same parental ML, and the child ML counts are collapsed into the parental ML. Figure 17 Molecular label correction is shown, where a pairwise Hamming distance of 1 is overrepresented. After molecular label correction, molecular labels that are one Hamming distance apart are clustered and collapsed into the same parental molecular label. Figure 18 A curve showing the number of corrected MLs compared to the number of reads covered is shown. Since all reads are retained, this method can be used to remove single-base PCR or sequencing errors.

[0307] Table 13. After duplicate removal of molecular labels, only a small number of unique molecular labels are considered errors given the entire transcriptome assay

[0308]

[0309] In summary, these data demonstrate a correction method that can be used to correct or adjust data from an entire transcriptome assay since all reads are retained.

[0310] Example 3

[0311] Molecular Marker Count for High-Input Samples

[0312] This example describes unique molecular labels used as input molecules.

[0313] When used with small sample inputs (such as single cells), BD Precise TM targeted assays can be most suitable for allowing random and unique labeling of mRNA. Since the number of transcripts relative to the barcode pool increases in high RNA / cell input experiments, the percentage of recycled MLs to label the same gene increases and is theoretically calculated using the Poisson distribution ( Figure 14 ). In these cases, without statistical correction, using ML to quantify gene expression will underestimate the number of molecules initially present without any Poisson correction or correction based on two negative binomial distributions.

[0314] In extremely high input samples, where the mRNA quantity of each gene exceeds the entire set of 6561 barcodes, Poisson correction or correction based on two negative binomial distributions is no longer possible. For example, regardless of 65,000 or 100,000 input molecules, at most 6561 saturated barcodes are expected in any case. Thus, genes and samples seemingly with high sample input can be altered, potentially underestimating the ML counts.

[0315] In summary, these data demonstrate that when using ML to quantify gene expression, the raw data needs to be adjusted.

[0316] Example 4

[0317] Recursive Substitution Error Correction (RSEC)

[0318] This example demonstrates recursive substitution error correction.

[0319] In BD Precise TM Two collaborative methods can be applied in the targeted assay analysis pipeline to remove ML errors. Briefly, using recursive substitution error correction (RSEC), ML errors originating from sequencing base call substitutions are identified and adjusted to true ML barcodes. Subsequently, distribution-based error correction (DBEC) is used to adjust ML errors originating from library preparation steps or sequencing base deletion errors.

[0320] The RSEC algorithm can adjust ML errors originating from PCR or sequencing substitutions. These rare error events have been observed when examining ML coverage. For example, the ML coverage of an incorrect ML can be significantly lower than that of the true ML in a fully sequenced sample ( Figure 15 ); in the case of using two very similar MLs during the initial Molecular Indexing TM (reverse transcription) step, they usually have similar ML coverage and do not need to be eliminated. As the sequencing depth increases, more ML errors occur, so RSEC is crucial for adjusting the ML counts of highly sequenced barcode libraries.

[0321] Briefly, RSEC considers two factors in error correction: 1) the similarity of ML sequences; and 2) their ML coverage. For each target gene, MLs are joined when their ML sequences are within 1 base of each other (Hamming distance = 1). For each join between ML x and y, if:

[0322] coverage(y) > 2 * coverage(x) + 1, Equation (5)

[0323] where y represents the "parental ML" and x represents the "child ML".

[0324] Based on this assignment, the child ML can be folded into its parent ML. This process is recursive until there are no more identifiable parent / child MLs for the gene.

[0325] Figure 19 A schematic diagram showing an example of the recursive substitution error correction outlined above. Before RSEC correction, the MLs in the raw data included nine unique MLs: GTCAAATT, GTCAAAAT, GTCAAAAA, TTCAAAAA, TTCAGAAA, CTCAAAAA, TTCAAACT, TTCAAAAT, and TTCAAACA (SEQ ID NO: 3 - 11). By applying RSEC, GTCAAA T T (SEQ ID NO: 3) can be folded into GTCAAA A T (SEQ ID NO: 4) because the two MLs have a one - nucleotide (underlined) difference and the ML count of ML GTCAAATT (SEQ ID NO: 3) is lower than that of GTCAAAAT (SEQ ID NO: 4). Further, ML GTCAAAA T (SEQ ID NO: 4) can be folded into ML GTCAAAA A (SEQ ID NO: 5) (the difference in the ML sequence is underlined), which has a higher ML count than GTCAAAAT (SEQ ID NO: 4). Similarly, ML TTCAGAAA (SEQ ID NO: 7) and CTCAAAAA (SEQ ID NO: 8) can be folded into ML TTCAAAAA (SEQ ID NO: 6). ML TTCAAACT (SEQ ID NO: 9) can be folded into ML TTCAAAAT (SEQ ID NO: 10), which can in turn be folded into ML TTCAAAAA (SEQ ID NO: 6). ML TTCAAACA (SEQ ID NO: 11) has more than one nucleotide difference compared to all other MLs and thus will not be folded into any of the other eight MLs. Before RSEC correction, the original ML count was nine. After RSEC correction, the ML count is two: ML TTCAAAAA (SEQ ID NO: 6) and TTCAAACA (SEQ ID NO: 11).

[0326] In summary, these data demonstrate the use of RSEC to correct the original ML count.

[0327] Example 5

[0328] ML Coverage Calculation

[0329] This example describes the ML coverage calculation.

[0330] After RSEC, the gene ML technology for each well is calculated to determine if they are suitable for further correction. Genes with low ML coverage (< 4 reads per ML) bypass subsequent correction steps and are reported in the final ML data table and recorded as "low depth" in the bioinformatics pipeline. For genes with extremely high inputs, where at least 6557 out of the possible 6561 barcodes are observed, where determining the number of molecules is challenging due to barcode diversity and the gene is labeled as "saturated". For gene MLs that do not meet either of the 2 decision points, they proceed to the subsequent DBEC algorithm and are marked as "ignored (Pass)" in the output log file. Additionally, genes with an average of more than 650 MLs per well are recorded as "high input" because > 5% of these MLs are recycled based on the Poisson distribution ( Figure 15 ).

[0331] In summary, this example describes the ML coverage calculation.

[0332] Example 6

[0333] Distribution-Based Error Correction (DBEC)

[0334] This example describes distribution-based error correction.

[0335] Unlike RSEC, the DBEC algorithm is a method for distinguishing whether an ML is an error or a true signal, regardless of its ML sequence. While RSEC can use both ML sequence and ML coverage information to correct errors, DBEC mainly relies on ML coverage only to correct non-substitution errors. As previously mentioned, error barcodes typically have a low ML coverage range, which is different from the true barcode ML coverage range; this difference in ML coverage can be observed in the histogram of ML coverage as different distributions ( Figure 15 ). Given this difference, DBEC fits two negative binomial distributions to statistically distinguish ML errors (with lower ML coverage) from true signals with higher ML coverage.

[0336] Remove Recycled MLs to Achieve Optimal Distribution Fitting

[0337] For a given gene, as the detected ML increases, the percentage of recycled MLs (i.e., the same ML is used to label 2 or more mRNAs from the same gene) increases and can be estimated. Using the Poisson distribution (λ 非独特 ), the ML recycling rate equation (Equation (6)) is used to estimate for well i (n 非独特,i) The number of recycled ML. If the estimated recycled ML is greater than 5% of the total ML of a given gene in well i, then that gene in well i is labeled "high input". For these "high input" data, the leading ML coverage ML will be removed from the distribution fitting - but retained for later counting steps - to obtain a better negative binomial distribution.

[0338] P(X>1|λ 非独特 ),λ 非独特

[0339] = The number of ML / 6561. Equation (6)

[0340]

[0341] Add Pseudo-Points for Low-Expression Genes

[0342] If the unique number of ML is less than 10, it is usually more challenging to fit the distribution due to the sparsity of the data. To alleviate this problem, DBEC adds pseudo points at 1% signal count for assisting in distribution fitting, but does not affect the data.

[0343] Parameter Estimation

[0344] To fit two negative binomial distributions to separate the errors from the signal ML, two sets of starting values for parameter estimation are approximated. Assume the error distribution is a negative binomial with a mean and dispersion of 1.

[0345] Error / Signal Probability Estimation

[0346] Assume the signal and error distributions are negative binomial (μ 信号 , size 信号 ) and negative binomial (μ 错误 , size 错误 ), respectively. To determine the number of signal ML (in ascending order), calculate the probabilities that the reads from a given ML are from the signal and error distributions until Equation (8) is satisfied, where all previous ML are considered error ML.

[0347] P(X = r|μ = μ 错误 , size = size 错误 ) < P(X = r|μ = μ 信号 , size = size 信号 . Equation (8)

[0348] In summary, this example shows the calculations for performing distribution-based error correction.

[0349] Example 7

[0350] Correct PCR and Sequencing Errors Based on DBEC

[0351] This example demonstrates the correction of PCR and sequencing errors based on two negative binomial distributions.

[0352] Figures 20A to 20C Exemplary results of correcting PCR and sequencing errors based on two negative binomial distributions of CD69 are shown. Figure 20A Two negative binomial distributions of CD69 are shown (D n is the noise negative binomial distribution, and D s is the signal binomial distribution) in the Figure 20B fitting on the ML count data shown in the ML depth histogram of. Figure 20B The dashed line in shows the separation of the ML signal and SL error determined by the two negative binomial distributions shown in Figure 20A . Figure 20C The vertical lines in show the local maxima of the second derivative determined as based on the cumulative sum plot of the reads. Similar to Figures 20A to 20C , Figures 21A to 21C Exemplary results of correcting PCR and sequencing errors based on two negative binomial distributions of CD3E are shown.

[0353] In summary, these data show that DBEC can be used to correct PCR and sequencing errors.

[0354] Example 8

[0355] ML Count Correction Using Two Negative Binomial Distributions

[0356] This example demonstrates the ML counts of ten targets corrected using two negative binomial distributions.

[0357] Figures 22A to 22J An exemplary non-limiting validation of a dataset corrected using two negative binomial distributions is shown. As Figures 22A to 22J shown, the ML counts of 10 targets are corrected. Figures 22A to 22J The vertical lines in each plot of show the separation of the ML signal and SL error of the targets determined using two negative binomial distributions.

[0358] In summary, these data validate the ML count correction using two negative binomial distributions.

[0359] Example 9

[0360] BD Precise from 96-well mixed Jurkat and breast cancer (BrCa) single cells TM t-Random neighbor of targeted assay Domain Embedding Visualization

[0361] This example demonstrates a method for correcting PCR and sequencing errors based on recursive substitution error correction and distribution-based error correction for mixed Jurkat and breast cancer (BrCa) single cells.

[0362] Figures 23A to 23D Shows Precise of Jurkat and breast cancer (BrCa) single cells from a 96-well mix (86 genes examined). TM Exemplary t-stochastic neighbor embedding (t-SNE) visualization of a targeted assay. Figure 23A Shows identification of cell clusters using DBScan with the same parameters before and after ML adjustment. Figures 23B to 23D Shows single marker expression scaled by both color and point size. Figure 23B Shows PSMB4 (housekeeping gene), which is present in both cell types, and after ML adjustment, the lack of PSMB4 signal is further highlighted in the "low signal" cluster. Figure 23C Shows CD3E (lymphocyte marker highlighting Jurkat cell clusters). Figure 23D Shows CDH1 (epithelial cell marker highlighting BrCa clusters).

[0363] In summary, these data demonstrate that ML regulation removes ML noise, which allows for clear differentiation of gene expression between cell clusters.

[0364] Example 10

[0365] Differential Expression Analysis between Cell Clusters

[0366] This example demonstrates a method for correcting PCR and sequencing errors based on recursive substitution error correction and distribution-based error correction for low-signal cells and breast cancer (BrCa) cells.

[0367] Figures 24A to 24B Is a non-limiting exemplary graph showing differential expression analysis of genes with >0 ML between cell clusters in two selected clusters calculated by DBScan and determined by gene marker levels in each cluster. Figure 24A Shows gene expression in the 'low signal' cluster compared to the remaining cells. Figure 24A The top graph shows the original ML comparison, which shows that ML noise is generally higher for genes with higher average expression in other cells. Figure 24A The bottom graph shows the reduction of ML noise detected in the "low signal" cluster after ML adjustment using RSEC and DBEC, which allows for a clearer differentiation of gene expression between clusters. Figure 24B Shows gene expression in the 'BrCa' cluster compared to the remaining cells. Figure 24B The top graph shows that the original ML in non-BrCa cells also has significant ML counts for BrCa markers such as KRT1, MUC1. Figure 24BThe bottom panel shows that the adjusted ML of BrCa markers is highly enriched in BrCa clusters compared to other cells.

[0368] In summary, these data demonstrate that for cells such as low-signal cells and breast cancer cells, PCR and sequencing errors can be corrected based on recursive substitution error correction and distribution-based error correction.

[0369] Example 11

[0370] Adjust Molecular Marker Counts for Mixed Jurkat and T47D Cells

[0371] This example demonstrates a method for adjusting the molecular marker counts of mixed Jurkat and T47D cells.

[0372] Figures 25A to 25D is an exemplary non-limiting t-stochastic neighbor embedding (t-SNE) visualization of the BD Precise targeted assay of mixed Jurkat and breast cancer (T47D) single cells (with 86 genes examined) from a 96-well plate. TM Figure 25A Shows the identification of cell clusters using DBScan with the same parameters before and after ML adjustment. Figures 25B to 25D Shows the expression of individual markers scaled by color and point size. Figure 25B Shows the scaling of PSMB4 (a housekeeping gene present in both cell types and after ML adjustment). The lack of PSMB4 signal is further highlighted in the no-template control (NTC) cluster. Figure 25C Shows the scaling of CD3E (a lymphocyte marker highlighting the Jurkat cell cluster). Figure 25D Shows the scaling of CDH1 (an epithelial cell marker highlighting the T47D cluster).

[0373] Figures 26A to 26B is a non-limiting exemplary heatmap showing differential gene expression of molecular marker counts between different cell clusters identified by Figure 26A before any error correction step ( Figure 26B the raw ML shown), and after RSEC and DBEC correction ( Figures 25A to 25D the adjusted ML shown). Genes with low expression are blue, and genes with high expression are orange. Genes with similar gene expression patterns between these cell types cluster together. Without error correction, the NTC has noise from highly expressed genes such as CD3E and KRT18 (which are Jurkat and T47D markers respectively). Additionally, error correction reveals different gene expression patterns between Jurkat and T47D.

[0374] In summary, these data demonstrate that ML regulation can remove MI noise, which allows for clear discrimination of gene expression between cell clusters.

[0375] Example 12

[0376] Immune Receptor Barcode Error Correction Using Recursive Substitution Error Correction

[0377] This example demonstrates immune receptor barcode error correction based on recursive substitution error correction.

[0378] Figures 27A to 27B A table showing a non-limiting example of immune receptor barcode error correction based on recursive substitution error correction is presented. Performing immune receptor barcode error correction involves adjusting the count of a target nucleotide sequence (NS) (e.g., the putative nucleotide sequence of CDR3) by recursive substitution error correction. Multiple clusters of putative sequences of CDR3 are identified. One cluster includes sequences that differ by one nucleotide (underlined) from the parental nucleotide sequence

[0379] TGTGTGGTGAACGGAGACGGCACTGCCAGTAAACTCACCTTT (SEQ ID NO:58), such as the sub-nucleotide sequences TGTGTGGTGAACGGAGACGGCACTGCCAGTAAACTCAC T TTT (SEQ ID NO:59) and

[0380] C GTGTGGTGAACGGAGACGGCACTGCCAGTAAACTCACCTTT (SEQ ID NO:77). The nucleotide sequence

[0381] TGTGTGGTGAACGGAGACGGCACTGCCAGTAAACTCACCTTT (SEQ ID NO:58) is the parental nucleotide sequence because it has the highest raw read count of 294 in this cluster, which is the number of molecular markers with different sequences associated with this nucleotide sequence in the sequencing data. Other clusters include

[0382] TGTGCTGTCCACCGAGGAAGCCAAGGAAATCTCATCTTT (SEQ ID NO:78) and TGTGCTGTCCACCGAGGAAGCCAAGGAAATCTCATC G TT (SEQ ID NO:79); TGTGCAGGAGAATCTGGGGATTACCAGAAAGTTACCTTT (SEQ ID NO:80) and

[0383] TGTGCAGGAGAATCTGGGG G TTACCAGAAAGTTACCTTT(SEQ ID NO:81); TGTGCAGCAACCGAGTCCTATGGTCAGAATTTTGTCTTT(SEQ ID NO:82) and TGTGCAGCAAC A GAGTCCTATGGTCAGAATTTTGTCTTT(SEQ ID NO:83) and

[0384] TGCCTCGTGGGGAGCCTTTCTGGTTCTGCAAGGCAACTGACCTTT(SEQ ID NO:84) and

[0385] TGCCTCGTGGGGAGCCTTTC C GGTTCTGCAAGGCAACTGACCTTT(SEQ ID NO:85). The occurrence of each sub-nucleotide sequence is attributed to the corresponding parental molecular marker sequence (as indicated by the arrow from the column labeled "Original Reads" to the column labeled "NS Adjusted Reads").

[0386] Performing immune receptor barcode error correction involves adjusting the count of molecular markers (ML) by recursive substitution error correction. Clusters of molecular marker sequences include the parental molecular marker sequence AGTGCGAG (SEQ ID NO:110) and the sub-molecular marker sequences AGTGCG G G and AGTGC N AG (SEQ ID Nos. 111 and 112 respectively), which differ from the parental molecular marker sequence by one nucleotide (underlined). In this example, the molecular marker sequence AGTGCGAG (SEQ ID NO:110) associated with the CDR3 sequence TGTGTGGTGAACGGAGACGGCACTGCCAGTAAACTCACCTTT (SEQ ID NO:58) has the highest nucleotide sequence adjusted reads for all CDR3 sequences associated with the molecular marker AGTGCGAG (SEQ ID NO:110). The occurrence of each sub-nucleotide sequence (sub-molecular marker sequences AGTGCG G G and AGTGC NAG (SEQ ID NO: 111 and 112 respectively) for two times and once respectively) is attributed to the corresponding parental molecular marker sequence with the highest nucleotide sequence adjusted read count (319 for the parental molecular marker AGTGCGAG (SEQ ID NO: 110)). Performing immune receptor barcode error correction includes simultaneously adjusting the nucleotide sequence and the count of the molecular marker based on RSEC. In this example, when applying RSEC, no adjustment of the molecular marker count is made when each nucleotide sequence and the corresponding molecular marker are considered as one sequence.

[0387] Multiple chimeras were identified, and the molecular marker counts corresponding to the chimeras were removed. For example, the molecular marker sequence AGTGCGAG (SEQ ID NO: 110) is associated with multiple CDR3 nucleotide sequences such as

[0388] TGTGTGGTGAACGGAGACGGCACTGCCAGTAAACTCACCTTT (SEQ ID NO: 58) and

[0389] TGTGCTGTCCACCGAGGAAGCCAAGGAAATCTCATCTTT (SEQ ID NO: 78). The adjusted occurrence of the nucleotide sequence

[0390] TGTGCTGTCCACCGAGGAAGCCAAGGAAATCTCATCTTT (SEQ ID NO: 78) (e.g., the number of occurrences observed in the sequencing data after the above adjustment) is 7, which is lower than the adjusted occurrence of the nucleotide sequence

[0391] TGTGTGGTGAACGGAGACGGCACTGCCAGTAAACTCACCTTT (SEQ ID NO: 58) (which is 322), which is the highest among all nucleotide sequences with the molecular marker sequence AGTGCGAG (SEQ ID NO: 110). The putative sequences of the targets corresponding to the identified chimeric sequences of the CDR3 are removed (as indicated by the arrows from the column labeled "NS and ML adjusted reads" to the column labeled "Removed chimeras"). By adjusting and removing chimeras, the CDR3 sequence was determined to be

[0392] TGTGTGGTGAACGGAGACGGCACTGCCAGTAAACTCACCTTT (SEQ ID NO.58), which is associated with the molecular marker sequence AGTGCGAG (SEQ ID NO.110) with an adjusted molecular marker count of 322.

[0393] In summary, these data demonstrate the adjustment of nucleotide sequence counts, molecular marker counts, and nucleotide sequences and molecular counts using RSEC, and the removal of chimeric CDR3 sequences.

[0394] Example 13

[0395] Immune Receptor Barcode Error Correction Using Recursive Substitution Error Correction and Distribution-Based Error Correction

[0396] This example demonstrates immune receptor barcode error correction using recursive substitution error correction and distribution-based error correction.

[0397] A 500-cell sample consisting of 75% healthy peripheral blood mononuclear cells (PBMCs) and 25% Jurkat cells was loaded into a Rhapsody TM cartridge and captured in microwells containing Rhapsody TM beads. Rhapsody TM beads are magnetic beads with barcodes attached to the bead surface. Each bead is attached with a barcode that has a cell marker (which has the same cell marker sequence) and a molecular marker (which is selected from a set of different molecular sequences). The barcodes of different beads have cell markers with different cell marker sequences. Each bead's barcode has a capture site for capturing TCR mRNA molecules. As described herein, the captured TCR mRNA molecules are barcoded and sequenced. The cell marker and the molecular marker are used to determine the cellular and molecular origin of each TCR molecule. The molecular marker is used to determine the number or count of occurrences of different TCR molecules.

[0398] Immune receptor barcode error correction (described with reference to Figure 7 , Figure 8 , Figure 27, and Example 12) was used to correct the sequencing data of different TCR molecules. Briefly, the counts of the putative nucleotide sequences of different TCR genes (e.g., TCRb) were adjusted based on recursive substitution error correction. The counts of the molecular markers were adjusted based on recursive substitution error correction. The counts of the nucleotide sequences (e.g., TCRb) and the molecular markers were adjusted based on recursive substitution error correction. The molecular marker counts corresponding to chimeras were removed. Subsequently, the molecular marker counts were adjusted using the distribution-based error correction described herein. Figure 28 is a histogram showing non-limiting exemplary results of immune receptor barcode error correction for TCRb followed by distribution-based error correction. Without error correction, TCR diversity (including TCRb diversity) would be overestimated.

[0399] In summary, these data demonstrate that replacing error correction with recursion adjusts TCR nucleotide sequence counts, molecular marker counts, and TCR nucleotide sequence and molecular counts, removes chimeric TCR sequences, and then performs distribution-based error correction to avoid overestimating TCR diversity.

[0400] In at least some of the previously described embodiments, one or more elements used in one embodiment may be used interchangeably in another embodiment, unless such substitution is technically infeasible. Those skilled in the art will understand that various other omissions, additions, and modifications may be made to the above methods and structures without departing from the scope of the claimed subject matter. All such modifications and changes are intended to fall within the scope of the subject matter defined by the appended claims.

[0401] Regarding the use of substantially any plural and / or singular terms herein, the skilled artisan can convert from the plural to the singular and / or from the singular to the plural as appropriate for the context and / or application. For clarity, various singular / plural permutations may be set forth explicitly herein. As used in this specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a / an" and "the" include plural referents. Unless otherwise stated, any reference to "or" herein is intended to cover "and / or".

[0402] Those skilled in the art will understand that, generally speaking, the terms used herein, particularly the terms in the appended claims (e.g., the body of the appended claims), are generally intended to be "open" terms (e.g., the term "including" should be interpreted as "including but not limited to", the term "having" should be interpreted as "having at least", the term "includes" should be interpreted as "includes but is not limited to", etc.). Those skilled in the art will further understand that if a specific number of the introduced claim recitations is intended, such an intention will be expressly stated in the claim, and in the absence of such a statement, no such intention exists. For example, as an aid to understanding, the following appended claims may contain the use of introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed to mean that the introduction of a claim recitation by the indefinite article "a" or "an" limits any particular claim containing such introduced claim recitation to an embodiment containing only one such recitation, even when the same claim includes the introductory phrases "one or more" or "at least one" and the indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be interpreted to mean "at least one" or "one or more"); this also applies to the use of the definite article to introduce a claim recitation. In addition, even if a specific number of the introduced claim recitations is expressly stated, those skilled in the art will recognize that such a statement should be interpreted to mean at least the stated number (e.g., merely stating "two recitations" without further modifiers means at least two recitations, or two or more recitations). Further, in those cases where a convention similar to "at least one of A, B, and C, etc." is used, generally such syntactic construction is intended in the sense that those skilled in the art will understand the convention (e.g., "a system having at least one of A, B, and C" will include, but not be limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In those cases where a convention similar to "at least one of A, B, or C, etc." is used, generally such syntactic construction is intended in the sense that those skilled in the art will understand the convention (e.g., "a system having at least one of A, B, or C" will include, but not be limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.).Those skilled in the art will further understand that, in fact, whether in the specification, claims or drawings, any disjunctive words and / or phrases presenting two or more alternative terms should be understood as contemplating the possibility of including one of the terms, either term, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B".

[0403] In addition, when a feature or aspect of this disclosure is described in terms of a Markush group, those skilled in the art will recognize that this disclosure also thereby describes any individual member or subgroup of members of the Markush group.

[0404] As those skilled in the art will understand, for any and all purposes, such as in providing a written description, all ranges disclosed herein also include any and all possible subranges and combinations of subranges thereof. Any listed range can be readily identified as fully describing and enabling the same range to be broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into lower thirds, middle thirds, upper thirds, etc. As those skilled in the art will also understand, all language such as "up to", "at least", "greater than", "less than", etc. includes the stated number and refers to ranges that can subsequently be broken down into subranges as discussed above. Finally, as those skilled in the art will understand, a range includes each individual member. Thus, for example, a group having 1 - 3 items refers to a group having 1, 2, or 3 items. Similarly, a group having 1 - 5 items refers to a group having 1, 2, 3, 4, or 5 items, and so on.

[0405] Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for illustrative purposes and are not intended to limit the true scope and spirit pointed out by the following claims. Sequence Listing <110> Celula Research, Inc. Erin Sham Jue Fan Jennifer Tsai <120> Immune Receptor Barcode Error Correction <130> BDCRI.035WO <150> 62 / 562,978 <151> 2017 - 09 - 25 <160> 113 <170> PatentIn version 3.5 <210> 1 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 1 aaaaaaaaaa aaaaaaaaaa 20 <210> 2 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 2 tttttttttt tttttttttt 20 <210> 3 <211> 8 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 3 gtcaaatt 8 <210> 4 <211> 8 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 4 gtcaaaat 8 <210> 5 <211> 8 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 5 gtcaaaaa 8 <210> 6 <211> 8 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 6 ttcaaaaa 8 <210> 7 <211> 8 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 7 ttcagaaa 8 <210> 8 <211> 8 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 8 ctcaaaaa 8 <210> 9 <211> 8 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 9 ttcaaact 8 <210> 10 <211> 8 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 10 ttcaaaat 8 <210> 11 <211> 8 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 11 ttcaaaca 8 <210> 12 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 12 Cys Val Val Asn Gly Asp Gly Thr Ala Ser Lys Leu Thr Phe 1 5 10 <210> 13 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 13 Cys Val Val Asp Gly Asp Gly Thr Ala Ser Lys Leu Thr Phe 1 5 10 <210> 14 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 14 Cys Val Val Ser Gly Asp Gly Thr Ala Ser Lys Leu Thr Phe 1 5 10 <210> 15 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 15 Cys Val Val Asn Gly Val Gly Thr Ala Ser Lys Leu Thr Phe 1 5 10 <210> 16 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 16 Cys Val Val Asn Gly Asp Gly Ala Ala Ser Lys Leu Thr Phe 1 5 10 <210> 17 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 17 Cys Val Val Asn Gly Asp Gly Thr Ala Gly Lys Leu Thr Phe 1 5 10 <210> 18 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 18 Cys Val Val Asn Gly Asp Gly Thr Ala Ser Arg Leu Thr Phe 1 5 10 <210> 19 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 19 Cys Val Val Asn Gly Asp Gly Thr Ala Ser Lys Leu Ala Phe 1 5 10 <210> 20 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 20 Cys Val Val Asn Gly Asp Gly Thr Ala Ser Lys Leu Thr Leu 1 5 10 <210> 21 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 21 Cys Val Val Asn Gly Asp Gly Thr Ala Ser Lys Pro Thr Phe 1 5 10 <210> 22 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Description=“Synthetic oligonucleotide or polypeptide” <400> 22 Cys Val Val Asn Gly Asp Gly Thr Thr Ser Lys Leu Thr Phe 1 5 10 <210> 23 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Description=“Synthetic oligonucleotide or polypeptide” <400> 23 Cys Val Val Asn Glu Asp Gly Thr Ala Ser Lys Leu Thr Phe 1 5 10 <210> 24 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Description=“Synthetic oligonucleotide or polypeptide” <400> 24 Cys Val Ala Asn Gly Asp Gly Thr Ala Ser Lys Leu Thr Phe 1 5 10 <210> 25 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Description=“Synthetic oligonucleotide or polypeptide” <400> 25 Cys Val Met Asn Gly Asp Gly Thr Ala Ser Lys Leu Thr Phe 1 5 10 <210> 26 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 26 Cys Ala Val Asn Gly Asp Gly Thr Ala Ser Lys Leu Thr Phe 1 5 10 <210> 27 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 27 Arg Val Val Asn Gly Asp Gly Thr Ala Ser Lys Leu Thr Phe 1 5 10 <210> 28 <211> 13 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 28 Cys Ala Val His Arg Gly Ser Gln Gly Asn Leu Ile Phe 1 5 10 <210> 29 <211> 13 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 29 Cys Ala Val His Arg Gly Ser Gln Gly Asn Leu Ile Val 1 5 10 <210> 30 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 30 Cys Ala Gly Glu Ser Gly Asp Tyr Gln Lys Val Thr Phe 1 5 10 <210> 31 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 31 Cys Ala Gly Glu Ser Gly Gly Tyr Gln Lys Val Thr Phe 1 5 10 <210> 32 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 32 Cys Ala Ala Thr Glu Ser Tyr Gly Gln Asn Phe Val Phe 1 5 10 <210> 33 <211> 15 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic Oligonucleotide or polypeptide <400> 33 Cys Leu Val Gly Ser Leu Ser Gly Ser Ala Arg Gln Leu Thr Phe 1 5 10 15 <210> 34 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic Oligonucleotide or polypeptide <400> 34 Cys Val Val Thr Ala Ser Gly Gly Tyr Gln Lys Val Thr Phe 1 5 10 <210> 35 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic Oligonucleotide or polypeptide <400> 35 Cys Ala Val Ala Pro Tyr Gly Asn Asn Arg Leu Ala Phe 1 5 10 <210> 36 <211> 15 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic Oligonucleotide or polypeptide <400> 36 Cys Ala Val Thr Arg Phe Ser Gly Gly Tyr Asn Lys Leu Ile Phe 1 5 10 15 <210> 37 <211> 17 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 37 Cys Ala Val Ser Lys Gly Ala Arg Ser Gly Asn Thr Gly Lys Leu Ile 1 5 10 15 Phe <210> 38 <211> 10 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 38 Cys Ala Leu Phe Asn Asn Asp Met Arg Phe 1 5 10 <210> 39 <211> 13 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 39 Cys Ala Leu Ser Pro Gly Gly Tyr Gln Lys Val Thr Phe 1 5 10 <210> 40 <211> 15 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 40 Cys Ala Gly Leu Lys Leu Glu Thr Ser Gly Ser Arg Leu Thr Phe 1 5 10 15 <210> 41 <211> 11 <212> PRT <213> Synthetic Sequence <220> <223> Note = "Description of synthetic sequence: synthetic oligonucleotide or polypeptide" <400> 41 Cys Ala Gly Gly Tyr Gly Asn Lys Leu Val Phe 1 5 10 <210> 42 <211> 16 <212> PRT <213> Synthetic Sequence <220> <223> Note = "Description of synthetic sequence: synthetic oligonucleotide or polypeptide" <400> 42 Cys Ala Gly Ala Arg Gly Ser Asn Phe Gly Asn Glu Lys Leu Thr Phe 1 5 10 15 <210> 43 <211> 12 <212> PRT <213> Synthetic Sequence <220> <223> Note = "Description of synthetic sequence: synthetic oligonucleotide or polypeptide" <400> 43 Cys Ala Ala Asn Asn Ala Gly Asn Val Leu Thr Phe 1 5 10 <210> 44 <211> 15 <212> PRT <213> Synthetic Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 44 Cys Ala Ala Pro Ser Leu Gly Gly Ser Ala Arg Gln Leu Thr Phe 1 5 10 15 <210> 45 <211> 15 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 45 Cys Ala Ala Ser Ile Arg Gly Asp Ser Ser Tyr Lys Leu Ile Phe 1 5 10 15 <210> 46 <211> 13 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <220> <221> Position corresponding to the stop codon <222> (7)..(8) <400> 46 Cys Ala Ala Ser Arg Ala Asp Gly Asn Gln Phe Tyr Phe 1 5 10 <210> 47 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 47 Cys Ala Ala Ser Pro Met Asn Arg Asp Asp Lys Ile Ile Phe 1 5 10 <210> 48 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 48 Cys Ala Ala Ser Ile Thr Asp Ser Trp Gly Lys Leu Gln Phe 1 5 10 <210> 49 <211> 13 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 49 Cys Val Val Ser Ala Lys Asn Thr Asp Lys Leu Ile Phe 1 5 10 <210> 50 <211> 12 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 50 Cys Ala Tyr Arg Ser Ser Asn Tyr Gln Leu Ile Trp 1 5 10 <210> 51 <211> 15 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic Oligonucleotide or polypeptide <400> 51 Cys Ala Val Val Pro Phe Gly Gly Gly Gly Asn Lys Leu Thr Phe 1 5 10 15 <210> 52 <211> 12 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic Oligonucleotide or polypeptide <400> 52 Cys Ala Gly Trp Ser Asn Asp Tyr Lys Leu Ser Phe 1 5 10 <210> 53 <211> 13 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic Oligonucleotide or polypeptide <400> 53 Cys Ala Ala Ser Gly Gly Ser Asn Tyr Lys Leu Thr Phe 1 5 10 <210> 54 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic Oligonucleotide or polypeptide <400> 54 Cys Ala Met Arg Glu Gly Gly Gly Ser Asn Asp Tyr Lys Leu Ser Phe 1 5 10 15 <210> 55 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 55 Cys Ile Arg Leu Pro Gly Asn Thr Gly Lys Leu Ile Phe 1 5 10 <210> 56 <211> 13 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 56 Cys Ala Tyr Val Ala Ala Ala Gly Asn Lys Leu Thr Phe 1 5 10 <210> 57 <211> 12 <212> PRT <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 57 Cys Ala Gly Ala Pro Gly Ser Tyr Ile Pro Thr Phe 1 5 10 <210> 58 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 58 tgtgtggtga acggagacgg cactgccagt aaactcacct tt 42 <210> 59 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 59 tgtgtggtgg acggagacgg cactgctagt aaactcacct tt 42 <210> 60 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 60 tgtgtggtga gcggagacgg cactgccagt aaactcacct tt 42 <210> 61 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 61 tgtgtggtga acggagtcgg cactgccagt aaactcacct tt 42 <210> 62 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 62 tgtgtggtga acggagacgg cgctgccagt aaactcacct tt 42 <210> 63 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 63 tgtgtggtga acggagacgg cactgccggt aaactcacct tt 42 <210> 64 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 64 tgtgtggtga acggagacgg cactgccagt agactcacct tt 42 <210> 65 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 65 tgtgtggtga acggagacgg cactgccagt aaactcgcct tt 42 <210> 66 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 66 tgtgtggtga acggagacgg cactgccagt aaactcactt tt 42 <210> 67 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Description = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 67 tgtgtggtga acggagacgg cactgccagt aaactcaccc tt 42 <210> 68 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Description = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 68 tgtgtggtga acggagacgg cactgccagt aaacccacct tt 42 <210> 69 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Description = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 69 tgtgtggtga acggagacgg cactgccagc aaactcacct tt 42 <210> 70 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Description = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 70 tgtgtggtga acggagacgg cactaccagt aaactcacct tt 42 <210> 71 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 71 tgtgtggtga acggagacgg cacagccagt aaactcacct tt 42 <210> 72 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 72 tgtgtggtga acgaagacgg cactgccagt aaactcacct tt 42 <210> 73 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 73 tgtgtggcga acggagacgg cactgccagt aaactcacct tt 42 <210> 74 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 74 tgtgtgatga acggagacgg cactgccagt aaactcacct tt 42 <210> 75 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic "Oligonucleotide or polypeptide" <400> 75 tgtgcggtga acggagacgg cactgccagt aaactcacct tt 42 <210> 76 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic" "Oligonucleotide or polypeptide" <400> 76 tgcgtggtga acggagacgg cactgccagt aaactcacct tt 42 <210> 77 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic" "Oligonucleotide or polypeptide" <400> 77 cgtgtggtga acggagacgg cactgccagt aaactcacct tt 42 <210> 78 <211> 39 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic" "Oligonucleotide or polypeptide" <400> 78 tgtgctgtcc accgaggaag ccaaggaaat ctcatcttt 39 <210> 79 <211> 39 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: Synthetic" "Oligonucleotide or polypeptide" <400> 79 tgtgctgtcc accgaggaag ccaaggaaat ctcatcgtt 39 <210> 80 <211> 39 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 80 tgtgcaggag aatctgggga ttaccagaaa gttaccttt 39 <210> 81 <211> 39 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 81 tgtgcaggag aatctggggg ttaccagaaa gttaccttt 39 <210> 82 <211> 39 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 82 tgtgcagcaa ccgagtccta tggtcagaat tttgtcttt 39 <210> 83 <211> 39 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 83 tgtgcagcaa cagagtccta tggtcagaat tttgtcttt 39 <210> 84 <211> 45 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 84 tgcctcgtgg ggagcctttc tggttctgca aggcaactga ccttt 45 <210> 85 <211> 45 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 85 tgcctcgtgg ggagcctttc cggttctgca aggcaactga ccttt 45 <210> 86 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 86 tgtgtggtga ccgcttctgg gggttaccag aaagttacct tt 42 <210> 87 <211> 39 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 87 tgtgctgtgg ccccctatgg gaacaacaga ctcgctttt 39 <210> 88 <211> 45 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 88 tgtgctgtga ctcggttttc tggtggctac aataagctga ttttt 45 <210> 89 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 89 tgtgctgtca gtaagggggc taggtctggc aacacaggca aactaatctt t 51 <210> 90 <211> 30 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 90 tgtgctctgt ttaacaatga catgcgcttt 30 <210> 91 <211> 39 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 91 tgtgctctgt cccctggggg ttaccagaaa gttaccttt 39 <210> 92 <211> 45 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 92 tgtgcagggt taaaactaga aaccagtggc tctaggttga ccttt 45 <210> 93 <211> 33 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 93 tgtgcagggg ggtatggaaa caaactggtc ttt 33 <210> 94 <211> 48 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 94 tgtgcaggag cgaggggatc taactttgga aatgagaaat taaccttt 48 <210> 95 <211> 36 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 95 tgtgcagcta ataatgcagg caacgtgctc accttt 36 <210> 96 <211> 45 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 96 tgtgcagccc cctccctggg gggttctgca aggcaactga ccttt 45 <210> 97 <211> 45 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 97 tgtgcagcaa gtataagggg ggatagcagc tataaattga tcttc 45 <210> 98 <211> 40 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 98 tgtgcagcaa gtagagccga ccggtaacca gttctatttt 40 <210> 99 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 99 tgtgcagcaa gcccaatgaa cagagatgac aagatcatct tt 42 <210> 100 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 100 tgtgcagcaa gcataactga cagctggggg aaattgcagt tt 42 <210> 101 <211> 39 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 101 tgtgtggtga gcgcgaaaaa caccgacaag ctcatcttt 39 <210> 102 <211> 36 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 102 tgtgcttata ggagtagcaa ctatcagtta atctgg 36 <210> 103 <211> 45 <212> DNA <213> Artificial sequence <220> <223> Note = "Description of artificial sequence: synthetic oligonucleotide or polypeptide" <400> 103 tgtgctgtgg tccccttcgg gggaggagga aacaaactca ccttt 45 <210> 104 <211> 36 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 104 tgtgcaggat ggtctaacga ctacaagctc agcttt 36 <210> 105 <211> 39 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 105 tgtgcagcaa gtggaggtag caactataaa ctgacattt 39 <210> 106 <211> 48 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 106 tgtgcaatga gagagggcgg tggttctaac gactacaagc tcagcttt 48 <210> 107 <211> 39 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 107 tgcatccgcc tgcctggcaa cacaggcaaa ctaatcttt 39 <210> 108 <211> 39 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 108 tgtgcttatg tcgcagctgc aggcaacaag ctaactttt 39 <210> 109 <211> 36 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 109 tgtgcaggag ccccaggaag ctacatacct acattt 36 <210> 110 <211> 8 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 110 agtgcgag 8 <210> 111 <211> 8 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 111 agtgcggg 8 <210> 112 <211> 8 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <220> <221> Hybrid Feature <222> (6)..(6) <223> n is a, c, g, or t <400> 112 agtgcnag 8 <210> 113 <211> 8 <212> DNA <213> Artificial Sequence <220> <223> Note = "Description of artificial sequence: Synthetic oligonucleotide or polypeptide" <400> 113 atagagat 8

Claims

1. A non-therapeutic method for correcting errors in sequencing data, the method comprising: (a) obtaining sequencing data of a plurality of randomly barcoded targets, each randomly barcoded target including a random barcode from among a plurality of random barcodes, each random barcode including a cell marker and a molecular marker, wherein the molecular markers of at least two of the plurality of random barcodes include different molecular marker sequences, and wherein at least two of the plurality of random barcodes include cell markers having the same cell marker sequence; and (b) for at least one of the plurality of targets: (i) identifying a putative sequence of the target in the sequencing data; (ii) counting the occurrences of the molecular marker sequences associated with the putative sequence of the target identified in (i) in the sequencing data; (iii) using directed adjacency to identify clusters of the putative sequences of the target, wherein the putative sequences of the targets in a cluster: (A) are within a first directed adjacency threshold of each other, and (B) include one or more parental sequences and one or more subsequences of the one or more parental sequences, wherein the occurrence of the parental sequence is greater than or equal to a first directed adjacency occurrence threshold; (iv) using the clusters of the putative sequences of the target identified in (iii) to collapse the obtained sequencing data by attributing the occurrences of the subsequences in the one or more subsequences to the parental sequences of the subsequences; (v) using directed adjacency to identify clusters of the molecular marker sequences associated with the putative sequences of the target, wherein the molecular marker sequences of one or more of the clusters of the molecular marker sequences associated with the putative sequences: (A) are within a second directed adjacency threshold of each other, and (B) contain one or more parental molecular marker sequences and sub-molecular marker sequences of the one or more parental molecular marker sequences, and wherein the occurrence of the parental molecular marker sequence is greater than or equal to a second directed adjacency occurrence threshold; (vi) using the clusters of the molecular marker sequences identified in (v) to collapse the sequencing data by attributing the occurrences of the sub-molecular marker sequences in the one or more sub-molecular marker sequences to the parental molecular marker sequences of the sub-molecular marker sequences; (vii) using directed adjacency to identify clusters of combined sequences, wherein each combined sequence includes a sequence in the sequence of the target and an associated molecular marker sequence in the molecular marker sequences, wherein the combined sequences in one or more of the clusters of the combined sequences: (A) are within a third directed adjacency threshold of each other; and (B) contain one or more parental combined sequences and one or more sub-combined sequences of the parental combined sequences, and wherein the occurrence of the parental combined sequence is greater than or equal to a third directed adjacency occurrence threshold; (viii) using the clusters of the combined sequences identified in (vii) to collapse the sequencing data by attributing the occurrences of the sub-combined sequences in the one or more sub-combined sequences to the parental combined sequences of the sub-combined sequences; (ix) Identify one or more putative sequences of the target corresponding to one or more chimeric sequences of the target, wherein the occurrence of one or more putative sequences of the target corresponding to one or more chimeric sequences of the target is less than the occurrence of the remaining one or more putative sequences of the target that do not correspond to one or more chimeric sequences of the target; (x) Remove from the sequencing data the one or more putative sequences of the target identified in (ix) that correspond to one or more chimeric sequences of the target; and (xi) Estimate the occurrence of the target, wherein the estimated occurrence of the target is related to the number of molecular marker sequences counted in (ii) after folding the sequencing data in (iv), (vi), and (viii) and removing the one or more putative sequences of the target corresponding to one or more chimeric sequences of the target in (x).

2. The method according to claim 1, wherein, the plurality of targets includes targets of the entire transcriptome of a cell.

3. The method according to claim 1, wherein, the plurality of targets includes mRNA targets encoding immune receptors.

4. The method according to claim 1, wherein, the plurality of targets includes genes, and the genes include variable sequences.

5. The method according to claim 4, wherein, the gene encodes a T cell receptor.

6. The method according to claim 1, wherein, the putative sequences of the target differ from each other by at least one nucleotide.

7. The method according to claim 1, wherein, the first directed adjacency threshold is a Hamming distance, and the Hamming distance is one.

8. The method according to claim 1, wherein, the first directed adjacency occurrence threshold is twice the occurrence of a subsequence less than one.

9. The method according to any one of claims 1-8, wherein, the second directed adjacency threshold is a Hamming distance, and the Hamming distance is one.

10. The method according to any one of claims 1-8, wherein, the second directed adjacency occurrence threshold is twice the occurrence of a sub-molecular marker sequence less than one.

11. The method according to any one of claims 1-8, wherein, the third directed adjacency threshold is a Hamming distance, and the Hamming distance is one.

12. The method according to any one of claims 1-8, wherein, the third directed adjacency occurrence threshold is twice the occurrence of a sub-combination sequence less than one.

13. The method according to any one of claims 1 to 8, wherein, identifying one or more putative sequences of the target corresponding to one or more chimeric sequences of the target: identifying a putative sequence of the target associated with one molecular marker sequence among the plurality of molecular marker sequences; identifying a putative sequence among the putative sequences of the target associated with the one molecular marker sequence, wherein the occurrence of the one molecular marker sequence is less than a chimeric occurrence threshold corresponding to a chimeric sequence in one or more chimeric sequences of the target.

14. The method according to claim 13, wherein, The value of the chimeric occurrence threshold is the occurrence of a putative sequence in the putative sequence of the target associated with the one molecular marker sequence, the occurrence being greater than the occurrence of any other sequence in the putative sequence of the target.

15. The method according to any one of claims 1 to 8, the method further comprising: Adjusting the sequencing data after folding the sequencing data in (iv), (vi), and (viii) and removing one or more putative sequences of the target corresponding to one or more chimeric sequences of the target in (x).

16. The method according to claim 15, wherein, Adjusting the sequencing data after folding the sequencing data in (iv), (vi), and (viii) and removing one or more putative sequences of the target corresponding to one or more chimeric sequences of the target in (x) comprises: After folding the sequencing data in (iv), (vi), and (viii) and removing one or more putative sequences of the target corresponding to one or more chimeric sequences of the target in (x), thresholding the molecular marker sequences associated with the putative sequences of the target to determine signal molecular marker sequences and noise molecular marker sequences associated with the sequences of the target in the sequencing data counted in (b).

17. The method according to claim 16, wherein, Thresholding the molecular marker sequences associated with the putative sequences of the target includes performing a statistical analysis on the molecular marker sequences of the target.

18. The method according to claim 17, wherein, Performing the statistical analysis includes: Fitting the molecular marker sequences associated with the putative sequences of the target and their occurrences to two negative binomial distributions; Use the two negative binomial distributions to determine the occurrence of the signaling molecule marker sequence n ; and After folding the sequencing data in (iv), (vi), and (viii) and removing one or more putative sequences of the target corresponding to one or more chimeric sequences of the target in (x), the noise molecular marker sequences are removed from the sequencing data obtained in (a), wherein the noise molecular marker sequences include molecular marker sequences whose occurrences are less than the occurrence of the n th most abundant molecular marker, and wherein the signal molecular marker sequences include molecular marker sequences whose occurrences are greater than or equal to the occurrence of the n th most abundant molecular marker.

19. The method according to claim 18, wherein, The two negative binomial distributions include a first negative binomial distribution corresponding to the signal molecular marker sequences and a second negative binomial distribution corresponding to the noise molecular marker sequences.

20. A computer system for determining the occurrence of a target, the computer system comprising: A hardware processor; and A non-transitory memory having instructions stored thereon, the instructions when executed by the hardware processor cause the processor to perform the method according to any one of claims 1 to 8.

21. A computer-readable medium, the computer-readable medium includes a software program, the software program includes code for performing the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Massively parallel single cell analysis

    US20150299784A1

  • Massively parallel single cell analysis

    WO2015031691A1

  • Methods of identifying a pair of binding partners

    CN102549427A

  • Large-scale biomolecular analysis with sequence tags

    CN105658812A