Systems and methods for sample preparation, sample sequencing, and sequencing data bias correction and quality control

By verifying and bias-correcting nucleic acid sequence data, the system addresses inaccuracies in cancer characterization, facilitating precise personalized cancer treatments.

JP2025128078APending Publication Date: 2025-09-02BOSTONGENE CORP
View PDF 42 Cites 0 Cited by

Patent Information

Application Number
JP2025071801
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-03-18
Filing Date
2025-04-23
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Current methods for processing biological samples for nucleic acid sequencing and analyzing sequencing data are inadequate, leading to biased and inaccurate characterization of cancers, which hampers personalized cancer treatment decisions.

Method used

A system and method for verifying nucleic acid sequence data by processing it to determine its source and completeness, and correcting for biases in gene expression data to identify appropriate cancer treatments.

Benefits of technology

Ensures accurate characterization of cancers by correcting biases in sequencing data, enabling personalized cancer treatment decisions based on reliable nucleic acid information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025128078000001_ABST
    Figure 2025128078000001_ABST
Patent Text Reader

Abstract

To provide a system, method, and non-transitory computer-readable medium for sample preparation, sample sequencing, bias correction of sequencing data, and quality control.SOLUTION: A system for sample preparation, sample sequencing, bias correction of sequencing data, and quality control obtains a tumor sample from a subject who has cancer, is suspected of having cancer, or is at risk of having cancer. Next, one or more quality control evaluations are performed on the primary sample. Furthermore, one or more quality control evaluations are performed on DNA, RNA, and / or library processes. In addition, one or more quality control evaluations are performed on nucleic acid sequence data.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Application No. 62 / 870,622, filed July 3, 2019, entitled "Compositions and Methods for Sample Preparation and Characterization of Cancer Therefrom," and U.S. Provisional Application No. 62 / 991,570, filed March 18, 2020, entitled "Nucleic Acid Data Quality Control," the entire disclosures of each of which are incorporated herein by reference.

[0002] Some aspects of the techniques described herein relate to collecting and processing tumor and / or healthy tissue samples to extract nucleic acid and perform nucleic acid sequencing. Some aspects of the techniques described herein relate to processing nucleic acid sequencing data to remove bias from nucleic acid sequencing data. Also described herein are various methods for evaluating the quality of nucleic acid sequence information obtained by sequencing. [Background technology]

[0003]

[0003] Properly characterizing one or more types of cancer a patient or subject has, and potentially selecting one or more effective therapies for the patient based on that characterization, can be crucial to the patient's survival and overall health. The manner in which biological samples from subjects are processed to obtain sequence data (e.g., RNA expression data) for characterizing one or more types of cancer, and the manner in which the data is processed, can adversely affect the characterization of one or more cancers. For example, high-throughput nucleic acid sequencing platforms (e.g., next-generation sequencing platforms) can generate large amounts of DNA and RNA sequence data from patient samples. Advances in sample preparation, data processing, and evaluation of sequence information from different NGS platforms using custom software are required to characterize cancers, predict prognosis, identify effective therapies, and otherwise support personalized care for patients with cancer. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] U.S. Provisional Patent Application No. 62 / 943,976 [Patent Document 2] International PCT Publication No. WO2018 / 231771 [Patent Document 3] PCT Application No. PCT / US20 / 037017 [Patent Document 4] International PCT Publication No. WO2018 / 231772 [Patent Document 5] International Patent Application No. PCT / US2018 / 037018 [Patent Document 6] International PCT Publication No. WO2018 / 231762 [Patent Document 7] International Patent Application No. PCT / US2018 / 037008 [Patent Document 8] International PCT Publication No. WO00 / 53211 [Patent Document 9] U.S. Patent No. 5,981,568 [Patent Document 10] International PCT Publication No. WO90 / 07936 [Patent Document 11] International PCT Publication No. WO94 / 03622 [Patent Document 12] International PCT Publication No. WO93 / 25698 [Patent Document 13] International PCT Publication No. WO93 / 25234 [Patent Document 14] International PCT Publication No. WO93 / 11230 [Patent Document 15] International PCT Publication No. WO93 / 10218 [Patent Document 16] International PCT Publication No. WO91 / 02805 [Patent Document 17] U.S. Patent No. 5,219,740 [Patent Document 18] U.S. Patent No. 4,777,127 [Patent Document 19] British Patent No. 2,200,651 [Patent Document 20] European Patent No. 0 345 242 [Patent Document 21] International PCT Publication No. WO94 / 12649 [Patent Document 22] International PCT Publication No. WO93 / 03769 [Patent Document 23] International PCT Publication No. WO93 / 19191 [Patent Document 24] International PCT Publication No. WO94 / 28938 [Patent Document 25] International PCT Publication No. WO95 / 11984 [Patent Document 26] International PCT Publication No. WO95 / 00655 [Patent Document 27] U.S. Patent No. 5,814,482 [Patent Document 28] International PCT Publication No. WO95 / 07994 [Patent Document 29] International PCT Publication No. WO96 / 17072 [Patent Document 30] International PCT Publication No. WO95 / 30763 [Patent Document 31] International PCT Publication No. WO97 / 42338 [Patent Document 32] International PCT Publication No. WO90 / 11092 [Patent Document 33] U.S. Patent No. 5,580,859 [Patent Document 34] U.S. Patent No. 5,422,120 [Patent Document 35] International PCT Publication No. WO95 / 13796 [Patent Document 36] International PCT Publication No. WO94 / 23697 [Patent Document 37] International PCT Publication No. WO91 / 14445 [Patent Document 38] European Patent No. 0524968 [Non-patent literature]

[0005] [Non-Patent Document 1] Nicolas L Bray, Harold Pimentel, Pall Melsted and Lior Pachter, "Near-optimal probabilistic RNA-seq quantification," Nature Biotechnology 34, pp. 525–527 (2016), doi:10.1038 / nbt.3519 [Non-patent document 2] Vaught et al., "Biospecimen and biorepositories: from afterthought to science" (Cancer Epidemiol Biomarkers Prev. 2012, February 21(2):253-5) [Non-patent document 3] Vaught and Henderson, "Biological sample collection, processing, storage and information management," IARC Sci Publ. 2011, (163):23-42 [Non-patent document 4] BioFiles: For Life Science Research, Issue 2, 2006, www.sigmaaldrich.com / content / dam / sigma-aldrich / docs / Sigma / General_Information / 2 / biofiles_issue2.pdf [Non-patent document 5] Quatromoni et al., “An optimized disaggregation method for human lung tumors that preserves the phenotype and function of the immune cell,” J Leukoc Biol. 2015 Jan, 97(1): pp. 201–209. [Non-patent document 6] Pennartz et al., “Generation of Single-Cell Suspensions from Mouse Neural Tissue,” JOVE Issue 29, doi:10.3791 / 1267, Published:7 / 07 / 2009 [Non-Patent Document 7] www.youtube.com / watch?v=N0jftyYqM38 [Non-patent document 8] Heng et al., Biol Proced Online. 2009, 11:161-169 [Non-Patent Document 9] hemberg-lab.github.io / scRNA.seq.course / introduction-to-single-cell-rna-seq.html [Non-Patent Document 10] Bagnoli et al. “Studying Cancer Heterogeneity by Single-Cell RNA Sequencing,” Methods Mol Biol. 2019, 1956: 305-319. [Non-Patent Document 11] Sun et al., "Single-cell RNA sequencing reveals gene expression signatures of breast cancer-associated endothelial cells," Oncotarget. February 16, 2018, 9(13): 10945-10961. [Non-Patent Document 12] Kulkarni et al., "Beyond bulk: a review of single cell transcriptomics methodologies and applications," Curr Opin Biotechnol. 2019, April 9, 58:129-136. [Non-Patent Document 13] Huang et al., “High Throughput Single Cell RNA Sequencing, Bioinformatics Analysis and Applications,” Adv Exp Med Biol. 2018, 1068:33–43. [Non-Patent Document 14] Zilionis et al. “Single-Cell Transcriptomics of Human and Mouse Lung Cancers Reveals Conserved Myeloid Populations across Individuals and Species”, Immunity. April 5, 2019, pii:S1074-7613(19)30126-8 [Non-Patent Document 15] Kashima "An Informative Approach to Single-Cell Sequencing Analysis", Adv Exp Med Biol. 2019, pages 1129:81~96, doi:10.1007 / 978-981-13-6037-4_6

Non-patent document 16

Non-patent document 17

Non-patent document 18

Non-patent document 19

Non-patent document 20

Non-patent document 21

Non-patent document 22

Non-patent document 23

Non-patent document 24

Non-patented document 25

Non-patent document 26

Non-patent document 27

Non-patent document 28

Non-patent document 36

Non-patent document 37

Non-patent document 38

Non-patent document 39

Non-patent document 40

Non-patent document 41

Non-patent document 42

Non-patent document 43

Non-patent document 44

Non-patented document 45

Non-patent document 46

[0006] Some embodiments include at least one computer hardware processor and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method, the method comprising: Provided is a system that includes obtaining nucleic acid data, the nucleic acid data including sequence data indicating at least 5 kilobases (kb) of nucleotide sequence of DNA and / or RNA from a previously obtained biological sample of a subject having, suspected of having, or at risk of having a disease, and claimed information indicating the claimed source and / or claimed completeness of the sequence data; and verifying the nucleic acid data by processing the sequence data to obtain determined information indicating the determined source and / or determined completeness of the sequence data; and determining whether the determined information matches the claimed information.

[0007] Some embodiments provide at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method, comprising obtaining nucleic acid data, the nucleic acid data representing at least 5 kilobases (kb) of nucleotide sequence of DNA and / or RNA from a previously obtained biological sample of a subject having, suspected of having, or at risk of having a disease; obtaining nucleic acid data, including claimed information indicating the claimed source and / or claimed completeness of the sequence data; verifying the nucleic acid data by processing the sequence data to obtain determined information indicating the determined source and / or determined completeness of the sequence data; and determining whether the determined information matches the claimed information.

[0008] Some embodiments use at least one hardware processor to: obtain nucleic acid data, the nucleic acid data including sequence data indicating at least 5 kilobases (kb) of nucleotide sequence of DNA and / or RNA from a previously obtained biological sample of a subject having, suspected of having, or at risk of having a disease, and claimed information indicating the claimed source and / or claimed completeness of the sequence data; and verify the nucleic acid data by processing the sequence data to obtain determined information indicating the determined source and / or determined completeness of the sequence data; and determining whether the determined information matches the claimed information.

[0009] In some embodiments, the sequence data may include raw DNA or RNA sequence data, DNA exome sequence data (e.g., from whole exome sequencing (WES)), DNA genomic sequence data (e.g., from whole genome sequencing (WGS)), RNA expression data, gene expression data, bias-corrected gene expression data, or any other suitable type of sequence data, including data obtained from a sequencing platform and / or data derived from data obtained from a sequencing platform.

[0010] In some embodiments, the method further includes, upon determining that the asserted information matches the determined information, processing the sequence data to determine whether the sequence data is indicative of one or more disease characteristics.

[0011] In some embodiments, the method further includes determining whether the determined information matches the purported information and processing the sequence data to determine whether it is indicative of one or more disease characteristics.

[0012] In some embodiments, the method further includes, when it is determined that the asserted information does not match the determined information, generating an indication that the determined information does not match the asserted information and not processing the sequence data in subsequent analysis and / or obtaining additional sequence data and / or biological sample and / or other information about the subject.

[0013] In some embodiments, the method further includes determining that the asserted information does not match the determined information and generating an indication that the determined information does not match the asserted information to not process the sequence data in subsequent analysis and / or obtain additional sequence data and / or biological sample and / or other information about the subject.

[0014] In some embodiments, the asserted information indicates a asserted source of the sequence data, and the method further includes processing the sequence data to obtain determined information indicating a determined source for the sequence data, and determining whether the determined source is consistent with the asserted source for the sequence data.

[0015] In some embodiments, the determined information indicating the determined source for the sequence data indicates the subject's MHC genotype, whether the nucleic acid data is RNA or DNA data, the tissue type of the biological sample, the tumor type of the biological sample, the sequencing platform used to generate the sequence data, the SNP match, and / or whether the RNA sample is polyA enriched.

[0016] In some embodiments, the determined information indicating the determined source for the sequence data indicates at least two of the subject's MHC genotype, whether the nucleic acid data is RNA data or DNA data, the tissue type of the biological sample, the tumor type of the biological sample, the sequencing platform used to generate the sequence data, the SNP match, and whether the RNA sample is polyA enriched.

[0017] In some embodiments, the determined information indicating the determined source for the sequence data indicates at least three of the subject's MHC genotype, whether the nucleic acid data is RNA or DNA data, the tissue type of the biological sample, the tumor type of the biological sample, the sequencing platform used to generate the sequence data, the SNP match, and whether the RNA sample is polyA enriched.

[0018] In some embodiments, the claimed information indicates a claimed completeness of the sequence data, and the method further includes processing the sequence data to obtain determined information indicative of a determined completeness for the sequence data, and determining whether the determined completeness is consistent with the claimed completeness for the sequence data.

[0019] In some embodiments, the determined information indicating the determined completeness indicates total sequence coverage, exon coverage, chromosomal coverage, the ratio of nucleic acids encoding two or more subunits of a multimeric protein, species contamination, single nucleotide polymorphisms (SNPs), complexity, and / or the percentage (%) of guanine (G) and cytosine (C) in the sequence data.

[0020] In some embodiments, the determined information indicative of the determined completeness indicates at least two of total sequence coverage, exon coverage, chromosome coverage, the ratio of nucleic acids encoding two or more subunits of a multimeric protein, species contamination, single nucleotide polymorphisms (SNPs), complexity, and the percentage (%) of guanine (G) and cytosine (C) in the sequence data.

[0021] In some embodiments, the determined information indicating the determined completeness indicates at least three of total sequence coverage, exon coverage, chromosome coverage, the ratio of nucleic acids encoding two or more subunits of a multimeric protein, species contamination, single nucleotide polymorphisms (SNPs), complexity, and the percentage (%) of guanine (G) and cytosine (C) in the sequence data.

[0022] In some embodiments, the asserted information for the sequence data includes the subject's MHC allele information.

[0023] In some embodiments, the method further includes determining one or more MHC allele sequences from the sequence data and determining whether the one or more MHC allele sequences match the claimed MHC allele information for the subject.

[0024] In some embodiments, determining one or more MHC allele sequences comprises determining MHC allele sequences for six MHC loci from the sequence data.

[0025] In some embodiments, the sequence data indicates the nucleotide sequence for the RNA, and the purported information indicates whether the RNA is polyA enriched.

[0026] In some embodiments, the sequence data is used to determine a therapy for a subject when the asserted information is determined to match the determined information.

[0027] In some embodiments, determining the therapy comprises determining a plurality of gene group expression levels, the plurality of gene group expression levels comprising a gene group expression level for each gene group in a set of genes, the set of genes comprising at least one gene group associated with cancer aggressiveness and at least one gene group associated with the cancer microenvironment, and identifying the therapy using the determined gene group expression levels.

[0028] In some embodiments, the method further comprises administering a therapy to the subject.

[0029] In some embodiments, the determined information is determined to match the purported information, the sequence data is processed to determine a therapy for the subject, and the therapy is administered to the subject.

[0030] In some embodiments, the disease is cancer and the therapy is a cancer treatment. In some embodiments, the subject is a human.

[0031] In some embodiments, processing the sequence data to obtain a determined source includes determining one or more single nucleotide polymorphisms (SNPs) in the sequence data and determining whether the one or more SNPs in the sequence data match one or more SNPs in a reference sequence.

[0032] In some embodiments, the reference sequence is the sequence of a nucleic acid in a second biological sample from the subject.

[0033] In some embodiments, processing the sequence data to obtain a determined completeness includes determining a first level of a first nucleic acid encoding a first subunit of the multimeric protein, determining a second level of a second nucleic acid encoding a second subunit of the multimeric protein, and determining whether the ratio between the first level and the second level matches an expected ratio. In some embodiments, the multimeric protein is a dimer. In some embodiments, the first subunit and the second subunit are a first and a second CD3 subunit, a first and a second CD8 subunit, or a first and a second CD79 subunit.

[0034] Some embodiments provide a system for identifying a cancer treatment for a subject having, suspected of having, or at risk of having cancer, the system including at least one sequencing platform configured to generate gene expression data from enriched RNA obtained from a first biological sample previously obtained from the subject, wherein the enriched RNA is obtained by: (i) extracting RNA from the first biological sample of a first tumor to obtain an extracted RNA; and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA, wherein the RNA expression data comprises at least 5 kilobases (kb). The at least one sequencing platform includes at least one computer hardware processor; and a processor-implemented and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to obtain RNA expression data using at least one sequencing platform; convert the RNA expression data into gene expression data; determine bias-corrected gene expression data from the gene expression data, at least in part by removing from the gene expression data expression data for at least one gene that introduces a bias in the gene expression data; and identify a cancer treatment for the subject using the bias-corrected gene expression data.

[0035] Some embodiments provide a system for identifying a cancer treatment for a subject having, suspected of having, or at risk of having cancer, the system comprising at least one computer hardware processor and at least one non-transitory computer readable storage medium storing processor-executable instructions, the processor-executable instructions, when executed by the at least one computer hardware processor, causing the at least one computer hardware processor to obtain RNA expression data from at least one sequencing platform, the RNA expression data comprising at least 5 kilobases (5 kb), the RNA expression data comprising a first sequence of a first tumor previously obtained from the subject. and at least one non-transitory computer-readable storage medium configured to perform the following steps: obtain RNA expression data from a biological sample, the RNA expression data being obtained at least in part by (i) extracting RNA from a first biological sample of a first tumor to obtain extracted RNA, and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA; convert the RNA expression data into gene expression data; determine bias-corrected gene expression data from the gene expression data, at least in part by removing expression data for at least one gene that introduces bias into the gene expression data from the gene expression data; and identify a cancer treatment for the subject using the bias-corrected gene expression data. In some embodiments, the system may further comprise at least one sequencing platform.

[0036] Some embodiments provide at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to: obtain RNA expression data from at least one sequencing platform, wherein the RNA expression data comprises at least five kilobases (5 kb), and the RNA expression data was obtained from a first biological sample of a first tumor previously obtained from a subject having, suspected of having, or at risk of having cancer, at least in part by (i) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA, and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA; converting the RNA expression data into gene expression data; determining bias-corrected gene expression data from the gene expression data, at least in part by removing from the gene expression data expression data for at least one gene that introduces a bias in the gene expression data; and identifying a cancer treatment for the subject using the bias-corrected gene expression data.

[0037] Some embodiments provide a method including: obtaining a first biological sample of a first tumor, the first biological sample previously obtained from a subject having, suspected of having, or at risk of having cancer; extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; enriching the extracted RNA for coding RNA to obtain enriched RNA; sequencing the enriched RNA using at least one sequencing platform to obtain RNA expression data comprising at least 5 kilobases (kb); obtaining the RNA expression data using the at least one sequencing platform using at least one hardware processor; converting the RNA expression data into gene expression data; determining bias-corrected gene expression data from the gene expression data, at least in part by removing expression data for at least one gene that introduces bias into the gene expression data from the gene expression data; and identifying a cancer treatment for the subject using the bias-corrected gene expression data.

[0038] In some embodiments, the method further comprises administering the identified cancer treatment to the subject.

[0039] In some embodiments, enriching the RNA for coding RNA comprises performing polyA enrichment.

[0040] In some embodiments, the at least one gene that introduces bias into the gene expression data comprises a gene having an average transcript length that is longer or shorter than the average length of the transcripts in the gene expression data, a gene having at least one threshold variation in average transcript expression level based on the transcript expression levels in a reference sample, and / or a gene having a poly-A tail that is less in length by at least a threshold amount compared to the average length of the poly-A tails of genes from the first biological sample from which the RNA expression data was obtained and / or the reference sample.

[0041] In some embodiments, at least one gene that introduces bias into the gene expression data belongs to a gene family selected from the group consisting of histone-encoding genes, mitochondrial genes, interleukin-encoding genes, collagen-encoding genes, B-cell receptor-encoding genes, and T-cell receptor-encoding genes.

[0042] In some embodiments, the at least one gene is selected from the group consisting of HIST1H1A, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H1T, HIST1H2AA, HIST1H2AB, HIST1H2AC, HIST1H2AD, HIST1H2AE, HIST1H2AG, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H2AK, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H2AK, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H2AK, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H2AK, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H2AK, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H2AI ... T1H2AL, HIST1H2AM, HIST1H2BA, HIST1H2BB, HIST1H2BC, HIST1H2BD, HIST1H2BE, HIST1H2BF, HIST1H2BG, HIST1 H2BH, HIST1H2BI, HIST1H2BJ, HIST1H2BK, HIST1H2BL, HIST1H2BM, HIST1H2BN, HIST1H2BO, HIST1H3A, HIST1H3B, HIST1H3C, HIST1H3D, HIST1H3E, HIST1H3F, HIST1H3G, HIST1H3H, HIST1H3I, HIST1H3J, HIST1H4A, HIST1H4B, HI ST1H4C, HIST1H4D, HIST1H4E, HIST1H4F, HIST1H4G, HIST1H4H, HIST1H4I, HIST1H4J, HIST1H4K, HIST1H4L, HIST and at least one histone-encoding gene selected from the group consisting of HIST2H2AA3, HIST2H2AA4, HIST2H2AB, HIST2H2AC, HIST2H2BE, HIST2H2BF, HIST2H3A, HIST2H3C, HIST2H3D, HIST2H3PS2, HIST2H4A, HIST2H4B, HIST3H2A, HIST3H2BB, HIST3H3, and HIST4H4.

[0043] In some embodiments, the at least one gene is selected from the group consisting of MT-ATP6, MT-ATP8, MT-CO1, MT-CO2, MT-CO3, MT-CYB, MT-ND1, MT-ND2, MT-ND3, MT-ND4, MT-ND4L, MT-ND5, MT-ND6, MT-RNR1, MT-RNR2, MT-TA, MT-TC, MT-TD, MT-TE, MT-TF, MT-TG, MT-TH, MT-TI, MT-TK, MT-TL1, MT-TL2, and M and at least one mitochondrial gene selected from the group consisting of MTRNR2L1, MTRNR2L10, MTRNR2L11, MTRNR2L12, MTRNR2L13, MTRNR2L3, MTRNR2L4, MTRNR2L5, MTRNR2L6, MTRNR2L7, and MTRNR2L8.

[0044] In some embodiments, determining the bias-corrected gene expression data further comprises renormalizing the gene expression data after removing expression data for at least one gene that introduces bias into the gene expression data.

[0045] In some embodiments, converting the RNA expression data to gene expression data includes removing non-coding transcripts from the RNA expression data to obtain filtered RNA expression data, and normalizing the filtered RNA expression data after removing the non-coding transcripts to obtain gene expression data in Transcripts Per Million (TPM) and / or any other suitable format.

[0046] In some embodiments, removing non-coding transcripts from RNA expression data removes non-coding transcripts from the RNA expression data, such as pseudogenes, polymorphic pseudogenes, processed pseudogenes, transcribed processed pseudogenes, unitary pseudogenes, non-processed pseudogenes, transcribed unitary pseudogenes, constant chain immunoglobulin (IG C) pseudogenes, joining chain immunoglobulin (IG J) pseudogenes, variable chain immunoglobulin (IG V) genes, transcribed non-processed genes, translated non-processed genes, joining chain T cell receptor (TR J) genes, variable chain T cell receptor (TR V) genes, small nuclear RNAs (snRNAs), small nucleolar RNAs (snoRNAs), microRNAs (miRNAs), ribozymes, ribosomal RNAs (rRNAs), mitochondrial tRNAs (Mt tRNAs), mitochondrial rRNAs (Mt The method includes removing non-coding transcripts belonging to a group selected from the list consisting of rRNA, Cajal body-specific RNA (scaRNA), residual introns, sense intron RNA, sense overlapping RNA, nonsense mutation-mediated decay RNA, non-stop decay RNA, antisense RNA, long intervening non-coding RNA (lincRNA), macro-long non-coding RNA (macro-lncRNA), processed transcripts, 3' overlapping non-coding RNA (3' overlapping ncrna), small RNA (sRNA), miscellaneous RNA (miscRNA), vault RNA (vaultRNA), and TEC RNA.

[0047] In some embodiments, information (e.g., sequence information) for one or more transcripts for one or more of these types of transcripts can be obtained in a nucleic acid database (e.g., a Gencode database, e.g., Gencode V23, a Genbank database, an EMBL database, or other database).

[0048] In some embodiments, the method further comprises aligning the RNA expression data to a reference and annotating the RNA expression data before performing the removal of non-coding transcripts.

[0049] In some embodiments, the RNA expression data comprises at least 25 million paired-end reads. In some embodiments, the RNA expression data comprises at least 50 million paired-end reads with an average read length of at least 100 bp.

[0050] In some embodiments, identifying a cancer treatment for a subject using the bias-corrected gene expression data comprises: using the bias-corrected gene expression data to determine a plurality of gene group expression levels, wherein the plurality of gene group expression levels comprises a gene group expression level for each gene group in a set of genes, wherein the set of genes comprises at least one gene group associated with cancer aggressiveness and at least one gene group associated with the cancer microenvironment; and identifying the cancer treatment using the determined gene group expression levels.

[0051] In some embodiments, the cancer treatment is selected from the group consisting of radiation therapy, surgery, chemotherapy, and immunotherapy.

[0052] In some embodiments, the method further comprises obtaining a second biological sample of a second tumor, wherein the second biological sample was previously obtained from the subject.

[0053] In some embodiments, the method further includes combining the first biological sample and the second biological sample to form a combined tumor sample, and extracting RNA includes extracting RNA from the combined tumor sample.

[0054] In some embodiments, the method further includes extracting RNA from a second biological sample and combining the RNA extracted from the second biological sample with the RNA extracted from the first biological sample to form a combined extracted RNA, wherein enriching the RNA for coding RNA includes enriching the combined extracted RNA for coding RNA.

[0055] In some embodiments, the extracted RNA comprises at least 1 μg of RNA after RNA extraction.

[0056] In some embodiments, the extracted RNA has a total mass of at least 1000-6000 ng and a purity corresponding to a ratio of absorbance at 260 nm to absorbance at 280 nm of at least 2.0.

[0057] In some embodiments, the method further includes performing a quality control assessment on the RNA expression data by, at least in part, obtaining claimed information indicative of the claimed source and / or claimed completeness of the RNA expression data, processing the RNA expression data to obtain determined information indicative of the determined source and / or determined completeness of the RNA expression data, and determining whether the determined information matches the claimed information.

[0058] In some embodiments, processing the RNA expression data includes processing the RNA expression data to determine a tissue type of the first biological sample, a tumor type of the first biological sample, and / or a percentage (%) of guanine (G) and / or cytosine (C).

[0059] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure, which may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein. The drawings are not necessarily drawn to scale. [Brief explanation of the drawings]

[0060] [Figure 1A]1 is an exemplary flowchart showing a sample preparation and quality control process, illustrating an example of a process pipeline that includes one or more quality control assessments during biopsy sample collection, DNA / RNA extraction and library construction, and / or nucleic acid bioinformatics analysis. [Figure 1B] FIG. 1 is an exemplary flowchart showing a sample preparation and quality control process, illustrating one example of a process for obtaining a subject's biopsy sample, extracting nucleic acid from the sample, performing nucleic acid sequencing, and processing the nucleic acid sequence to identify one or more appropriate cancer therapies for the subject. [Figure 2A] 1 is a graphical representation showing the levels and distribution of RNA transcripts depending on the type of RNA enrichment method used and whether stranded or non-stranded RNA was used for sequencing; FIG. 2 is a graphical representation of the distribution of RNA after RNA enrichment by either ribosomal RNA (r-RNA) depletion or polyA enrichment. [Figure 2B] 1 is a graphical representation showing the levels and distribution of RNA transcripts depending on the type of RNA enrichment method used and whether stranded or non-stranded RNA was used for sequencing; and FIG. 2 is a graphical representation of the levels of RNA measured after RNA sequencing of either stranded or non-stranded RNA for IL24, ICAM4, and GAPDH. [Figure 3] Graphical representation showing the distribution of different RNA transcripts for different types of RNA as indicated in the legend. Each column represents a unique sample. All samples were prepared from the same tissue type, using the same RNA enrichment method, and the same sequencing service. The bottom panel shows the data in the top panel in Transcripts per Kilobase Million (TPM). [Figure 4A] FIG. 1 shows the distribution of polyA tails of RNA transcripts from HeLa cell samples and some examples of polyA tails for histone family genes. [Figure 4B]FIG. 1 shows a comparison of mitochondrial RNA expression for samples that were polyA enriched or enriched by rRNA depletion (represented by total RNA). [Figure 4C] FIG. 1 shows a comparison of the expression of histone-coding RNAs for samples that were polyA enriched or enriched by rRNA depletion (represented by total RNA). [Figure 5] Principal component analysis (PCA) of RNA expression in cell samples containing different proportions of HeLa cells, with or without polyA enrichment, and with or without data filtering. Data filtering included removing non-coding RNA transcripts, histone-encoding transcripts, and mitochondrial transcripts. The PCA0 component describes the significant differences between polyA sequencing and total RNA sequencing. The PCA1 component describes the different cell line ratios. Samples were prepared as a mixture of two different cell lines at five different ratios. [Figure 6A] 1 is an exemplary flowchart illustrating a process 200 for obtaining enriched RNA-seq data from tumors of subjects having, suspected of having, or at risk of having cancer. [Figure 6B] 2 is an exemplary flowchart showing a process 210 for obtaining bias-corrected gene expression data from RNA expression data to identify cancer treatments for subjects with, suspected of having, or at risk of having cancer. [Figure 6C] 2 is an exemplary flowchart showing a process 220 for processing RNA obtained from a tumor sample to identify a cancer treatment for a subject having, suspected of having, or at risk of having cancer. [Figure 7A] 3 is a flowchart of an exemplary process pipeline 300 including a bioinformatics quality control process for evaluating nucleic acid sequence data obtained from tumor samples and using the nucleic acid sequence data to identify cancer treatments for subjects with, suspected of having, or at risk of having cancer. [Figure 7B] 3 is a flowchart of an exemplary process pipeline 300 including a bioinformatics quality control process for evaluating nucleic acid sequence data obtained from tumor samples and using the nucleic acid sequence data to identify cancer treatments for subjects with, suspected of having, or at risk of having cancer. [Figure 8] 8 is an exemplary flowchart illustrating a process 800 showing a computerized process for processing and validating sequence data and related information. [Figure 9] FIG. 5 is a block diagram of an exemplary computer system 500 that may be used to implement one or more embodiments of a process pipeline for preparing, evaluating, and / or analyzing sequence data. [Figure 10] 6 is a block diagram of an example environment 600 in which one or more embodiments of the techniques described in this application may be implemented. [Figure 11] FIG. 1 shows the results of MHC allele analysis on sequence information obtained from three nucleic acids (RNA-Seq, WES tumor, and WES normal) for two subjects (103 and 105). [Figure 12] An example of a bar graph showing the probability that sequence information is from a particular type of tumor (eg, BRCA, which is associated with breast cancer). [Figure 13A] 1 is a graph depicting an example of the relationship between expression levels of protein subunits. [Figure 13B] 1 is a graph depicting an example of the relationship between expression levels of protein subunits. [Figure 14A] 1 is an example of a bar graph showing the probability of obtaining sequence information from a sample containing only polyadenylated RNA or from a sample containing total or whole RNA (total RNA). [Figure 14B] 1 is an example of a bar graph showing the probability of obtaining sequence information from a sample containing only polyadenylated RNA or from a sample containing total or whole RNA (total RNA). [Figure 15]FIG. 1 illustrates the analysis of three batches of gene expression data containing tumor and normal samples, showing an example of principal component analysis. DETAILED DESCRIPTION OF THE INVENTION

[0061] Recent advances in personalized genome sequencing and cancer genome sequencing technology have made it possible to obtain patient-specific information about cancer cells (e.g., tumor cells) and the cancer microenvironment from one or more biological samples obtained from individual patients. The inventors have realized that this information can be used to characterize the type of cancer a patient has and potentially select one or more effective therapies for the patient. This information can also be used to determine how a patient is responding to treatment over time and, if necessary, to select one or more new therapies for the patient. This information can also be used to determine whether a patient should be included or excluded from participation in a clinical trial.

[0062] The inventors recognize that the workflow used to obtain sequence data for a patient strongly influences the inferences that can be drawn about the patient's cancer, including, but not limited to, determining whether the patient will respond to one or more particular therapies, whether the patient will have an adverse reaction to one or more particular therapies, whether the patient is a candidate for enrollment in a clinical trial, whether the patient has one or more particular biomarkers (e.g., biomarkers indicative of potential response to a therapy, biomarkers indicative of survival, etc.), whether the patient will experience disease progression (e.g., from early-stage cancer to late-stage cancer, relapse from remission, etc.), whether a different therapy or therapies should be selected for the patient, and / or other suitable prognostic, diagnostic, and / or clinical inferences.

[0063] If the workflow used to obtain sequence data contains errors, suboptimal processing, sources of bias in the data, and the like, it is often not possible to make inferences about a subject's cancer with the desired or necessary confidence, or even at all. Worse yet, errors in the workflow for generating sequence data can lead to incorrect inferences about the patient, potentially resulting in inappropriate treatment or missed opportunities for better treatment. Furthermore, workflow errors can lead to wasted laboratory resources (e.g., having to reprocess samples) and wasted computing resources (e.g., performing expensive computations on megabytes or gigabytes of sequence data, occupying processor and networking resources only to later discard the results and / or have to repeat the process).

[0064] A conventional workflow used to obtain sequence data for a patient involves multiple steps, including obtaining a biological sample from the patient (e.g., by performing a biopsy, obtaining a blood sample, a saliva sample, or any other suitable biological sample from the patient), preparing the biological sample for sequencing using a sequencing platform (e.g., a next-generation sequencing (NGS) platform), and obtaining raw data output by the sequencing platform. Various conventional bioinformatics processing pipelines and other algorithms can then use the raw data output by the sequencing platform in an attempt to make one or more of the above-mentioned inferences.

[0065] However, such a conventional workflow for obtaining sequencing data is prone to errors at every stage. For example, errors may occur in laboratories when handling samples from multiple patients. In fact, it is not uncommon for laboratories to receive biological samples claimed to be from one patient, when in fact the sample is from another patient. As another example, the biological sample may not be properly processed in the laboratory and may not have the nucleic acid concentration and / or quality required for subsequent analysis. As yet another example, errors may be introduced by the sequencing platform itself and / or subsequent post-processing steps (e.g., alignment and variant calling). As yet another example, the raw sequencing data generated by the sequencing platform may contain artifacts and undesired sequences and / or transcripts. Other examples of various errors are described herein.

[0066] In some embodiments, the sequence data or sequencing data may include raw DNA or RNA sequence data, DNA exome sequence data (e.g., from whole exome sequencing (WES)), DNA genomic sequence data (e.g., from whole genome sequencing (WGS)), RNA expression data, gene expression data, bias-corrected gene expression data, or any other suitable type of sequence data, including data obtained from a sequencing platform and / or data derived from data obtained from a sequencing platform.

[0067] To address the shortcomings of conventional workflows for obtaining patient sequencing data, the inventors have developed techniques that address various sources of errors that may be present in the sequencing data. These techniques developed by the inventors include: (1) novel sample preparation techniques for preparing biological samples for sequencing using one or more sequencing platforms; (2) novel techniques for post-processing the raw data output by the sequencing platforms to remove irrelevant data and sources of bias (e.g., non-coding transcripts and gene-associated expression data that introduce bias into the sequence data); and (3) novel quality control techniques that facilitate the detection and correction of errors in the sequence data. In some embodiments, techniques from each of these three categories may be utilized in a workflow for obtaining patient sequence data; however, it should be understood that this is not a limitation of the techniques described herein, and that in some embodiments, any one or more (but not necessarily all) of these techniques may be used in a workflow.

[0068] By way of example, in some embodiments, novel sample preparation and post-processing techniques may obtain sequencing data and remove sources of bias from the sequencing data by: (1) obtaining a first biological sample of a first tumor, where the first biological sample was previously obtained from a subject having, suspected of having, or at risk of having cancer; (2) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; (3) enriching the extracted RNA for coding RNA to obtain enriched RNA; and (4) sequencing the enriched RNA using at least one sequencing platform. (4) performing sequencing to obtain RNA expression data comprising at least 5 kilobases (kb); and (5) using at least one hardware processor to: (a) obtain the RNA expression data using at least one sequencing platform; (b) convert the RNA expression data into gene expression data; (c) determine bias-corrected gene expression data from the gene expression data, at least in part by removing expression data for at least one gene that introduces bias into the gene expression data; and (d) identify a cancer treatment for the subject using the bias-corrected gene expression data.

[0069] Removing bias from gene expression data in this way improves sequencing technology for many reasons. First, it removes artifacts and sources of bias from sequencing data, resulting in fewer errors in any downstream processing and higher fidelity output. Second, the inventors recognize that removing sources of bias in this way allows for a more accurate and faithful representation of a patient's molecular functional characteristics (for example, via the molecular functional expression signature described herein). The inventors recognize that bias-corrected gene expression data can be used to identify more effective therapies for patients, improve the ability to determine whether one or more cancer therapies are effective when administered to patients, improve the ability to identify clinical trials in which subjects can participate, and / or identify improvements to many other prognostic, diagnostic, and clinical applications.

[0070] As another example, in some embodiments, a novel quality control technique includes using at least one computer hardware processor to: (a) obtain nucleic acid data, the nucleic acid data including (i) sequence data representing at least 5 kilobases (kb) of nucleotide sequence of DNA and / or RNA from a previously obtained biological sample of a subject having, suspected of having, or at risk of having a disease, and (ii) claimed information indicating the claimed source and / or claimed completeness of the sequence data; and (b) verifying the nucleic acid data by (i) processing the sequence data to obtain determined information indicating the determined source and / or determined completeness of the sequence data, and (ii) determining whether the determined information matches the claimed information. Examples of various such verification techniques are described herein and are important examples of quality control techniques developed by the inventors and described herein.

[0071] Adopting such quality control techniques also leads to improvements in sequencing and computer technologies. First, sequencing data that fail one or more quality control checks are not used for some or all of the downstream processing that reduces or eliminates errors in downstream applications (e.g., identifying biomarkers, tumor microenvironment types, potential patient therapies, etc.). Often, such downstream processing requires expensive (often cloud-based) computational processing of large datasets (e.g., sequencing data contains tens of millions of reads, which must be aligned, annotated, and processed in many other ways). Utilizing quality control to prevent computationally expensive processes from being performed reduces or eliminates the wasteful use of computing resources, conserving processing power, memory, and networking resources (this is an improvement in computing technology in addition to sequencing technology). Identifying errors can also reduce resource waste in laboratories that process multiple samples by freeing up equipment to process biological samples that pass initial quality control checks. Additionally, using sequence data for downstream processing that passes various quality control checks may identify more effective therapies for patients, improve the ability to determine whether one or more cancer therapies will be effective when administered to a patient, improve the ability to identify clinical trials in which subjects may participate, and / or identify improvements to many other prognostic, diagnostic, and clinical applications.

[0072] 1A and 1B show an example of a process pipeline for sample preparation and quality control as described herein. The process pipeline of FIG. 1 is illustrative of embodiments of the methods and systems provided in this disclosure and should not be construed as limiting its scope in any way. This disclosure provides that a process pipeline need not include all of the process steps or the order of process steps illustrated in FIG. 1. One or more processes may be omitted, repeated, or performed in a different order depending on the application.

[0073] FIG. 1A shows a non-limiting process pipeline 100 that includes one or more quality control assessments. In activity 101, a biological sample (e.g., a tumor biopsy) is obtained from a subject (e.g., a subject who has, is suspected of having, or is at risk for having cancer). In some embodiments, the sample is obtained from a physician, hospital, clinic, or other healthcare provider. One or more sample quality control assessments in quality control activity 102 may be performed on the biological sample. In some embodiments, the quality control assessment on the biological sample (e.g., biopsy material) includes determining whether the sample is in an appropriate form (e.g., fresh frozen or FFPE) and / or whether it is accompanied by sufficient information to identify the nature and source of the sample. Nucleic acids (e.g., DNA and / or RNA) may then be extracted from biological samples that meet the criteria of sample quality control activity 102. One or more nucleic acid quality control assessments may then be performed in activity 103, which may, for example, evaluate one or more physical attributes of the extracted nucleic acids, a nucleic acid library prepared from the extracted nucleic acids, and / or the pooled nucleic acids or libraries. Nucleic acids (e.g., DNA and / or RNA) that meet the criteria of nucleic acid quality control activity 103 can then be processed (e.g., enriched for polyA RNA) and / or sequenced to obtain raw DNA and / or RNA sequence data (e.g., RNA expression data). In some embodiments, the RNA expression data is processed to obtain gene expression data and, optionally, to remove data for one or more types of genes that may interfere with (e.g., bias) subsequent analysis of the gene expression data. In some embodiments, the gene expression data is normalized (e.g., after removing data for one or more interfering genes).In some embodiments, one or more sequence quality control assessments may be performed on the DNA and / or RNA sequence data (e.g., on the processed, e.g., normalized, gene expression data) for bioinformatics quality control activities 104. In some embodiments, the one or more bioinformatics quality control assessments are performed to determine whether the sequence data is from the expected source (e.g., patient, tissue, tumor, etc.) and / or has sufficient completeness for further analysis. In some embodiments, sequence data that meets the conditions of bioinformatics quality control activities 104 is further processed, e.g., to determine a diagnosis, prognosis, and / or therapy for the subject, to evaluate and / or monitor the subject, and / or for one or more clinical applications (e.g., to evaluate a therapy).

[0074] In some embodiments, the sequence data may include raw DNA or RNA sequence data, DNA exome sequence data (e.g., from whole exome sequencing (WES)), DNA genomic sequence data (e.g., from whole genome sequencing (WGS)), RNA expression data, gene expression data, bias-corrected gene expression data, or any other suitable type of sequence data, including data obtained from a sequencing platform and / or data derived from a sequencing platform including such data, for example, but not limited to, data described herein. FIG. 1B illustrates a non-limiting process pipeline 110 for preparing nucleic acids from a biological sample (e.g., a tumor biopsy) and obtaining and processing nucleic acid sequence data for subsequent analysis (e.g., for diagnostic, prognostic, therapeutic, and / or other clinical applications). Process pipeline 110 executes in activity 111 by obtaining a biological sample (e.g., a tumor sample) from a subject having, suspected of having, or at risk of having cancer. Nucleic acids (e.g., DNA and / or RNA) are obtained (e.g., extracted) from a sample in activity 112. One or more quality control assessments of the nucleic acids are performed in activity 113. One or more nucleic acid libraries are prepared in activity 114, e.g., using nucleic acids that satisfy the criteria of at least one quality control assessment of activity 113. The nucleic acid libraries are sequenced (e.g., to obtain RNA expression data for the RNA) using at least one sequencing platform in sequencing activity 115. In some embodiments, the RNA expression data is converted to gene expression data in activity 116, and the gene expression data is optionally bias-corrected, at least in part, by removing expression data for at least one gene that introduces bias in the gene expression data.One or more bioinformatics quality control assessments are performed on the DNA or RNA sequence data and / or RNA sequence data from activity 115 (e.g., bias-corrected gene expression data from activity 116) in bioinformatics quality control activity 117. In some embodiments, the nucleic acid data (e.g., satisfying the conditions of at least one bioinformatics quality control assessment of activity 117) is further processed in activity 118 (e.g., determining one or more signatures of disease from the gene expression data) and performing a diagnostic, prognostic, therapeutic, and / or other clinical evaluation of the subject (e.g., identifying a treatment, e.g., a cancer treatment, for the subject) in activity 119. In some embodiments, a treatment (e.g., a cancer treatment) is administered to the subject.

[0075] In some embodiments, activity 111 includes obtaining a bulk biopsy tissue from a subject or patient. In some embodiments, activity 111 includes obtaining a blood sample from a subject or patient. In some embodiments, activity 111 includes obtaining a single-cell suspension. In some embodiments, activity 111 includes obtaining any type of sample suitable for preparing nucleic acids for subsequent sequencing analysis. In some embodiments, activity 111 includes obtaining multiple types of samples.

[0076] In some embodiments, when bulk biopsy tissue is obtained, the tissue is processed (e.g., homogenized in the presence of TriZol) to extract nucleic acids, such as DNA or RNA, in activity 112. In some embodiments, when a single-cell suspension is obtained, the suspension is processed to extract nucleic acids, such as DNA or RNA, in activity 112. In some embodiments, nucleic acids suitable for germline whole exome sequencing (WES) may be extracted in activity 112. In some embodiments, nucleic acids suitable for tumor whole exome sequencing (WES) may be extracted in activity 112. In some embodiments, nucleic acids suitable for tumor RNA sequencing may be extracted in activity 112. In some embodiments, nucleic acids suitable for CYTOF (mass cytometry) may be extracted in activity 112. In some embodiments, nucleic acids suitable for any type of sequencing known in the art may be extracted in activity 112.

[0077] In activity 113, one or more quality control assessments may be performed. Acceptable and / or target thresholds may be determined and used as references. In some embodiments, the total amount of extracted DNA or RNA may be used for the quality control assessment. In some embodiments, a spectrophotometer, e.g., a small-volume full-spectrum UV-visible spectrophotometer (e.g., a NanoDrop spectrophotometer available from ThermoFisher Scientific, www.thermofisher.com), may be used for the quality control assessment of DNA or RNA. In some embodiments, a fluorometer (e.g., a Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com), e.g., for DNA or RNA quantification, may be used for the quality control assessment of DNA or RNA. In some embodiments, an automated electrophoresis system (e.g., a TAPESTATION) may be used for the quality control assessment of DNA or RNA. In some embodiments, a real-time PCR system (e.g., a LIGHTCYCLER®) may be used for the quality control assessment of DNA or RNA.

[0078] In some embodiments, activity 114 includes preparing a library for the extracted nucleic acids that met at least one quality control threshold in activity 113. In some embodiments, activity 114 includes one or more methods described in Example 2.

[0079] In some embodiments, activity 115 includes sequencing nucleic acids (e.g., the DNA, RNA, or related library of activity 114) using at least one nucleic acid sequencing platform (e.g., a next-generation nucleic acid sequencing platform) to obtain DNA sequence data and / or RNA sequence data (e.g., RNA expression data). The sequence data obtained in activity 115 may be stored in any suitable format (e.g., in the form of one or more FASTQ files).

[0080] In some embodiments, the RNA expression data is converted to gene expression data at activity 116. In some embodiments, the RNA expression data is aligned to known genes in a database, e.g., a known assembled genome (e.g., the human genome) or a transcriptome in a database. In some embodiments, a program for quantifying transcripts using high-throughput sequencing reads, e.g., from bulk and single-cell RNA-Seq data (e.g., Kallisto (hg38), available from Github, www.github.com, e.g., as described in Nicolas L Bray, Harold Pimentel, Pall Melsted, and Lior Pachter, “Near-optimal probabilistic RNA-seq quantification,” Nature Biotechnology 34, 525-527 (2016), doi:10.1038 / nbt.3519), and / or Gencode (e.g., Gencode V23) is used for sequence alignment and / or annotation. In some embodiments, activity 116 includes gene aggregation. In some embodiments, activity 116 includes removing expression data for one or more non-coding transcripts from the gene expression data. In some embodiments, activity 116 includes removing expression data for one or more genes that may bias the gene expression data. In some embodiments, activity 116 includes removing expression data for histone-encoding genes and / or mitochondrial-encoding genes. In some embodiments, activity 116 includes performing normalization (e.g., TPM normalization) after removing the expression data for non-coding and / or bias-associated genes from the gene expression data. This normalization may be referred to herein as "re-normalization."

[0081] In activity 117, one or more bioinformatics quality control assessments are performed on the nucleic acid sequence data, e.g., DNA sequence data and / or RNA sequence data (e.g., bias-corrected and / or normalized gene expression data). In some embodiments, one or more bioinformatics quality control assessments may be performed to assess the source and / or completeness of the nucleic acid sequence data. In some embodiments, one or more bioinformatics quality control assessments described herein are performed.

[0082] In some embodiments, the method includes all of the processes illustrated in FIG. 1 . However, in some embodiments, a subset of the processes is performed, and any one or more of those processes may be omitted, overlapping, and / or performed in a different order than illustrated in FIG. 1 . In some embodiments, the method includes a process for preparing nucleic acids from a biological sample, optionally including one or more quality control steps, where the nucleic acids are sequenced on at least one sequencing platform. In some embodiments, the method includes processing nucleic acid information obtained (e.g., received) from the sequencing platform to generate DNA or RNA sequence data for subsequent analysis (e.g., generating bias-corrected, optionally normalized gene expression data for subsequent analysis). In some embodiments, one or more of the processes in FIG. 1 are implemented on a computer. In some embodiments, the method includes identifying a treatment (e.g., a cancer treatment) for a subject (e.g., a subject having, suspected of having, or at risk of having cancer). In some embodiments, the method includes administering the treatment to the subject.

[0083] biological samples Any of the methods, systems, or elements described in other claims may use or be used to analyze a biological sample from a subject. In some embodiments, the biological sample is obtained from a subject who has or is suspected of having cancer. One or more biological samples from the subject may be analyzed as described herein to obtain information about the subject's cancer. The biological sample may be any type of biological sample, including, for example, a bodily fluid (e.g., blood, urine, or cerebrospinal fluid), one or more cells (e.g., from a scraping or brushing, such as an oral mucosal specimen or tracheal brushing), a tissue section (e.g., cheek tissue, muscle tissue, lung tissue, heart tissue, brain tissue, or skin tissue), or part or all of an organ (e.g., brain, lung, liver, bladder, kidney, pancreas, intestine, or muscle), or other type of biological sample (e.g., feces or hair).

[0084] In some embodiments, the biological sample is a tumor sample from the subject. In some embodiments, the biological sample is a blood sample from the subject. In some embodiments, the biological sample is a tissue sample from the subject.

[0085] A tumor sample, in some embodiments, refers to a sample comprising cells from a tumor. In some embodiments, a tumor sample comprises cells from a benign tumor, e.g., non-cancerous cells. A tumor sample comprises cells from a pre-malignant tumor, e.g., pre-cancerous cells. In some embodiments, a tumor sample comprises cells from a malignant tumor, e.g., cancerous cells.

[0086] Examples of tumors include, but are not limited to, adenoma, fibroma, hemangioma, lipoma, cervical dysplasia, pulmonary metaplasia, leukoplakia, carcinoma, sarcoma, germ cell tumor, and blastoma.

[0087] A blood sample, in some embodiments, refers to a sample containing cells, e.g., cells from a blood sample. In some embodiments, the blood sample contains non-cancerous cells. In some embodiments, the blood sample contains pre-cancerous cells. In some embodiments, the blood sample contains cancerous cells. In some embodiments, the blood sample contains blood cells. In some embodiments, the blood sample contains red blood cells. In some embodiments, the blood sample contains white blood cells. In some embodiments, the blood sample contains platelets. Examples of cancerous blood cells include, but are not limited to, leukemia, lymphoma, and myeloma. In some embodiments, the blood sample is taken to obtain cell-free nucleic acid (e.g., cell-free DNA) in the blood.

[0088] The blood sample may be a whole blood sample or a fractionated blood sample. In some embodiments, the blood sample comprises whole blood. In some embodiments, the blood sample comprises fractionated blood. In some embodiments, the blood sample comprises a buffy coat. In some embodiments, the blood sample comprises serum. In some embodiments, the blood sample comprises plasma. In some embodiments, the blood sample comprises a blood clot.

[0089] A tissue sample, in some embodiments, refers to a sample comprising cells from the tissue. In some embodiments, a tumor sample comprises non-cancerous cells from the tissue. In some embodiments, a tumor sample comprises pre-cancerous cells from the tissue. In some embodiments, a tumor sample comprises pre-cancerous cells from the tissue.

[0090] The disclosed methods encompass a variety of tissues, including organ or non-organ tissues, including, but not limited to, muscle tissue, brain tissue, lung tissue, liver tissue, epithelial tissue, connective tissue, and nerve tissue. In some embodiments, the tissue may be normal, diseased, or suspected diseased. In some embodiments, the tissue may be a tissue section or whole, intact tissue. In some embodiments, the tissue may be animal or human tissue. Animal tissue includes, but is not limited to, tissue obtained from rodents (e.g., rats or mice), primates (e.g., monkeys), dogs, cats, and livestock.

[0091] The biological sample may be from any source within the subject's body, including, but not limited to, any bodily fluid (such as blood (e.g., whole blood, serum, or plasma), saliva, tears, synovial fluid, cerebrospinal fluid, pleural fluid, pericardial fluid, peritoneal fluid, and / or urine), hair, skin (including portions of the epidermis, dermis, and / or hypodermis), oropharynx, laryngopharynx, esophagus, stomach, bronchi, salivary glands, tongue, oral cavity, nasal cavity, vaginal cavity, anal cavity, bone, bone marrow, brain, thymus, spleen, small intestine, appendix, colon, rectum, anus, liver, biliary tract, pancreas, kidney, ureter, bladder, urethra, uterus, vagina, vulva, ovaries, cervix, scrotum, penis, prostate, testicles, seminal vesicles, and / or any type of tissue (e.g., muscle tissue, epithelial tissue, connective tissue, or nervous tissue).

[0092] Any biological sample described herein can be obtained from a subject using any known technique. For example, regarding the collection, processing, and storage of biological samples, see the publications "Biospecimens and biorepositories: from afterthought to science" by Vaught et al. (Cancer Epidemiol Biomarkers Prev. 2012 February, 21(2):253-5) and "Biological sample collection, processing, storage and information management" by Vaught and Henderson (IARC Sci Publ. 2011, (163):23-42), each of which is incorporated herein in its entirety.

[0093] In some embodiments, the biological sample may be obtained from surgery (e.g., laparoscopic, microsurgical, or endoscopic), bone marrow biopsy, punch biopsy, endoscopic biopsy, or needle biopsy (e.g., fine needle aspiration, core needle biopsy, vacuum-assisted biopsy, or image-guided biopsy). In some embodiments, the biological sample may be obtained from an autopsy.

[0094] In some embodiments, one or more cells (i.e., a cellular biological sample) may be obtained from a subject using a scraping or brushing method. A cellular biological sample may be obtained from any area in or from the subject's body, including, for example, one or more areas of the cervix, esophagus, stomach, bronchi, or oral cavity. In some embodiments, one or more tissue pieces (e.g., tissue biopsies) from a subject may be used. In some embodiments, a tissue biopsy may include one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) biological samples from one or more tumors or tissues known to have or suspected of having cancerous cells.

[0095] Any biological sample from a subject described herein can be stored using any method that maintains the stability of the biological sample. In some embodiments, maintaining the stability of a biological sample means preventing the components of the biological sample (e.g., DNA, RNA, protein, or tissue structure or morphology) from deteriorating until measured so that the measurements, when measured, represent the state of the sample when it was obtained from the subject. In some embodiments, the biological sample is stored in a composition that can permeate it and protect the components of the biological sample (e.g., DNA, RNA, protein, or tissue structure or morphology) from deteriorating. As used herein, degradation is the conversion of one component to another such that the original form is no longer detectable at the same level as before degradation.

[0096] In some embodiments, biological samples are stored using cryopreservation. Non-limiting examples of cryopreservation include, but are not limited to, step-down freezing, rapid freezing, direct plunge freezing, snap freezing, slow freezing using a programmable freezer, and vitrification. In some embodiments, biological samples are stored using freeze-drying. In some embodiments, biological samples are placed in a container already containing a preservative (e.g., RNALater for preserving RNA) after collection from a subject and then frozen (e.g., by snap freezing). In some embodiments, such storage in a frozen state occurs immediately after collection of the biological sample. In some embodiments, the biological sample may be kept in a preservative or in a preservative-free buffer at either room temperature or 4°C for a period of time (e.g., up to 1 hour, up to 8 hours, or up to 1 day, or several days) before being frozen.

[0097] Non-limiting examples of preservatives include formalin solution, formaldehyde solution, RNALater or other equivalent solutions, TriZol or other equivalent solutions, DNA / RNA Shield or equivalent solutions, EDTA (e.g., Buffer AE (10 mM Tris-Cl, 0.5 mM EDTA, pH 9.0)) and other coagulants, and Acids Citrate Dextrons (e.g., for blood specimens).

[0098] In some embodiments, specialized containers may be used to collect and / or store biological samples. For example, vacutainers may be used to store blood. In some embodiments, vacutainers may contain preservatives (e.g., coagulants or anticoagulants). In some embodiments, the container in which the biological sample is stored may be housed in a secondary container for better preservation or to avoid contamination.

[0099] Any biological sample from a subject described herein may be stored under any conditions that preserve the stability of the biological sample. In some embodiments, the biological sample is stored at a temperature that preserves the stability of the biological sample. In some embodiments, the sample is stored at room temperature (e.g., 25°C). In some embodiments, the sample is stored under refrigeration (e.g., 4°C). In some embodiments, the sample is stored under frozen conditions (e.g., -20°C). In some embodiments, the sample is stored under ultra-low temperature conditions (e.g., -50°C to -800°C). In some embodiments, the sample is stored under liquid nitrogen (e.g., -1700°C). In some embodiments, biological samples are stored at -60°C to -8°C (e.g., -70°C) for up to 5 years (e.g., up to 1 month, up to 2 months, up to 3 months, up to 4 months, up to 5 months, up to 6 months, up to 7 months, up to 8 months, up to 9 months, up to 10 months, up to 11 months, up to 1 year, up to 2 years, up to 3 years, up to 4 years, or up to 5 years). In some embodiments, biological samples are stored for up to 20 years (e.g., up to 5 years, up to 10 years, up to 15 years, or up to 20 years) as described by any of the methods described herein.

[0100] The methods of the present disclosure involve obtaining one or more biological samples from a subject for analysis. In some embodiments, one biological sample is taken from a subject for analysis. In some embodiments, multiple (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more) biological samples are taken from a subject for analysis. In some embodiments, one biological sample is analyzed from a subject. In some embodiments, multiple (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more) biological samples are analyzed. When multiple biological samples from a subject are analyzed, the biological samples may be procured simultaneously (e.g., multiple biological samples may be taken at the same procedure), or the biological samples may be taken at different time points (e.g., at different procedures, including a procedure 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 days after the initial procedure; a procedure 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 weeks after the initial procedure; a procedure 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 months after the initial procedure; a procedure 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 months after the initial procedure; or a procedure 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 years after the initial procedure).

[0101] The second or subsequent biological sample may be taken or obtained from the same region (e.g., from the same tumor or tissue region) or from a different region (e.g., including a different tumor). The second or subsequent biological sample may be taken or obtained from the subject after one or more treatments, and may be taken from the same region or a different region. As a non-limiting example, the second or subsequent biological sample may be useful in determining whether the cancer in each biological sample has different characteristics (e.g., in the case of biological samples taken from two physically separate tumors in a patient's body) or whether the cancer has responded to one or more treatments (e.g., in the case of two or more biological samples taken from the same tumor or different tumors before and after treatment). In some embodiments, each of the at least one biological sample is a bodily fluid sample, a cell sample, or a tissue biopsy sample.

[0102] In some embodiments, one or more biological specimens are combined (e.g., placed in the same container for storage) before further processing. For example, a first sample of a first tumor obtained from a subject may be combined with a second sample of a second tumor obtained from the subject, where the first and second tumors may or may not be the same tumor. In some embodiments, the first tumor and the second tumor are similar but not the same (e.g., two tumors in the brain of a subject). In some embodiments, the first biological sample and the second biological sample from a subject are samples of different types of tumors (e.g., a tumor in muscle tissue and a tumor in brain tissue).

[0103] In some embodiments, the sample (e.g., a tumor sample or a blood sample) from which RNA and / or DNA is extracted is sufficiently large that at least 2 μg (e.g., at least 2 μg, at least 2.5 μg, at least 3 μg, at least 3.5 μg, or more) of RNA can be extracted therefrom. In some embodiments, the sample from which RNA and / or DNA is extracted may be peripheral blood mononuclear cells (PBMCs). In some embodiments, the sample from which RNA and / or DNA is extracted may be any type of cell suspension. In some embodiments, the sample (e.g., a tumor sample or a blood sample) from which RNA and / or DNA is extracted is sufficiently large that at least 1.8 μg of RNA can be extracted therefrom. In some embodiments, at least 50 mg (e.g., at least 1 mg, at least 2 mg, at least 3 mg, at least 4 mg, at least 5 mg, at least 10 mg, at least 12 mg, at least 15 mg, at least 18 mg, at least 20 mg, at least 22 mg, at least 25 mg, at least 30 mg, at least 35 mg, at least 40 mg, at least 45 mg, or at least 50 mg) of tissue sample is collected and RNA and / or DNA is extracted therefrom. In some embodiments, at least 20 mg of tissue sample is collected and RNA and / or DNA is extracted therefrom. In some embodiments, at least 30 mg of tissue sample is collected. In some embodiments, at least 10-50 mg (e.g., 10-50 mg, 10-15 mg, 10-30 mg, 10-40 mg, 20-30 mg, 20-40 mg, 20-50 mg, or 30-50 mg) of tissue sample is collected and RNA and / or DNA is extracted therefrom. In some embodiments, at least 30 mg of tissue sample is collected, hi some embodiments, at least 20-30 mg of tissue sample is collected and RNA and / or DNA is extracted therefrom.In some embodiments, the sample (e.g., a tumor sample or a blood sample) from which RNA and / or DNA is extracted is large enough that at least 0.2 μg (e.g., at least 200 ng, at least 300 ng, at least 400 ng, at least 500 ng, at least 600 ng, at least 700 ng, at least 800 ng, at least 900 ng, at least 1 μg, at least 1.1 μg, at least 1.2 μg, at least 1.3 μg, at least 1.4 μg, at least 1.5 μg, at least 1.6 μg, at least 1.7 μg, at least 1.8 μg, at least 1.9 μg, or at least 2 μg) of RNA can be extracted therefrom. In some embodiments, the sample (e.g., a tumor sample or a blood sample) from which RNA and / or DNA is extracted is large enough such that at least 0.1 μg (e.g., at least 100 ng, at least 200 ng, at least 300 ng, at least 400 ng, at least 500 ng, at least 600 ng, at least 700 ng, at least 800 ng, at least 900 ng, at least 1 μg, at least 1.1 μg, at least 1.2 μg, at least 1.3 μg, at least 1.4 μg, at least 1.5 μg, at least 1.6 μg, at least 1.7 μg, at least 1.8 μg, at least 1.9 μg, or at least 2 μg) of RNA can be extracted.

[0104] subject Aspects of the present disclosure relate to biological samples obtained from a subject. In some embodiments, the subject is a mammal (e.g., a human, mouse, cat, dog, horse, hamster, cow, pig, or other domestic animal). In some embodiments, the subject is a human. In some embodiments, the subject is an adult (e.g., 18 years of age or older). In some embodiments, the subject is a child (e.g., under 18 years of age). In some embodiments, the human subject is a person who has at least one form of cancer or has been diagnosed with at least one form of cancer. In some embodiments, the cancer the subject is suffering from is carcinoma, sarcoma, myeloma, leukemia, lymphoma, or a mixed type cancer comprising more than one of carcinoma, sarcoma, myeloma, leukemia, and lymphoma. Carcinoma refers to a malignant neoplasm of epithelial origin or a cancer of the lining or adventitia of the body. Sarcoma refers to a cancer originating in supportive and connective tissues such as bone, tendon, cartilage, muscle, and fat. Myeloma is a cancer originating in the plasma cells of the bone marrow. Leukemia ("liquid cancer" or "blood cancer") is cancer of the bone marrow (site of blood cell production). Lymphoma develops in the glands or nodes of the lymphatic system, a network of blood vessels, nodes, and organs (particularly the spleen, tonsils, and thymus) that produce white blood cells, or lymphocytes, which cleanse the body's fluids and fight infection. Non-limiting examples of mixed-type cancers include adenosquamous carcinoma, mixed mesodermal tumor, carcinosarcoma, and teratocarcinoma. In some embodiments, the subject has a tumor. The tumor can be benign or malignant. In some embodiments, the cancer is any one of skin cancer, lung cancer, breast cancer, prostate cancer, colon cancer, rectal cancer, cervical cancer, and uterine cancer. In some embodiments, the subject is at risk of developing cancer because, for example, the subject has one or more genetic risk factors or has been or is being exposed to one or more carcinogens (e.g., tobacco smoke or chewing tobacco).

[0105] Single-cell suspension In some embodiments, methods for characterizing cancers a subject has or is suspected of having (e.g., RNA sequencing, DNA sequencing, or multiplexed flow cytometry) are performed at the single-cell level to capture the heterogeneity of a single tumor or cancer tissue, or multiple tumors or cancer tissues. That is, measurement and evaluation of single cells in a tumor sample provides information that is not confounded by the genotypic or phenotypic heterogeneity of the bulk sample. In some embodiments, single-cell suspensions are prepared from one or more biological samples obtained from a subject for use in methods such as single-cell RNA or DNA sequencing or mass cytometry.

[0106] Thus, some embodiments of any one of the methods described herein include forming a single-cell suspension of cells from a sample of the tumor (e.g., a first sample of the tumor). In some embodiments, forming a single-cell suspension of cells from the sample of the tumor includes dissecting the tumor sample to obtain tumor sample fragments. Curved scissors may be used to dissect the tumor tissue sample. In some embodiments, the tumor sample fragments are 0.5-3 mm. 3 (For example, 1 to 2 mm 3 In some embodiments, the tumor tissue sample or fragment thereof is kept moist during dissection.

[0107] Methods for preparing single-cell suspensions from tumor samples can include any one or more of the following steps, in any order: mincing, enzymatic and / or non-enzymatic digestion, vigorous pipetting, cell straining, washing, and counting. In some embodiments, one or more of these steps are repeated (e.g., 1, 2, 3, 4, or 5 or more times).

[0108] In some embodiments, the tumor sample or tumor sample fragment is incubated with an enzyme cocktail. Any number and combination of enzymes can be used, see, e.g., BioFiles: For Life Science Research, Issue 2, 2006, www.sigmaaldrich.com / content / dam / sigma-aldrich / docs / Sigma / General_Information / 2 / biofiles_issue2.pdf, which is incorporated herein by reference in its entirety, and which specifically incorporates herein any of the enzymes or other components (e.g., media) described therein.

[0109] Quatromoni et al., "An optimized disaggregation method for human lung tumors that preserves the phenotype and function of the immune cell," J Leukoc Biol. 2015 Jan, 97(1):201-209, provides a comparison of different enzyme cocktails and is incorporated herein by reference in its entirety. In some embodiments, the enzyme cocktail includes as components one or more of the following: medium (e.g., L-15 medium), antibacterial agent (e.g., penicillin and / or streptomycin), antifungal agent (e.g., amphotericin), collagenase (e.g., collagenase I, collagenase II, collagenase IV), DNAse (e.g., DNAse I), elastase, hyaluronidase, protease (e.g., protease XIV, trypsin, papain, thermolysin). Coll I has an inherent balance of collagenase, caseinase, clostripain, and trypsin activity, Coll II contains higher relative levels of protease activity, particularly clostripain, and Coll IV is designed to have particularly low tryptic activity (Quatromoni et al., J Leukoc Biol. 2015 Jan, 97(1):201-209). In some embodiments, only collagenase I, collagenase II, or collagenase IV is used. In some embodiments, a mixture of two collagenases is used (e.g., collagenase I and collagenase II, collagenase I and collagenase IV, or collagenase II and collagenase IV). In some embodiments, more than two collagenases are used (e.g., collagenase I, collagenase II, and collagenase IV).

[0110] In some embodiments, the enzyme cocktail includes one or more of the following components: medium (e.g., complete medium), penicillin, streptomycin, and collagenase (e.g., collagenase I or collagenase IV). The concentration of the enzymes in the cocktail can be adjusted. A non-limiting example of an enzyme cocktail is collagenase I (0.2 mg / ml), collagenase IV (1 mg / ml), complete medium, penicillin (0.001%), and DNAse.

[0111] In some embodiments, at least 25 ml (e.g., at least 25 ml, at least 26 ml, at least 27 ml, at least 28 ml, at least 29 ml, or at least 30 ml) of enzyme cocktail is added per 0.5 gm of tumor tissue. In some embodiments, a sample of tumor or a fragment thereof is incubated in the enzyme cocktail while the sample is shaken or agitated (e.g., spun at 85 RPM and / or with vigorous pipetting). In some embodiments, a sample of tumor or a fragment thereof is incubated in the enzyme cocktail at a temperature between 20-50°C (e.g., 20-50°C, 20-25°C, 25-30°C, 25-35°C, 30-40°C, 35-45°C, 40-50°C, or 30-50°C). In some embodiments, a method for preparing a single-cell suspension includes filtering the enzyme cocktail, for example, through a cell strainer (e.g., 50 μm, 70 μm, or 100 μm). In some embodiments, a filter that is too fine may result in a cell composition with a high concentration of fibroblasts. In some embodiments, a filter that is too coarse may result in cell clumps. In some embodiments, cell clumps are broken up using mechanical force (e.g., vigorous pipetting, application of pressure using a syringe).

[0112] In some embodiments, the filtered cells are lysed using an RBC lysis buffer to lyse the red blood cells. RBC lysis buffer is commercially available (see, e.g., www.abcam.com / red-blood-cell-rbc-lysis-buffer-ab204733.html).

[0113] In some embodiments, methods for preparing single-cell suspensions include enzymatic and mechanical disaggregation. Examples of methods for disaggregating cells from tissues are described in the following publications: Quatromoni et al., "An optimized disaggregation method for human lung tumors that preserves the phenotype and function of the immune cell," J Leukoc Biol. January 2015, 97(1):201-209; Pennartz et al., "Generation of Single-Cell Suspensions from Mouse Neural Tissue," JOVE Issue 29, doi:10.3791 / 1267, Published: 7 / 07 / 2009; and www.youtube.com / watch?v=N0jftyYqM38.

[0114] In some embodiments, enzyme-free cell dissociation buffer is used.See, for example, ThermoFisher Scientific catalog numbers 13151014 and 13150016, or Millipore Sigma Aldrich catalog number S-014-B.Heng et al., Biol Proced Online. 2009, 11:161-169, presents a comparison between enzymatic and non-enzymatic means of cell dissociation, and is incorporated herein by reference in its entirety.

[0115] In some embodiments, the number of cells in the single cell suspension is counted and their viability is tested. The following example provides an example of the overall process for forming a single cell suspension from a sample of tumor tissue.

[0116] In some embodiments, the method includes forming a single-cell suspension of cells from a tumor sample and dividing it into at least first and second portions. The first and second portions of the single-cell suspension can be of equal size or different sizes (e.g., containing different numbers of cells). In some embodiments, all portions of the single-cell suspension (e.g., the first portion, the second portion, etc.) are stored in separate containers and stored under the same or similar conditions (e.g., in liquid nitrogen or at -80°C). In some embodiments, different portions of the single-cell suspension are stored under different conditions before or after any further processing (e.g., labeling with antibodies for protein expression studies). In some embodiments, cells isolated from a biological sample are cultured and expanded before storage. In some embodiments, cells isolated from a biological sample are cultured and expanded after storage.

[0117] In some embodiments, any one of the methods described herein further comprises forming a lysate from at least a portion (e.g., a first or second portion) of the single cell suspension. In some embodiments, different portions of the single cell suspension comprise different types of cells. In some embodiments, the portion of the single cell suspension from which the lysate is formed comprises at least 1 x 10 6 cells (e.g., at least 1 × 10 6 cells, at least 2 x 10 6 cells, at least 3 x 10 6 cells, at least 4 x 10 6 cells, or at least 5 x 10 6 In some embodiments, the portion of the single cell suspension from which the lysate is formed contains at least 2 x 10 6The single cell suspension comprises 100 cells. The lysate may be stored in a storage medium that prevents degradation of DNA and / or RNA (e.g., RNALater). In some embodiments, the method includes extracting RNA from the lysate from the single cell suspension or portions of the single cell suspension and performing RNA sequencing on the extracted RNA to obtain RNA expression data. These RNA expression data can be used to determine tumor heterogeneity.

[0118] An overview of single-cell RNA sequencing is provided at hemberg-lab.github.io / scRNA.seq.course / introduction-to-single-cell-rna-seq.html, Figure 2.1 of which is incorporated herein by reference. In some embodiments, methods for performing RNA sequencing of single-cell suspensions include isolating single-cell RNA, pre-amplifying reverse-transcribed cDNA, preparing a cDNA library (e.g., using the Fluidigm C1 protocol), and sequencing the resulting library using a platform such as the Illumina HiSeq 2500.

[0119] Methods for performing single-cell RNA sequencing are described in Bagnoli et al., "Studying Cancer Heterogeneity by Single-Cell RNA Sequencing," Methods Mol Biol. 2019;1956:305-319; Sun et al., "Single-cell RNA sequencing reveals gene expression signatures of breast cancer-associated endothelial cells," Oncotarget. 2018;2;16;9(13):10945-10961; Kulkarni et al., "Beyond bulk: a review of single cell transcriptomics methodologies and applications," Curr Opin Biotechnol. 2019;4;9;58:129-136; and Huang et al., "High Throughput Single Cell RNA Sequencing, Bioinformatics Analysis and Applications," Adv Exp Med Biol., each of which is incorporated herein by reference in its entirety. 2018, pp. 1068:33-43, Zilionis et al. "Single-Cell Transcriptomics of Human and Mouse Lung Cancers Reveals Conserved Myeloid Populations across Individuals and Species", Immunity. April 5, 2019, pii:S1074-7613(19)30126-8), Kashima et al. "Single-Cell Sequencing Analysis", Adv Exp Med Biol. 2019, pp. 1129:81-96, doi:10.1007 / 978-981-13-6037-4_6, Seki et al. "An Informative Approach to Single-Cell Sequencing Analysis", Adv Exp Med Biol.2019, 1129:81-96, "Single-Cell DNA-Seq and RNA-Seq in Cancer Using the C1 System," Adv Exp Med Biol. 2019, 1129:27-50, doi:10.1007 / 978-981-13-6037-4_3), and See et al., "A Single-Cell Sequencing Guide for Immunologists," Front Immunol. 2018, 9:2425.

[0120] Gan et al., "Identification of cancer subtypes from single-cell RNA-seq data using a consensus clustering method," BMC Med Genomics. 2018, 11(Suppl 6):117, describes a clustering method for single-cell RNA sequencing data, which is incorporated herein by reference in its entirety.

[0121] In some embodiments, the single-cell RNA sequencing method used is any one of Fluidigm C1 System (SMART-seq), Fluidigm C1 System (mRNA Seq HT), SMART-seq2, 10X Genomics Chromium System, and MARS-seq. See et al., Front Immunol. 2018, 9:2425, provides a comparison of these methods and is incorporated herein by reference in its entirety.

[0122] In some embodiments, any one of the methods described herein further comprises performing measurements on a single cell suspension. In some embodiments, different measurements are performed in parallel on the same cells. Macaulay et al., Trends Genet. 2017 Feb;33(2):155-168, describes a method for performing multiple measurements from a single cell, and is incorporated herein by reference in its entirety.

[0123] In some embodiments, any one of the methods described herein further comprises performing mass cytometry on at least a first portion of the single-cell suspension. Mass cytometry is a mass spectrometry technique based on inductively coupled plasma mass spectrometry and time-of-flight mass spectrometry used to characterize cells. In some embodiments, mass cytometry involves conjugating an antibody to an isotopically pure element and then using it to label cellular molecules (e.g., proteins). In some embodiments, the cells are nebulized and passed through an argon plasma to ionize the metal antibody. The metal signal is then analyzed by a time-of-flight mass spectrometer to identify and quantify cellular molecules within the cells. In some embodiments, the single-cell suspension or portion thereof on which mass cytometry is performed contains at least 1×10 6 cells (e.g., at least 1 × 10 6 cells, at least 2 x 10 6 cells, at least 3 x 10 6 cells, at least 4 x 10 6 cells, at least 5 x 10 6 cells, at least 6 x 10 6 cells, at least 7 x 10 6 cells, at least 8 x 10 6 cells, at least 9 x 10 6 cells, or at least 10 x 10 6 In some embodiments, the single cell suspension or portion thereof on which mass cytometry is performed contains at least 5×10 6 Contains cells.

[0124] Methods for performing mass cytometry are described in Galli et al., "The end of omics? High dimensional single cell analysis in precision medicine," Eur J Immunol. February 2019, 49(2):212-220; Brodin, "The biology of the cell - insights from mass cytometry," FEBS J. November 3, 2018, doi:10.1111 / febs.14693; Olsen et al., "The anatomy of single cell mass cytometry data," Cytometry A. February 2019, 95(2):156-172; and Behbehani, "Applications of Mass Cytometry in Clinical Medicine: The Promise and Perils of Clinical CyTOF," Clin Lab Med., each of which is incorporated herein by reference in its entirety. 2017; 37(4):945-964; Gondhalekar et al., "Alternatives to current flow cytometry data analysis for clinical and research studies," Methods. 2018; 134-135:113-129; and Soares et al., "Go with the flow: advances and trends in magnetic flow cytometry," Anal Bioanal Chem. 2019; 411(9):1839-1862, doi: 10.1007 / s00216-019-01593-9. Epub 2019; 19 February 2019.

[0125] Other assays Any of the biological samples described herein can be used to obtain expression data using conventional assays or those described herein. In some embodiments, the expression data includes gene expression levels. Gene expression levels can be detected by detecting the products of gene expression, such as mRNA and / or protein.

[0126] In some embodiments, the gene expression level is determined by detecting the level of a protein in a sample and / or by detecting the activity level of a protein in a sample. As used herein, the phrases "determining" or "detecting" can include assessing the presence, absence, quantity and / or amount (which may be an effective amount) of a substance in a sample, including deriving a qualitative or quantitative concentration level of such substance, or otherwise assessing the value and / or classification of such substance in a sample from a subject.

[0127] The level of the protein can be measured using an immunoassay. Examples of immunoassays include any known assay (but are not limited to), including immunoblot analysis (e.g., Western blot), immunohistochemistry, flow cytometry, immunofluorescence analysis (IF), enzyme-linked immunosorbent assay (ELISA) (e.g., sandwich ELISA), radioimmunoassay, electrochemiluminescence-based detection assay, magnetic immunoassay, lateral flow assay, and any related techniques. Additional suitable immunoassays for detecting the level of the proteins provided herein will be apparent to those skilled in the art.

[0128] Such immunoassays may involve the use of an agent (e.g., an antibody) specific for a target protein. The phrase "specifically binds" an agent, such as an antibody, to a target protein is well understood in the art, and methods for determining such specific binding are also well known in the art. An antibody is said to exhibit "specific binding" if it reacts with or binds to a particular target protein more frequently, more rapidly, for a longer period of time, and / or with higher affinity than alternative proteins. It is also understood by reading this definition that, for example, an antibody that specifically binds to a first target peptide may or may not specifically or selectively bind to a second target peptide. As such, "specific binding" or "selective binding" does not necessarily require (although it may include) exclusive binding. Generally, but not necessarily, reference to binding implies selective binding. In some instances, an antibody that "specifically binds" to a target peptide or epitope thereof may not bind to other peptides or other epitopes in the same antigen. In some embodiments, a sample can be contacted with multiple binding agents that bind different proteins simultaneously or sequentially (eg, multiplex analysis).

[0129] As used herein, the term "antibody" refers to a protein that contains at least one immunoglobulin variable domain or immunoglobulin variable domain sequence. For example, an antibody can contain a heavy (H) chain variable region (abbreviated herein as VH) and a light (L) chain variable region (abbreviated herein as VL). In another example, an antibody contains two heavy (H) chain variable regions and two light (L) chain variable regions. The term "antibody" encompasses antigen-binding fragments of antibodies (e.g., single-chain antibodies, Fab and sFab fragments, F(ab')2, Fd fragments, Fv fragments, scFv, and domain antibody (dAb) fragments (de Wildt et al., Eur J Immunol. 1996, 26(3):629-39), as well as complete antibodies. Antibodies can have structural characteristics of IgA, IgG, IgE, IgD, IgM (and their subtypes). Antibodies can be from any source, including, but not limited to, primates (human and non-human primates) and primatized (such as humanized) antibodies.

[0130] In some embodiments, an antibody as described herein can be conjugated to a detectable label, and binding of the detection reagent to the peptide of interest can be determined based on the intensity of the signal emitted from the detectable label. Alternatively, a secondary antibody specific to the detection reagent can be used. One or more antibodies can be conjugated to a detectable label. Any suitable label known in the art can be used in the assay methods described herein. In some embodiments, the detectable label comprises a fluorophore. As used herein, the term "fluorophore" (also called "fluorescent label" or "fluorescent dye") refers to a moiety that absorbs light energy at a defined excitation wavelength and emits light energy at a different wavelength. In some embodiments, the detection moiety is or comprises an enzyme. In some embodiments, the enzyme produces a colored product from a colorless substrate (e.g., β-galactosidase).

[0131] It will be clear to those skilled in the art that the present disclosure is not limited to immunoassay.Detection assays that are not based on antibody, such as mass spectrometry, are also useful for detecting and / or quantifying the level of protein and / or protein as provided herein.The assays that rely on chromogenic substrates are also useful for detecting and / or quantifying the level of protein and / or protein as provided herein.

[0132] Alternatively, the level of nucleic acid encoding a gene in a sample can be measured by conventional methods. In some embodiments, measuring the expression level of nucleic acid encoding a gene includes measuring mRNA. In some embodiments, the expression level of mRNA encoding a gene can be measured using real-time reverse transcriptase (RT) Q-PCR or nucleic acid microarray. Methods for detecting nucleic acid sequences include, but are not limited to, polymerase chain reaction (PCR), reverse transcriptase PCR (RT-PCR), in situ PCR, quantitative PCR (Q-PCR), real-time quantitative PCR (RT Q-PCR), in situ hybridization, Southern blot, Northern blot, sequence analysis, microarray analysis, reporter gene detection, or other DNA / RNA hybridization platforms.

[0133] In some embodiments, the level of nucleic acid encoding a gene in a sample can be measured via a hybridization assay. In some embodiments, the hybridization assay comprises at least one binding partner. In some embodiments, the hybridization assay comprises at least one oligonucleotide binding partner. In some embodiments, the hybridization assay comprises at least one labeled oligonucleotide binding partner. In some embodiments, the hybridization assay comprises at least one pair of oligonucleotide binding partners. In some embodiments, the hybridization assay comprises at least one pair of labeled oligonucleotide binding partners.

[0134] Any binding agent that specifically binds to a desired nucleic acid or protein can be used in the methods and kits described herein to measure expression levels in a sample. In some embodiments, the binding agent is an antibody or aptamer that specifically binds to a desired protein. In other embodiments, the binding agent can be one or more oligonucleotides that are complementary to a nucleic acid or a portion thereof. In some embodiments, a sample can be contacted with multiple binding agents that bind different proteins or different nucleic acids, simultaneously or sequentially (e.g., multiplex analysis).

[0135] To measure the expression level of a protein or nucleic acid, the sample may be contacted with a binder under suitable conditions. Generally, the term "contacting" refers to exposing the binder to the sample or cells taken therefrom for a suitable period of time sufficient to allow complexes to form between the binder and the target protein or nucleic acid, if any, in the sample. In some embodiments, the contacting is performed by capillary action, in which the sample is moved across the surface of a support membrane.

[0136] In some embodiments, the assay may be performed on a low-throughput platform, including a single-assay format. In some embodiments, the assay may be performed on a high-throughput platform. Such high-throughput assays may include using binding agents immobilized on a solid support (e.g., one or more chips). Methods for immobilizing the binding agent depend on factors such as the nature of the binding agent and the material of the solid support, and may require specific buffers. Such methods will be apparent to those skilled in the art.

[0137] DNA and / or RNA extraction In any one of the methods described herein, RNA is extracted from the biological sample to prevent degradation of the RNA and / or to prevent enzyme inhibition in downstream processing, e.g., preparation of DNA (i.e., a cDNA library from the RNA). In any one of the methods described herein, DNA is extracted from the biological sample to prevent degradation of the DNA and / or to prevent enzyme inhibition in downstream processing, e.g., preparation of DNA. In some embodiments, the term "extraction" in the context of obtaining DNA or RNA from a biological sample is used interchangeably with the term "isolation."

[0138] The methods described herein involve the extraction of RNA and / or DNA from a biological sample (e.g., a tumor sample or a blood sample). As described above, a biological sample can be composed of multiple samples from one or more tissues (e.g., one or more different tumors). In some embodiments, RNA and / or DNA is extracted from combined samples. In some embodiments, RNA and / or DNA is extracted from multiple biological samples from a subject and then combined before further processing (e.g., storage or DNA library preparation). In some embodiments, multiple samples of extracted RNA and / or DNA are combined with each other after retrieval from storage. In some embodiments, at least tumor DNA is extracted from one or more tumor tissues. In some embodiments, at least tumor RNA is extracted from one or more tumor tissues. In some embodiments, at least normal DNA is extracted from one or more normal tissues and used as a control. In some embodiments, at least normal RNA is extracted from one or more normal tissues and used as a control. DNA / RNA extraction protocols are described at least in Example 2.

[0139] Methods for extracting DNA and / or RNA from biological samples are known in the art, and reagents and kits for this purpose are commercially available. Gomez-Akata et al., "Methods for extracting 'omes from microbialites," J. Microbiol. Methods. March 12, 2019, 160:1-10, describes extraction methods applicable to DNA and RNA extraction from microorganisms, as well as their advantages and disadvantages, and is incorporated herein by reference in its entirety. The methods described in Gomez-Akata et al. are generally applicable to RNA and / or DNA extracted from tissues. Moore, Curr. Protoc. Immunol. May 2001, Chapter 10: Unit 10.1, describes the purification and concentration of DNA from aqueous solutions, and is incorporated herein by reference in its entirety.

[0140] In some embodiments, extracting DNA and / or RNA involves lysing cells of a biological sample and isolating the DNA and / or RNA from other cellular components. Examples of methods for lysing cells include, but are not limited to, mechanical lysis, liquid homogenization, sonication, freeze-thaw, chemical lysis, alkaline lysis, and manual grinding.

[0141] Methods for extracting DNA and / or RNA include, but are not limited to, solution-phase extraction and solid-phase extraction. In some embodiments, solution-phase extraction comprises organic extraction, such as phenol-chloroform extraction. In some embodiments, solution-phase extraction comprises high-salt extraction, such as guanidinium thiocyanate (GuTC) or guanidinium chloride (GuCl) extraction. In some embodiments, solution-phase extraction comprises ethanol precipitation. In some embodiments, solution-phase extraction comprises isopropanol precipitation. In some embodiments, solution-phase extraction comprises ethidium bromide (EtBr)-cesium chloride (CsCl) gradient centrifugation. In some embodiments, extracting DNA and / or RNA comprises non-ionic detergent extraction, such as cetyltrimethylammonium bromide (CTAB) extraction.

[0142] In some embodiments, extracting DNA and / or RNA comprises solid-phase extraction.Any solid phase that binds DNA and / or RNA can be used to extract DNA and / or RNA in the methods and systems described herein.Examples of solid phases that bind DNA and / or RNA include, but are not limited to, silica matrices, ion-exchange matrices, glass particles, magnetizable cellulose beads, polyamide matrices, and nitrocellulose membranes.

[0143] In some embodiments, the solid phase extraction method comprises a spin column-based extraction method, hi some embodiments, the solid phase extraction method comprises a bead-based extraction method, hi some embodiments, the solid phase extraction method comprises a cation exchange resin, e.g., a styrene divinyl benzene copolymer resin.

[0144] The systems and methods described herein encompass extracting DNA and / or RNA from a single biological sample or multiple biological samples. In some embodiments, extracting DNA comprises extracting DNA from a single sample. In some embodiments, extracting DNA comprises extracting DNA from multiple samples. In some embodiments, extracting DNA comprises extracting DNA from a first sample and a second sample. In some embodiments, extracting DNA comprises extracting DNA from one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more samples.

[0145] In some embodiments, extracting RNA comprises extracting RNA from a single sample. In some embodiments, extracting RNA comprises extracting RNA from multiple samples. In some embodiments, extracting RNA comprises extracting RNA from a first sample and a second sample. In some embodiments, extracting RNA comprises extracting RNA from one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more samples.

[0146] DNA and / or RNA extracted from a biological sample can be combined with DNA and / or RNA extracted from another biological sample. This can be achieved by combining one or more biological samples to extract nucleic acids, or by combining nucleic acids extracted from one or more biological samples. In some embodiments, a first biological sample is combined with a second biological sample to form a combined sample, and DNA and / or RNA is extracted from the combined sample. In some embodiments, DNA and / or RNA extracted from a first biological sample can be combined with DNA and / or RNA extracted from a second biological sample.

[0147] The systems and methods described herein encompass extracting any type of DNA and / or RNA from a biological sample. In some embodiments, extracting DNA comprises extracting genomic DNA (gDNA). In some embodiments, extracting DNA comprises extracting mitochondrial DNA (gDNA). In some embodiments, extracting RNA comprises extracting messenger RNA (mRNA). In some embodiments, extracting RNA comprises extracting precursor mRNA (pre-mRNA). In some embodiments, extracting RNA comprises extracting ribosomal RNA (rRNA). In some embodiments, extracting RNA comprises extracting transfer RNA (tRNA).

[0148] In some embodiments, a single kit is used to purify DNA and RNA from the same sample.A non-limiting example of a kit for doing so is the Qiagen AllPrep DNA / RNA kit.In some embodiments, a robot is employed to perform DNA and / or RNA extraction.

[0149] In some embodiments, if the extracted RNA sample does not have sufficient yield and / or quality, one of the following results may occur: First, common transcripts may be overrepresented in the RNA sequencing data, and low-abundance transcripts may be underrepresented. Second, low RNA quality may result in insufficient read length (i.e., short reads), and / or inadequate read quality may result in misidentification of RNA.

[0150] For whole-exome sequencing, low DNA quantity and quality can lead to misidentification of base pairs, resulting in false variant discovery (e.g., false positives) or false occurrences where variants are not identified (e.g., false negatives). Another problem that can result from low DNA quantity and quality is insufficient coverage of the exome (e.g., missing sequences).

[0151] In some embodiments, the quality and / or quantity of the extracted RNA and / or DNA is checked before it is further processed for RNA sequencing or whole exome sequencing (WES). In some embodiments, the extracted RNA sample has a total mass of at least 1,000-6,000 ng. In some embodiments, the extracted RNA sample has a total mass of at least 100-60,000 ng (e.g., 100-60,000 ng, 500-30,000 ng, 800-20,000 ng, 1,000-15,000 ng, 1,000-10,000 ng, 1,000-8,000 ng, 1,000-6,000 ng, 10,000-20,000 ng, or 20,000-60,000 ng). In some embodiments, the acceptable amount of total RNA for further sequencing is at least 100-1,000 ng (e.g., 100-1,000 ng, 500-1,000 ng, or 300-900 ng). In some embodiments, the target amount of total RNA for further sequencing is greater than 200-1,000 ng (e.g., 200-1,000 ng, 500-1,000 ng, or 300-1,000 ng). In some embodiments, the purity of the extracted RNA sample corresponds to a ratio of absorbance at 260 nm to absorbance at 280 nm of at least 1 (e.g., at least 1, at least 1.2, at least 1.4, at least 1.6, at least 1.8, or at least 2). In some embodiments, the purity of the extracted RNA sample corresponds to a ratio of absorbance at 260 nm to absorbance at 280 nm of at least 2. The ratio of absorbance at 260 nm to absorbance at 280 nm is used to assess the purity of DNA and RNA. A ratio of ~1.8 is generally accepted as "pure" for DNA, and a ratio of ~2.0 is generally accepted as "pure" for RNA. In either case, if this ratio is significantly lower, it may indicate the presence of protein, phenol, or other contaminants that absorb strongly around 280 nm. Absorbance may be measured using a spectrophotometer.

[0152] In some embodiments, the purity or integrity of extracted RNA or DNA (e.g., a DNA fragment library) by any one of the methods described herein corresponds to an RNA Integrity Number (RIN) of at least 4 (e.g., at least 4, at least 5, at least 6, at least 7, at least 8, or at least 9). In some embodiments, the purity of extracted nucleic acid (e.g., RNA or DNA) by any one of the methods described herein corresponds to an RNA Integrity Number (RIN) of at least 7. RIN has been demonstrated to be robust and reproducible in studies compared to other RNA integrity calculation algorithms, solidifying its position as the preferred method for determining the quality of the RNA being analyzed (Imbeaud et al., "Towards standardization of RNA quality assessment using user-independent classifiers of microcapillary electrophoresis traces," Nucleic Acids Research. 33(6):e56).

[0153] In some embodiments, the extracted DNA sample has a total mass of at least 100-20,000 ng (e.g., 100-20,000 ng, 500-15,000 ng, 800-10,000 ng, 1,000-15,000 ng, 1,000-10,000 ng, 1,000-8,000 ng, 1,000-6,000 ng, or 1,000-2,000 ng). In some embodiments, the extracted DNA sample has a total mass of at least 1,000-2,000 ng. In some embodiments, the acceptable total DNA amount for further sequencing is at least 20-200 ng (e.g., 20-200 ng, 30-200 ng, or 50-150 ng). In some embodiments, the target total DNA amount for further sequencing is greater than 30-200 ng (e.g., 30-200 ng, 50-200 ng, or 100-200 ng). In some embodiments, the target purity of the extracted DNA sample is a value corresponding to a ratio of absorbance at 260 nm to absorbance at 280 nm of at least 1.8-2 (e.g., at least 1.8-2, at least 1.8-1.9). In some embodiments, the purity of the extracted DNA sample is a value corresponding to a ratio of absorbance at 260 nm to absorbance at 280 nm of at least 1 (e.g., at least 1, at least 1.2, at least 1.4, at least 1.6, at least 1.8, or at least 2). In some embodiments, an acceptable purity of the extracted DNA sample is a value corresponding to a ratio of absorbance at 260 nm to absorbance at 280 nm of at least 1.5 (e.g., at least 1.5, at least 1.7, at least 2). In some embodiments, the target purity for a sample of extracted DNA corresponds to a range of absorbance at 260 nm to absorbance at 230 nm of at least 2 to 2.2 (e.g., at least 2 to 2.2, at least 2 to 2.1). In some embodiments, an acceptable purity for a sample of extracted DNA corresponds to a range of absorbance at 260 nm to absorbance at 230 nm of at least 1.5 (e.g., at least 1.5, at least 1.7, at least 2).In some embodiments, the purity of a sample of extracted DNA as described herein is analyzed by a spectrophotometer, such as a small-volume full-spectrum UV-visible spectrophotometer (e.g., a NanoDrop spectrophotometer available from ThermoFisher Scientific, www.thermofisher.com).

[0154] In some embodiments, the extracted DNA sample has a target concentration of at least 4.5 ng / μl (e.g., 4.5 ng / μl, 5.5 ng / μl, 6.5 ng / μl). In some embodiments, the extracted DNA sample has an acceptable concentration of at least 3 ng / μl (e.g., 3 ng / μl, 5 ng / μl, 10 ng / μl). In some embodiments, the extracted DNA concentration determination is performed, for example, by a fluorometer for DNA or RNA quantification (e.g., a Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com).

[0155] In some embodiments, the extracted DNA sample has a target concentration of at least 4 ng / μl (e.g., 4 ng / μl, 6 ng / μl, 8 ng / μl). In some embodiments, the extracted DNA sample has an acceptable concentration of at least 2.5 ng / μl (e.g., 2.5 ng / μl, 4.5 ng / μl, 5.5 ng / μl). In some embodiments, the extracted DNA concentration determination is performed by Tapestation.

[0156] In some embodiments, the extracted RNA sample has a target concentration of at least 2 ng / μl (e.g., 2 ng / μl, 4 ng / μl, 6 ng / μl). In some embodiments, the extracted RNA sample has an acceptable concentration of at least 4 ng / μl (e.g., 4 ng / μl, 6 ng / μl, 10 ng / μl). In some embodiments, the extracted DNA concentration determination is performed, for example, by a fluorometer for DNA or RNA quantification (e.g., a Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com).

[0157] In some embodiments, the extracted RNA sample has a target concentration of at least 4 ng / μl (e.g., 4 ng / μl, 6 ng / μl, 8 ng / μl). In some embodiments, the extracted RNA sample has an acceptable concentration of at least 1.5 ng / μl (e.g., 1.5 ng / μl, 3.5 ng / μl, 5.5 ng / μl). In some embodiments, the extracted RNA concentration determination is performed by Tapestation. In some embodiments, an acceptable RNA Integrity Number (RIN) is at least 5 (e.g., 5, 6, 7). In some embodiments, the target RNA Integrity Number (RIN) is at least 8 (e.g., 8, 9, 10). In some embodiments, the RIN is performed by Tapestation.

[0158] In some embodiments, the target purity for a sample of extracted RNA is a value that corresponds to a range of ratios of absorbance at 260 nm to absorbance at 280 nm of at least 1.8 to 2 (e.g., at least 1.8 to 2, at least 1.8 to 1.9). In some embodiments, the purity for a sample of extracted RNA is a value that corresponds to a range of ratios of absorbance at 260 nm to absorbance at 280 nm of at least 1.8. In some embodiments, an acceptable purity for a sample of extracted RNA is a value that corresponds to a range of ratios of absorbance at 260 nm to absorbance at 280 nm of at least 1.5 (e.g., at least 1.5, at least 1.7, at least 2). In some embodiments, the target purity for a sample of extracted RNA is a value that corresponds to a range of ratios of absorbance at 260 nm to absorbance at 230 nm of at least 2 to 2.2 (e.g., at least 2 to 2.2, at least 2 to 2.1). In some embodiments, acceptable purity of a sample of extracted RNA is a value corresponding to a ratio of absorbance at 260 nm to absorbance at 230 nm of at least 1.5 (e.g., at least 1.5, at least 1.7, at least 2). In some embodiments, the purity of a sample of extracted RNA as described herein is analyzed by a spectrophotometer, such as a small-volume full-spectrum UV-visible spectrophotometer (e.g., a NanoDrop spectrophotometer available from ThermoFisher Scientific, www.thermofisher.com). In some embodiments, the concentration of extracted DNA is at least 10-2000 ng / μl (e.g., 10-2000 ng / μl, 10-1000 ng / μl, 10-200 ng / μl, 1-200 ng / μl, 0.5-400 ng / μl, 0.5-200 ng / μl, 100-200 ng / μl, 100-400 ng / μl, 100-500 ng / μl, 50-500 ng / μl, or 50-250 ng / μl).

[0159] Protocols for quality control of extracted RNA or DNA samples are described at least in Example 6. In some embodiments, the purity of extracted DNA and / or RNA samples as described herein can be analyzed by any other suitable technique or tool. In some embodiments, extracted RNA or DNA samples are not further processed if they do not meet certain quantity or purity criteria, as described above. In some embodiments, extracted RNA or DNA samples are combined with another sample if they do not meet certain quantity or purity criteria.

[0160] Library preparation for RNA sequencing Methods for preparing cDNA libraries from RNA samples are known in the art. For example, www.illumina.com / content / dam / illumina-marketing / documents / applications / ngs-library-prep / for-all-you-seq-rna.pdf provides illustrations of different methods for preparing cDNA libraries for RNA sequencing. Non-limiting examples of cDNA library preparation include ClickSeq, 3Seq, and cP-RNA-Seq. In some embodiments, preparing a cDNA library from RNA includes purifying mRNA from the RNA sample (RNA enrichment). In some embodiments, the enriched RNA is fragmented. In some embodiments, after selection of the appropriate RNA fraction is complete, the molecules are fragmented into smaller pieces between 50 and 1000 bp (e.g., 50-100 bp, 100-800 bp, 100-500 bp, or 200-500 bp), depending on the sequencing platform being used. This fragmentation can be achieved either by fragmenting double-stranded (ds) cDNA or by fragmenting RNA, both methods resulting in the same final product: a double-stranded cDNA library with adapters attached to each fragment.

[0161] In some embodiments, the library preparation method includes one or more amplification steps to add functional elements (e.g., sample indexes, molecular barcodes, or flow cell oligo binding sites), enrich for sequencing-competent DNA fragments, and / or generate sufficient amounts of library DNA for downstream processing. In some embodiments, the enriched RNA (e.g., fragmented and enriched RNA) is amplified using random primers (e.g., random hexamers). In some embodiments, the enriched RNA (e.g., fragmented and enriched RNA) is amplified using oligo-dTs. In some embodiments, the RNA is then removed from the formed cDNA. In some embodiments, the cDNA is amplified to include sequencing adaptors and indexes (i.e., multiple indexes). Adapters are 10-100 bp (e.g., 10-20, 10-100, 20-80, 30-70, 40-60, 20-100, 40-100, 40-80, 30-60, or 45-65 bp) DNA sequences that can be attached to a flow cell for sequencing. Adapters also allow PCR enrichment of adapter-ligated DNA fragments. Adapters can enable indexing or barcoding of samples so that multiple cDNA libraries can be mixed together in one sequencing sample (or lane), i.e., multiplexing. In some embodiments, the index or barcode is 4-20 bp in length (e.g., 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 4-20, 5-15, 6-12, or 4-12 bp). tucf-genomics.tufts.edu / documents / protocols / TUCF_Understanding_Illumina_TruSeq_Adapters.pdf provides an exemplary protocol for preparing cDNA libraries using adapters and indexing, which is incorporated herein by reference in its entirety.Protocols for constructing DNA or RNA libraries are described in at least Examples 3 and 5.

[0162] RNA enrichment Methods of RNA enrichment (also referred to herein as "RNA enrichment") for enriching mRNA during cDNA library preparation are known in the art. RNA enrichment can be targeted or non-targeted. Targeted methods of RNA enrichment include the use of sequence-specific capture probes. A non-limiting example of targeted mRNA enrichment includes CaptureSeq (sapac.illumina.com / science / sequencing-method-explorer / kits-and-arrays / aptureseq.html), which utilizes capture probes specific to sequences of interest. Other platforms and tools suitable for targeted mRNA enrichment can also be used.

[0163] Examples of non-targeted mRNA enrichment methods include polyA capture using oligo-dT (e.g., conjugated to beads) and rRNA depletion. Petrova et al., Scientific Reports volume 7, Article number: 41114 (2017), provides a comparison of various rRNA depletion methods, the entire contents of which are incorporated herein by reference. In some embodiments, rRNA depletion can be performed using an enzymatic approach (e.g., using an exonuclease that does not process mRNA). In some embodiments, rRNA depletion methods include subtractive hybridization, whereby rRNA is captured using sequence-specific probes (see, e.g., www.sciencedirect.com / topics / immunology-and-microbiology / subtractive-hybridization).

[0164] In some embodiments, polyA capture involves capturing polyA-tailed mRNA using a polyA-specific capture probe (oligo dT). In some embodiments, the capture probe is immobilized to facilitate purification. In some embodiments, the capture probe is immobilized on beads (e.g., magnetic beads). In some embodiments, commercially available kits are used to prepare DNA libraries from RNA samples. In some embodiments, the Illumina TruSeq RNA Library Prep kit is used.

[0165] The choice of mRNA enrichment can have a significant impact on the selection of sequenced transcripts. For example, in some embodiments, cDNA libraries prepared using polyA enrichment, compared to rRNA depletion methods, result in libraries containing a higher fraction (e.g., greater than 80%, greater than 90%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, or greater than 99.9%) of protein-coding transcripts relative to non-coding transcripts (e.g., rRNA, miRNA, and IncRNA).

[0166] In some embodiments, prepared cDNA libraries are inspected for quality. In some embodiments, quantification of libraries for use in sequencing is typically performed before the libraries are pooled for target enrichment or amplification to ensure equal representation of indexed libraries in multiplexed applications. In some embodiments, quantification is also used to ensure optimal dilution of individual libraries or library pools before sequencing. Accurate and reproducible quantification of adapter-ligated library molecules contributes to obtaining consistent and reproducible results and maximizing sequencing yields. Loading more DNA than recommended can saturate the flow cell or result in high cluster density, while loading too little DNA can result in low cluster density and reduced coverage and depth.

[0167] Methods for quantifying DNA libraries include electrophoresis, fluorometry, spectrophotometry, digital PCR, droplet digital PCR, and qPCR. Various instruments exist for measuring the quantity and / or quality of DNA libraries, such as the Agilent High Sensitivity D1000 ScreenTape System.

[0168] Aspects of the present disclosure provide quality control of nucleic acids to be analyzed for sequencing. Aspects of the present disclosure provide quality control of DNA to be analyzed for sequencing. Aspects of the present disclosure provide quality control of RNA to be analyzed for sequencing. In some embodiments, the nucleic acids can include any suitable type of DNA or RNA. In some embodiments, quality control of nucleic acids includes verifying biopsy conditions and documentation. In some embodiments, verifying biopsy conditions and documentation can include, but is not limited to, inventorying and registering nucleic acid material. In some embodiments, verifying biopsy conditions and documentation includes receiving nucleic acid material. For example, patient samples received from a healthcare provider can be verified to determine whether the patient tissue is fresh-frozen or formalin-fixed, paraffin-embedded. Laboratory personnel verify biopsy compliance of registered entities. Laboratory personnel verify proper storage of biopsy samples during transport. Laboratory personnel verify the physical condition of the biopsy sample. If laboratory personnel identify any errors with the biopsy sample, the source of the biopsy sample (e.g., a healthcare provider) can be notified. In some embodiments, if the received biopsy sample is a patient tissue cell line, the sample is prepared for extraction. In some embodiments, if the received biopsy sample is extracted DNA or RNA, the sample is stored at -80°C for further sequencing. In some embodiments, the extracted DNA can be reference gDNA. In some embodiments, the extracted RNA can be reference RNA.

[0169] In some embodiments, the quality control procedure defines a target range. The target range may represent the most ideal quality for a given step (e.g., extraction). In some embodiments, the quality control procedure defines an acceptable range. The acceptable range may represent the ideal or acceptable quality for a given step. In some embodiments, nucleic acid quality control includes ensuring quality in the process of constructing a DNA library. In some embodiments, nucleic acid quality control includes ensuring quality in the process of constructing an RNA library. As shown in FIG. 7 and Example 6, DNA or RNA library preparation includes extracting DNA or RNA from patient tissue samples. In some embodiments, a spectrophotometer, such as a small-volume full-spectrum UV-visible spectrophotometer (e.g., the NanoDrop spectrophotometer available from ThermoFisher Scientific, www.thermofisher.com), can be used to determine the quality of the DNA or RNA extraction. By way of example, extracted DNA of >100 ng / μl indicates that the extracted DNA passes the quality control test. Extracted RNA >500ng / μl indicates that the extracted RNA has passed quality control testing. Alternatively, a 260nm to 280nm absorbance ratio (260 / 280) of 1.8-2.0 indicates that the extracted DNA has passed quality control testing. A 260nm to 280nm absorbance ratio (260 / 280) of 2.0 indicates that the extracted RNA has passed quality control testing. Alternatively, a 260nm to 230nm absorbance ratio (260 / 230) of 2.0-2.2 indicates that the extracted DNA has passed quality control testing. A 260nm to 230nm absorbance ratio (260 / 230) of 2.0-2.2 indicates that the extracted RNA has passed quality control testing.In some embodiments, a fluorometer (e.g., a Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com), for example, for DNA or RNA quantification, may be used to determine the quality of the DNA or RNA extraction. In some embodiments, an electrophoresis device, for example, an automated electrophoresis device (e.g., a TapeStation System available from Agilent, www.agilent.com), may be used to determine the quality of the DNA or RNA extraction. In some embodiments, any suitable technique or tool may be used to determine the quality of the DNA or RNA extraction.

[0170] In some embodiments, the acceptable total DNA amount for further DNA library construction is at least 200-1,000 ng (e.g., 200-1,000 ng, 300-1,000 ng, or 300-1,000 ng). In some embodiments, the target total DNA amount for further sequencing is more than 500-1,000 ng (e.g., 500-1,000 ng, 600-1,000 ng, or 800-1,000 ng). In some embodiments, the acceptable total RNA amount for further RNA library construction is at least 0.5-4 nmol / L (e.g., 200-1,000 ng, 300-1,000 ng, or 300-1,000 ng). In some embodiments, the target amount of total RNA for further RNA library construction is at least 0.5 to 4 nmol / L (eg, 500 to 1,000 ng, 600 to 1,000 ng, or 800 to 1,000 ng).

[0171] In some embodiments, the acceptable DNA concentration for further DNA library construction is at least 17 ng / μl (e.g., 17 ng / μl, 25 ng / μl, 35 ng / μl). In some embodiments, the target DNA concentration for further DNA library construction is at least 42 ng / μl (e.g., 42 ng / μl, 50 ng / μl, 80 ng / μl). In some embodiments, the acceptable RNA concentration for further RNA library construction is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 1 ng / μl, 3 ng / μl). In some embodiments, the target RNA concentration for further RNA library construction is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 1 ng / μl, 3 ng / μl). In some embodiments, DNA and RNA concentrations are detected, for example, by a fluorometer for DNA or RNA quantification (e.g., a Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com).

[0172] In some embodiments, the acceptable DNA concentration for further DNA library construction is at least 15 ng / μl (e.g., 15 ng / μl, 25 ng / μl, 35 ng / μl). In some embodiments, the target DNA concentration for further DNA library construction is at least 402 ng / μl (e.g., 40 ng / μl, 50 ng / μl, 80 ng / μl). In some embodiments, the acceptable RNA concentration for further RNA library construction is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 1 ng / μl, 3 ng / μl). In some embodiments, the target RNA concentration for further RNA library construction is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 1 ng / μl, 3 ng / μl). In some embodiments, the acceptable RNA concentration for further RNA library construction is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 1 nmol / l, 5 nmol / l). In some embodiments, the target RNA concentration for further RNA library construction is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 1 nmol / l, 5 nmol / l). In some embodiments, DNA and RNA concentrations are detected by Tapestation.

[0173] In some embodiments, the acceptable RNA concentration for further RNA library construction is at least 0.5 nmol / L (e.g., 0.5 nmol / L, 1 nmol / L, 5 nmol / L). In some embodiments, the target RNA concentration for further RNA library construction is at least 0.5 nmol / L (e.g., 0.5 nmol / L, 1 nmol / L, 5 nmol / L). In some embodiments, the DNA and RNA concentrations are detected by a nucleic acid amplification device (e.g., a PCR system), for example, a real-time PCR system (e.g., the LightCycler Instrument available from Roche, www.lifescience.roche.com). In some embodiments, the DNA and RNA concentrations can be detected by any suitable technique or tool.

[0174] In some embodiments, if RNA is extracted, reverse transcription may be performed. In some embodiments, an RNA library may be constructed after reverse transcription is performed. In some embodiments, a fluorometer (e.g., a Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com), for example, for DNA or RNA quantification, may be used to determine the quality of the DNA or RNA library. In some embodiments, any suitable method may be used to determine the quality of the DNA or RNA library. In some embodiments, an electrophoresis device, for example, an automated electrophoresis device (e.g., a TapeStation System available from Agilent, www.agilent.com), may be used to determine the quality of the DNA or RNA library. In some embodiments, a fluorometer (e.g., a Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com), for example, for DNA or RNA quantification, may be used to determine the quality of the DNA or RNA extraction. In some embodiments, a nucleic acid amplification device (e.g., a PCR system), such as a real-time PCR system (e.g., the LightCycler Instrument available from Roche, www.lifescience.roche.com), can be used to determine the quality of the RNA library. In some embodiments, one or more RNA libraries can be pooled. In some embodiments, if DNA is extracted, the extracted DNA can be used for DNA library construction. In some embodiments, DNA fragments in the constructed DNA library can be hybridized and / or captured.In some embodiments, a fluorometer (e.g., Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com), for example, for DNA or RNA quantification, may be used to determine the quality of the DNA hybridization and capture steps. In some embodiments, an electrophoresis device, for example, an automated electrophoresis device (e.g., TapeStation System available from Agilent, www.agilent.com), may be used to determine the quality of the DNA hybridization and capture steps. In some embodiments, a nucleic acid amplification device (e.g., PCR system), for example, a real-time PCR system (e.g., LightCycler Instrument available from Roche, www.lifescience.roche.com), may be used to determine the quality of the DNA hybridization and capture steps. In some embodiments, any suitable method may be used to determine the quality of the DNA hybridization and capture steps. In some embodiments, one or more DNA libraries may be pooled. In some embodiments, an electrophoresis device, such as an automated electrophoresis device (e.g., the TapeStation System available from Agilent, www.agilent.com), can be used to determine the quality of DNA or RNA library pooling. In some embodiments, any suitable method can be used to determine the quality of DNA or RNA library pooling.

[0175] In some embodiments, the acceptable and / or target final DNA concentration range for pooling is at least 0.5-4 nmol / L (e.g., 0.5-4 nmol / L, 0.5-3 nmol / L, 2-4 nmol / L). In some embodiments, the acceptable DNA concentration for pooling is at least 0.1 ng / μL (e.g., 0.1 ng / μL, 0.8 ng / μL, 4 ng / μL), for example, when a fluorometer for DNA or RNA quantification (e.g., Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com) is used. In some embodiments, the target DNA concentration for pooling is at least 0.1 ng / μL (e.g., 0.1 ng / μL, 0.8 ng / μL, 4 ng / μL), for example, when a fluorometer for DNA or RNA quantification (e.g., Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com) is used.

[0176] In some embodiments, the acceptable DNA concentration for pooling is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 0.8 ng / μl, 4 ng / μl) when an electrophoresis device, such as an automated electrophoresis device (e.g., the TapeStation System available from Agilent, www.agilent.com), is used. In some embodiments, the target DNA concentration for pooling is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 0.8 ng / μl, 4 ng / μl) when an electrophoresis device, such as an automated electrophoresis device (e.g., the TapeStation System available from Agilent, www.agilent.com), is used. In some embodiments, the acceptable DNA concentration for pooling is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 0.8 nmol / l, 3 nmol / l) when an electrophoresis device, such as an automated electrophoresis device (e.g., the TapeStation System available from Agilent, www.agilent.com), is used. In some embodiments, the target DNA concentration for pooling is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 0.8 nmol / l, 3 nmol / l) when an electrophoresis device, such as an automated electrophoresis device (e.g., the TapeStation System available from Agilent, www.agilent.com), is used. In some embodiments, the acceptable concentration and / or density of DNA is in the range of 380-440 ng (e.g., 380-440 ng, 400-440 ng, 420-440 ng) when an electrophoresis device, such as an automated electrophoresis device (e.g., TapeStation System available from Agilent, www.agilent.com), is used.In some embodiments, the acceptable DNA concentration for pooling is at least 0.5 nmol / L (e.g., 0.5 nmol / L, 0.8 nmol / L, 3 nmol / L) when a nucleic acid amplification device (e.g., a PCR system), such as a real-time PCR system (e.g., the LightCycler Instrument available from Roche, www.lifescience.roche.com) is used. In some embodiments, the target DNA concentration for pooling is at least 0.5 nmol / L (e.g., 0.5 nmol / L, 0.8 nmol / L, 3 nmol / L) when a LightCycler is used.

[0177] In some embodiments, nucleic acid quality control includes ensuring the quality of DNA or RNA libraries after construction, such as during the sequencing process. In some embodiments, cluster density may be a parameter for sample run quality control (Example 6). Cluster density is an important factor in optimizing sequencing data quality and yield. Without wishing to be bound by any theory, an optimal cluster density indicates that at least the DNA or RNA library is balanced. In some embodiments, quality score and signal-to-noise ratio may be parameters for sample run quality control.

[0178] In some embodiments, nucleic acid quality control includes ensuring sequencing quality. In some embodiments, sequencing quality control includes bioinformatics quality control. In some embodiments, sequencing can be DNA sequencing. In some embodiments, sequencing can be RNA sequencing. In some embodiments, sequencing can be any type of sequencing technique known in the art for determining the DNA or RNA expression profile of a given biological sample. By way of example, sequencing can be whole-exome sequencing. Sequencing can be transcriptome sequencing. Sequencing can be Sanger sequencing.

[0179] In some embodiments, up to 1 ng (e.g., up to 0, up to 0.1 μl, up to 0.5 μl, up to 0.8 μl, up to 0.9 μl, up to 1 μl, up to 1.2 μl, up to 1.4 μl, up to 1.5 μl, up to 1.8 μl, or up to 2 μl) of a library of up to 2 μl (e.g., up to 0.1, up to 0.2, up to 0.3, up to 0.4, up to 0.5, up to 0.6, up to 0.7, up to 0.8, up to 0.9, or up to 1 ng) of solution is used for quality control testing. In some embodiments, parameters examined include size and size distribution of DNA molecules, and purity.

[0180] In some embodiments, standard methods for preparing libraries of cDNA fragments from RNA fail to preserve information related to which DNA strand was the original template during transcription and subsequent synthesis of mRNA transcripts. Because antisense transcripts likely have regulatory roles distinct from their protein-coding counterparts, this loss of strand information results in an incomplete understanding of the transcriptome. Strand-specific RNA-Seq can be performed to preserve this strandness. Methods for preserving strandness and preparing cDNA fragment libraries for this purpose are known in the art (e.g., Mills et al., "Strand-Specific RNA-Seq Provides Greater Resolution of Transcriptome Profiling," Curr Genomics. May 2013, 14(3):173-181). In some embodiments, library preparation for stranded RNA-Seq utilizes strand-specific adapters of known orientation. In some embodiments, the strand is chemically modified to remember its origin.

[0181] In some embodiments, the adapter-based method involves strand-specific 3'-end RNA-seq. In some embodiments, strand-specific 3'-end RNA-seq involves anchored oligo(dT) primers, which are first used to select mRNA, resulting in double-stranded cDNA molecules. Then, adapters for paired-end sequencing are ligated to both ends of the cDNA molecules. The fragments are then sequenced to generate paired-end reads aligned to the reference genome. Aligned reads containing adenine stretches at the end of the transcript must be transcripts derived from the DNA antisense strand, while any reads that align with a thymine stretch at the front must be transcripts derived from the DNA sense strand.

[0182] In some embodiments, the adapter-based method utilizes single-stranded (ss) cDNA and Illumina adapters and four DNA ligases that allow for ligation of 3' and 5' adapters to the ssDNA. The second strand is never synthesized or sequenced, so strand information is preserved.

[0183] In some embodiments, any suitable technique or tool may be used to preserve strandedness. For example, flow cell reverse transcription sequencing (FRT-Seq) may be used to preserve strandedness. In some embodiments, FRT-Seq or an equivalent technique involves ligating adapters to either end of fragmented, purified polyadenylated mRNA. In some embodiments, each adapter contains two regions: one to which a sequencing primer anneals and one complementary to an oligonucleotide present on the flow cell. The complementary regions allow the mRNA fragments to hybridize to the flow cell. The mRNA fragments are then reverse transcribed on the flow cell surface.

[0184] Other non-limiting adapter-based methods that preserve strandedness include direct strand-specific sequencing (DSSS) and the SOLiDR Total RNA-Seq Kit (tools.thermofisher.com / content / sfs / manuals / cms_078610.pdf), which preserves strand specificity through the addition of directional adapters.

[0185] In some embodiments, chemical modification of the strand to remember its origin involves marking the original RNA template using bisulfite treatment. In some embodiments, dUTP is incorporated into the reverse transcription reaction, resulting in a ds cDNA in which the original strand has deoxythymidine residues and the complementary strand has deoxyuridine residues. Uracil-DNA-glycosylase (UDG) treatment can then be used to degrade the complementary strand.

[0186] WES library preparation The "exome" is the sum of all regions in the genome consisting of exons. Exons are regions of DNA that are transcribed into messenger RNA, as opposed to introns, which are removed by splicing proteins. Exome sequencing is a capture-based method developed to identify variants present in the coding regions of genes that affect protein function. Because the coding portion of the genome comprises only 1-2% of the entire genome, this approach represents a cost-effective strategy for detecting DNA modifications that may alter protein function compared to whole genome sequencing. In some embodiments, whole exome sequencing (WES) involves preparing a library of DNA fragments for sequencing from a sample of DNA. In some embodiments, the DNA is first fragmented to the appropriate size (depending on the sequencing platform used), and then sequencing platform-specific adapters are added. In some embodiments, the library is amplified before the next step in the process (target enrichment or sequencing).

[0187] Kits for library preparation are commercially available, non-limiting examples of which include KAPA HyperPrep Kits, Agilent HaloPlex, Agilent SureSelect QXT, IDT xGEN Exome, Illumina Nextera Rapid Capture Exome, Roche Nimblegen SeqCap, and MYcroarray MYbaits. In some embodiments, any kit capable of preparing a DNA library for WES can be used. For example, the Agilent Human All Exon V6 Capture Kit (www.agilent.com / cs / library / datasheets / public / SureSelect%20V6%20DataSheet%205991-5572EN.pdf) is used to prepare a DNA library for WES. In some embodiments, the Clinical Research Exome kit (www.agilent.com / en / promotions / clinical-research-exome-v2) is used. The amount of DNA required depends on the specific reagents used to prepare the library. For example, 100 ng of genomic DNA is sufficient for the Agilent SureSelect XT2 V6 Exome, while 500 ng of genomic DNA is required for the IDT xGEN Exome Panel. A comparison of various capture kits is provided at www.genohub.com / exome-sequencing-library-preparation / .

[0188] In some embodiments, library preparation methods include one or more amplification steps to add functional elements (e.g., sample indexes, molecular barcodes, or flow cell oligo binding sites), enrich for sequencing-competent DNA fragments, and / or generate sufficient amounts of library DNA for downstream processing. Exemplary library preparation methods are shown in Examples 3 and 5.

[0189] In some embodiments, prepared DNA libraries are inspected for quality. In some embodiments, quantification of libraries for use in sequencing is typically performed before the libraries are pooled for target enrichment or amplification to ensure equal representation of indexed libraries in multiplexed applications. In some embodiments, quantification is also used to ensure optimal dilution of individual libraries or library pools before sequencing. Accurate and reproducible quantification of adapter-ligated library molecules contributes to obtaining consistent and reproducible results and maximizing sequencing yields. Loading more DNA than recommended can saturate the flow cell or result in high cluster density, while loading too little DNA can result in low cluster density and reduced coverage and depth.

[0190] Methods for quantifying DNA libraries include electrophoresis, fluorometry, spectrophotometry, digital PCR, droplet digital PCR, and qPCR. Various instruments exist for measuring the quantity and / or quality of DNA libraries, such as the Agilent High Sensitivity D1000 ScreenTape System.

[0191] In some embodiments, the prepared DNA library is inspected for quality. In some embodiments, up to 1 ng (e.g., up to 0, up to 0.1 μl, up to 0.5 μl, up to 0.8 μl, up to 0.9 μl, up to 1 μl, up to 1.2 μl, up to 1.4 μl, up to 1.5 μl, up to 1.8 μl, or up to 2 μl) of a solution of up to 2 μl (e.g., up to 0.1, up to 0.2, up to 0.3, up to 0.4, up to 0.5, up to 0.6, up to 0.7, up to 0.8, up to 0.9, or up to 1 ng) of the library is used for quality control testing. In some embodiments, the parameters inspected include the size and size distribution of the DNA molecules, and purity.

[0192] RNA sequencing RNA sequencing is a tool for measuring the transcriptome. The transcriptome consists of a distinct population of RNA molecules, including mRNA, rRNA, tRNA, and other non-coding RNAs (e.g., microRNA, lncRNA, etc.). In some embodiments, RNA sequencing is used to profile the transcriptome (e.g., coding and / or non-coding regions). In some embodiments, it is used to identify differentially expressed genes in different biological samples (e.g., cells, tissues, or body fluids). In some embodiments, RNA sequencing is used to determine the genetic effects of splicing events, identify novel transcripts, detect structural variations (e.g., gene fusions and isoforms), and / or detect single nucleotide variants.

[0193] In some embodiments, the term "RNA sequencing" can be used interchangeably with "RNA seq," "RNA-seq," or variations thereof, which are known in the art to refer to any technology, tool, or platform that interrogates the transcriptome. Note that when "RNA sequencing," "RNA seq," "RNA-seq," or variations thereof are referenced in this disclosure, they do not refer to specific technologies or tools associated with a particular platform or company, unless otherwise indicated using non-limiting examples to demonstrate the processes or systems as described herein. In some embodiments, RNA sequencing can be performed using any suitable sequencing platform and / or sequencing method. Non-limiting examples of high-throughput sequencing platforms include mRNA-seq, total RNA-seq, targeted RNA-seq, single-cell RNA-seq, RNA exome capture platforms, or small RNA-seq (e.g., Illumina, www.illumina.com), SMRT (single molecule, real-time) sequencing (e.g., Pacific Biosciences, https: / / www.pacb.com), and RNA sequencing (e.g., ThermoFisher, https: / / www.thermofisher.com).

[0194] As described above, RNA sequencing can be targeted or non-targeted. Targeted approaches involve using sequence-specific probes or oligonucleotides to sequence one or more specific regions of the transcriptome. In some embodiments, targeted RNA sequencing involves methods such as mRNA enrichment (e.g., by polyA enrichment or rRNA depletion).

[0195] In some embodiments, the RNA sequencing is whole transcriptome sequencing. Whole transcriptome sequencing involves measuring the full complement of transcripts in a sample. In some embodiments, whole transcriptome sequencing is used to determine the global expression level of each transcript (e.g., both coding and non-coding) and identify exons, introns, and / or their junctions.

[0196] In some embodiments, RNA is sequenced directly without preparing cDNA from a sample of RNA. In some embodiments, direct RNA sequencing comprises single molecule RNA sequencing (DRSTM).

[0197] In some embodiments, the RNA sequencing is mRNA sequencing. In some embodiments, the mRNA sequencing is sequencing of only coding transcripts, with the goal of excluding non-coding regions. In some embodiments, the mRNA sequencing is independent of polyA enrichment. In some embodiments, the mRNA sequencing is dependent on polyA enrichment.

[0198] In some embodiments, RNA is extracted from a biological sample, mRNA is enriched from the extracted RNA, and a cDNA library is constructed from the enriched mRNA. In some embodiments, single pieces of cDNA from the cDNA library are attached to a solid matrix. In some embodiments, single pieces of cDNA from the cDNA library are attached to a solid matrix by limiting dilution. In some embodiments, the cDNA pieces attached to the matrix are then sequenced (e.g., using Pacbio or Pacifbio technology). In some embodiments, the cDNA pieces attached to the matrix are amplified and sequenced (e.g., using SOLiD, 454 Pyrosequencing, Ion Torrent, or dedicated emulsion PCR (emPCR) in a connector-based bridge reaction (Illumina) platform).

[0199] In some embodiments, cDNA transcripts can be sequenced in parallel by measuring the incorporation of fluorescent nucleotides (e.g., Illumina), fluorescent short linkers (e.g., SOLiD), the release of by-products from the incorporation of normal nucleotide (454), measuring fluorescence emission, or measuring pH changes (e.g., Ion Torrent). In some embodiments, cDNA transcripts can be sequenced using any known sequencing platform. Jazayeri et al., "RNA-seq: a glance at technologies and methodologies," Acta biol. Colomb., vol. 20 no. 2, Bogota, May / Aug. 2015, provides a comparison of different RNA-seq platforms, which is incorporated herein by reference in its entirety, including Tables 3 and 4. Mestan et al., "Genomic sequencing in clinical trials," Journal of Translational Medicine, 2011, 9:222, performs a similar analysis of sequencing in clinical trials.

[0200] In some embodiments, RNA sequencing is strand-specific. cDNA synthesis from RNA results in loss of strandness. In some embodiments, strandness is preserved by chemically labeling either or both of the RNA strand and the cDNA strand formed by reverse transcription or antisense transcription, as described above, or by using adapter-based technology to distinguish the original RNA strand from the complementary DNA strand.

[0201] In some embodiments, non-stranded RNA sequencing is performed. In some embodiments, stranded RNA sequencing should be avoided for clinical samples. In some embodiments, non-stranded RNA sequencing is used to compare data obtained from biological samples with RNA sequencing data from established datasets (e.g., The Cancer Genome Atlas (TCGA) and the International Cancer Genome Consortium (ICGC)).

[0202] In some embodiments, RNA sequencing obtains paired-end reads. Paired-end reads are reads of the same nucleic acid fragment, and are reads that start from either end of the fragment. In some embodiments, RNA sequencing is performed with at least 2x25 paired-end reads (2x25, 2x50, 2x75, 2x100, 2x125, 2x150, 2x175, 2x200, 2x225, 2x250, 2x275, 2x300, 2x325, or 2x350). In some embodiments, RNA sequencing is performed with at least 2x75 paired-end reads. RNA sequencing with 2x75 paired-end reads means that, on average, each paired-end read reads 75 base pairs. In some embodiments, RNA sequencing is performed with a total of at least 20 million paired-end reads (e.g., at least 20 million, at least 30 million, at least 40 million, at least 50 million, at least 60 million, at least 70 million, at least 80 million, at least 90 million, at least 100 million, at least 120 million, at least 140 million, at least 150 million, at least 160 million, at least 180 million, at least 200 million, at least 250 million, at least 300 million, at least 350 million, or at least 400 million). In some embodiments, RNA sequencing is performed with a total of at least 50 million paired-end reads. In some embodiments, RNA sequencing is performed with a total of at least 100 million paired-end reads.

[0203] In some embodiments, quality control is performed for RNA sequencing. In some embodiments, cluster density or cluster PF% is a parameter for determining the quality of a sample run. In some embodiments, the target range for cluster density or cluster PF% is at least 170-220 (e.g., 170-220, 190-220, 210-220). In some embodiments, the acceptable range for cluster density or cluster PF% is at least 280 (e.g., 280, 300, 450).

[0204] In some embodiments, %≧Q30 is a parameter for determining the quality of a sample run. In some embodiments, the target %≧Q30 is at least 85% (e.g., 85%, 90%, 95%). In some embodiments, the acceptable %≧Q30 is at least 75% (e.g., 75%, 85%, 95%).

[0205] In some embodiments, the % error rate is a parameter for determining the quality of a sample run. In some embodiments, the target % error rate is at least less than 0.7% (e.g., 0.6%, 0.5%, 0.4%). In some embodiments, the acceptable % error rate is at least less than 1% (e.g., 0.9%, 0.8%, 0.7%).

[0206] Whole exome sequencing (WES) Whole exome sequencing (WES) is a genomic technology for sequencing all of the protein-coding regions of genes in the genome. In some embodiments, WES is performed to identify genetic mutations that alter protein sequences. In some embodiments, WES is performed to identify genetic mutations that alter protein sequences at a cost lower than that of whole genome sequencing.

[0207] In some embodiments, whole exome sequencing (WES) is performed on a sample of DNA extracted from a biological sample. In some embodiments, a library of DNA fragments is prepared from the extracted DNA sample. In some embodiments, any one of the methods described herein includes performing whole exome sequencing (WES) on the library of DNA fragments. Preparation of a DNA library from a DNA sample for WES is as described above.

[0208] In some embodiments, the DNA library is quantified before sequencing (e.g., using next generation sequencing (NGS)). In some embodiments, the DNA library is pooled before sequencing. In some embodiments, the DNA library is amplified before sequencing. In some embodiments, the DNA library is indexed before sequencing to track the origin of the DNA fragments.

[0209] In some embodiments, WES involves target enrichment, which allows for selective capture of genomic regions of interest prior to sequencing. In some embodiments, array-based capture is used (e.g., using microarrays). In some embodiments, in-solution capture is used.

[0210] Any high-throughput DNA sequencing platform and / or method can be used in any one of the methods described herein.In some embodiments, DNA sequencing can be carried out by using any suitable platform and / or method. Non-limiting examples of high throughput sequencing methods include single molecule real-time sequencing, ion semiconductor (Ion Torrent sequencing), pyrosequencing (i.e., 454), sequencing by synthesis (Illumina), Illumina (Solexa) sequencing, combinatorial probe anchor synthesis (cPAS- BGI / MGI), sequencing by ligation (SOLiD sequencing), nanopore sequencing (e.g., using an Oxford Nanopore Technologies instrument), chain termination (Sanger sequencing), massively parallel signature sequencing (MPSS) polony sequencing, Heliscope single molecule sequencing, and single molecule real-time (SMRT) sequencing (e.g., using a Pacific Biosciences instrument). Other non-limiting examples of high-throughput sequencing techniques include tunneling current DNA sequencing, sequencing by hybridization, sequencing using mass spectrometry, microfluidic Sanger sequencing, and RNAP sequencing.

[0211] In some embodiments, DNA sequencing obtains paired-end reads. Paired-end reads are reads of the same nucleic acid fragment, and are reads that start from either end of the fragment. In some embodiments, DNA sequencing is performed with at least 2x25 paired-end reads (2x25, 2x50, 2x75, 2x100, 2x125, 2x150, 2x175, 2x200, 2x225, 2x250, 2x275, 2x300, 2x325, or 2x350). In some embodiments, DNA sequencing is performed with at least 2x75 paired-end reads. DNA sequencing with 2x75 paired-end reads means that, on average, each paired-end read reads 75 base pairs. In some embodiments, DNA sequencing is performed with a total of at least 20 million paired-end reads (e.g., at least 20 million, at least 30 million, at least 40 million, at least 50 million, at least 60 million, at least 70 million, at least 80 million, at least 90 million, at least 100 million, at least 120 million, at least 140 million, at least 150 million, at least 160 million, at least 180 million, at least 200 million, at least 250 million, at least 300 million, at least 350 million, or at least 400 million). In some embodiments, DNA sequencing is performed with a total of at least 50 million paired-end reads. In some embodiments, DNA sequencing is performed with a total of at least 100 million paired-end reads. In some embodiments, DNA sequencing is performed to obtain at least 20-fold coverage (e.g., at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, at least 1000-fold, at least 120-fold, at least 125-fold, at least 150-fold, at least 175-fold, at least 200-fold, at least 250-fold, at least 300-fold, or at least 400-fold).Coverage, also referred to as depth, is the number of times that a single base pair in a sample of nucleic acid is read or sequenced on average. In some embodiments, the portion of the genome targeted for capture and sequencing is at least 10 Mb (e.g., at least 10 Mb, at least 20 Mb, at least 30 Mb, at least 40 Mb, at least 50 Mb, at least 60 Mb, at least 70 Mb, at least 80 Mb, at least 90 Mb, at least 100 Mb, at least 120 Mb, at least 150 Mb, at least 200 Mb, at least 250 Mb, at least 300 Mb, or at least 350 Mb). In some embodiments, the portion of the genome targeted for capture and sequencing is at least 48 Mb (e.g., after using the Agilent Human All Exon V6 Capture system). In some embodiments, the portion of the genome targeted for capture and sequencing is at least 54 Mb (eg, after using the Clinical Research Exome Capture System (Agilent)).

[0212] In some embodiments, quality control is performed for whole-exome sequencing. In some embodiments, cluster density or cluster PF% is a parameter for determining the quality of a sample run. In some embodiments, the target range for cluster density or cluster PF% is at least 170-220 (e.g., 170-220, 190-220, 210-220). In some embodiments, the acceptable range for cluster density or cluster PF% is at least 280 (e.g., 280, 300, 450).

[0213] In some embodiments, the actual yield is a parameter for determining the quality of a sample run. In some embodiments, the target actual yield is at least 15 Gbp (e.g., 15 Gbp, 20 Gbp, 30 Gbp).

[0214] In some embodiments, %≧Q30 is a parameter for determining the quality of a sample run. In some embodiments, the target %≧Q30 is at least 85% (e.g., 85%, 90%, 95%). In some embodiments, the acceptable %≧Q30 is at least 75% (e.g., 75%, 85%, 95%).

[0215] In some embodiments, the % error rate is a parameter for determining the quality of a sample run. In some embodiments, the target % error rate is at least less than 0.7% (e.g., 0.6%, 0.5%, 0.4%). In some embodiments, the acceptable % error rate is at least less than 1% (e.g., 0.9%, 0.8%, 0.7%).

[0216] Reagents and Kits Contemplated herein are reagents and kits comprising the reagents for carrying out any one of the methods described herein. In some embodiments, kits as provided herein include reagents (e.g., buffers, preservatives, inhibitors, or enzymes) and / or laboratory equipment (e.g., pipettes, filters, tubes, storage containers such as vacutainers, or dissection tools) for storing biological samples obtained from a subject.

[0217] In some embodiments, kits as provided herein include reagents (e.g., buffers, preservatives, inhibitors, or enzymes) and / or labware (e.g., pipettes, filters, or tubes) for extracting RNA and / or DNA from a biological sample or a sample derived from a biological sample (e.g., a single-cell solution). In some embodiments, kits as provided herein include reagents (e.g., buffers, preservatives, inhibitors, enzymes, or dyes) and / or labware (e.g., pipettes, filters, tubes, storage containers, or electrophoresis paper) for measuring the quality and quantity of RNA and / or DNA extracted from a biological sample. In some embodiments, kits as provided herein include reagents (e.g., buffers, preservatives, inhibitors, enzymes, or dyes) and / or labware (e.g., pipettes, filters, tubes, storage containers, or electrophoresis paper) for measuring the quality and quantity of a DNA library for sequencing (e.g., RNA-seq or WES).

[0218] In some embodiments, kits as provided herein include reagents (e.g., buffers, preservatives, inhibitors, or enzymes) and / or laboratory equipment (e.g., pipettes, filters, tubes, storage containers such as vacutainers, or dissection tools) for preparing single-cell solutions from biological samples.

[0219] In some embodiments, kits as provided herein include reagents (e.g., buffers, inhibitors, or enzymes such as reverse transcriptase) and / or labware (e.g., pipettes, filters, tubes, storage containers) for preparing DNA libraries for sequencing.

[0220] In some embodiments, kits as provided herein comprise reagents (e.g., buffers, preservatives, inhibitors, or enzymes) and / or labware (e.g., pipettes, filters, tubes, storage containers such as vacutainers, or dissection tools) for any combination of two or more of the following operations: storing a biological sample; extracting RNA and / or DNA from the biological sample; testing the quality and quantity of the extracted RNA and / or DNA sample and / or DNA library prepared therefrom; preparing a single-cell solution from the biological sample; and preparing a DNA library from the extracted RNA and / or DNA.

[0221] In some embodiments, any one of the kits described herein includes components for making a cell dissociation cocktail. The cell dissociation cocktail may be enzymatic or non-enzymatic. In some embodiments, the kit includes one or more enzyme cocktails. In some embodiments, the kit includes as components one or more of the following: medium (e.g., L-15 medium), antibacterial agent (e.g., penicillin and / or streptomycin), antifungal agent (e.g., amphotericin), collagenase (e.g., collagenase I, collagenase II, collagenase IV), DNAse (e.g., DNAse I), elastase, hyaluronidase, protease (e.g., protease XIV, trypsin, papain, thermolysin). In some embodiments, any one of the kits described herein includes as enzymes one or more of collagenase I and collagenase IV. In some embodiments, these enzymes are packaged in separate containers. In some embodiments, these enzymes are packaged in a single container.

[0222] In some embodiments, the kit includes a smaller instrument such as a spectrophotometer. In some embodiments, the kit includes instructions for storing a biological sample, extracting RNA and / or DNA from the biological sample, testing the quality and quantity of the extracted RNA and / or DNA sample and / or a DNA library prepared therefrom, preparing a single-cell solution from the biological sample, and preparing a DNA library from the extracted RNA and / or DNA, or a combination of any two or more of these. In some embodiments, the kit includes instructions for performing any one of the methods described herein. In some embodiments, the kit is tailored or customized for a specific tissue type, such as a solid tumor biopsy, a liquid biopsy, a blood sample, or urine.

[0223] Data Processing Aspects of the present disclosure relate to processing data obtained from RNA sequencing. In some embodiments, a method for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) includes aligning and annotating genes in the RNA expression data with known sequences of the human genome to obtain annotated RNA expression data, removing non-coding transcripts from the annotated RNA expression data, converting the annotated RNA expression data to gene expression data in transcripts per kilobase million (TPM) format, identifying at least one gene that introduces bias into the gene expression data, and removing the at least one gene from the gene expression data to obtain bias-corrected gene expression data. In some embodiments, the method for processing RNA expression data includes obtaining RNA expression data from a subject having or suspected of having cancer.

[0224] In some embodiments, the non-coding transcript is a pseudogene, a polymorphic pseudogene, a processed pseudogene, a transcribed processed pseudogene, a unitary pseudogene, a non-processed pseudogene, a transcribed unitary pseudogene, a constant chain immunoglobulin (IG C) pseudogene, a joining chain immunoglobulin (IG J) pseudogene, a variable chain immunoglobulin (IG V) gene, a transcribed non-processed gene, a translated non-processed gene, a joining chain T cell receptor (TR J) gene, a variable chain T cell receptor (TR V) gene, a small nuclear RNA (snRNA), a small nucleolar RNA (snoRNA), a microRNA (miRNA), a ribozyme, a ribosomal RNA (rRNA), a mitochondrial tRNA (Mt tRNA), a mitochondrial rRNA (Mt The gene expression profiles include genes belonging to groups selected from the list consisting of rRNA, Cajal body-specific RNA (scaRNA), residual intron, sense intron RNA, sense overlap RNA, nonsense mutation-mediated decay RNA, nonstop decay RNA, antisense RNA, long intervening noncoding RNA (lincRNA), macro-long noncoding RNA (macro-lncRNA), processed transcript, 3' overlapping noncoding RNA (3' overlapping ncrna), small RNA (sRNA), miscellaneous RNA (miscRNA), vault RNA (vaultRNA), and TEC RNA.

[0225] In some embodiments, information (e.g., sequence information) for one or more transcripts for one or more of these types of transcripts can be obtained in a nucleic acid database (e.g., a Gencode database, e.g., Gencode V23, a Genbank database, an EMBL database, or other database).

[0226] In some embodiments, the methods for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) include using the bias-corrected gene expression data to identify a cancer treatment (also referred to herein as an anti-cancer therapy) for a subject. In some embodiments, any one of the methods for processing RNA expression data is further combined with administering one or more anti-cancer or cancer therapies to the subject. In some embodiments, any one of the methods for processing RNA expression data is further combined with prescribing or recommending administration of one or more anti-cancer or cancer therapies to the subject.

[0227] Acquisition of RNA expression data In some embodiments, a method for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) includes obtaining RNA expression data from a subject (e.g., a subject having or diagnosed with cancer). In some embodiments, obtaining the RNA expression data includes obtaining a biological sample and processing it to perform RNA sequencing using any one of the RNA sequencing methods described herein. In some embodiments, the RNA expression data is obtained from a laboratory or center that performed the experiments to obtain the RNA expression data (e.g., the laboratory or center that performed RNA-seq). In some embodiments, the laboratory or center is a clinical laboratory or center.

[0228] In some embodiments, the RNA expression data is obtained by obtaining a computer storage medium (e.g., a data storage drive) on which the data resides. In some embodiments, the RNA expression data is obtained via a secure server (e.g., an SFTP server or Illumina BaseSpace). In some embodiments, the data is obtained in the form of a text-based file (e.g., a FASTQ file). In some embodiments, the file in which the sequencing data is stored also includes a quality score for the sequencing data. In some embodiments, the file in which the sequencing data is stored also includes sequence identifier information.

[0229] Alignment and annotation In some embodiments, a method for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) includes aligning and annotating genes in the RNA expression data with known sequences of the human genome to obtain annotated RNA expression data.

[0230] In some embodiments, aligning the RNA expression data involves aligning the data to a known assembled genome (e.g., the human genome) or transcriptome database for the subject's particular species. A variety of sequence alignment software is available and can be used to align the data to an assembled genome or transcriptome database. Non-limiting examples of alignment software include short (unspliced) aligners (e.g., BLAT, BFAST, Bowtie, Burrows-Wheeler Aligner, Short Oligonucleotide Analysis package, or Mosaik), spliced ​​aligners, aligners based on known splice junctions (e.g., Errange, IsoformEx, or SpliceSeq), or de novo splice aligners (e.g., ABMapper, BBMap, CRAC, or HiSAT). In some embodiments, any suitable tool can be used for data alignment and annotation. For example, Kallisto (github.com / pachterlab / kallisto) is used for data alignment and annotation. In some embodiments, a known genome is referred to as a reference genome. A reference genome (also called a reference assembly) is a digital nucleic acid sequence database assembled as a representative example of a species set of genes. In some embodiments, the human and mouse reference genomes used in any one of the methods described herein are maintained and improved by the Genome Reference Consortium (GRC). Non-limiting examples of human reference releases are GRCh38, GRCh37, NCBI Build 36.1, NCBI Build 35, and NCBI Build 34. Non-limiting examples of transcriptome databases include transcriptome shotgun assemblies (TSA).

[0231] In some embodiments, annotating the RNA expression data includes identifying the locations of genes and / or coding regions in the data to be processed by comparing with an assembled genome or transcriptome database. Non-limiting examples of data sources for annotation include GENCODE (www.gencodegenes.org), RefSeq (see, e.g., www.ncbi.nlm.nih.gov / refseq / ), and Ensembl. In some embodiments, annotating genes in the RNA expression data is based on the GENCODE database (e.g., GENCODE V23 annotation, www.gencodegenes.org).

[0232] Consea et al., "A survey of best practices for RNA-seq data analysis," Genome Biology 201617:13, provides best practices for analyzing RNA-seq data that are applicable to any one of the methods described herein and are incorporated by reference in their entirety. Also, Pereira and Rueda, bioinformatics-core-shared-training.github.io / cruk-bioinf-sschool / Day2 / rnaSeq_align.pdf, describes a method for analyzing RNA sequencing data that is applicable to any one of the methods described herein and is incorporated by reference in its entirety.

[0233] Removal of non-coding transcripts In some embodiments, methods for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) include removing non-coding transcripts from annotated RNA expression data. Aligning and annotating the RNA expression data allows for the identification of coding and non-coding reads. In some embodiments, non-coding reads to transcripts are removed to focus analytical efforts on the expression of proteins (e.g., those that may be involved in cancer pathology). In some embodiments, removing reads to non-coding transcripts from the data reduces data variance, for example, across replicates of the same or similar samples (e.g., nucleic acids from the same cell or cell type).In some embodiments, non-limiting examples of expression data that may be removed include pseudogenes, polymorphic pseudogenes, processed pseudogenes, transcribed processed pseudogenes, unitary pseudogenes, non-processed pseudogenes, transcribed unitary pseudogenes, constant chain immunoglobulin (IG C) pseudogenes, joining chain immunoglobulin (IG J) pseudogenes, variable chain immunoglobulin (IG V) genes, transcribed non-processed genes, translated non-processed genes, joining chain T cell receptor (TR J) genes, variable chain T cell receptor (TR V) genes, small nuclear RNAs (snRNAs), small nucleolar RNAs (snoRNAs), microRNAs (miRNAs), ribozymes, ribosomal RNAs (rRNAs), mitochondrial tRNAs (Mt tRNAs), mitochondrial rRNAs (Mt The non-coding transcripts may include one or more non-coding transcripts (e.g., 10-50, 50-100, 100-1,000, 1,000-2,500, 2,500-5,000 or more non-coding transcripts) belonging to a group selected from the list consisting of rRNA, Cajal body-specific RNA (scaRNA), residual intron, sense-intron RNA, sense-overlapping RNA, nonsense-mediated decay RNA, non-stop decay RNA, antisense RNA, long intervening non-coding RNA (lincRNA), macro-long non-coding RNA (macro-lncRNA), processed transcript, 3' overlapping non-coding RNA (3' overlapping ncrna), small RNA (sRNA), miscellaneous RNA (miscRNA), vault RNA, and TEC RNA.

[0234] In some embodiments, information (e.g., sequence information) for one or more transcripts for one or more of these types of transcripts may be obtained in a nucleic acid database (e.g., the Gencode database, e.g., Gencode V23, the Genbank database, the EMBL database, or other databases). In some embodiments, a portion (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 99.5% or more) of the non-coding transcripts, histone-encoding genes, mitochondrial genes, interleukin-encoding genes, collagen-encoding genes, and / or T-cell receptor-encoding genes described herein are removed from the aligned and annotated RNA expression data.

[0235] Conversion to TPM and gene aggregation In some embodiments, methods for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) include normalizing the RNA expression data with respect to the length of the transcripts read (e.g., into transcripts per kilobase million (TPM) format). In some embodiments, the RNA expression data that has been normalized with respect to transcript length is first aligned and annotated. Converting the data to TPM allows expression to be expressed in terms of concentrations rather than counts, which in turn allows for comparison of samples with different total read counts and / or read lengths.

[0236] In some embodiments, RNA expression data normalized for transcript read length is analyzed to obtain gene expression data (expression data for genes). This is also called gene aggregation. Gene aggregation involves combining expression data for reads for transcripts of all isoforms of a gene to obtain expression data for that gene. In some embodiments, gene aggregation to obtain gene expression data is performed after TPM normalization but before identifying genes that introduce bias. In some embodiments, gene aggregation is performed before converting the data to TPM.

[0237] Wagner et al., Theory Biosci. (2012) 131:281-285, provides a description of how TPM can be calculated, which is incorporated herein by reference in its entirety. In some embodiments, to calculate TPM, the formula

[0238]

number

[0239] is used.

[0240] Removing bias Because converting RNA expression data to obtain expression in TPM format requires dividing the number of reads for a given transcript by the length of the transcript's reads, bias may be introduced into the data for various reasons (as described below). Accordingly, some embodiments of any one of the methods described herein include identifying at least one gene that introduces bias into the gene expression data. Some embodiments of any one of the methods described herein include identifying at least one gene that introduces bias into the gene expression data, and removing the expression data for the at least one gene from the gene expression data to obtain bias-corrected gene expression data.

[0241] In some embodiments, removing data from a dataset may involve deleting the data from the dataset, marking the data so that it is not used in subsequent processing of some or all of the dataset, and / or performing any other suitable processing to prevent the data from being used in subsequent processing of some or all of the dataset. For example, removing particular expression data (e.g., expression data for at least one gene that introduces bias) from gene expression data may involve deleting the particular expression data from the gene expression data, marking the particular expression data, and / or performing any other suitable processing to prevent the particular expression data from being used in subsequent processing of some or all of the gene expression data. As another example, removing non-coding transcripts from RNA expression data (as described above) may involve deleting the non-coding transcripts, marking the non-coding transcripts, and / or performing any other suitable subsequent processing to prevent the non-coding transcripts from being used in subsequent processing of some or all of the RNA expression data. As yet another example, removing sequence data that is determined to fail one or more quality control checks during performance of the quality control techniques described herein may involve deleting the sequence data, marking the sequence data, and / or performing any other suitable processing such that the sequence data that did not pass the quality control checks is not used in some or all further processing.

[0242] In some embodiments, bias in the expression data converted to TPM format is due to transcripts with an average length that is at least a threshold amount higher or lower than the average length of the transcripts as read across the expression dataset. For example, for genes in which one or more transcripts of one or more isoforms have a length that is a threshold amount (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more) lower than the mean or median transcript length across the expression dataset, the expression of the gene in TPM format will appear artificially high. Conversely, for genes where one or more reads of one or more isoforms have a length that is a threshold (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 13, or 15 or more standard deviations) higher than the mean or median read length across the expression dataset, the expression of the gene in TPM format will appear artificially low. In some embodiments, the threshold is set in terms of standard deviations (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 13, or 15 or more standard deviations). In some embodiments, the threshold is set based on transcript length and / or read length, e.g., less than 5 bp, less than 10 bp, less than 15 bp, less than 20 bp, less than 25 bp, less than 50 bp, less than 75 bp, less than 100 bp, or less than 150 bp or more.

[0243] In some embodiments, the bias is due to the length of the poly-A tail on the transcript. In some embodiments, RNA transcripts with poly-A tails that are, on average, smaller or larger than the average length of the poly-A tails for RNA transcripts in the sample have a higher or lower enrichment than the average enrichment of all RNA transcripts in the sample. Thus, genes may be associated with poly-A tails whose lengths are smaller by at least a threshold amount compared to the average length of the poly-A tails of genes from the sample from which the RNA expression data was obtained. In some embodiments, such expression data for such genes is also removed from the gene expression data, thereby obtaining bias-corrected gene expression data. Removing expression data associated with one or more genes from a dataset to reduce bias may be considered a type of data filtering. In some embodiments, "filtering" may refer to one or more of removing expression data for genes that appear artificially high or low (e.g., due to the length of the transcript or the length of the poly-A tail associated with the transcript) and removing expression data for non-coding RNAs from the data.

[0244] In some embodiments, identifying at least one gene that introduces bias into the gene expression data comprises analyzing the length of transcripts in the dataset being analyzed. In some embodiments, removing the expression data of the at least one gene that introduces bias from the gene expression data reduces variability and improves the overall accuracy of subsequent gene expression-based analyses.

[0245] In some embodiments, identifying at least one gene that introduces bias into the gene expression data involves using knowledge gained from analyzing data outside the expression dataset in question, e.g., using a reference dataset. The inventors have recognized that filtering out (the expression data) of genes with poly-A tail lengths that fall outside the average range of poly-A tails in the RNA expression dataset effectively removes bias and / or outliers in the gene expression data. For example, knowledge that a particular gene family introduces bias can be known a priori (from previously performed experiments or previously performed processing of the data) for processing the RNA expression data and used to filter the data for that family of genes.

[0246] In some embodiments, genes that introduce bias into an expression dataset may belong to a family of genes that have poly-A tails that are, on average, small or large relative to the average length of the poly-A tails of genes from the sample from which the RNA expression data was obtained (or another reference sample). In some embodiments, "small or large" may refer to a small or large number relative to a known average threshold for one or more genes.

[0247] In some embodiments, the genes that introduce bias into the expression dataset belong to a gene family selected from the group consisting of histone-encoding genes, mitochondrial genes, interleukin-encoding genes, collagen-encoding genes, B-cell receptor-encoding genes, and T-cell receptor-encoding genes. In some embodiments, the genes that introduce bias into the expression dataset may be any other genes that have, on average, small or large poly-A tails compared to the average length of poly-A tails of genes from the sample from which the RNA expression data was obtained (or another reference sample).

[0248] In some embodiments, the histone-encoding genes, mitochondrial genes, interleukin-encoding genes, collagen-encoding genes, B cell receptor-encoding genes, and / or T cell receptor-encoding genes are genes in the human sample that contain a poly-A tail that is small on average or large on average relative to the average length of the poly-A tails of genes from the sample from which the RNA expression data was obtained. For example, the histone-encoding genes contain a poly-A tail that is small on average relative to the average length of the poly-A tails of genes from the sample from which the RNA expression data was obtained. In some embodiments, the histone-encoding genes do not contain a poly-A tail. In some embodiments, a poly-A tail is minimally or not detected in the histone-encoding genes.

[0249] In some embodiments, abbreviations or acronyms of one or more genes or proteins are used in this application to refer to genes (or protein-encoding genes) using their recognized scientific names. Additional information about genes and / or encoded proteins can be found in one or more gene sequence databases, such as the NIH Gene Sequence Database (GenBank, www.ncbi.nlm.nih.gov), the EMBL database (the European Molecular Biology Laboratory Nucleotide Sequence Database, www.ebi.ac.uk / embl / index.html), the EMBL European Bioinformatics Institute database (EMBL-EBI European Nucleotide Archive, www.ebi.ac.uk / ena), the GENCODE database (www.gencodegenes.org), or other suitable databases, the contents of which are incorporated by reference herein for the different types of genes and gene names referenced herein. In some embodiments, gene or protein abbreviations or acronyms refer to human genes (or human protein-encoding genes).

[0250] In some embodiments, the histone encoding genes are HIST1H1A, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H1T, HIST1H2AA, HIST1H2AB, HIST1H2AC, HIST1H2AD, HIST1H2AE, HIST1H2AG, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H 2AK, HIST1H2AL, HIST1H2AM, HIST1H2BA, HIST1H2BB, HIST1H2BC, HIST1H2BD, HIST1H2BE, HIST1H2BF, HIST 1H2BG, HIST1H2BH, HIST1H2BI, HIST1H2BJ, HIST1H2BK, HIST1H2BL, HIST1H2BM, HIST1H2BN, HIST1H2BO, HIS T1H3A, HIST1H3B, HIST1H3C, HIST1H3D, HIST1H3E, HIST1H3F, HIST1H3G, HIST1H3H, HIST1H3I, HIST1H3J, H IST1H4A, HIST1H4B, HIST1H4C, HIST1H4D, HIST1H4E, HIST1H4F, HIST1H4G, HIST1H4H, HIST1H4I, HIST1H4J, HIST1H4K, HIST1H4L, HIST2H2AA3, HIST2H2AA4, HIST2H2AB, HIST2H2AC, HIST2H2BE, HIST2H2BF, HIST2H3A, HIST2H3C, HIST2H3D, HIST2H3PS2, HIST2H4A, HIST2H4B, HIST3H2A, HIST3H2BB, HIST3H3, or HIST4H4.In some embodiments, the mitochondrial gene is MT-ATP6, MT-ATP8, MT-CO1, MT-CO2, MT-CO3, MT-CYB, MT-ND1, MT-ND2, MT-ND3, MT-ND4, MT-ND4L, MT-ND5, MT-ND6, MT-RNR1, MT-RNR2, MT-TA, MT-TC, MT-TD, MT-TE, MT-TF, MT-TG, MT-TH, MT-TI, MT-TK, MT-TL1, MT-TL2, MT-TM, MT-TN, MT-TP, MT-TQ, MT-TR, MT-TS1, MT-TS2, MT-TT, MT-TV, MT-TW, MT-TY, MTRNR2L1, MTRN R2L10, MTRNR2L11, MTRNR2L12, MTRNR2L13, MTRNR2L3, MTRNR2L4, MTRNR2L5, MTRNR2L6, MTRNR2L7, or MTRNR2L8.

[0251] In some embodiments, removing expression data for at least one gene that introduces bias into the gene expression data comprises removing expression data for one or more (e.g., at least 2, at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, a number between 2 and 1000, or any suitable number of genes within these ranges) genes in each of one or more (2, 3, 4, 5, or all) gene families comprising histone-encoding genes, mitochondrial genes, interleukin-encoding genes, collagen-encoding genes, B-cell receptor-encoding genes, and T-cell receptor-encoding genes. In some embodiments, removing expression data for at least one gene that introduces bias into the gene expression data comprises removing expression data for one or more genes that have either a small or large poly-A tail on average compared to the average length of the poly-A tails of genes from the sample from which the RNA expression data was obtained (or a reference sample).

[0252] In some embodiments, after expression data for at least one gene that introduces bias has been removed from the gene expression data, the remaining gene expression data may be normalized again ("re-normalized") (e.g., to TPM or any other suitable units, such as Reads Per Kilobase Million (RPKM) or Fragments Per Kilobase Million (FPKM)) so that the normalized expression values ​​are not biased by the expression data of the removed bias gene. In some embodiments, the expression data for the remaining genes may include expression data for at least 1,000 genes, at least 5,000 genes, at least 10,000 genes, between 500 and 5,000 genes, between 1,000 and 10,000 genes, between 5,000 and 15,000 genes, or any suitable number of genes within these ranges.

[0253] Post-sequencing nucleic acid data quality control As presented in this disclosure, quality control is performed periodically during the sample preparation process. For example, the purity of extracted nucleic acids or the size distribution of DNA libraries is detected. If one or more quality control issues occur and cannot be corrected in the laboratory, the provider of the biological sample (e.g., a medical service provider) is notified before proceeding to subsequent steps. After the quality issues are resolved, the sample preparation process is completed and bioinformatics analysis (e.g., post-sequencing processing) is performed.

[0254] Embodiments of the methods and systems described herein provide for quality control to be performed on gene expression data to improve the accuracy and reliability of subsequent expression analysis (e.g., to determine a patient or subject's diagnosis, prognosis, and / or treatment) and resulting recommendations.

[0255] In some embodiments, bioinformatics quality control of sequence data can be performed as a standalone process (e.g., based on nucleic acid data received from a healthcare provider) or in conjunction with a prior sample preparation process (e.g., when a patient sample is provided by a healthcare provider as opposed to nucleic acid sequence data). As illustrated in FIG. 7 , activities 301 through 310 illustrate a non-limiting sample preparation process as described in this disclosure, while activities 311 through 315 illustrate a non-limiting quality control process as described in this disclosure. In some embodiments, one or more of activities 301 through 310 can be performed independently (e.g., without one or more of activities 311 through 315). In some cases, one or more of activities 301 through 310 can be skipped or delayed. Activities 311 through 315 can be performed independently (e.g., without activities 301 through 310). In some cases, one or more of activities 311 through 315 may be skipped or delayed. In some cases, one or more of the sample preparation (activities 301 through 310) and quality control (activities 311 through 315) processes may both be performed. In some cases, one or more of the sample preparation processes and one or more of the quality control processes may be performed.

[0256] In some embodiments, process pipeline 300 includes obtaining a first tumor sample from a subject having, suspected of having, or at risk of having cancer in activity 301; extracting RNA from the first sample of the first tumor in activity 302; enriching the extracted RNA for coding RNA to obtain enriched RNA in activity 303; preparing a first library of cDNA fragments from the enriched RNA for non-stranded RNA sequencing in activity 304; obtaining RNA expression data for the subject having, suspected of having, or at risk of having cancer in activity 305; aligning and annotating genes in the RNA expression data with known sequences of the human genome to obtain annotated RNA expression data in activity 306; removing non-coding transcripts from the annotated RNA expression data in activity 307; and analyzing the annotated RNA expression data in activity 308 based on Transcripts Per Kilobase. In activity 309, identifying at least one gene that introduces bias into the gene expression data; in activity 310, removing at least one gene from the gene expression data to obtain bias-corrected gene expression data; in activity 311, obtaining sequence information and claimed information; in activity 312, determining one or more features from the sequence information; in activity 313, determining whether the one or more features match the claimed information; in activity 314, making an additional determination of at least one of the features; and in activity 315, identifying a cancer treatment for the subject using the bias-corrected gene expression data.

[0257] In some embodiments, activity 305 may include obtaining RNA expression data by using a sequencing platform or by receiving it from a healthcare provider or laboratory. In some embodiments, activity 306 may include converting the RNA expression data into gene expression data. As described herein, "known sequence of the human genome" may refer to a reference. In some embodiments, activity 307 may include converting the RNA expression data into gene expression data. In some embodiments, activity 307 may include obtaining filtered RNA expression data. In some embodiments, activity 308 may include normalizing the filtered RNA expression data to obtain gene expression data in terms of transcripts per kilobase million (TPM). In some embodiments, the asserted information of activity 311 may indicate the asserted source and / or asserted completeness of the sequence data. In some embodiments, activity 312 may include determining one or more disease characteristics. In some embodiments, activity 312 may include processing the sequence information or data to obtain determined information indicating the determined source and / or determined completeness of the sequence information or data. In some embodiments, activity 313 may include determining whether the determined information matches the asserted information. In some embodiments, the at least one additional determination of characteristics in process 314 may include determining a disease characteristic or a characteristic not directly related to a disease.

[0258] Aspects of the methods and systems described herein provide approaches for validating nucleic acid sequence data by obtaining both the sequence data and claimed information related to one or more features of the sequence data (e.g., source, nucleic acid type, expected completeness, etc.), determining one or more features from the sequence data, and verifying that the one or more features determined from the sequence data match the claimed information related to those features. In some embodiments, the claimed information may be information about the patient, tissue type, tumor type, nucleic acid type (RNA, DNA, WES, polyA, etc.), sequencing protocol used, etc., or combinations thereof. In some embodiments, the claimed information may be an expected and / or acceptable (e.g., acceptable for subsequent analysis of the sequence data) completeness threshold of the sequence information, including, for example, expected and / or acceptable levels of GC content, contamination, coverage (e.g., genomic, exome, exon, protein-coding, or other coverage), or other measures of completeness.

[0259] Nucleic acid sequencing, particularly next-generation sequencing (NGS), allows for the generation of large amounts of information about a given nucleic acid (such as DNA, RNA, genome, exome, or transcriptome). However, due to the large number of different sequencing platforms available, the wide variety of sample preparation and sequencing protocols and techniques used, and the variability and inconsistency between platforms and protocols, the content and coverage of the resulting nucleic acid sequencing information varies substantially. Furthermore, when evaluating large sets of sequence information from several sequencing runs, or from multiple sequencing runs (e.g., including historical data from different medical visits for one or more patients), or from different studies (e.g., studies to generate prognostic or diagnostic assessments, or studies to evaluate the effects of drugs or treatments on disease progression), combining sequence information from different sources can be difficult. In addition, it can be difficult to detect misidentified sequence data when large amounts of information from different sources are combined.

[0260] Currently, there are no robust methods for validating (e.g., increasing confidence, reducing uncertainty, correcting for or omitting low-quality sequence information, providing a signal to verify or re-examine questionable sequence information or outliers, etc.) the source and / or completeness (e.g., which may also be referred to herein as quality) of sequence information that may be subject to further use (e.g., being used for analysis beyond the initial sequencing step), e.g., for diagnostic, prognostic, and / or clinical applications.

[0261] This disclosure recognizes the proliferation of next-generation sequencing technologies and platforms across various sectors of the scientific community. This disclosure also recognizes the various protocols and methodologies associated with the different technologies and platforms employed. Variations in platforms and protocols for using different platforms result in variability in the data and sequence information realized from their use, which poses a significant hurdle to using the sequence information for substantive analysis, particularly when such sequence information is used for analysis beyond the original data performed by the original user of the sample (e.g., by a secondary user beyond the user who procured and performed the original sequencing, a third party to the sequencing, etc.).

[0262] Accordingly, the present disclosure presents various methods and processes for assessing the quality of sequence information (e.g., for correct identification of the sequence information, sample identification, subject identification, etc.) and even for assessing the integrity of the sequence information (e.g., creating checkpoints to screen for various integrity issues, e.g., contamination or degradation). For example, in some embodiments, described herein are methods for assessing sequence information by obtaining sequence information from nucleic acids of a subject's sample, obtaining claimed information, determining characteristics of the sequence information (e.g., source, identity, status, characteristics), and comparing the claimed information to the determined information. Sequence information can be obtained (e.g., acquired) from any source or through any means known in the art. Thus, sequence information can be generated using any suitable sequencing technology. Alternatively, sequence information can be obtained electronically from a third party that generated the sequence information. In some embodiments, sequence information (e.g., reference sequence information) is obtained from an existing databank of sequences. In some embodiments, sequence information is obtained from a company, non-profit organization, academic institution, or medical institution.

[0263] In some embodiments, the sample may be any specimen, biopsy, or biological component obtained (e.g., procured, collected, received) from a subject. For example, in some embodiments, the sample may be a blood sample, a hair sample, a tissue sample, a bodily fluid sample, a cell sample, a blood component sample, or any other cell or tissue sample from which nucleic acids can be obtained for sequencing.

[0264] In some embodiments, the subject may be any living organism in need of treatment or diagnosis using the methods or systems of the present disclosure. For example, without limitation, the subject may include mammals and non-mammals. As used herein, "mammal" refers to any animal that comprises the class Mammalia (e.g., human, mouse, rat, cat, dog, sheep, rabbit, horse, cow, goat, pig, guinea pig, hamster, chicken, turkey, or non-human primate (e.g., marmoset, macaque)). In some embodiments, the mammal is a human. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human.

[0265] In some embodiments, the sample may be a biological sample obtained from a subject, e.g., a patient. In some embodiments, the sample may be blood, serum, sputum, urine, or a tissue biopsy (e.g., from any tissue including, but not limited to, the heart, liver, pancreas, CNS, gastrointestinal tract, mouth, colon, kidney, and skin). In some embodiments, the sample may be suspected to be a disease sample (e.g., a cancer sample). In some embodiments, the sample may be a healthy sample (e.g., used as a reference).

[0266] In some embodiments, the sequence information is obtained from a next-generation sequencing platform (e.g., Illumina™, Roche™, Ion Torrent™, etc.), or any high-throughput or massively parallel sequencing platform. In some embodiments, these methods may be automated, and in some embodiments, there may be manual intervention. In some embodiments, the sequence information may be the result of non-next-generation sequencing (e.g., Sanger sequencing). In some embodiments, sample preparation may be according to a manufacturer's protocol. In some embodiments, sample preparation may be a custom protocol or other protocol for research, diagnostic, prognostic, and / or clinical purposes. In some embodiments, the protocol may be experimental. In some embodiments, the origin or preparation method of the sequence information may be unknown.

[0267] In some embodiments, the size of the obtained RNA and / or DNA sequence data comprises at least 5 kilobases (kb). In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 10 kb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 100 kb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 500 kb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 1 megabase (Mb). In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 10 Mb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 100 Mb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 500 Mb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 1 gigabase (Gb). In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 10 Gb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 100 Gb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 500 Gb.

[0268] In some embodiments, sequence information can be generated using nucleic acids from a sample from a subject. In some embodiments, the sequence information can be sequence data representing the nucleotide sequence of DNA and / or RNA from a previously obtained biological sample from a subject with, suspected of, or at risk of having a disease. In some embodiments, the nucleic acid is deoxyribonucleic acid (DNA). In some embodiments, the nucleic acid is prepared so that the entire genome is present in the nucleic acid. In some embodiments, the nucleic acid is processed so that only the protein-coding region of the genome remains (e.g., the exome). When the nucleic acid is prepared so that only the exome is sequenced, this is called whole exome sequencing (WES). In various methods for isolating exomes for sequencing, known in the art, for example, solution-based separation, labeled probes are used to hybridize target regions (e.g., the exome), which can then be further separated from other regions (e.g., unbound oligonucleotides). These labeled fragments can then be prepared and sequenced.

[0269] In some embodiments, the nucleic acid is ribonucleic acid (RNA). In some embodiments, the sequenced RNA includes both coding and non-coding transcribed RNA found in the sample. When such RNA is used for sequencing, the sequencing is said to be generated from "total RNA" and is sometimes referred to as whole transcriptome sequencing. Alternatively, nucleic acids can be prepared such that coding RNA (e.g., mRNA) is isolated and used for sequencing. This can be done through any means known in the art, for example, by isolating or screening RNA for polyadenylated sequences. This is sometimes referred to as mRNA-Seq.

[0270] Sequence information can include sequence data generated by a nucleic acid sequencing protocol (e.g., a series of nucleotides in a nucleic acid molecule identified by next generation sequencing, Sanger sequencing, etc.), as well as information contained therein (e.g., information indicating source, tissue type, etc.), which can also be considered information that can be inferred or determined from the sequence data. For example, in some embodiments, RNA sequence information can be analyzed to determine whether the nucleic acid is predominantly polyadenylated.

[0271] The claimed information may refer to the sequence data and, therefore, information about the nucleic acid, sample, and / or subject from which the sequence data was obtained. In some embodiments, the claimed information is provided along with the sequence data and can be verified by analyzing the sequence data as described herein. The claimed information may relate to characteristics of the nucleic acid, sample, or subject and can be used to assess the quality of the nucleic acid (e.g., the source or integrity of the nucleic acid). The claimed information can refer to the claimed source and / or claimed completeness of the sequence data or information.

[0272] In some embodiments, a third party may provide the sequence data as well as the associated claimed information. In some embodiments, the claimed information is obtained from the same entity from which the sequencing data was obtained. In some embodiments, the claimed information and the sequencing data are obtained from different parties. In some embodiments, the claimed information is obtained from a database. In some embodiments, the claimed information is a reference value or property. In some embodiments, the claimed information may assert the identity of the sequence information, the identity of the nucleic acid(s) in the sequence information, the identity of the sample from which the sequence information was generated, or the identity of the subject from whom the sample was obtained. In some embodiments, the claimed information may identify the sequence data as having been obtained from polyadenylated RNA, as derived from whole transcriptome sequencing, or from WES. In some embodiments, the claimed information may identify the cell or tissue type for the sample from which the nucleic acid was obtained. In some embodiments, the claimed information may assert the tumor type for the sample from which the nucleic acid was obtained. In some embodiments, the claimed information may identify the MHC profile for the subject from which the sample was obtained (e.g., the sequence of MHC alleles of the subject from whom the nucleic acid was obtained). In some embodiments, the claimed information may identify an expected protein subunit ratio for the sample. In some embodiments, the claimed information may provide an expected complexity value for the sequence information. In some embodiments, the claimed information may provide an expected contamination value for the sequence information. In some embodiments, the claimed information may provide an expected coverage value for the sequence information. In some embodiments, the claimed information may provide an expected exon coverage value for the sequence information. In some embodiments, the claimed information may provide an expected read composition value for the sequence information. In some embodiments, the claimed information may provide an expected Phred score for the sequence information.In some embodiments, the claimed information may provide predicted single nucleotide polymorphism (SNP) values ​​for the sequence information. In some embodiments, the claimed information may relate to GC content values ​​for the sequence information. In some embodiments, the claimed information may include additional information. In some embodiments, the claimed information may include information relating to multiple or more than one features of the sequence information. In some embodiments, the claimed information is any combination of the aforementioned features (e.g., determined values, properties, characteristics, etc.).

[0273] As used herein, a "feature" may be a property or characteristic determined from an analysis of sequence information that goes beyond the sequence of nucleotides in the sequence information and provides a user with information about the sequence information, the sample from which the information was derived, and / or the subject from whom the sample was obtained. The sequence information may be associated with gene expression data obtained from a healthcare provider or laboratory. For example, a feature may indicate the source (e.g., patient, subject, nucleic acid type), patient or subject identity, tissue type, tumor type, polyadenylation status, MHC sequence, protein subunit, complexity, contamination, coverage (e.g., full sequence, exons, etc.), read composition, quality and / or Phred score, location of single nucleotide polymorphisms (SNPs), and / or GC content. Features of sequence information may indicate whether the sequence information is potentially consistent or inconsistent with claimed information from a healthcare provider or laboratory.

[0274] When considering identity or source, it is important to recognize that this term not only refers to the identification of a particular subject or patient as a particular individual, but also that one or more features of sequence information for a sample can be identified as being the same as one or more features of sequence information obtained from another sample. For example, sequence information A can be compared to sequence information B, which is presented and claimed to be from the same nucleic acid, subject or patient, tissue, or tumor. Identity can be confirmed or challenged by the methods herein without knowing the actual identity of the subject, but the finding that identity is consistent with another given sequence information can be supported. In some embodiments, the identity of sequence information is used to compare claimed information for a given sample, subject, tissue, or tumor. In some embodiments, the identity of sequence information is used to compare claimed information for another nucleic acid or reference value.

[0275] In some embodiments, these determined features of the sequence information are then evaluated (e.g., determined, matched, aligned, measured, assessed) against the claimed information. This evaluation can be done to increase confidence that the sequence information is of a particular origin (e.g., source), has been correctly identified, or has particular or specific characteristics (e.g., is from a polyadenylated nucleic acid). In this regard, these methods can be used to provide checkpoints and means for highlighting potential problems (e.g., discordant values ​​(e.g., for the determined features and claimed information), or determined values ​​that fall outside accepted or established ranges). Such problems may indicate or suggest problems with the integrity (e.g., corrupted, contaminated) or source (e.g., misidentified, mislabeled, etc.) of the sequence data. Using the methods and processes herein to match determined features with the claimed information for given sequence information reduces the likelihood that fraudulent or low-quality sequence data will be used in an analysis and increases confidence that the sequence data is of sufficient quality to be used for diagnostic, prognostic, and / or clinical analysis.

[0276] In some embodiments, evaluating whether the determined information matches the asserted information involves determining whether the determined information matches the asserted information exactly or within a specified threshold. More generally, in some embodiments, evaluating whether two values ​​"match" may involve determining whether the two values ​​match exactly or within a specified threshold. The threshold may be 0, which in some embodiments requires an exact match. The threshold may be greater than 0, such that when numerical values ​​are being compared, the values ​​are said to "match" if they are within the threshold of one another (e.g., when the absolute difference between the numerical values ​​is less than or equal to the threshold). In some embodiments, the threshold may be set as a function of a standard deviation (or a multiple thereof), a quantile, a percentile, or any other suitable statistical measure. In some embodiments, evaluating whether two values ​​"match" may involve determining, when there is a difference between the two values, whether the difference is statistically significant. Such a determination may be performed using statistical hypothesis testing, thresholds, or any other suitable statistical or mathematical technique, as aspects of the technology described herein are not limited in this respect.

[0277] In some embodiments, one or more quality control parameters are checked against the bioinformatics data. In some embodiments, tumor purity may be checked. Tumor purity, as described herein, may refer to the percentage of cancer cells in the admixture. In some embodiments, the target tumor purity for WES is ≧20% (e.g., 20%, 40%, 60%). In some embodiments, the target tumor purity for RNA-seq is ≧20% (e.g., 20%, 40%, 60%).

[0278] In some embodiments, the depth of coverage can be checked. In some embodiments, the depth of coverage for WES is ≧150× average coverage of tumor samples (e.g., 150×, 180×, 200×). In some embodiments, the target depth of coverage for RNA-seq is ≧100× (e.g., 100×, 150×, 200×).

[0279] In some embodiments, the alignment rate may be checked. In some embodiments, the target alignment rate for WES is greater than 90% (e.g., 91%, 95%, 99%). In some embodiments, the target alignment rate for RNA-seq is greater than 90% (e.g., 91%, 95%, 99%).

[0280] In some embodiments, base call quality scores, such as Phred scores, can be checked. In some embodiments, the target Phred score for WES is greater than 30 (e.g., 35, 40, 50). In some embodiments, the target Phred score for RNA-seq is greater than 30 (e.g., 35, 40, 50).

[0281] In some embodiments, coverage uniformity can be checked. In some embodiments, the target coverage uniformity for WES is 85% of base pairs in the target region being covered at ≥ 20-fold for tumor tissue (e.g., 85%, 95%, 99%). In some embodiments, the target coverage uniformity for WES is 85% of base pairs in the target region being covered at ≥ 20-fold for normal tissue (e.g., 85%, 95%, 99%). In some embodiments, the target region for determining coverage uniformity can be the ExonV7 target region using the coding region from the CCDS (consensus coding sequence) gene.

[0282] In some embodiments, GC bias may be checked. In some embodiments, the target GC bias for WES is at least 50 (e.g., 50, 60, 70). In some embodiments, the acceptable range of GC bias for WES is at least 45-65 (e.g., 45-65, 50-65, 55-65). In some embodiments, the target GC bias for RNA-seq is at least 50 (e.g., 50, 60, 70). In some embodiments, the acceptable range of GC bias for RAN-seq is at least 45-65 (e.g., 45-65, 50-65, 55-65).

[0283] In some embodiments, the mapping quality may be checked. In some embodiments, the mapping quality for the WES is ≧10 (e.g., 10, 20, 30).

[0284] In some embodiments, the overlap rate can be checked. In some embodiments, the overlap rate for WES is less than 30% (e.g., 29.9%, 25%, 15%). In some embodiments, the overlap rate for RAN-seq is less than 85% (e.g., 84.99%, 80%, 70%).

[0285] In some embodiments, insert size may be checked. In some embodiments, the acceptable median insert size for tumor tissue for WES is about 150 (e.g., 150, 280, 250). In some embodiments, the target median insert size for tumor tissue for WES is about 200 (e.g., 200, 250, 350). In some embodiments, the acceptable median insert size for normal tissue for WES is about 150 (e.g., 150, 280, 250). In some embodiments, the target median insert size for normal tissue for WES is about 200 (e.g., 200, 250, 350). In some embodiments, the acceptable median insert size for tumor tissue for RNA-seq is about 150 (e.g., 150, 280, 250). In some embodiments, the target median insert size for tumor tissue for RNA-seq is about 200 (e.g., 200, 250, 350).

[0286] In some embodiments, contamination can be checked. In some embodiments, acceptable contamination for WES is less than 0.05% (e.g., 0.04%, 0.03%, 0.01%). In some embodiments, acceptable contamination for RNA-seq is less than 0.05% (e.g., 0.04%, 0.03%, 0.01%).

[0287] In some embodiments, SNP concordance between pairs of tumor and normal samples from the same patient may be checked. In some embodiments, the target SNP concordance for WES is greater than 90% (e.g., 91%, 95%, 98%). In some embodiments, the acceptable SNP concordance for WES is greater than 85% (e.g., 86%, 90%, 98%). In some embodiments, the target SNP concordance for RNA-seq is greater than 90% (e.g., 91%, 95%, 98%). In some embodiments, the acceptable SNP concordance for RNA-seq is greater than 85% (e.g., 86%, 90%, 98%).

[0288] In some embodiments, HLA allele matching of pairs of tumor and normal samples from the same patient can be checked. In some embodiments, the threshold for normal versus tumor tissue for WES is less than 5 (e.g., 4.5, 3, 2.5). In some embodiments, the threshold for tumor RNA-seq tumor versus normal WES tissue for RNA-seq is less than 5 (e.g., 4.5, 3, 2.5).

[0289] In some embodiments, sequence information can be evaluated for genomic contamination (e.g., non-human genomic contamination). In some embodiments, samples or sequence information are evaluated to determine whether they are contaminated by determining whether they contain sequences from other species or reference genomes, such as mouse, zebrafish, Drosophila, Caenorhabditis elegans, Saccharomyces, Arabidopsis, microbiome, mycoplasma, adapter, UniVec, and phiX rRNA. In some embodiments, the target threshold for ADA genomic contamination for WES is greater than 60 (e.g., 65, 70, 80). In some embodiments, the acceptable threshold for ADA genomic contamination for WES is greater than 40 (e.g., 45, 60, 80). In some embodiments, the target threshold for ADA genomic contamination for RNA-seq is greater than 40 (e.g., 50, 60, 80). In some embodiments, the acceptable threshold for ADA genomic contamination for RNA-seq is greater than 20 (e.g., 30, 50, 70).

[0290] In some embodiments, only one feature is evaluated against the claimed information. In some embodiments, multiple features are evaluated against the claimed information. In some embodiments, at least two or more features are evaluated against the claimed information. In some embodiments, at least three or more features are evaluated against the claimed information. In some embodiments, at least four or more features are evaluated against the claimed information. In some embodiments, at least five or more features are evaluated against the claimed information. In some embodiments, at least six or more features are evaluated against the claimed information. In some embodiments, at least seven or more features are evaluated against the claimed information. In some embodiments, at least eight or more features are evaluated against the claimed information. In some embodiments, at least nine or more features are evaluated against the claimed information. In some embodiments, at least ten or more features are evaluated against the claimed information. In some embodiments, at least eleven or more features are evaluated against the claimed information. In some embodiments, at least 12 or more features are evaluated against the claimed information. In some embodiments, at least 13 or more features are evaluated against the claimed information. In some embodiments, at least 14 or more features are evaluated against the claimed information. In some embodiments, at least 15 or more features are evaluated against the claimed information.

[0291] In some embodiments, if a feature or a determined value is found to be inconsistent with or inconsistent with the claimed information, additional steps are performed. In some embodiments, if a feature or a determined value is found to be inconsistent with or inconsistent with the claimed information, the sequence information is rejected (e.g., not used in subsequent analysis). In some embodiments, if a feature or a determined value is found to be inconsistent with or inconsistent with the claimed information, the sequence information is re-examined, i.e., any evaluation of the feature or determination is performed at least once more, or a second or subsequent time (e.g., a third, fourth, fifth, sixth, etc.). In some embodiments, if a feature or a determined value is found to be inconsistent with or inconsistent with the claimed information, another, second, or subsequent time (e.g., a third, fourth, fifth, sixth, etc.) of sequence information is obtained and then examined, i.e., any evaluation of the feature or determination is performed at least once more, or a second or subsequent time (e.g., a third, fourth, fifth, sixth, etc.) regardless of the initial determination and evaluation performed on the first sequence information. In some embodiments, if a feature or determined value is found to not match or correspond to the claimed information, the sequence information is reported as such to the user. In some embodiments, if a feature or determined value is found to not match or correspond to the claimed information, any combination of these steps may be performed.

[0292] In some embodiments, if a feature or determined value is found to not fit or match the claimed information, the sequence information may still be evaluated for characteristics related to the disease (e.g., cancer), but information regarding the quality (e.g., the extent and nature of one or more features of the determined sequence information that do not match the claimed information) may be provided to a user (e.g., a physician or other healthcare professional). In some embodiments, the characteristic relates to the type of cancer, its environment, its stage, its location, its tissue of origin, its statistical likelihood of responding to various treatments or therapies, or other properties that may aid a practitioner in treating a subject.

[0293] In some embodiments, if the features or determined values ​​are found to match or correspond to the claimed information (e.g., match, exceed, or otherwise satisfy a reference value or threshold), additional steps may be performed. In some embodiments, if the features or determined values ​​are found to match or correspond to the claimed information (e.g., match, exceed, or otherwise satisfy a reference or threshold), additional steps may be performed. In some embodiments, if the features or determined values ​​are found to match or correspond to the claimed information (e.g., match, exceed, or otherwise satisfy a reference or threshold), the sequence information is evaluated for characteristics related to cancer. In some embodiments, the characteristics relate to the type of cancer, its environment, its stage, its location, its tissue of origin, its statistical likelihood of responding to various treatments or therapies, or other properties that may aid a practitioner in treating a subject.

[0294] In some embodiments, after one or more quality control steps are performed, a report is generated for the user along with the results of the quality control steps that were performed.

[0295] Therefore, in one aspect, the present disclosure relates to a method for evaluating the sequence information of at least one nucleic acid to determine at least one characteristic thereof.At least one characteristic can be used to evaluate the quality or completeness of sequence information, to interrogate the source of sequence information, or to enable the analysis of other sequence information, which may or may not be from the same sequencing platform or from the same or different sample preparation protocol.Furthermore, at least one characteristic can be used as a quality control measure to ensure that threshold quality and low-quality sequence information are omitted from subsequent analysis.

[0296] Accordingly, in one aspect, the present disclosure relates to a method of evaluating sequence information by: (a) obtaining sequence information, the sequence information comprising (1) sequence data from a first ribonucleic acid (RNA) or (2) sequence data from a first whole exome sequence (WES); and (b) determining one or more characteristics of the sequence data, the sequence data being selected from the group consisting of: (i) the identity of the subject from whom the nucleic acid was obtained; (ii) the tissue of origin from which the nucleic acid was obtained; (iii) the tumor type from which the nucleic acid was obtained; (iv) a quality measure of the first RNA-seq data; (v) whether the RNA-seq data was obtained from polyadenylated (polyA) RNA or total RNA; (vi) whether the first sequence dataset is first WES sequence data; (vii) the sequencing platform used to generate the first sequence dataset; and (viii) a quality measure of the first sequence dataset.

[0297] In some embodiments, the method further comprises obtaining additional sequence information if one or more features of the sequence information are below a quality control threshold suitable for further analysis.

[0298] In some embodiments, the characteristic being evaluated is the identity of the subject. In some embodiments, the identity of the subject is determined by performing one or more of the following assessments: a major histocompatibility complex assessment and a SNP match assessment; and the results of the assessment are compared to the claimed value for the subject or a second sequence dataset from the subject.

[0299] In some embodiments, the feature being evaluated is the tissue of origin. In some embodiments, the tissue of origin is determined by performing one or more of the evaluations from a group including protein expression and biomarker analysis. In another aspect, the present disclosure relates to a method for evaluating a feature, which includes assigning the tissue of origin of the sample from which the sequence information was generated. In some embodiments, the method includes evaluating the sequence information for marker or gene expression indicative of the tissue type from which the sequence information was generated. In some embodiments, the method includes evaluating the marker or gene expression against a database that is the same for different tissue types. Different tissues throughout a subject's body express different proteins, which create a profile of such tissue. Therefore, by evaluating the protein expression profile and matching it to the tissue type, it is possible to identify the sample, and therefore the tissue from which the sequence information was obtained. This can be done through various methods known in the art. For example, assessing the number of a given messenger RNA (mRNA) transcript (e.g., used as a surrogate for assessing protein expression) can be assessed against a database of known tissue markers (e.g., protein expression profiles), can be assessed against a provided set of markers for the subject, or can be assessed against second sequence information or a set of tissue markers obtained from the subject. In some embodiments, the tissue of origin is determined by assessing sequence information for the markers (e.g., protein expression) and matching the markers against a database of tissues. In some embodiments, the tissue of origin is determined by assessing sequence information for the markers (e.g., protein expression) and matching the markers against a set of markers from the subject's tissues. In some embodiments, the tissue of origin is determined by assessing sequence information for the markers (e.g., protein expression) and matching the markers against second sequence information obtained from a subject whose tissue of origin is known.

[0300] In some embodiments, the feature evaluated is a measure of the completeness of sequence information. In some embodiments, the completeness measure of the first RNA-seq data is determined by performing one or more evaluations from a group including: determining the coverage of one or more genes in the RNA-seq data; determining the relative coverage of two or more exons for at least one gene in the RNA-seq data; determining the expression ratio of two known reference genes from the RNA-seq data; or other features, or a combination of two or more of them. In some embodiments, the completeness measure of DNA-seq data is determined by performing one or more evaluations from a group including the total coverage and / or chromosome coverage of the DNA-seq data; or other features, or a combination of two or more of them.

[0301] In some embodiments, the RNA-seq data is analyzed to determine whether it was obtained from polyA RNA or total RNA. In some embodiments, the RNA-seq data is analyzed by assessing the expression levels of one or more mitochondrial or histone genes from the RNA-seq data, and / or other characteristics typical of polyA or total RNA.

[0302] In some embodiments, the feature evaluated is the sequencing platform used to generate the sequence. In some embodiments, the sequencing platform used to generate the WES sequence data is determined by performing one or more evaluations from a group including determining the percent variance of one or more reference genes in the WES sequence data, or other properties of the sequencing data typical of the sequencing platform used to generate the sequence data.

[0303] In some embodiments, the method includes evaluating at least one of the features described herein. In some embodiments, the method includes evaluating at least two of the features described herein. In some embodiments, the method includes evaluating at least three of the features described herein. In some embodiments, the method includes evaluating at least four of the features described herein. In some embodiments, the method includes evaluating at least five of the features described herein. In some embodiments, the method includes evaluating at least six of the features described herein. In some embodiments, the method includes evaluating at least seven of the features described herein.

[0304] In some embodiments, the quality (e.g., source or integrity) of sequence information from one or more nucleic acid samples (e.g., at least two nucleic acid samples) is evaluated by: (a) determining the sequences of two or more (e.g., 2, 3, 4, 5, 6 or more) major histocompatibility complex (MHC); and (b) determining whether the MHCs from one or more samples match. In some embodiments, if the MHCs do not match (e.g., if the calculated match value is less than a statistically significant threshold), the sequence information from each of the nucleic acids is considered likely to be from different sources, of insufficient quality, removed, discarded, retested, and / or reported to the user as such. In some embodiments, if the calculated match value (x) between WES normal / tumor / RNAseq is 0 < x ≦ 2 (e.g., 1, 1.5, 2), this represents "warning" as acceptable. A warning means that the calculated match value is within the range considered acceptable but is thought to be close to unacceptable. In some embodiments, if the calculated match value (x) between WES normal / tumor / RNAseq is > 5, this represents unacceptable or poor quality. In some embodiments, if the calculated match value (x) between WES normal / tumor / RNAseq is 0, this represents good quality. In some embodiments, if the MHCs match (e.g., if the match value is greater than or equal to a statistically significant threshold), the sequence information from each of the nucleic acid samples is considered likely to be from the same source, of sufficient quality, retained for further analysis, and / or reported to the user as such.

[0305] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is assessed by determining a match value for single nucleotide polymorphisms (SNPs) within the sequence information. In some embodiments, the method further comprises assessing the match value. In some embodiments, if the match value is less than 85%, less than 80%, or less than 75%, the sequence information from each of the nucleic acid samples is deemed to be of insufficient quality, likely to be from a different source, and is removed, discarded, retested, and / or reported to a user as such. In some embodiments, if the match value is less than 75%, the sequence information is deemed unacceptable. In some embodiments, if the match value is greater than 80% but less than 95%, the sequence information is deemed to be in a range close to being unacceptable. In some embodiments, if the match value is greater than 95%, the sequence information is deemed acceptable. In some embodiments, if the match value is at least 75%, at least 80%, or at least 85%, the sequence information from each of the nucleic acid samples is deemed to be of sufficient quality, likely to be from the same source, and is retained and / or reported to a user as such. In some embodiments, at least 5,000 SNPs may be evaluated for concordance values. In some embodiments, at least 6,000 SNPs may be evaluated for concordance values. In some embodiments, at least 7,000 SNPs may be evaluated for concordance values. In some embodiments, at least 8,000 SNPs may be evaluated for concordance values.

[0306] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is assessed by determining a contamination value for the sequence information. In some embodiments, if the contamination value exceeds a statistically significant threshold, the sequence information is removed, discarded, retested, and / or reported as such to the user. In some embodiments, if the contamination value is greater than 0.05% (e.g., 0.06%, 1%, 2%), the sequence information is deemed close to unacceptable (e.g., a warning). In some embodiments, if the contamination value is greater than 0.1% (e.g., 0.1%, 0.5%, 1%), the sequence information is deemed unacceptable for blood samples and fresh frozen tissues. In some embodiments, if the contamination value is less than the threshold, the sequence information is retained and / or reported as such to the user.

[0307] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is assessed by analyzing the sequence information from one or more nucleic acid samples against a set of tumor types, determining a predicted tumor type from the sequence information, and determining whether the predicted tumor type matches the tumor type provided (e.g., asserted) for one or more nucleic acid samples. In some embodiments, determining the predicted tumor type as a quality control step may be performed using a computerized system or process as described herein. In some embodiments, determining the predicted tumor type as a quality control step may be performed by using machine learning techniques to determine cancer grade from sequence data, as described herein and in U.S. Provisional Patent Application No. 62 / 943,976, filed December 5, 2019, entitled "Machine Learning Techniques for Gene Expression Analysis," which is incorporated herein by reference in its entirety. In some embodiments, if there is a discrepancy between the tumor type (e.g., cancer grade) obtained from sequence assessment and the claimed information, the sequence information is identified as being of suspect or insufficient quality and is removed, discarded, retested, and / or reported as such to the user. In some embodiments, if there is a concordance between the predicted tumor type and the expected tumor type for one or more nucleic acid samples, the sequence information is deemed to be of sufficient quality and is retained and / or reported as such to the user.

[0308] In some embodiments, matching the predicted tumor type with the provided tumor type involves using a set of reference genes from a training dataset that includes multiple signature genes that are up- or down-regulated in a particular tumor type relative to normal healthy samples. For example, if the predicted tumor type is prostate cancer (e.g., claimed information), the sample is checked against known reference genes for prostate cancer. In some embodiments, the predicted tumor type is evaluated against its tumor grade, which can help determine the signature genes of the claimed cancer grade at different stages of cancer.

[0309] As explained above, in some embodiments, determining a predicted tumor type as a quality control step may be performed by determining cancer grade from sequence information by using a machine learning approach that employs a statistical model trained using training data.

[0310] For example, in some embodiments, a statistical model can be used to predict characteristics of a biological sample using gene expression data based on an input ranking of genes, ranked based on their respective expression levels for a sequencing platform. Using input rankings instead of specific values ​​for expression levels allows the same or similar data processing pipeline to be used across different expression data regardless of the specific method by which the expression levels were obtained (e.g., regardless of the sequencing platform, sequencing conditions, sample preparation, data processing to obtain the expression levels, etc.). In some embodiments, a statistical model can be used to predict the cancer grade of a biological sample. In some embodiments, a statistical model can be used to predict the tissue of origin of a biological sample, which may also be used to perform quality control as described herein.

[0311] For example, in some embodiments, a ranking of genes based on gene expression levels (in a biological sample) as determined by a sequencing platform can be provided as input to a statistical model trained to predict a tissue of origin for a biological sample. The predicted tissue of origin can be compared against a tissue of origin claimed as part of the quality control techniques described herein. As another example, in some embodiments, a ranking of genes based on gene expression levels (in a biological sample) as determined by a sequencing platform can be provided as input to a statistical model trained to predict a cancer grade for a biological sample. The predicted cancer grade can be compared against a cancer grade claimed as part of the quality control techniques described herein.

[0312] In some embodiments, the set of genes that are ranked depends on the particular biological characteristic of interest, for example, one set of genes may be used to determine tissue of origin, and another set of genes may be used to determine cancer grade.

[0313] In some embodiments, the expression data may be obtained for cells in a biological sample, and the subject has, is suspected of having, or is at risk for having cancer. In the context where tissue of origin is the characteristic being determined, the tissue of origin is for cells in the biological sample. The tissue of origin may refer to the particular tissue type from which the cells originate, such as lung, pancreas, stomach, colon, liver, bladder, kidney, thyroid, lymph node, adrenal gland, skin, breast, ovary, prostate, etc.

[0314] For example, in some embodiments, this involves using a gene set to predict the tissue of origin, which may include the cell of origin, for diffuse large B-cell lymphoma (DLBCL), such as germinal center B-cell (GCB) and activated B-cell (ABC). The genes in the gene set may be selected from the group consisting of ITPKB, MYBL1, LMO2, BATF, IRF4, LRMP, CCND2, SLA, SP140, PIM1, CSTB, BCL2, TCF4, P2RX5, SPINK2, VCL, PTPN1, REL, FUT8, RPL21, PRKCB1, CSNK1E, GPR18, IGHM, ACP1, SPIB, HLA-DQA1, KRT8, FAM3C, and HLA-DMB.

[0315] In the context of cancer grade being a characteristic being determined, the cancer grade is for cells in a biological sample. The cancer grade may refer to the proliferation and differentiation characteristics of the cells in the biological sample, and generally refers to a numerical grade determined by visual observation of the cells using a microscope, such as grade 1, grade 2, grade 3, and grade 4.

[0316] For example, some embodiments involve using a gene set to predict breast cancer grade. The genes in the gene set include UBE2C, MYBL2, PRAME, LMNB1, CXCL9, KPNA2, TPX2, PLCH1, CCL18, CDK1, MELK, CCNB2, RRM2, CCNB1, NUSAP1, SLC7A5, TYMS, GZMK, SQLE, C1orf106, CDC25B, ATAD2, QPRT, CCNA2, NEK2, IDO1, NDC80, ZWINT, ABCA12, TOP2A, TDO2, S100A8, LAMP3, and MMP1. , GZMB, BIRC5, TRIP13, RACGAP1, ASPM, ESRP1, MAD2L1, CENPF, CDC20, MCM4, MKI67, PBK, CKS2, KIF2C, MRPL13, TTK, BUB1, TK1, FOX M1, CEP55, EZH2, ECT2, PRC1, CENPU, CCNE2, AURKA, HMGB3, APOBEC3B, LAGE3, CDKN3, DTL, ATP6V1C1, KIAA0101, CD2, KIF11, KIF20A , CDCA8, NCAPG, CENPN, MTFR1, MCM2, DSCC1, WDR19, SEMA3G, KCND3, SETBP1, KIF13B, NR4A2, NAV3, PDZRN3, MAGI2, CACNA1D, STC2, CHAD, PDGFD, ARMCX2, FRY, AGTR1, MARCH8, ANG, ABAT, THBD, RAI2, HSPA2, ERBB4, ECHDC2, FST, EPHX2, FOSB, STARD13, ID4, FAM129 A, FCGBP, LAMA2, FGFR2, PTGER3, NME5, LRRC17, OSBPL1A, ADRA2A, LRP2, C1orf115, COL4A5, DIXDC1, KIAA1324, HPN, KLF4, SCUBE2, FMO5, SORBS2, CARD10, CITED2, MUC1, BCL2, RGS5, CYBRD1, OMD, IGFBP4, LAMB2, DUSP4, PDLIM5, IRS2, and CX3CR1.

[0317] As another example, some embodiments involve using a gene set to predict renal clear cell carcinoma grade. The genes in the gene set include PLTP, C1S, LY96, TSKU, TPST2, SERPINF1, SRPX2, SAA1, CTHRC1, GFPT2, CKAP4, SERPINA3, CFH, PLAU, BASP1, PTTG1, MOCOS, LEF1, SLPI, PRAME, STEAP3, LGALS2, CD44, FLNC, UBE2C, CTSK, SULF2, TMEM45A, FCGR1A, PLOD2, C19orf80, PDGFRL, IGF2BP3, SLC7A5, PRRX1, RARRES1, LHFPL2, KDELR3, TRIB3, IL20RB, It may be selected from the group consisting of FBLN1, KMO, C1R, CYP1B1, KIF2A, PLAUR, CKS2, CDCP1, SFRP4, HAMP, MMP9, SLC3A1, NAT8, FRMD3, NPR3, NAT8B, BBOX1, SLC5A1, GBA3, EMCN, SLC47A1, AQP1, PCK1, UGT2A3, BHMT, FMO1, ACAA2, SLC5A8, SLC16A9, TSPAN18, SLC17A3, STK32B, MAP7, MYLIP, SLC22A12, LRP2, CD34, PODXL, ZBTB42, TEK, FBP1, and BCL2.

[0318] Aspects of using statistical models to predict tissue of origin, cancer grade, and / or other characteristics of biological samples are described in U.S. Provisional Patent Application No. 62 / 943,976, filed December 5, 2019, entitled "Machine Learning Techniques for Gene Expression Analysis," which is incorporated herein by reference in its entirety.

[0319] Returning to the aspect of assessing the quality of sequence information, in some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is assessed by determining the presence or absence of polyadenylated RNA genes to predict whether the sequence information was obtained from polyA RNA. In some embodiments, if there is a discrepancy between the predicted polyA status and the expected (e.g., asserted) polyA status of one or more samples, the sequence information for those samples is deemed to be of insufficient quality, suspect, and is removed, discarded, retested, and / or reported as such to the user. In some embodiments, if there is a concordance between the predicted polyA status and the expected polyA status for one or more nucleic acid samples, the sequence information is deemed to be of sufficient quality, is retained, and / or is reported as such to the user.

[0320] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is assessed by determining a complexity value for the sequence information. In some embodiments, determining the complexity value includes determining the number of duplicates. In some embodiments, the % duplicate rate can be determined for DNA or RNA libraries. In some embodiments, if a large percentage of libraries are duplicated, this indicates either a library of low complexity or over-amplification of DNA or cDNA fragments. In some cases, differences in complexity or amplification between libraries indicate the introduction of some bias in the data (e.g., different % GC content). In some embodiments, if the complexity value is less than 75% or less than 80%, the sequence information is deemed to be of insufficient quality or of questionable quality and is removed, discarded, retested, and / or reported to the user as such. In some embodiments, if the complexity value is at least less than 80% or at least 85%, the sequence information is deemed to be of sufficient quality for further analysis and is retained and / or reported to the user as such.

[0321] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is assessed by predicting a tissue source for the nucleic acid. In some embodiments, if there is a match between the predicted tissue source for the nucleic acid and the claimed tissue source, the sequence information is deemed to be of insufficient quality, suspect, and is removed, discarded, retested, and / or reported as such. In some embodiments, if there is a match between the predicted tissue source and the claimed tissue source, the sequence information is deemed to be of sufficient quality for further analysis and is retained and / or reported to the user as such.

[0322] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is assessed by (a) determining gene expression levels for two different subunits of a known protein and (b) determining an expression ratio for the two different subunits. In some embodiments, if the determined expression ratio does not match the expected expression ratio for the protein subunit, the sequence information is identified as being of insufficient quality, suspect, removed, discarded, retested, and / or reported as such to a user. In some embodiments, if the determined expression ratio matches the expected expression ratio for the protein subunit, the sequence information is deemed to be of sufficient quality for further analysis, retained, and / or reported as such to a user.

[0323] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is assessed by determining a Phred score for the sequence information. In some embodiments, if the Phred score is less than 27, the sequence information is deemed to be of insufficient quality, suspect, and is removed, discarded, retested, and / or reported to a user as such. In some embodiments, if the Phred score is less than 20, the sequence information is removed, discarded, retested, and / or reported to a user as such. In some embodiments, if the Phred score is greater than 20 but less than 27, the sequence information is deemed close to being removed, discarded, retested, and / or reported to a user as such. In some embodiments, if the Phred score is at least 27, the sequence information is deemed to be of sufficient quality for further analysis and is retained and / or reported to a user as such.

[0324] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is assessed by determining the GC content of the sequence information. In some embodiments, if the GC content is at least 30% but not more than 55%, the sequence information is deemed sufficient for further analysis, retained, and / or reported to the user as such. In some embodiments, if the GC content is within the range of 45-65%, the sequence information is deemed sufficient for further analysis, retained, and / or reported to the user as such (i.e., acceptable). In some embodiments, a GC content of at least 50% (e.g., 50%, 51%, 60%) is a target value, at least for human samples.

[0325] In some embodiments, at least two (eg, 3, 4, 5, 6, 7, 8, 9, 10, or more) different methods for assessing the quality (eg, source and / or completeness) of sequence information are performed.

[0326] In some embodiments, the methods performed herein evaluate sequence information from a mammal. In some embodiments, the mammal is a human.

[0327] In some embodiments, the sample from which the sequence information was generated is a subject having, suspected of having, or at risk of having a disease, hi some embodiments, the disease is cancer.

[0328] In some embodiments, a report is generated that includes one or more features or results of the methods described herein. In some embodiments, the report further includes an analysis of the results of the methods described herein.

[0329] In some embodiments, the disclosed methods or processes may be executed on a system or computer processor (e.g., a laptop, desktop, server, or other computerized machine). Components of the system may be located remotely and communicate over a network, such as a local area network or wide area network, or by Internet Protocol. The system may interface with a user via a web-enabled browser and a graphical user interface (GUI). In some embodiments, the system is under the control of a user at a single location. In some embodiments, the system is composed of components that are not located at a single location and may not be under the direct control of a user. In some embodiments, system information is stored locally.

[0330] As described herein, the terms "process," "activity," "step," or variations thereof when used in a computerized process or flowchart therein may be used interchangeably unless otherwise indicated.

[0331] As described herein, the terms "patient," "subject," "human subject," or variations thereof, can be used interchangeably unless otherwise indicated.

[0332] 6A is a flowchart illustrating an exemplary computerized process 200 for performing non-stranded RNA sequencing with coding RNA enrichment. Process 200 begins with activity 201, where a first sample of a first tumor is obtained from a subject having, suspected of having, or at risk of having cancer. Further aspects related to obtaining a first sample of a first tumor from a subject having, suspected of having, or at risk of having cancer are presented under the heading "Biological Sample."

[0333] Process 200 then proceeds to activity 202, where RNA from the first sample of the first tumor is extracted. Aspects related to extracting RNA from the first sample of the first tumor are described under the heading "DNA and / or RNA Extraction."

[0334] Process 200 then proceeds to activity 203, where the extracted RNA is enriched for coding RNA to obtain enriched RNA. Aspects related to enriching extracted RNA for coding RNA to obtain enriched RNA are described under the heading "RNA Enrichment."

[0335] Process 200 then proceeds to activity 204, where a first library of cDNA fragments is prepared from the enriched RNA for non-stranded RNA sequencing. Aspects related to preparing a first library of DNA fragments from the enriched RNA for non-stranded RNA sequencing are described under the heading "Library Preparation for RNA Sequencing."

[0336] Process 200 then proceeds to activity 205, where non-stranded RNA sequencing is performed on the first library of cDNA fragments prepared from the enriched RNA. Aspects related to performing non-stranded DNA sequencing on the first library of DNA fragments prepared from the enriched RNA are described under the heading "RNA Sequencing." It should be understood that one or more activities of process 200 may be optional.

[0337] 6B is a flow chart illustrating a computerized process 210 for identifying cancer treatments by obtaining bias-corrected gene expression data. Process 210 begins with activity 211, where RNA expression data is obtained for a subject who has, is suspected of having, or is at risk for having cancer. Aspects related to obtaining RNA expression data are described under the heading "Obtaining RNA Expression Data."

[0338] Process 210 then proceeds to activity 212, where genes in the RNA expression data are aligned to a reference and the RNA expression data is annotated. Aspects related to aligning and annotating genes in the RNA expression data with known sequences of the human genome to obtain annotated RNA expression data are described under the heading "Alignment and Annotation."

[0339] Process 210 then proceeds to activity 213, where non-coding transcripts from the annotated RNA expression data are removed to obtain filtered RNA expression data. Aspects related to removing non-coding transcripts from the annotated RNA expression data are described under the heading "Removing Non-coding Transcripts."

[0340] Process 210 then proceeds to activity 214, where the filtered RNA expression data is normalized to obtain gene expression data. The gene expression data may be in Transcripts Per Kilobase Million (TPM) format. Normalizing the filtered RNA expression data to gene expression data in Transcripts Per Kilobase Million (TPM) format is described in the section entitled "Conversion to TPM and Gene Aggregation."

[0341] Process 210 then proceeds to activity 215, where at least one gene that introduces bias into the gene expression data is identified. Aspects of identifying at least one gene that introduces bias into the gene expression data are described under the heading "Removing Bias."

[0342] Process 210 then proceeds to activity 216, where expression data associated with at least one gene that introduces bias is removed from the gene expression data to obtain bias-corrected gene expression data. Aspects of removing expression data associated with at least one gene that introduces bias into the gene expression data from the gene expression data to obtain bias-corrected gene expression data are described under the heading "Bias Removal."

[0343] Process 210 then proceeds to activity 217, where a cancer treatment is identified for the subject using the bias-corrected gene expression data. Aspects related to identifying a cancer treatment for the subject using bias-corrected gene expression data are described under the heading "Cancer Treatment Identification."

[0344] 6C is a flowchart illustrating a computerized process 220 for identifying a cancer treatment for a subject having, suspected of having, or at risk of having cancer using bias-corrected gene expression data. Process 220 begins with activity 221, where RNA is enriched for coding RNA in a sample of extracted RNA from a first tumor sample from a subject having, suspected of having, or at risk of having cancer. Aspects related to enriching RNA for coding RNA in a sample of extracted RNA are described under the heading "DNA and / or RNA Extraction."

[0345] Process 220 then proceeds to activity 222, where non-stranded RNA sequencing is performed on the first library of cDNA fragments prepared from the enriched RNA to obtain RNA expression data. Aspects related to performing non-stranded RNA sequencing on the first library of cDNA fragments prepared from the enriched RNA to obtain RNA expression data are described under the heading "RNA Sequencing."

[0346] Process 220 then proceeds to activity 223, where the RNA expression data is converted into gene expression data. Process 220 then proceeds to activity 224, where at least one gene that introduces bias into the gene expression data is identified. Process 220 then proceeds to activity 225, where expression data associated with the at least one gene that introduces bias is removed from the gene expression data to obtain bias-corrected gene expression data. Aspects related to activities 223, 224, and 225 are described under the heading "Bias Removal."

[0347] Process 220 then proceeds to activity 226, where a cancer treatment is identified for the subject using the bias-corrected gene expression data. Aspects related to identifying a cancer treatment for the subject using bias-corrected gene expression data are described under the heading "Cancer Treatment Identification."

[0348] FIG. 7 is an exemplary flowchart showing a computerized process 300 for preparing patient samples for sequencing analysis and performing bioinformatics quality control, whereby appropriate cancer treatments may be obtained for patients or subjects whose nucleic acids were extracted for sequencing analysis.

[0349] In the illustrated embodiment, process 300 includes obtaining a first sample of a first tumor from a subject having, suspected of having, or at risk of having cancer in activity 301; extracting RNA from the first sample of the first tumor in activity 302; enriching the RNA for coding RNA to obtain enriched RNA in activity 303; preparing a first library of cDNA fragments from the enriched RNA for non-stranded RNA sequencing in activity 304; obtaining RNA expression data for the subject in activity 305; aligning and annotating genes in the RNA expression data with known sequences of the human genome to obtain annotated RNA expression data in activity 306; removing non-coding transcripts from the annotated RNA expression data in activity 307; and synthesizing the annotated RNA expression data (e.g., Transcripts Per Kilobase) in activity 308. In activity 309, identifying at least one gene that introduces bias into the gene expression data; in activity 310, removing the expression data for the at least one gene that introduces bias from the gene expression data to obtain bias-corrected gene expression data; in activity 311, obtaining sequence information and claimed information; in activity 312, determining one or more features from the sequence information; in activity 313, determining whether the one or more features match the claimed information; in activity 314, making an additional determination of at least one of the features; and in activity 315, identifying a cancer treatment for the subject using the bias-corrected gene expression data.

[0350] It should be understood that one or more activities in process 300 may be optional. For example, in some embodiments, activities 301 and 303 may be performed, with activity 303 being optional. In some embodiments, activities 301, 302, and 303 are all performed. In some embodiments, activities 301, 302, and 303 are all omitted, but the remaining activities are performed. This is useful when extracted, enriched RNA from a patient sample is already available before the start of process 300. In some embodiments, the one or more features in activity 312 include one or more features of source, patient, tissue type, tumor type, polyA status, MHC sequence, protein subunit ratio, complexity, contamination, coverage, exon coverage, read composition, Phred score, SNP match, and GC content. In some embodiments, the one or more features in activity 312 further include the strandedness of the RNA-seq analysis. In some embodiments, any one or more of the features in activity 312 may be determined. In some embodiments, the additional determination of features in activity 314 can include, but is not limited to, SNP match value, contamination value, polyA status, complexity value, Phred score, and GC content. In some embodiments, the additional determination of any one or more of the features can be performed in activity 314. In some embodiments, any one or more of activity 303, process 307, and process 314 can be omitted. In some embodiments, all of the activities of computerized process 300 can be performed.

[0351] FIG. 8 illustrates a non-limiting process pipeline 800 for processing and validating sequence data and claimed information associated with the sequence data for subsequent analysis (e.g., for diagnosis, prognosis, therapy, and / or other clinical applications). Activity 801 is performed by obtaining nucleic acid data including sequence data and claimed information indicating a claimed source for the sequence data. In some embodiments, the nucleic acid data is obtained from a previously processed biological sample. In some embodiments, the biological sample was previously obtained from a subject having, suspected of having, or at risk for having cancer. In some embodiments, activity 801 is performed by obtaining nucleic acid data including the claimed completeness of the sequence data. In some embodiments, activity 801 is performed by obtaining nucleic acid data including sequence data and claimed information indicating the claimed source and claimed completeness of the sequence data. In some embodiments, the claimed information indicates the claimed completeness of the sequence data. In some embodiments, the claimed information indicates the subject from whom the nucleic acids were obtained. For example, in some embodiments, the claimed information includes MHC allele information and / or SNP information for one or more loci for the subject. After activity 801, process 800 proceeds to activities 802 and 803, where the validity of the nucleic acid data obtained at activity 801 is confirmed. The validation includes processing the sequence data at activity 802 to obtain a determined completeness and / or a determined source, and determining whether the determined completeness and / or the determined source match the claimed completeness and / or the claimed source, respectively. The sequence data is processed at activity 802 to obtain determined information indicative of the determined source of the sequence data at activity 802a and / or determined information indicative of the determined completeness of the sequence data at activity 802b.In some embodiments, activity 802a may include determining information indicative of at least one, two, or three of the subject's MHC genotype, whether the nucleic acid data is RNA or DNA, the tissue type of the biological sample, the tumor type of the biological sample, the sequencing platform used to generate the sequence data, SNP matching (e.g., determining whether one or more SNPs in the sequence data match one or more SNPs in a reference sequence), and / or whether the RNA sample is polyA enriched. In some embodiments, activity 802b may include determining a first level of a first nucleic acid encoding a first subunit of a multimeric protein, determining a second level of a second nucleic acid encoding a second subunit of the multimeric protein, and determining whether a ratio between the first level and the second level matches an expected ratio. In some embodiments, the first subunit and the second subunit are first and second CD3 subunits, first and second CD8 subunits, or first and second CD79 subunits. In some embodiments, the determined information indicating the determined completeness indicates at least one, two, or three of total sequence coverage, exon coverage, chromosome coverage, the ratio of nucleic acids encoding two or more subunits of a multimeric protein, species contamination, complexity, and / or the percentage (%) of guanine (G) and cytosine (C) in the sequence data. In some embodiments, activity 803 includes determining one or more MHC allele sequences from the sequence data and determining whether the one or more MHC allele sequences match claimed MHC allele information for the subject. In some embodiments, determining the MHC alleles includes determining sequences for six MHC loci from the sequence data.

[0352] In activity 803, the determined completeness and / or source is evaluated by determining whether the determined source of the sequence data matches the claimed source of the sequence data and / or whether the determined completeness of the sequence data matches the claimed completeness of the sequence data.

[0353] If the information asserted and determined at activity 803 match (i.e., yes), process 800 proceeds to activity 804, where the sequence data is further evaluated to determine whether the sequence data is indicative of a diagnosis, prognosis, therapy, or other clinical outcome. For example, in some embodiments, the sequence data is further processed at activity 804 to provide a cancer treatment recommendation for a subject who has, is suspected of having, or is at risk of having cancer. In some embodiments, activity 804 executes by determining a therapy for the subject, which is then administered to the subject.

[0354] In some embodiments, the process may further include administering a therapy to the subject. In some embodiments, the therapy is a cancer therapy.

[0355] In some embodiments, determining a therapy for the subject may include determining a plurality of gene group expression levels, including a gene group expression level for each gene group in a set of genes. In some embodiments, the set of genes includes at least one gene group associated with cancer aggressiveness and at least one gene group associated with the cancer microenvironment. A therapy for the subject is identified using the determined gene group expression levels.

[0356] If the asserted information and the determined information do not match at activity 803 (i.e., no), process 800 proceeds to 805, where one or more corrective actions are taken. In some embodiments, the corrective action includes generating an indication that the determined information does not match the asserted information, generating an indication not to process the sequence data in subsequent analyses, and / or generating an indication to obtain additional sequence data and / or biological samples and / or other information about the subject.

[0357] In some embodiments, the method includes all of the activities illustrated in FIG. 8. However, in some embodiments, a subset of these activities is performed, and any one or more of the activities may be omitted, duplicated, and / or performed in a different order than illustrated in FIG. 8. For example, either activity 802a or activity 802b is performed in activity 802. For example, activity 803 may be performed twice to confirm a decision. For example, one or more activities in process 800 may be performed after one or more corrective actions in activity 805. In some embodiments, one or more activities of FIG. 8 are implemented on a computer.

[0358] In some embodiments, the expression level of one or more genes in a sample is analyzed to assess the origin and / or quality of the sample. For example, the expression of one or more genes known to be expressed in a particular cell, tissue, or tumor type is assessed to determine whether the expression level is expected based on the expected cell, tissue, or tumor being analyzed. Similarly, the expression of one or more genes known not to be expressed (or not highly expressed) in a particular cell, tissue, or tumor type is assessed to determine whether the expression level is expected based on the expected cell, tissue, or tumor being analyzed.

[0359] In some embodiments, the expression levels of one or more genes are analyzed for each of a plurality of samples (e.g., 2, 3, 4, 5, 4-10, 1-50, 50-500, or more samples). If expression of one or more genes is lower or higher than expected, this may indicate that the quality and / or source / origin of the analyzed data is not as expected. In some embodiments, data from samples with unexpected levels of expression (e.g., lower or higher than expected) for one or more genes are excluded from further analysis. In some embodiments, new sequence information is obtained for samples with unexpected levels of expression for one or more genes, e.g., to confirm whether the initial data is correct. In some embodiments, samples with unexpected levels of expression for one or more genes may be further analyzed, e.g., to determine whether the samples were from a different source than originally indicated.

[0360] In some embodiments, expression levels for one or more genes are analyzed (e.g., using tSNE, PCA, or other techniques) to determine whether gene expression or patterns of gene expression are similar or different in separate samples. In some embodiments, if datasets containing the same cell type or tissue type do not cluster within a group, or if one or more datasets are identified as statistically different from other datasets containing the same cells or tissues, the datasets identified as different may be excluded and further analyzed or flagged as potentially suspicious. In some embodiments, additional sequence data may be obtained for samples identified as potentially suspicious.

[0361] An exemplary embodiment of a computer system 500 that may be used in connection with any of the embodiments of the technology described herein is shown in FIG. 9. Computer system 500 includes one or more processors 510 and one or more articles of manufacture that include non-transitory computer-readable storage media (e.g., memory 520 and one or more non-volatile storage media 530). Processor 510 may control the writing of data to and reading of data from memory 520 and non-volatile storage device 530 in any suitable manner, and the aspects of the technology described herein are not limited in this respect. To perform any of the functions described herein, processor 510 may execute one or more processor-executable instructions stored in one or more non-transitory computer-readable storage media (e.g., memory 520), which may act as non-transitory computer-readable storage media that store processor-executable instructions for execution by processor 510.

[0362] Computing device 500 may also include a network input / output (I / O) interface 540 that may be used to communicate between the computing device and other computing devices (e.g., over a network), and may also include one or more user I / O interfaces 550 that may be used by the computing device to provide output to and receive input from users. User I / O interfaces may include devices such as a keyboard, a mouse, a microphone, a display device (e.g., a monitor or touchscreen), speakers, a camera, and / or various other types of I / O devices.

[0363] The embodiments described above may be implemented in numerous ways. For example, these embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code may be executed on any suitable processor (e.g., a microprocessor) or collection of processors, whether provided on a single computing device or distributed across multiple computing devices. It should be understood that any component or collection of components that performs the functions described above may be generally considered to be one or more controllers that control the functions described above. The one or more controllers may be implemented in various ways, such as dedicated hardware or general-purpose hardware (e.g., one or more processors) that are programmed using microcode or software to perform the functions described above.

[0364] In this regard, one implementation of the embodiments described herein includes at least one computer-readable storage medium (e.g., RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or other tangible, non-transitory computer-readable storage medium) encoded with a computer program (i.e., a plurality of executable instructions) that, when executed on one or more processors, performs the above-described functions of one or more embodiments. The computer-readable medium may be portable such that the program stored thereon can be loaded onto any computing device to implement aspects of the technology described herein. In addition, it should be understood that reference to a computer program that, when executed, performs any of the above-described functions is not limited to an application program running on a host computer. Rather, the terms computer program and software are used generically herein to refer to any type of computer code (e.g., application software, firmware, microcode, or other form of computer instructions) that may be employed to program one or more processors to implement aspects of the technology described herein.

[0365] Aspects of the technology described herein provide computer-implemented methods for evaluating, generating, visualizing, and / or classifying biological characteristics (e.g., cancer grade, tissue of origin) of sequence information of a subject (e.g., a cancer patient) or an individual having, suspected of having, or at risk of having a disease (e.g., cancer).

[0366] In some embodiments, the software program may provide a user with a visual representation of subject (e.g., patient) characteristics and / or other information related to the subject's (e.g., patient's) cancer using an interactive graphical user interface (GUI). Such software programs may run in any suitable computing environment, including, but not limited to, a cloud computing environment, a device co-located with the user's location (e.g., the user's laptop, desktop, smartphone, etc.), one or more devices remote from the user (e.g., one or more servers), etc.

[0367] For example, in some embodiments, the techniques described herein may be implemented in an exemplary environment 600 shown in Figure 10. As shown in Figure 10, within exemplary environment 600, one or more biological samples from a subject 680 may be provided to a laboratory 670. The laboratory 670 may process the biological sample to obtain expression data (e.g., DNA, RNA, and / or protein expression data) and / or sequence information, which may be provided via a network 610 to at least one database 660 that stores information about the subject (e.g., patient) 680.

[0368] Network 610 may be a wide area network (e.g., the Internet), a local area network (e.g., a corporate intranet), and / or any other suitable type of network. Any of the devices shown in Figure 10 may connect to network 610 using one or more wired links, one or more wireless links, and / or any suitable combination thereof.

[0369] In the illustrated embodiment of FIG. 10 , the at least one database 620 may store expression data and / or sequence information for a subject (e.g., a patient), medical history data for the subject (e.g., a patient), lab result data for the subject (e.g., a patient), and / or any other suitable information regarding the subject 680. Examples of stored lab result data for a subject (e.g., a patient) include biopsy test results, imaging test results (e.g., MRI results), and blood test results. The information stored in the at least one database 620 may be stored in any suitable format and / or using any suitable data structure, and aspects of the technology described herein are not limited in this respect. The at least one database 620 may store data in any suitable manner (e.g., in one or more databases, one or more files). The at least one database 620 may be a single database or multiple databases.

[0370] As shown in FIG. 10 , exemplary environment 600 includes one or more external databases 620, which may store information for patients other than patient 680. For example, external database 660 may store expression data and / or sequence information (of any suitable type) for one or more patients, medical history data for one or more patients, test result data for one or more patients (e.g., imaging results, biopsy results, blood test results), demographic and / or personal information for one or more patients, and / or any other suitable type of information. In some embodiments, external database 660 may store information available in one or more publicly accessible databases, such as TCGA (The Cancer Genome Atlas), one or more databases of clinical trial information, and / or one or more databases maintained by commercial sequencing suppliers. External database 660 may store such information in any suitable manner using any suitable hardware, and the aspects of the technology described herein are not limited in this respect.

[0371] In some embodiments, at least one database 620 and external database 660 may be the same database, may be part of the same database system, or may be physically co-located, although aspects of the technology described herein are not limited in this respect.

[0372] For example, in some embodiments, server 640 may access information stored in databases 620 and / or 660 and use this information to perform the processes described herein with reference to FIG. 10 to determine one or more characteristics of the biological sample and / or sequence information.

[0373] In some embodiments, server 640 may comprise one or more computing devices. When server 640 comprises multiple computing devices, the devices may be physically co-located (e.g., in a single room) or distributed across multiple physical locations. In some embodiments, server 640 may be part of a cloud computing infrastructure. In some embodiments, one or more servers 640 may be co-located within a facility operated by an entity (e.g., a hospital, a research institute) to which physician 650 belongs. In such embodiments, it may be easier to enable server 640 to access private medical data of patient 880.

[0374] 10 , in some embodiments, the results of the analysis performed by server 640 may be provided to physician 650 via computing device 630 (which may be a portable computing device such as a laptop or smartphone, or a fixed computing device such as a desktop computer). The results may be provided via a written report, email, a graphical user interface, and / or other suitable means. It should be understood that although in the embodiment of FIG. 10 the results are provided to physician 650, in other embodiments the analysis results may be provided to patient 680 or the patient's 680 caregiver, a healthcare provider such as a nurse, or someone involved in the clinical trial.

[0375] In some embodiments, the results may be part of a graphical user interface (GUI) presented to physician 650 via computing device 630. In some embodiments, the GUI may be presented to the user as part of a web page displayed by a web browser running on computing device 630. In some embodiments, the GUI may be presented to the user using an application program (different from the web browser) running on computing device 630. For example, in some embodiments, computing device 630 may be a mobile device (e.g., a smartphone), and the GUI may be presented to the user via an application program (e.g., an “app”) running on the mobile device.

[0376] The GUI presented on the computing device 630 may provide a wide range of oncological data related to both the patient and the patient's cancer in a compact, informative, and novel way. Previously, oncological data was obtained in batches from multiple sources of data, making the process of obtaining such information costly both in terms of time and money. Using the techniques and graphical user interface illustrated herein, a user can access the same amount of information at once, with fewer demands on the user and fewer demands on the computing resources required to provide such information. The fewer demands on the user help reduce clinician errors associated with searching various sources of information. The fewer demands on computing resources helps reduce the processor power, network bandwidth, and memory required to provide the wide range of oncological data, leading to improved computing technologies. In some embodiments, the reports of the present disclosure are presented to the user by the system or using a GUI.

[0377] Therefore, in one aspect, the present disclosure relates to a method for evaluating sequence information to determine at least one characteristic. This evaluation can be performed on a computer or other automated machine that can execute programmable instructions, or can be performed manually by an evaluator. The characteristic can be used to generate a report that informs the evaluator of at least one characteristic of the sequence information. In some embodiments, the characteristic is the sequence of the MHC alleles of the sequence information.

[0378] The major histocompatibility complex (MHC) (called human leukocyte antigens (HLA) in humans) is a mechanism by which the immune system can distinguish between self and non-self cells. It is a collection of glycoproteins (carbohydrate-containing proteins) present on the plasma membrane of almost all somatic cells. The MHC is a highly polymorphic gene important to the immune system of living organisms, derived from 20 genes, with over 50 variations per gene between individuals and allowing codominance between alleles. These glycoproteins are part of a pathway that allows the immune system to distinguish between self and non-self cells through abnormalities in the MHC displayed on the plasma membrane.

[0379] These properties, such as the highly polymorphic nature and codominance of the MHC, and the large number of alleles that may exist in a given species, make a subject's MHC profile highly specific and unique. Thus, it is highly unlikely that two individuals, except for identical twins, will possess cells with the same set of MHC molecules. Thus, by evaluating the sequence of the MHC profile of sequence information, this can be used to confirm or disqualify the information from distinguishing between sequence information, claimed information, other sequence information, or combinations thereof.

[0380] In some embodiments, one MHC allele is used for the assessment. In some embodiments, at least two MHC alleles are used for the assessment. In some embodiments, at least three MHC alleles are used for the assessment. In some embodiments, at least four MHC alleles are used for the assessment. In some embodiments, at least five MHC alleles are used for the assessment. In some embodiments, at least six MHC alleles are used for the assessment.

[0381] In some embodiments, the feature evaluated is a single nucleotide polymorphism (SNP) concordance value. "SNP" or "single nucleotide polymorphism," as used herein, refers to a difference in a nucleic acid sequence (e.g., a genome, a sequence dataset) at a single nucleotide (e.g., adenine (A), thymine (T), cytosine (C), and / or guanine (G)) shared between subjects of a species or within individual subjects on paired chromosomes. A SNP may be, or represent, an altered nucleotide (e.g., A changed to T, G changed to A, etc.), referred to as a substitution; a removed nucleotide, in which a nucleotide is completely absent from the sequence, referred to as a deletion; or an added nucleotide, in which an additional nucleotide is added to the sequence. A SNP may (e.g., a nonsynonymous SNP) or may not (e.g., synonymous) cause a change in the encoded protein. Furthermore, when a SNP is nonsynonymous, it may cause a change in the encoded amino acid (e.g., missense) or cause a premature stop codon (e.g., nonsense). Synonymous SNPs can also alter the message of a nucleic acid sequence by affecting or altering splice sites, transcription factor binding, and / or messenger RNA (mRNA) binding. These mutations (e.g., changes to the protein-coding capacity of the sequence) can cause many effects, including differences in phenotype and even various disease types. Furthermore, SNPs occur in large numbers within a subject's genome; it is estimated that a typical genome differs from a reference human genome at 4 to 5 million positions, of which more than 99.9% are SNPs.

[0382] Because SNPs are encoded in nucleic acids that are part of the genome, they are passed down from parents to offspring (both within the subject and within the subject when the nucleic acid is replicated). Therefore, because this inheritance is stable and numerous, SNPs can be used as genetic markers of subject relatedness and as a measure of the identity of two nucleic acid sequences as originating from the same subject. In some embodiments, SNP match values ​​are determined between sequence information and a reference sequence. In some embodiments, SNP match values ​​are determined between sequence information and a claimed value. In some embodiments, the SNP match value must be equal to or greater than a threshold value to be acceptable (e.g., deemed to have sufficient quality and completeness) for use in further analysis. In some embodiments, the threshold value is 80%. In some embodiments, a SNP concordance value is determined between the sequence dataset and the subject, and the SNP concordance value is at least 70% (e.g., at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 95%). A match value is considered sufficiently likely to be from the subject and identified as being from the subject if the match value is greater than or equal to 99.5%, at least 96%, at least 96.5%, at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, at least 99.999%, or more. As described herein, in some embodiments, determining the match value is as described in the present disclosure.SNP match detection can be performed by any means available or known in the art, for example, SNP match detection can be performed by various online tools such as Compair (github.com / nygenome / Compair) or GATK GenotypeConcordance (software.broadinstitute.org / gatk / documentation / tooldocs / 3.8-0 / org_broadinstitute_gatk_tools_walkers_variantutils_GenotypeConcordance.php), or can be calculated manually. In other examples, SNP match detection can be performed by tools such as those described on publicly available websites (genome.sph.umich.edu / wiki / VerifyBamID or software.broadinstitute.org / cancer / cga / contest).

[0383] In some embodiments, the feature evaluated is a quality score, such as, for example, a Phred score. As used herein, a "Phred score" (also known or sometimes referred to herein as a "Phred quality score") refers to a measure of quality for the identification of nucleotides sequenced by a nucleic acid sequencing system or platform (e.g., NGS). Phred scores are known in the art and are often generated from sequencing platforms based on several parameters (e.g., peak shape, resolution, etc.), and a score (Q) is assigned to each nucleotide base call. (For a detailed review of the calculations, see Ewing B, Hillier L, Wendl MC, Green P. "Base-calling of automated sequencer traces using phred. I. Accuracy assessment." Genome Res. 1998 March;8(3):175-85; and Ewing B, Green P. "Base-calling of automated sequencer traces using phred. II. Error probabilities." Genome Res. 1998 March;8(3):186-94.) The Phred score for each base indicates the likelihood that a nucleotide base call is incorrect (base call error probability (P)), and is expressed as Q = -10log 10The Phred score is determined by P. Thus, the score (e.g., Q) indicates the base calling accuracy; for example, a Phred score of 10 indicates 90% calling accuracy for the base of interest, and a Phred score of 40 indicates 99.99% calling accuracy for the same base. In some embodiments, the Phred score of the sequence information is determined and compared to a reference value. In some embodiments, the reference value is at least 27, at least 28, at least 29, at least 30, or greater than 30. In some embodiments, the Phred score is determined and compared to the Phred scores of other sequence information. In some embodiments, the Phred score is determined and compared to a claimed score. In some embodiments, the Phred score is used as a base-level determination of quality. In some embodiments, a sequence is compared to claimed information to compare identity, since if the Phred scores differ, it is unlikely that they are the same sequence information or from the same sample or subject.

[0384] In some embodiments, the characteristic evaluated is tumor type.

[0385] In some embodiments, the characteristic evaluated is tissue type.

[0386] In some embodiments, the feature evaluated is the polyadenylation status of the sequence information. As used herein, "polyadenylation" or "polyA" refers to a series of multiple adenosine monophosphate nucleotides attached to the 3' end of messenger RNA (mRNA) after transcription and cleavage of the 3' end of the transcript, liberating a hydroxyl group. The "polyA tail," as it is often referred to, is characteristic of fully processed mRNA and aids in various cellular processes. For example, the polyA tail is a binding site for proteins (such as polyA-binding proteins) that facilitate transport from the cell nucleus for translation, as well as affect mRNA translation and stability. If only protein-coding mRNA transcripts are present, it is likely (e.g., indicates) that the sample from which the sequence information was generated was generated using mRNA-Seq. In some embodiments, the polyA status indicates that mRNA-Seq was not used (e.g., whole transcriptome was used). In some embodiments, the polyA status is evaluated against the claimed information. In some embodiments, the polyA status is evaluated against a reference sequence. In some embodiments, the probability that sequence information will be generated using either mRNA-Seq or whole transcriptome must be higher than a threshold. In some embodiments, the threshold is a reference value. In some embodiments, the threshold level is 90%. In some embodiments, the threshold is asserted information. In some embodiments, the sequence information is from a sample that contains primarily polyadenylated nucleic acids. In some embodiments, the sequence information is from a sample that contains both polyadenylated and non-polyadenylated nucleic acids.

[0387] In some embodiments, the feature evaluated is the GC content of the sequence information. "G / C content" or "guanine (G)-cytosine (C) content," as used herein, refers to the percentage of nucleotides in a nucleic acid sample that are either G or C. This can be calculated by summing all of the G and C reads of a given sequence information and dividing by the total number of sequenced nucleotides. In some embodiments, the sequence information is evaluated and the GC content is calculated by summing the number of base calls that result in a G or C (e.g., G+C) in the sequence information and dividing by the total number of base calls in the sequence information (e.g., the number of nucleotides in the sequence dataset), i.e., (G+C) / (the number of nucleotides in the sequence dataset).

[0388] GC content can also be used as a quality measure of sequence information. Many known genomes have been sequenced, along with their respective exomes, transcriptomes, and various other portions (e.g., as a measure of specific RNA components). Furthermore, many of these sequences have been sequenced multiple times, generating average values ​​and ranges for their various components, such as the GC content of the human genome. The GC content of the human genome is known to vary from approximately 35% to 60%, with an average (e.g., mean) value of approximately 41%. Therefore, if, as a quality measure, sequence information identified as a human genome were assessed to have a GC content of 75%, there would be a problem with the quality of the sequence information (or the sample from which the sequence information was derived). As a result, the assessed GC content can be compared with known ranges of expected GC content for the sequence information, values ​​provided by the sequence information provider, values ​​from a database of such values, additional sequence information, or reference ranges for a given type of sequence information, to determine whether they are consistent or whether the GC content indicates problems, which may be due to degradation, residual primers, contamination, or other disruptions to the sequence information.

[0389] Thus, in some embodiments, GC content features are used in methods for assessing the integrity of sequence information by determining the GC content for each sequence dataset being evaluated; if the GC content is (i) less than 30% or greater than 55%, the nucleic acid sample is deemed likely to be of insufficient quality and is removed, discarded, retested, or reported as insufficient; if the GC content is at least 30% but less than 55%, the nucleic acid sample is deemed to be of sufficient quality and is retained. Additionally, in some embodiments, GC content can be calculated and compared to claimed information. GC content can, in some embodiments, be used to verify or challenge the sequence information as being from a given sample, subject, or specific sample or subject, the same sequence information claimed.

[0390] In some embodiments, the evaluated feature is the ratio of expression of protein subunits (e.g., the expression ratio of nucleic acids encoding different subunits of a protein). Protein expression can be measured by any means known in the art; for example, expression can be determined (e.g., quantified) by counting the number of reads mapped to each locus in the transcriptome. By evaluating the expression of protein subunits (e.g., a protein having multiple subunits expressed by different coding regions) and then calculating the ratio, it is possible to compare the ratio to a known value, reference value, threshold, other sequence information, or asserted information. In some embodiments, the evaluated protein subunits are from proteins present in human samples, regardless of the presence or absence of certain cancers. Without wishing to be bound by any theory, such proteins are encoded by housekeeping genes (e.g., positive or negative controls for human samples). In some embodiments, the known value, reference value, or threshold is a fixed ratio. For example, subunit A and subunit B of a known protein have a ratio of 1:1 or 2:1. In some embodiments, the protein subunits evaluated are from proteins present in human samples having or suspected of having certain types of cancer.

[0391] In some embodiments, this ratio is compared to a known value. In some embodiments, it is compared to claimed information. In some embodiments, it is compared to other sequence information. In some embodiments, proteins, and their subunits, are selected for analysis due to their properties. For example, they degrade quickly and therefore serve as a surrogate for the stability and / or quality of the sample from which the sequence information was generated. In some embodiments, they are selected for variability between subjects or samples, thereby allowing comparisons to confirm or disqualify the identity of the sequence information.

[0392] In some embodiments, the feature evaluated is...

Claims

1. 1. A method comprising: Use at least one hardware processor A step of acquiring nucleic acid data, wherein the nucleic acid data is sequence data representing at least 5 kilobases (kb) of nucleotide sequence of DNA and / or RNA from a previously obtained biological sample of a subject having, suspected of having, or at risk of having a disease; claimed information indicating the claimed source and / or claimed completeness of said sequence data; The nucleic acid data processing the sequence data to obtain determined information indicative of a determined source and / or a determined completeness of the sequence data; and determining whether the determined information matches the asserted information.

2. 2. The method of claim 1, further comprising processing the sequence data to determine whether the sequence data is indicative of one or more disease characteristics when the asserted information is determined to match the determined information.

3. determining whether the determined information is consistent with the purported information; and processing the sequence data to determine whether it is indicative of one or more disease characteristics.

4. When it is determined that the asserted information is inconsistent with the determined information, generating an indication that the determined information does not match the purported information; without processing said sequence data in subsequent analyses; and / or 4. The method of claim 1, further comprising obtaining additional sequence data and / or other information about the biological sample and / or the subject.

5. determining that the purported information does not match the determined information; generating an indication that the determined information does not match the purported information; without processing said sequence data in subsequent analyses; and / or 5. The method of any one of claims 1 or 2 to 4, further comprising obtaining additional sequence data and / or other information about the biological sample and / or the subject.

6. the claimed information indicates the claimed source of the sequence data, and the method processing the sequence data to obtain determined information indicative of a determined source for the sequence data; and determining whether the determined source is consistent with the purported source for the sequence data.

7. 6. The method of claim 6 or any one of claims 1 to 5, wherein the determined information indicating the determined source for the sequence data indicates the subject's MHC genotype, whether the nucleic acid data is RNA data or DNA data, the tissue type of the biological sample, the tumor type of the biological sample, the sequencing platform used to generate the sequence data, SNP matching, and / or whether the RNA sample is polyA enriched.

8. 7. The method of claim 7 or any one of claims 1 to 6, wherein the determined information indicating the determined source for the sequence data indicates at least two of the subject's MHC genotype, whether the nucleic acid data is RNA data or DNA data, the tissue type of the biological sample, the tumor type of the biological sample, the sequencing platform used to generate the sequence data, the SNP match, and whether the RNA sample is polyA enriched.

9. 8. The method of claim 8 or any one of claims 1 to 7, wherein the determined information indicating the determined source for the sequence data indicates at least three of the subject's MHC genotype, whether the nucleic acid data is RNA data or DNA data, the tissue type of the biological sample, the tumor type of the biological sample, the sequencing platform used to generate the sequence data, SNP match, and whether the RNA sample is polyA enriched.

10. the claimed information indicates the claimed completeness of the sequence data, and the method processing the sequence data to obtain determined information indicative of the determined completeness of the sequence data; and determining whether the determined completeness is consistent with the claimed completeness for the sequence data.

11. 10. The method of claim 1 or any one of claims 1 to 9, wherein the determined information indicating the determined completeness indicates total sequence coverage, exon coverage, chromosome coverage, the ratio of nucleic acids encoding two or more subunits of a multimeric protein, species contamination, single nucleotide polymorphisms (SNPs), complexity, and / or the percentage (%) of guanine (G) and cytosine (C) in the sequence data.

12. 11. The method of claim 1 or any one of claims 1 to 10, wherein the determined information indicating the determined completeness indicates at least two of the following: total sequence coverage, exon coverage, chromosome coverage, the ratio of nucleic acids encoding two or more subunits of a multimeric protein, species contamination, single nucleotide polymorphisms (SNPs), complexity, and the percentage (%) of guanine (G) and cytosine (C) in the sequence data.

13. 12. The method of claim 1 or any one of claims 1 to 11, wherein the determined information indicating the determined completeness indicates at least three of the following: total sequence coverage, exon coverage, chromosome coverage, the ratio of nucleic acids encoding two or more subunits of a multimeric protein, species contamination, single nucleotide polymorphisms (SNPs), complexity, and the percentage (%) of guanine (G) and cytosine (C) in the sequence data.

14. 14. The method of any one of claims 1 or 2 to 13, wherein the asserted information for the sequence data includes MHC allele information for the subject.

15. 14. The method of any one of claims 1 to 13, further comprising determining one or more MHC allele sequences from the sequence data and determining whether the one or more MHC allele sequences match the claimed MHC allele information for the subject.

16. 15. The method of claim 14, wherein determining one or more MHC allele sequences comprises determining MHC allele sequences for six MHC loci from the sequence data.

17. 17. The method of any one of claims 1 or 2 to 16, wherein the sequence data indicates the nucleotide sequence for an RNA and the asserted information indicates whether the RNA is polyA enriched.

18. 18. The method of any one of claims 1 or 2 to 17, further comprising using the sequence data to determine a therapy for the subject when it is determined that the asserted information matches the determined information.

19. The step of determining the therapy comprises: determining expression levels of a plurality of genes, wherein the plurality of gene expression levels comprises a gene expression level of each gene in a set of genes, the set of genes comprising at least one gene associated with cancer malignancy and at least one gene associated with the cancer microenvironment; and using said determined gene group expression levels to identify said therapy.

20. 20. The method of claim 19 or any one of claims 1 to 18, further comprising administering said therapy to said subject.

21. 21. The method of any one of claims 1 or 2 to 20, wherein the determined information is determined to match the purported information, and the sequence data is processed to thereby determine a therapy for the subject, and the therapy is administered to the subject.

22. 22. The method of any one of claims 1 or 2 to 21, wherein the disease is cancer and the therapy is cancer therapy.

23. 23. The method of any one of claims 1 or 2 to 22, wherein the subject is a human.

24. The step of processing the sequence data to obtain the determined source comprises: determining one or more single nucleotide polymorphisms (SNPs) within the sequence data; and determining whether the one or more SNPs in the sequence data match one or more SNPs in a reference sequence.

25. 24. The method of claim 24 or any one of claims 1 to 23, wherein the reference sequence is a sequence of a nucleic acid in a second biological sample of the subject.

26. processing the sequence data to obtain a determined completeness, determining a first level of a first nucleic acid encoding a first subunit of a multimeric protein; determining a second level of a second nucleic acid encoding a second subunit of the multimeric protein; and determining whether the ratio between the first level and the second level matches an expected ratio.

27. 26. The method of claim 26 or any one of claims 1 to 25, wherein the multimeric protein is a dimer.

28. 27. The method of any one of claims 1 to 26, wherein the first subunit and the second subunit are a first and second CD3 subunit, a first and second CD8 subunit, or a first and second CD79 subunit.

29. 1. A system comprising: at least one computer hardware processor; and at least one non-transitory computer-readable storage medium storing processor-executable instructions, said instructions, when executed by said at least one computer hardware processor, causing said at least one computer hardware processor to perform a method, said method comprising: A step of acquiring nucleic acid data, wherein the nucleic acid data is sequence data representing at least 5 kilobases (kb) of nucleotide sequence of DNA and / or RNA from a previously obtained biological sample of a subject having, suspected of having, or at risk of having a disease; claimed information indicating the claimed source and / or claimed completeness of said sequence data; The nucleic acid data processing the sequence data to obtain determined information indicative of a determined source and / or a determined completeness of the sequence data; and determining whether the determined information matches the asserted information.

30. At least one non-transitory computer-readable storage medium storing processor-executable instructions, the instructions, when executed by the at least one computer hardware processor, causing the at least one computer hardware processor to perform a method, the method comprising: A step of acquiring nucleic acid data, wherein the nucleic acid data is sequence data representing at least 5 kilobases (kb) of nucleotide sequence of DNA and / or RNA from a previously obtained biological sample of a subject having, suspected of having, or at risk of having a disease; claimed information indicating the claimed source and / or claimed completeness of said sequence data; The nucleic acid data processing the sequence data to obtain determined information indicative of a determined source and / or a determined completeness of the sequence data; and determining whether the determined information matches the purported information.

31. 1. A method comprising: obtaining a first biological sample of a first tumor, said first biological sample having been previously obtained from a subject having, suspected of having, or at risk of having cancer; extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; enriching the extracted RNA for coding RNA to obtain enriched RNA; sequencing the enriched RNA using at least one sequencing platform to obtain RNA expression data comprising at least 5 kilobases (kb); Use at least one hardware processor obtaining said RNA expression data using said at least one sequencing platform; converting the RNA expression data into gene expression data; determining bias-corrected gene expression data from said gene expression data at least in part by removing from said gene expression data expression data for at least one gene that introduces a bias into said gene expression data; and identifying a cancer treatment for said subject using said bias-corrected gene expression data.

32. 32. The method of claim 31, further comprising administering the identified cancer treatment to the subject.

33. 33. The method of claim 31 or 32, wherein enriching the RNA for coding RNA comprises performing polyA enrichment.

34. The at least one gene that introduces bias into the gene expression data is genes having an average transcript length that is longer or shorter than the average length of the transcripts in said gene expression data; genes having at least one threshold variation in mean transcript expression level based on transcript expression levels in a reference sample; and / or genes having a poly-A tail length that is at least a threshold amount less than the average length of the poly-A tails of genes from the first biological sample from which the RNA expression data was obtained and / or a reference sample.

35. 35. The method of any one of claims 31 or 32 to 34, wherein the at least one gene that introduces bias into the gene expression data belongs to a gene family selected from the group consisting of histone-encoding genes, mitochondrial genes, interleukin-encoding genes, collagen-encoding genes, B-cell receptor-encoding genes, and T-cell receptor-encoding genes.

36. The at least one gene is selected from the group consisting of HIST1H1A, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H1T, HIST1H2AA, HIST1H2AB, HIST1H2AC, HIST1H2AD, HIST1H2AE, HIST1H2AG, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H2AK, HIST1H2AL, and HIST1H2AM , HIST1H2BA, HIST1H2BB, HIST1H2BC, HIST1H2BD, HIST1H2BE, HIST1H2BF, HIST1H2BG, HIST1H2BH, HIST1H2BI, HIS T1H2BJ, HIST1H2BK, HIST1H2BL, HIST1H2BM, HIST1H2BN, HIST1H2BO, HIST1H3A, HIST1H3B, HIST1H3C, HIST1H3D, H IST1H3E, HIST1H3F, HIST1H3G, HIST1H3H, HIST1H3I, HIST1H3J, HIST1H4A, HIST1H4B, HIST1H4C, HIST1H4D, HIST1 H4E, HIST1H4F, HIST1H4G, HIST1H4H, HIST1H4I, HIST1H4J, HIST1H4K, HIST1H4L, HIST2H2AA3, HIST2H2AA4, HIST2 35. The method of any one of claims 31 to 34, comprising at least one histone encoding gene selected from the group consisting of H2AB, HIST2H2AC, HIST2H2BE, HIST2H2BF, HIST2H3A, HIST2H3C, HIST2H3D, HIST2H3PS2, HIST2H4A, HIST2H4B, HIST3H2A, HIST3H2BB, HIST3H3, and HIST4H4.

37. The at least one gene is selected from the group consisting of MT-ATP6, MT-ATP8, MT-CO1, MT-CO2, MT-CO3, MT-CYB, MT-ND1, MT-ND2, MT-ND3, MT-ND4, MT-ND4L, MT-ND5, MT-ND6, MT-RNR1, MT-RNR2, MT-TA, MT-TC, MT-TD, MT-TE, MT-TF, MT-TG, MT-TH, MT-TI, MT-TK, MT-TL1, MT-TL2, MT-TM, MT-TN, MT-TP, and MT-TQ. , MT-TR, MT-TS1, MT-TS2, MT-TT, MT-TV, MT-TW, MT-TY, MTRNR2L1, MTRNR2L10, MTRNR2L11, MTRNR2L12, MTRNR2L13, MTRNR2L3, MTRNR2L4, MTRNR2L5, MTRNR2L6, MTRNR2L7, and MTRNR2L8.

38. The step of determining bias-corrected gene expression data comprises:

38. The method of any one of claims 31 or 32-37, further comprising re-normalizing the gene expression data after removing the expression data for the at least one gene that introduces bias into the gene expression data.

39. converting the RNA expression data into gene expression data, removing non-coding transcripts from the RNA expression data to obtain filtered RNA expression data; and normalizing the filtered RNA expression data after removing the non-coding transcripts to obtain Transcripts Per Million (TPM) gene expression data.

40. The step of removing the non-coding transcripts from the RNA expression data may include removing pseudogenes, polymorphic pseudogenes, processed pseudogenes, transcribed processed pseudogenes, unitary pseudogenes, non-processed pseudogenes, transcribed unitary pseudogenes, constant chain immunoglobulin (IG C) pseudogenes, joining chain immunoglobulin (IG J) pseudogenes, variable chain immunoglobulin (IG V) genes, transcribed non-processed genes, translated non-processed genes, joining chain T cell receptor (TR J) genes, variable chain T cell receptor (TR V) genes, small nuclear RNAs (snRNAs), small nucleolar RNAs (snoRNAs), microRNAs (miRNAs), ribozymes, ribosomal RNAs (rRNAs), mitochondrial tRNAs (Mt tRNAs), mitochondrial rRNAs (Mt 39. The method of any one of claims 39 or 31 to 38, comprising removing non-coding transcripts belonging to a group selected from the list consisting of: rRNA, Cajal body-specific RNA (scaRNA), residual introns, sense intron RNA, sense overlapping RNA, nonsense-mediated decay RNA, non-stop decay RNA, antisense RNA, long intervening non-coding RNA (lincRNA), macro-long non-coding RNA (macro-lncRNA), processed transcripts, 3' overlapping non-coding RNA (3' overlapping ncrna), small RNA (sRNA), miscellaneous RNA (miscRNA), vault RNA, and TEC RNA.

41. Before performing the removal of non-coding transcripts, aligning the RNA expression data to a reference; 41. The method of any one of Claims 39 or 31 to 38 or 40, further comprising the step of annotating the RNA expression data.

42. 42. The method of any one of claims 31 or 32-41, wherein the RNA expression data comprises at least 25 million paired-end reads.

43. 42. The method of any one of claims 42 or 31 to 41, wherein the RNA expression data comprises at least 50 million paired-end reads, with an average read length of at least 100 bp.

44. identifying the cancer treatment for the subject using the bias-corrected gene expression data, using the bias-corrected gene expression data to determine a plurality of gene group expression levels, wherein the plurality of gene group expression levels comprises a gene group expression level for each gene group in a set of genes, the set of genes comprising at least one gene group associated with cancer malignancy and at least one gene group associated with the cancer microenvironment; and using said determined gene group expression levels to identify said cancer treatment.

45. 45. The method of any one of claims 34 or 31-33 or 35-44, wherein the cancer treatment is selected from the group consisting of radiation therapy, surgical therapy, chemotherapy, and immunotherapy.

46. 46. ​​The method of any one of claims 31 or 32-45, further comprising obtaining a second biological sample of a second tumor, the second biological sample having been previously obtained from the subject.

47. further comprising combining the first biological sample and the second biological sample to form a combined tumor sample; 46. ​​The method of any one of claims 46 or 31 to 45, wherein extracting the RNA comprises extracting the RNA from the combined tumor sample.

48. extracting RNA from the second biological sample; combining the RNA extracted from the second biological sample with the RNA extracted from the first biological sample to form a combined extracted RNA; 48. The method of any one of claims 46 or 31 to 45 or 47, wherein enriching the RNA for coding RNA comprises enriching the combined extracted RNA for coding RNA.

49. 49. The method of any one of claims 31 or 32 to 48, wherein the extracted RNA comprises at least 1 μg of RNA after RNA extraction.

50. 49. The method of any one of claims 49 or 31 to 48, wherein the extracted RNA has a total mass of at least 1000-6000 ng and a purity corresponding to an absorbance at 260 nm to absorbance at 280 nm ratio of at least 2.

0.

51. A quality control assessment of the RNA expression data is performed at least in part by obtaining claimed information indicating the claimed source and / or claimed completeness of said RNA expression data; processing the RNA expression data to obtain determined information indicative of a determined source and / or a determined completeness of the RNA expression data; 51. The method of any one of claims 31 or 32-50, further comprising the step of executing by determining whether the determined information matches the purported information.

52. 51. The method of any one of claims 31 to 50, wherein processing the RNA expression data comprises processing the RNA expression data to determine a tissue type of the first biological sample, a tumor type of the first biological sample, and / or a percentage (%) of guanine (G) and / or cytosine (C).

53. 1. A system for identifying a cancer treatment for a subject having, suspected of having, or at risk of having cancer, said system comprising: at least one sequencing platform configured to generate gene expression data from enriched RNA obtained from a first biological sample previously obtained from the subject, wherein the enriched RNA was obtained by: (i) extracting RNA from the first biological sample of a first tumor to obtain extracted RNA; and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA, wherein the RNA expression data comprises at least 5 kilobases (kb); at least one computer hardware processor; at least one non-transitory computer-readable storage medium storing processor-executable instructions, the processor-executable instructions, when executed by the at least one computer hardware processor, causing the at least one computer hardware processor to: obtaining said RNA expression data using said at least one sequencing platform; converting the RNA expression data into gene expression data; determining bias-corrected gene expression data from said gene expression data at least in part by removing from said gene expression data expression data for at least one gene that introduces a bias into said gene expression data; and identifying a cancer treatment for the subject using the bias-corrected gene expression data.

54. 1. A system for identifying a cancer treatment for a subject having, suspected of having, or at risk of having cancer, said system comprising: at least one computer hardware processor; at least one non-transitory computer-readable storage medium storing processor-executable instructions, the processor-executable instructions, when executed by the at least one computer hardware processor, causing the at least one computer hardware processor to: obtaining RNA expression data from at least one sequencing platform, the RNA expression data comprising at least 5 kilobases (5 kb), the RNA expression data having been obtained from a first biological sample of a first tumor previously obtained from the subject, at least in part by (i) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA, and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA; converting the RNA expression data into gene expression data; determining bias-corrected gene expression data from said gene expression data at least in part by removing from said gene expression data expression data for at least one gene that introduces a bias into said gene expression data; and identifying a cancer treatment for the subject using the bias-corrected gene expression data.

55. 55. The system of claim 54, further comprising the at least one sequencing platform.

56. at least one non-transitory computer-readable storage medium storing processor-executable instructions, the processor-executable instructions, when executed by the at least one computer hardware processor, causing the at least one computer hardware processor to: obtaining RNA expression data from at least one sequencing platform, the RNA expression data comprising at least 5 kilobases (5 kb), the RNA expression data obtained from a first biological sample of a first tumor previously obtained from a subject having, suspected of having, or at risk of having cancer, at least in part by (i) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA, and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA; converting the RNA expression data into gene expression data; determining bias-corrected gene expression data from said gene expression data at least in part by removing from said gene expression data expression data for at least one gene that introduces a bias into said gene expression data; and identifying a cancer treatment for the subject using the bias-corrected gene expression data.

Citation Information

Patent Citations

  • Personalized cancer vaccines

    JP2014523406A

  • Means and Methods for Analyzing a Sample by Means of Chromatography-Mass Spectrometry

    US20080234945A1

  • Automated sample quality assessment

    US20180321267A1

  • Neoantigen identification using hotspots

    WO2019075112A1

  • Expression of gag proteins from retroviruses in eucaryotic cells

    EP0345242A2