Systems and methods for sample preparation, sample sequencing, and bias correction and quality control of sequencing data - Patents.com

Novel sample preparation and quality control techniques for nucleic acid sequencing address errors in conventional workflows, enhancing data accuracy and therapeutic identification by reducing biases and improving molecular characterization.

JP7716384B2Active Publication Date: 2025-07-31BOSTONGENE CORP
View PDF 41 Cites 0 Cited by

Patent Information

Application Number
JP2022500543
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-18
Filing Date
2020-07-03
Publication Date
2025-07-31
Estimated Expiration
2040-07-03

AI Technical Summary

Technical Problem

Conventional workflows for obtaining nucleic acid sequencing data for cancer characterization are prone to errors, leading to inaccurate patient inferences and resource wastage due to improper sample handling, processing, and data bias, which can result in inappropriate treatments and missed therapeutic opportunities.

Method used

Developed techniques include novel sample preparation, post-processing to remove data bias, and quality control methods to verify the integrity and source of sequencing data, ensuring accurate representation of molecular characteristics for improved therapeutic identification.

Benefits of technology

These techniques enhance the accuracy of nucleic acid sequencing data by reducing errors and biases, enabling more faithful representation of patient molecular characteristics, leading to improved therapeutic identification and reduced resource wastage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007716384000015
    Figure 0007716384000015
  • Figure 0007716384000016
    Figure 0007716384000016
  • Figure 0007716384000017
    Figure 0007716384000017
Patent Text Reader

Abstract

Described herein are various methods for collecting and processing tumor and / or healthy tissue samples to extract nucleic acids and perform nucleic acid sequencing. Also described herein are various methods for processing nucleic acid sequencing data to remove bias from the nucleic acid sequencing data. Also described herein are various methods for assessing the quality of nucleic acid sequence information. The identity and / or completeness of nucleic acid sequence data is assessed before using the sequence information for subsequent analysis (e.g., for diagnostic, prognostic, or clinical purposes). These methods allow a subject, physician, or user to accurately characterize or classify various types of cancer, thereby determining a therapy or combination of therapies that may be effective in treating the subject's cancer based on accurate characterization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Application No. 62 / 870,622, filed Jul. 3, 2019, entitled "Compositions and Methods for Sample Preparation and Characterization of Cancer Therefrom", and U.S. Provisional Application No. 62 / 991,570, filed Mar. 18, 2020, entitled "Nucleic Acid Data Quality Control", the entire disclosure of each of which is incorporated herein by reference.

[0002] Some aspects of the technology described herein relate to collecting and processing tumor and / or healthy tissue samples to extract nucleic acids and perform nucleic acid sequencing. Some aspects of the technology described herein relate to processing nucleic acid sequencing data to remove bias from the nucleic acid sequencing data. Also described herein are various methods for assessing the quality of nucleic acid sequence information obtained by sequencing.

Background Art

[0003] Properly characterizing one or more types of cancer that a patient or subject has, and potentially selecting one or more effective therapies for the patient based on that characterization, can be extremely important for that patient's survival and overall health. The manner in which biological samples from a subject are processed to obtain array data (e.g., RNA expression data) for characterizing one or more types of cancer, and the manner in which the data are processed, can have a detrimental effect on the characterization of one or more cancers. For example, high-throughput nucleic acid sequencing platforms (e.g., next-generation sequencing platforms) can generate large amounts of DNA and RNA sequence data from patient samples. To characterize cancer, predict prognosis, identify effective therapies, and otherwise support the individualized care of patients suffering from cancer, there is a need for advancements in sample preparation, data processing, and the evaluation of sequence information from different NGS platforms by custom software.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Patent Document 5

Patent Document 6

Patent Document 7

Patent Document 8

Patent Document 9

Patent Document 10

Patent Document 11

Patent Document 12

Patent Document 13

Patent Document 14

Patent Document 15

Patent Document 16

Patent Document 17

Patent Document 18

Patent Document 19

Patent Document 20

Patent Document 21

Patent Document 22

Patent Document 23

Patent Document 24

Patent Document 25

Patent Document 26

Patent Document 27

Patent Document 28

[0005] [Non-Patent Document 1] Nicolas L Bray, Harold Pimentel, Pall Melsted and Lior Pachter, "Near-optimal probabilistic RNA-seq quantification", Nature Biotechnology 34, pp. 525 - 527 (2016), doi:10.1038 / nbt.3519 [Non-Patent Document 2] Vaught et al., "Biospecimen and biorepositories: from afterthought to science" (Cancer Epidemiol Biomarkers Prev. February 2012, 21(2):253 - 5)

Non-Patent Document 3

Non-Patent Document 4

Non-Patent Document 5

Non-Patent Document 6

Non-Patent Document 7

Non-Patent Document 8

Non-Patent Document 9

Non-Patent Document 10

Non-Patent Document 11

Non-Patent Document 12

Non-Patent Document 13

Non-Patent Document 14

Non-Patent Document 15

Non-Patent Document 16

Non-Patent Document 17

Non-Patent Document 18

Non-Patent Document 19

Non-Patent Document 20

Non-Patent Document 21

Non-Patent Document 22

Non-Patent Document 23

Non-Patent Document 24

Non-Patent Document 25

Non-Patent Document 26

Non-Patent Document 27

Non-Patent Document 28

Non-Patent Document 29

Non-Patent Document 30

Non-Patent Document 31

Non-Patent Document 32

Non-Patent Document 33

Non-Patent Document 34

Non-Patent Document 35

Non-Patent Document 36

Non-Patent Document 37

Non-Patent Document 38

Non-Patent Document 39

Non-Patent Document 40

Non-Patent Document 41

Non-Patent Document 42

Non-Patent Document 43

Non-Patent Document 44

Non-Patent Document 45

Non-Patent Document 46

Non-Patent Document 47

Non-Patent Document 48

Non-Patent Document 49

Non-Patent Document 50

Non-Patent Document 51

Summary of the Invention

Means for Solving the Problems

[0006] Some embodiments include at least one computer hardware processor and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method, the method comprising Obtaining nucleic acid data, wherein the nucleic acid data comprises sequence data showing at least 5 kilobases (kb) of nucleotide sequences of DNA and / or RNA from a previously obtained biological sample of a subject having a disease, suspected of having a disease, or at risk of having a disease, and claimed information indicating the claimed source and / or claimed completeness of the sequence data; obtaining nucleic acid data; processing the sequence data of the nucleic acid data to obtain determined information indicating the determined source and / or determined completeness of the sequence data; and verifying by determining whether the determined information matches the claimed information.

[0007] Some embodiments provide at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method. The method includes obtaining nucleic acid data, wherein the nucleic acid data comprises sequence data showing at least 5 kilobases (kb) of nucleotide sequences of DNA and / or RNA from a previously obtained biological sample of a subject having a disease, suspected of having a disease, or at risk of having a disease, and claimed information indicating the claimed source and / or claimed completeness of the sequence data; obtaining nucleic acid data; processing the sequence data of the nucleic acid data to obtain determined information indicating the determined source and / or determined completeness of the sequence data; and verifying by determining whether the determined information matches the claimed information.

[0008] Some embodiments involve using at least one hardware processor to obtain nucleic acid data, the nucleic acid data comprising sequence data showing at least 5 kilobases (kb) of nucleotide sequence of DNA and / or RNA from a previously obtained biological sample of a subject having a disease, suspected of having a disease, or at risk of having a disease, and claimed information indicating a claimed source and / or claimed integrity of the sequence data, obtaining the nucleic acid data, processing the nucleic acid data to obtain determined information indicating a determined source and / or determined integrity of the sequence data, and verifying by determining whether the determined information matches the claimed information.

[0009] In some embodiments, the sequence data can include raw DNA or RNA sequence data, DNA exome sequence data (e.g., from whole exome sequencing (WES)), DNA genomic sequence data (e.g., from whole genome sequencing (WGS)), RNA expression data, gene expression data, bias-corrected gene expression data, or data obtained from a sequencing platform and / or data derived from data obtained from a sequencing platform, and can include any other suitable type of sequence data.

[0010] In some embodiments, the method further includes determining whether the sequence data shows characteristics of one or more diseases when it is determined that the claimed information matches the determined information.

[0011] In some embodiments, the method further includes determining whether the determined information matches the claimed information and determining whether the sequence data shows characteristics of one or more diseases.

[0012] In some embodiments, the method generates an indication that the determined information does not match the claimed information when it is determined that the claimed information does not match the determined information, does not process the array data in subsequent analysis, and / or further includes obtaining additional array data and / or biological samples and / or other information about the subject.

[0013] In some embodiments, the method further includes determining that the claimed information does not match the determined information, generating an indication that the determined information does not match the claimed information, not processing the array data in subsequent analysis, and / or obtaining additional array data and / or biological samples and / or other information about the subject.

[0014] In some embodiments, the claimed information indicates a claimed source of the array data, and the method further includes processing the array data to obtain determined information indicating a determined source for the array data, and determining whether the determined source matches the claimed source for the array data.

[0015] In some embodiments, the determined information indicating a determined source for the array data indicates the subject's MHC genotype, whether the nucleic acid data is RNA data or DNA data, the tissue type of the biological sample, the tumor type of the biological sample, the sequencing platform used to generate the array data, SNP matches, and / or whether the RNA sample is polyA-enriched.

[0016] In some embodiments, the determined information indicating a determined source for the array data indicates at least two of the subject's MHC genotypes, whether the nucleic acid data is RNA data or DNA data, the tissue type of the biological sample, the tumor type of the biological sample, the sequencing platform used to generate the array data, SNP matches, and whether the RNA sample is polyA-enriched.

[0017] In some embodiments, the determined information indicating the determined source for the array data includes at least three of the subject's MHC genotypes, whether the nucleic acid data is RNA data or DNA data, the tissue type of the biological sample, the tumor type of the biological sample, the sequencing platform used to generate the array data, SNP matches, and whether the RNA sample is polyA-enriched.

[0018] In some embodiments, the asserted information indicates the asserted completeness of the array data, and the method further includes processing the array data to obtain determined information indicating the determined completeness of the array data, and determining whether the determined completeness matches the asserted completeness for the array data.

[0019] In some embodiments, the determined information indicating the determined completeness includes total sequence coverage, exon coverage, chromosomal coverage, the ratio of nucleic acids encoding two or more subunits of a multimeric protein, species contamination, single nucleotide polymorphisms (SNPs), complexity, and / or the percentage (%) of guanine (G) and cytosine (C) in the array data.

[0020] In some embodiments, the determined information indicating the determined completeness includes at least two of total sequence coverage, exon coverage, chromosomal coverage, the ratio of nucleic acids encoding two or more subunits of a multimeric protein, species contamination, single nucleotide polymorphisms (SNPs), complexity, and the percentage (%) of guanine (G) and cytosine (C) in the array data.

[0021] In some embodiments, the determined information indicating determined completeness includes at least three of total array coverage, exon coverage, chromosomal coverage, the ratio of nucleic acids encoding two or more subunits of a multimeric protein, species contamination, single nucleotide polymorphisms (SNPs), complexity, and the percentage (%) of guanine (G) and cytosine (C) in the sequence data.

[0022] In some embodiments, the claimed information for the sequence data includes the subject's MHC allele information.

[0023] In some embodiments, the method further includes determining one or more MHC allele sequences from the sequence data and determining whether the one or more MHC allele sequences match the claimed MHC allele information for the subject.

[0024] In some embodiments, determining one or more MHC allele sequences includes determining MHC allele sequences for six MHC loci from the sequence data.

[0025] In some embodiments, the sequence data represents a nucleotide sequence for RNA, and the claimed information indicates whether the RNA is polyA-enriched.

[0026] In some embodiments, the sequence data is used to determine a therapy for a subject when it is determined that the claimed information matches the determined information.

[0027] In some embodiments, determining the therapy includes determining a plurality of gene group expression levels, where the plurality of gene group expression levels includes the gene group expression level of each gene group in a set of gene groups, and the set of gene groups includes at least one gene group related to cancer malignancy and at least one gene group related to the cancer microenvironment, and using the determined gene group expression levels to identify a therapy.

[0028] In some embodiments, the method further includes administering a therapy to a subject.

[0029] In some embodiments, it is determined that the determined information matches the asserted information, the array data is processed to determine a subject's therapy, and the therapy is administered to the subject.

[0030] In some embodiments, the disease is cancer and the therapy is cancer treatment. In some embodiments, the subject is a human.

[0031] In some embodiments, obtaining the source determined by processing the array data includes determining one or more single nucleotide polymorphisms (SNPs) in the array data and determining whether one or more SNPs in the array data match one or more SNPs in a reference sequence.

[0032] In some embodiments, the reference sequence is the sequence of nucleic acids in a second biological sample of the subject.

[0033] In some embodiments, obtaining the integrity determined by processing the array data includes determining a first level of a first nucleic acid encoding a first subunit of a multimeric protein, determining a second level of a second nucleic acid encoding a second subunit of the multimeric protein, and determining whether the ratio between the first level and the second level matches an expected ratio. In some embodiments, the multimeric protein is a dimer. In some embodiments, the first subunit and the second subunit are the first and second CD3 subunits, the first and second CD8 subunits, or the first and second CD79 subunits.

[0034] Some embodiments provide a system for identifying cancer treatments for a subject having cancer, suspected of having cancer, or at risk of having cancer, the system comprising at least one sequencing platform configured to generate gene expression data from enriched RNA obtained from a first biological sample previously obtained from the subject, the enriched RNA being obtained by: (i) extracting RNA from a first biological sample of a first tumor to obtain extracted RNA; and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA, wherein the RNA expression data comprises at least 5 kilobases (kb), the enriched RNA being obtained by obtaining enriched RNA; at least one computer hardware processor; and at least one non-transitory computer-readable storage medium storing processor-executable instructions, the processor-executable instructions, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to: obtain RNA expression data using the at least one sequencing platform; convert the RNA expression data to gene expression data; determine bias-corrected gene expression data by removing expression data for at least one gene that, at least in part, introduces bias into the gene expression data from the gene expression data; and identify a cancer treatment for the subject using the bias-corrected gene expression data.

[0035] Some embodiments provide a system for identifying cancer treatments for a subject having cancer, suspected of having cancer, or at risk of having cancer, the system comprising at least one computer hardware processor and at least one non-transitory computer-readable storage medium storing processor-executable instructions, the processor-executable instructions, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to obtain RNA expression data from at least one sequencing platform, the RNA expression data comprising at least 5 kilobases (5 kb), the RNA expression data being obtained from a first biological sample of a first tumor previously obtained from the subject, at least in part by (i) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA, converting the RNA expression data to gene expression data, determining bias-corrected gene expression data from the gene expression data, at least in part by removing expression data for at least one gene that introduces bias into the gene expression data from the gene expression data, and identifying a cancer treatment for the subject using the bias-corrected gene expression data. The system may further comprise at least one sequencing platform in some embodiments.

[0036] Some embodiments are at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to obtain RNA expression data from at least one sequencing platform, where the RNA expression data includes at least 5 kilobases (5 kb), and the RNA expression data is from a first biological sample of a first tumor previously obtained from a subject having cancer, suspected of having cancer, or at risk of having cancer, and at least a portion of the obtaining includes (i) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA, and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA; obtaining RNA expression data; converting the RNA expression data to gene expression data; determining bias-corrected gene expression data from the gene expression data, at least in part by removing expression data for at least one gene that introduces bias into the gene expression data from the gene expression data; and identifying cancer treatment for the subject using the bias-corrected gene expression data.

[0037] Some embodiments include obtaining a first biological sample of a first tumor, the first biological sample having been previously obtained from a subject having cancer, suspected of having cancer, or at risk of having cancer; extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; enriching the extracted RNA for coding RNA to obtain enriched RNA; performing sequencing of the enriched RNA using at least one sequencing platform to obtain RNA expression data including at least 5 kilobases (kb); using at least one hardware processor to obtain the RNA expression data using at least one sequencing platform; converting the RNA expression data to gene expression data; determining bias-corrected gene expression data by removing, from the gene expression data, expression data for at least one gene that, at least in part, introduces bias into the gene expression data from the gene expression data; and using the bias-corrected gene expression data to identify cancer treatment for the subject.

[0038] In some embodiments, the method further includes administering the identified cancer treatment to the subject.

[0039] In some embodiments, enriching the RNA for coding RNA includes performing polyA enrichment.

[0040] In some embodiments, at least one gene that introduces bias into the gene expression data includes a gene having an average transcript length that is longer or shorter than the average transcript length of transcripts in the gene expression data, a gene having at least one threshold variation in average transcript expression level based on transcript expression levels in a reference sample, and / or a gene having a polyA tail length that is at least a threshold amount smaller compared to the average length of polyA tails of genes from the first biological sample and / or reference sample from which the RNA expression data was obtained.

[0041] In some embodiments, at least one gene that introduces bias into the gene expression data belongs to a gene family selected from the group consisting of histone-coding genes, mitochondrial genes, interleukin-coding genes, collagen-coding genes, B-cell receptor-coding genes, and T-cell receptor-coding genes.

[0042] In some embodiments, at least one gene includes at least one histone-coding gene selected from the group consisting of HIST1H1A, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H1T, HIST1H2AA, HIST1H2AB, HIST1H2AC, HIST1H2AD, HIST1H2AE, HIST1H2AG, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H2AK, HIST1H2AL, HIST1H2AM, HIST1H2BA, HIST1H2BB, HIST1H2BC, HIST1H2BD, HIST1H2BE, HIST1H2BF, HIST1H2BG, HIST1H2BH, HIST1H2BI, HIST1H2BJ, HIST1H2BK, HIST1H2BL, HIST1H2BM, HIST1H2BN, HIST1H2BO, HIST1H3A, HIST1H3B, HIST1H3C, HIST1H3D, HIST1H3E, HIST1H3F, HIST1H3G, HIST1H3H, HIST1H3I, HIST1H3J, HIST1H4A, HIST1H4B, HIST1H4C, HIST1H4D, HIST1H4E, HIST1H4F, HIST1H4G, HIST1H4H, HIST1H4I, HIST1H4J, HIST1H4K, HIST1H4L, HIST2H2AA3, HIST2H2AA4, HIST2H2AB, HIST2H2AC, HIST2H2BE, HIST2H2BF, HIST2H3A, HIST2H3C, HIST2H3D, HIST2H3PS2, HIST2H4A, HIST2H4B, HIST3H2A, HIST3H2BB, HIST3H3, and HIST4H4.

[0043] In some embodiments, at least one gene comprises at least one mitochondrial gene selected from the group consisting of MT-ATP6, MT-ATP8, MT-CO1, MT-CO2, MT-CO3, MT-CYB, MT-ND1, MT-ND2, MT-ND3, MT-ND4, MT-ND4L, MT-ND5, MT-ND6, MT-RNR1, MT-RNR2, MT-TA, MT-TC, MT-TD, MT-TE, MT-TF, MT-TG, MT-TH, MT-TI, MT-TK, MT-TL1, MT-TL2, MT-TM, MT-TN, MT-TP, MT-TQ, MT-TR, MT-TS1, MT-TS2, MT-TT, MT-TV, MT-TW, MT-TY, MTRNR2L1, MTRNR2L10, MTRNR2L11, MTRNR2L12, MTRNR2L13, MTRNR2L3, MTRNR2L4, MTRNR2L5, MTRNR2L6, MTRNR2L7, and MTRNR2L8.

[0044] In some embodiments, determining bias-corrected gene expression data further comprises re-normalizing the gene expression data after removing expression data for at least one gene that introduces bias into the gene expression data.

[0045] In some embodiments, converting RNA expression data to gene expression data comprises obtaining filtered RNA expression data by removing non-coding transcripts from the RNA expression data, and after removing the non-coding transcripts, normalizing the filtered RNA expression data to obtain gene expression data in Transcripts Per Million (TPM) and / or any other suitable format.

[0046] In some embodiments, removing non-coding transcripts from RNA expression data involves removing non-coding transcripts belonging to a group selected from the list consisting of pseudogenes, polymorphic pseudogenes, processed pseudogenes, transcribed processed pseudogenes, unitary pseudogenes, non-processed pseudogenes, transcribed unitary pseudogenes, constant chain immunoglobulin (IG C) pseudogenes, joining chain immunoglobulin (IG J) pseudogenes, variable chain immunoglobulin (IG V) genes, transcribed non-processed genes, translated non-processed genes, joining chain T cell receptor (TR J) genes, variable chain T cell receptor (TR V) genes, small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), microRNA (miRNA), ribozymes, ribosomal RNA (rRNA), mitochondrial tRNA (Mt tRNA), mitochondrial rRNA (Mt rRNA), Cajal body-specific RNA (scaRNA), retained introns, sense intron RNA, sense overlapping RNA, nonsense-mediated decay RNA, non-stop decay RNA, antisense RNA, long intergenic non-coding RNA (lincRNA), macro long non-coding RNA (macro lncRNA), processed transcripts, 3' overlapping non-coding RNA (3' overlapping ncrna), small RNA (sRNA), other RNA (miscRNA), vault RNA, and TEC RNA.

[0047] In some embodiments, information (e.g., sequence information) regarding one or more of these types of transcripts can be obtained from a nucleic acid database (e.g., the Gencode database, e.g., Gencode V23, the Genbank database, the EMBL database, or other databases).

[0048] In some embodiments, the method further includes aligning the RNA expression data to a reference and annotating the RNA expression data before performing the removal of non-coding transcripts.

[0049] In some embodiments, the RNA expression data comprises at least 25 million paired-end reads. In some embodiments, the RNA expression data comprises at least 50 million paired-end reads with an average read length of at least 100 bp.

[0050] In some embodiments, identifying cancer treatment for a subject using bias-corrected gene expression data comprises determining a plurality of gene group expression levels using the bias-corrected gene expression data, wherein the plurality of gene group expression levels comprises the gene group expression level of each gene group in a set of gene groups, and the set of gene groups comprises at least one gene group related to cancer malignancy and at least one gene group related to the cancer microenvironment, and identifying cancer treatment using the determined gene group expression levels.

[0051] In some embodiments, the cancer treatment is selected from the group consisting of radiation therapy, surgical therapy, chemotherapy, and immunotherapy.

[0052] In some embodiments, the method further comprises obtaining a second biological sample of a second tumor, the second biological sample having been previously obtained from the subject.

[0053] In some embodiments, the method further comprises combining the first biological sample and the second biological sample to form a combined tumor sample, and extracting RNA comprises extracting RNA from the combined tumor sample.

[0054] In some embodiments, the method further comprises extracting RNA from the second biological sample and combining the RNA extracted from the second biological sample with the RNA extracted from the first biological sample to form combined extracted RNA, and enriching RNA for coding RNA comprises enriching the combined extracted RNA for coding RNA.

[0055] In some embodiments, the extracted RNA contains at least 1 μg of RNA after RNA extraction.

[0056] In some embodiments, the extracted RNA has a total mass of at least 1000 - 6000 ng and a purity corresponding to a ratio of absorbance at 260 nm to absorbance at 280 nm that is at least 2.0.

[0057] In some embodiments, the method further includes performing a quality control assessment on the RNA expression data by obtaining at least some information that indicates the claimed source and / or claimed completeness of the RNA expression data, processing the RNA expression data to obtain determined information that indicates the determined source and / or determined completeness of the RNA expression data, and determining whether the determined information matches the claimed information.

[0058] In some embodiments, processing the RNA expression data includes processing the RNA expression RNA to determine the tissue type of the first biological sample, the tumor type of the first biological sample, and / or the percentage (%) of guanine (G) and / or cytosine (C).

[0059] The following drawings form a part of this specification and are included to further illustrate some aspects of the present disclosure and can be better understood by referring to one or more of these drawings in combination with the detailed description of the specific embodiments presented herein. The drawings are not necessarily drawn to scale.

Brief Description of the Drawings

[0060]

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 7A

Figure 7B

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13A

Figure 13B

Figure 14A

Figure 14B

Figure 15

BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Recent advances in personalized genomic sequencing and cancer genomic sequencing technologies have made it possible to obtain patient-specific information regarding cancer cells (e.g., tumor cells) and the cancer microenvironment from one or more biological samples obtained from individual patients. The inventors have understood that this information can be used to characterize the type of cancer a patient has and potentially to select one or more effective therapies for the patient. This information can also be used to determine how a patient is responding to treatment over time and, if necessary, to select new one or more therapies for the patient as needed. This information can also be used to determine whether a patient should be included or excluded from participation in a clinical trial.

[0062] The inventors recognize that the workflow used to obtain sequence data for a patient strongly influences the inferences that can be drawn about the patient's cancer. Such inferences include, but are not limited to, whether a patient will respond to a particular one or more therapies, whether a patient has a refractory response to a particular one or more therapies, whether a patient is a candidate for enrollment in a clinical trial, whether a patient has one or more particular biomarkers (e.g., biomarkers indicating potential response to a therapy, biomarkers indicating survival rate, etc.), whether there is disease progression in the patient (e.g., from early-stage cancer to late-stage cancer, recurrence from remission, etc.), whether one or more different therapies should be selected for the patient, and / or other suitable prognostic, diagnostic, and / or clinical inferences.

[0063] If the workflows used to obtain array data contain errors, suboptimal processing, sources of bias in the data, and the like, it may not be possible to make inferences about a subject's cancer while maintaining the desired or necessary level of reliability, or it may not even be possible to make such inferences at all. Even worse, if there are errors in the workflows used to generate array data, inferences about the patient may not be made correctly, potentially resulting in inappropriate treatment or missed opportunities for better treatment. Additionally, workflow errors can lead to waste of resources in the laboratory (e.g., having to reprocess samples) and waste of computing resources (e.g., performing costly computational processing on megabytes or gigabytes of array data, occupying processor and networking resources but ultimately discarding the results and / or having to repeat the process).

[0064] Conventional workflows used to obtain array data for a patient include multiple steps of obtaining a biological sample from the patient (e.g., by performing a biopsy, obtaining a blood sample, a saliva sample, or any other suitable biological sample from the patient), preparing the biological sample for sequencing using a sequencing platform (e.g., a next-generation sequencing (NGS) platform), and obtaining the raw data output by the sequencing platform. Various conventional bioinformatics processing pipelines and other algorithms may then use the raw data output by the sequencing platform in an attempt to perform one or more of the above inferences.

[0065] However, such conventional workflows for obtaining sequencing data are prone to errors at all stages. For example, in a laboratory, errors can occur when handling samples from multiple patients. In fact, it is not uncommon for a laboratory to receive a biological sample claimed to be from one patient when the sample is actually from another patient. As another example, a biological sample may not be properly processed in the laboratory and may not have the nucleic acid concentration and / or quality required for subsequent analysis. As yet another example, errors can be introduced by the sequencing platform itself and / or subsequent post-processing steps (e.g., alignment and variant calling). As yet another example, raw sequencing data generated by the sequencing platform may contain artifacts and unwanted sequences and / or transcripts. Other examples of various errors are described herein.

[0066] In some embodiments, the sequence data or sequencing data may include raw DNA or RNA sequence data, DNA exome sequence data (e.g., from whole exome sequencing (WES)), DNA genomic sequence data (e.g., from whole genome sequencing (WGS)), RNA expression data, gene expression data, bias-corrected gene expression data, or data obtained from a sequencing platform and / or data derived from data obtained from a sequencing platform, and may include any other suitable type of sequence data.

[0067] To address the drawbacks of conventional workflows for obtaining sequencing data for patients, the inventors have developed techniques to address the various sources of error that may be present in sequencing data. These techniques developed by the inventors include: (1) novel sample preparation techniques for preparing biological samples for sequencing using one or more sequencing platforms; (2) novel techniques for post-processing the raw data output by a sequencing platform to remove sources of irrelevant data and bias (e.g., transcripts of non-coding regions that introduce bias into sequence data and expression data related to genes); and (3) novel quality control techniques that facilitate the detection and correction of errors within sequence data. In some embodiments, techniques from each of these three categories can be utilized in a workflow for obtaining sequence data for a patient, although this is not a limitation of the techniques described herein, and it should be understood that in some embodiments, any one or more (but not necessarily all) of these techniques can be used in a workflow.

[0068] As an example, in some embodiments, the novel sample preparation techniques and post-processing techniques include obtaining sequencing data and removing sources of bias from the sequencing data by: (1) obtaining a first biological sample of a first tumor, the first biological sample having been previously obtained from a subject having cancer, suspected of having cancer, or at risk of having cancer; (2) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; (3) enriching the extracted RNA for coding RNA to obtain enriched RNA; (4) performing sequencing of the enriched RNA using at least one sequencing platform to obtain RNA expression data including at least 5 kilobases (kb); and (5) using at least one hardware processor to: (a) obtain RNA expression data using at least one sequencing platform; (b) convert the RNA expression data to gene expression data; (c) determine bias-corrected gene expression data by removing expression data for at least one gene that, at least in part, introduces bias into the gene expression data from the gene expression data; and (d) identify cancer treatment for the subject using the bias-corrected gene expression data.

[0069] Removing bias from gene expression data in this way provides improvements to sequencing technology for a number of reasons. First, this removes sources of artifacts and bias from the sequencing data, resulting in fewer errors in any downstream processing and a more faithful output. Second, the inventors recognize that removing sources of bias in this way enables the more accurate and faithful representation of a patient's molecular functional characteristics (e.g., via the molecular function expression signatures described herein). The inventors recognize that bias-corrected gene expression data can be used to identify more effective therapies for patients, to determine whether one or more cancer therapies are effective when administered to a patient, to identify clinical trials in which a subject may participate, and / or to identify improvements to many other prognostic, diagnostic, and clinical applications.

[0070] As another example, in some embodiments, a novel quality control technique uses at least one computer hardware processor to (a) obtain nucleic acid data, the nucleic acid data comprising (i) sequence data showing at least 5 kilobases (kb) of nucleotide sequence of DNA and / or RNA from a previously obtained biological sample of a subject having, suspected of having, or at risk of having a disease, and (ii) claimed information indicating a claimed source and / or claimed completeness of the sequence data, and (b) verify the nucleic acid data by (i) processing the sequence data to obtain determined information indicating a determined source and / or determined completeness of the sequence data, and (ii) determining whether the determined information matches the claimed information. Examples of various such verification techniques are described herein and are important examples of the quality control techniques developed by and described by the inventors herein.

[0071] Adopting such quality control techniques also leads to improvements in sequencing technology and computer technology. First, sequencing data that fails one or more quality control checks is not used in part or in whole for downstream processing that reduces or eliminates errors in downstream applications (e.g., identification of biomarkers, tumor microenvironment types, possible therapies for patients, etc.). Often, such downstream processing requires performing computationally expensive (often cloud-based) processing of large datasets (e.g., sequencing data contains tens of millions of reads that must be aligned, annotated, and processed in many other ways). By preventing computationally expensive processes from being executed using quality control, wasteful use of computing resources is reduced or eliminated, saving processing power, memory, and networking resources (which is an improvement in computing technology in addition to being an improvement in sequencing technology). Also, by identifying errors, waste of resources in a laboratory that processes multiple samples can be reduced by freeing up equipment to process biological samples that passed the initial quality control checks. In addition, using array data for downstream processing that passed various quality control checks improves the ability to identify more effective therapies for patients, determine whether one or more cancer therapies are effective when administered to a patient, identify clinical trials that a subject can participate in, and / or identify many other improvements to prognosis, diagnosis, and clinical applications.

[0072] Figures 1A and 1B show an example of a process pipeline for sample preparation and quality control as described herein. The process pipeline of FIG. 1 illustrates embodiments of the methods and systems provided in the present disclosure and should not be construed as limiting the scope in any way. The present disclosure stipulates that the process pipeline need not include all of the process steps illustrated in FIG. 1 or the order of the process steps. One or more processes may be omitted, repeated, or executed in a different order depending on the application.

[0073] Figure 1A shows a non-limiting process pipeline 100 that includes one or more quality control evaluations. In activity 101, a biological sample (e.g., a tumor biopsy) is obtained from a subject (e.g., a subject having cancer, suspected of having cancer, or at risk of having cancer). In some embodiments, the sample is obtained from a physician, hospital, clinic, or other healthcare service provider. One or more sample quality control evaluations in quality control activity 102 can be performed on the biological sample. In some embodiments, the quality control evaluation of the biological sample (e.g., biopsy material) includes determining whether the sample is in the appropriate form (e.g., fresh frozen or FFPE) and / or whether it is accompanied by sufficient information to identify the nature and source of the sample. Thereafter, nucleic acids (e.g., DNA and / or RNA) can be extracted from the biological sample that meets the conditions of sample quality control activity 102. Next, one or more nucleic acid quality control evaluations are performed in activity 103, thereby enabling the evaluation of one or more physical attributes of, for example, the extracted nucleic acids, nucleic acid libraries prepared from the extracted nucleic acids, and / or pooled nucleic acids or libraries. Thereafter, nucleic acids (e.g., DNA and / or RNA) that meet the conditions of nucleic acid quality control activity 103 are processed (e.g., concentrated for polyA RNA) and / or sequenced to obtain raw DNA and / or RNA sequence data (e.g., RNA expression data). In some embodiments, the RNA expression data is processed to obtain gene expression data, and optionally, data for one or more types of genes that may interfere with (e.g., bias) subsequent analysis of the gene expression data can be removed. In some embodiments, the gene expression data is normalized (e.g., after removing data for one or more interfering genes).In some embodiments, one or more array quality control assessments can be performed on DNA and / or RNA sequence data (e.g., processed, e.g., normalized, gene expression data) for bioinformatics quality control activities 104. In some embodiments, one or more bioinformatics quality control assessments are performed to determine whether the sequence data is from a predicted source (e.g., patient, tissue, tumor, etc.) and / or whether it has sufficient integrity for further analysis. In some embodiments, sequence data that meets the conditions of bioinformatics quality control activities 104 is further processed, for example, to evaluate and / or monitor a subject, and / or for one or more clinical applications (e.g., to evaluate a therapy) to determine a diagnosis, prognosis, and / or therapy for the subject.

[0074] In some embodiments, the array data includes raw DNA or RNA sequence data, DNA exome sequence data (e.g., from whole exome sequencing (WES)), DNA genomic sequence data (e.g., from whole genome sequencing (WGS)), RNA expression data, gene expression data, bias-corrected gene expression data, or data obtained from a sequencing platform, and / or without limitation, data derived from data obtained from such a sequencing platform including, for example, such data as described herein, and may include any other suitable type of array data. FIG. 1B illustrates a non-limiting process pipeline 110 for preparing nucleic acids from a biological sample (e.g., a tumor biopsy) and obtaining and processing nucleic acid sequence data for subsequent analysis (e.g., for diagnosis, prognosis, treatment, and / or other clinical applications). Process pipeline 110 is executed by obtaining a biological sample (e.g., a tumor sample) from a subject having cancer, suspected of having cancer, or at risk of having cancer in activity 111. Nucleic acids (e.g., DNA and / or RNA) are obtained (e.g., extracted) from the sample in activity 112. In activity 113, one or more quality control assessments of the nucleic acids are performed. One or more nucleic acid libraries are prepared in activity 114 using nucleic acids that meet the conditions of at least one quality control assessment of activity 113, for example. The nucleic acid library is sequenced in sequencing activity 115 using at least one sequencing platform (e.g., to obtain RNA expression data for RNA). In some embodiments, the RNA expression data is converted to gene expression data in activity 116, and the gene expression data is optionally bias-corrected by removing expression data for at least one gene that, at least in part, introduces bias into the gene expression data.One or more bioinformatics quality control evaluations are performed on DNA sequence data and / or RNA sequence data from activity 115 and / or RNA sequence data (e.g., bias-corrected gene expression data from activity 116) in bioinformatics quality control activity 117. In some embodiments, the nucleic acid data (e.g., meeting the conditions of at least one bioinformatics quality control evaluation of activity 117) is further processed in activity 118 (e.g., determining one or more signs of a disease from gene expression data), and a diagnosis, prognosis, treatment, and / or other clinical evaluation of the subject is performed in activity 119 (e.g., identifying a treatment for the subject, such as a cancer treatment). In some embodiments, a treatment (e.g., a cancer treatment) is administered to the subject.

[0075] In some embodiments, activity 111 includes obtaining a bulk biopsy tissue of a subject or patient. In some embodiments, activity 111 includes obtaining a blood sample of a subject or patient. In some embodiments, activity 111 includes obtaining a single cell suspension. In some embodiments, activity 111 includes obtaining any type of sample suitable for preparing nucleic acids for subsequent sequencing analysis. In some embodiments, activity 111 includes obtaining multiple types of samples.

[0076] In some embodiments, when bulk biopsy tissue is obtained, the tissue is processed (e.g., homogenized in the presence of TriZol), and nucleic acids such as DNA or RNA are extracted in activity 112. In some embodiments, when a single cell suspension is obtained, the suspension is processed to extract nucleic acids such as DNA or RNA in activity 112. In some embodiments, nucleic acids suitable for germline whole exome sequencing (WES) may be extracted in activity 112. In some embodiments, nucleic acids suitable for tumor whole exome sequencing (WES) may be extracted in activity 112. In some embodiments, nucleic acids suitable for tumor RNA sequencing may be extracted in activity 112. In some embodiments, nucleic acids suitable for CYTOF (mass cytometry) may be extracted in activity 112. In some embodiments, nucleic acids suitable for any type of sequencing known in the art may be extracted in activity 112.

[0077] In Activity 113, one or more quality control evaluations may be performed. Acceptable thresholds and / or target thresholds may be determined and used as references. In some embodiments, the total amount of extracted DNA or RNA may be used for quality control evaluations. In some embodiments, a spectrophotometer, e.g., a small-volume full-spectrum ultraviolet-visible spectrophotometer (e.g., NanoDrop spectrophotometer available from ThermoFisher Scientific, www.thermofisher.com), can be used for quality control evaluations of DNA or RNA. In some embodiments, a fluorometer (e.g., Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com), for example, for quantification of DNA or RNA, may be used for quality control evaluations of DNA or RNA. In some embodiments, an automated electrophoresis system (e.g., TAPESTATION) may be used for quality control evaluations of DNA or RNA. In some embodiments, a real-time PCR system (e.g., LIGHTCYCLER®) may be used for quality control evaluations of DNA or RNA.

[0078] In some embodiments, Activity 114 includes preparing a library for the extracted nucleic acids that meet the conditions of at least one quality control threshold in Activity 113. In some embodiments, Activity 114 includes one or more methods described in Example 2.

[0079] In some embodiments, Activity 115 includes sequencing nucleic acids (e.g., DNA, RNA, or related libraries of Activity 114) using at least one nucleic acid sequencing platform (e.g., a next-generation nucleic acid sequencing platform) to obtain DNA sequence data and / or RNA sequence data (e.g., RNA expression data). The sequence data obtained in Activity 115 may be stored in any suitable format (e.g., in the form of one or more FASTQ files).

[0080] In some embodiments, the RNA expression data is converted to gene expression data in activity 116. In some embodiments, the RNA expression data is aligned to known genes in a database, such as a known assembled genome (e.g., the human genome) or a transcriptome in the database. In some embodiments, programs for quantifying transcripts using high-throughput sequencing reads (e.g., Kallisto (hg38) available from Github, www.github.com, e.g., as described in Nicolas L Bray, Harold Pimentel, Pall Melsted and Lior Pachter, "Near-optimal probabilistic RNA-seq quantification", Nature Biotechnology 34, pages 525-527 (2016), doi:10.1038 / nbt.3519) and / or Gencode (e.g., Gencode V23) are used for sequence alignment and / or annotation, for example, from bulk and single cell RNA-Seq data. In some embodiments, activity 116 includes gene aggregation. In some embodiments, activity 116 includes removing expression data for one or more non-coding transcripts from the gene expression data. In some embodiments, activity 116 includes removing expression data for one or more genes that can bias the gene expression data. In some embodiments, activity 116 includes removing expression data for histone-coded genes and / or mitochondrial-coded genes. In some embodiments, activity 116 includes performing normalization (e.g., TPM normalization) after removing expression data for non-coding and / or bias-related genes from the gene expression data. This normalization may be referred to herein as "re-normalization".

[0081] In activity 117, one or more bioinformatics quality control evaluations are performed on nucleic acid sequence data, such as DNA sequence data and / or RNA sequence data (e.g., bias-corrected and / or normalized gene expression data). In some embodiments, one or more bioinformatics quality control evaluations may be performed to evaluate the source and / or integrity of the nucleic acid sequence data. In some embodiments, one or more bioinformatics quality control evaluations described in this application are performed.

[0082] In some embodiments, the method includes all of the processes illustrated in FIG. 1. However, in some embodiments, a subset of the processes may be performed, and any one or more of those processes may be omitted, repeated, and / or performed in an order different from that illustrated in FIG. 1. In some embodiments, the method includes a process that optionally includes one or more quality control steps for preparing nucleic acids from a biological sample, and the nucleic acids are sequenced on at least one sequencing platform. In some embodiments, the method includes processing nucleic acid information obtained (e.g., received) from a sequencing platform to generate DNA or RNA sequence data (e.g., bias-corrected, optionally normalized gene expression data for subsequent analysis). In some embodiments, one or more of the processes of FIG. 1 are implemented on a computer. In some embodiments, the method includes identifying a treatment (e.g., cancer treatment) for a subject (e.g., a subject having cancer, suspected of having cancer, or at risk of having cancer). In some embodiments, the method includes administering the treatment to the subject.

[0083] Biological sample The method, system, or other element recited in the claims can all be used to use or analyze a biological sample from a subject. In some embodiments, the biological sample is obtained from a subject having or suspected of having cancer. One or more biological samples from a subject can be analyzed as described herein to obtain information regarding the subject's cancer. The biological sample can be, for example, a body fluid (e.g., blood, urine, or cerebrospinal fluid), one or more cells (e.g., from scraping or brushing such as buccal mucosal specimen collection or tracheal brushing), a tissue piece (buccal tissue, muscle tissue, lung tissue, heart tissue, brain tissue, or skin tissue), or an organ (such as the brain, lung, liver, bladder, kidney, pancreas, intestine, or muscle), or any other type of biological sample including a part or all of them, or other types of biological samples (e.g., feces or hair).

[0084] In some embodiments, the biological sample is a sample of a tumor from a subject. In some embodiments, the biological sample is a sample of blood from a subject. In some embodiments, the biological sample is a sample of tissue from a subject.

[0085] The sample of a tumor, in some embodiments, refers to a sample containing cells from the tumor. In some embodiments, the sample of a tumor contains cells from a benign tumor, e.g., non-cancerous cells. The sample of a tumor contains cells from a pre-malignant tumor, e.g., pre-cancerous cells. In some embodiments, the sample of a tumor contains cells from a malignant tumor, e.g., cancerous cells.

[0086] Examples of tumors include, but are not limited to, adenoma, fibroma, hemangioma, lipoma, cervical dysplasia, pulmonary metaplasia, leukoplakia, carcinoma, sarcoma, germ cell tumor, and blastoma.

[0087] A blood sample, in some embodiments, refers to a sample containing cells, e.g., cells from a blood sample. In some embodiments, the blood sample contains non-cancerous cells. In some embodiments, the blood sample contains pre-cancerous cells. In some embodiments, the blood sample contains cancer cells. In some embodiments, the blood sample contains blood cells. In some embodiments, the blood sample contains red blood cells. In some embodiments, the blood sample contains white blood cells. In some embodiments, the blood sample contains platelets. Examples of cancerous blood cells include, but are not limited to, leukemia, lymphoma, and myeloma. In some embodiments, a blood sample is taken to obtain cell-free nucleic acids (e.g., cell-free DNA) in the blood.

[0088] The blood sample may be a sample of whole blood or a fractionated blood sample. In some embodiments, the blood sample contains whole blood. In some embodiments, the blood sample contains fractionated blood. In some embodiments, the blood sample contains meninges. In some embodiments, the blood sample contains serum. In some embodiments, the blood sample contains plasma. In some embodiments, the blood sample contains a blood clot.

[0089] A tissue sample, in some embodiments, refers to a sample containing cells from a tissue. In some embodiments, a tumor sample contains non-cancerous cells from a tissue. In some embodiments, a tumor sample contains pre-cancerous cells from a tissue. In some embodiments, a tumor sample contains pre-cancerous cells from a tissue.

[0090] The methods of the present disclosure encompass a variety of tissues, including but not limited to organ tissues or non-organ tissues, including muscle tissue, brain tissue, lung tissue, liver tissue, epithelial tissue, connective tissue, and nerve tissue. In some embodiments, the tissue may be normal tissue, diseased tissue, or tissue suspected of being diseased. In some embodiments, the tissue may be a tissue section or intact tissue. In some embodiments, the tissue may be animal tissue or human tissue. Animal tissues include, but are not limited to, tissues obtained from rodents (e.g., rats or mice), primates (e.g., monkeys), dogs, cats, and livestock.

[0091] A biological sample may be from any source within a subject and may include, but is not limited to, any body fluid [such as blood (e.g., whole blood, serum, or plasma), saliva, tears, synovial fluid, cerebrospinal fluid, pleural fluid, pericardial fluid, ascites, and / or urine, etc.], hair, skin (including a portion of the epidermis, dermis, and / or hypodermis), oropharynx, larynx, esophagus, stomach, bronchi, salivary glands, tongue, oral cavity, nasal cavity, vaginal cavity, anal cavity, bone, bone marrow, brain, thymus, spleen, small intestine, appendix, colon, rectum, anus, liver, biliary tract, pancreas, kidney, ureter, bladder, urethra, uterus, vagina, vulva, ovaries, neck, scrotum, penis, prostate, testicles, seminal vesicles, and / or any type of tissue (e.g., muscle tissue, epithelial tissue, connective tissue, or nerve tissue).

[0092] Any of the biological samples described herein can be obtained from a subject using any known technique. For example, for the collection, processing, and storage of biological samples, see the publications, Vaught et al., "Biospecimens and biorepositories: from afterthought to science" (Cancer Epidemiol Biomarkers Prev. February 2012, 21(2):253-5), and Vaught and Henderson, "Biological sample collection, processing, storage and information management" (IARC Sci Publ. 2011, (163):23-42), each of which is incorporated herein by reference in its entirety.

[0093] In some embodiments, the biological sample can be obtained from a surgical procedure (e.g., laparoscopic surgery, microsurgery, or endoscopic surgery), a bone marrow biopsy, a punch biopsy, an endoscopic biopsy, or a needle biopsy (e.g., fine needle aspiration, core needle biopsy, vacuum-assisted biopsy, or image-guided biopsy). In some embodiments, the biological sample can be obtained from an autopsy.

[0094] In some embodiments, one or more cells (i.e., a cellular biological sample) can be obtained from a subject using a scraping or brushing method. The cellular biological sample can be obtained from any region within or from the body of the subject, including, for example, one or more regions of the neck, esophagus, stomach, bronchus, or mouth. In some embodiments, one or more tissue pieces (e.g., a tissue biopsy) from the subject can be used. In some embodiments, the tissue biopsy can include one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or 10 or more) biological samples from one or more tumors or tissues known or suspected to have cancerous cells.

[0095] Any of the biological samples from a subject described herein can be stored using any method that maintains the stability of the biological sample. In some embodiments, maintaining the stability of a biological sample means preventing the components of the biological sample (e.g., DNA, RNA, protein, or the structure or morphology of the tissue) from degrading until measured such that when measured, the measured value represents the state of the sample when the sample was obtained from the subject. In some embodiments, the biological sample is stored in a composition that can penetrate it and protect the components of the biological sample (e.g., DNA, RNA, protein, or the structure or morphology of the tissue) from degrading. As used herein, degradation is a conversion of one component to another such that the initial form is no longer detected at the same level as before degradation.

[0096] In some embodiments, the biological sample is stored using cryopreservation. Non-limiting examples of cryopreservation include, but are not limited to, step-down freezing, rapid freezing, direct plunge freezing, snap freezing, slow freezing using a programmable freezer, and vitrification. In some embodiments, the biological sample is stored using lyophilization. In some embodiments, after the biological sample is collected from the subject, it is placed into a container that already contains a preservative (e.g., RNALater for preserving RNA) and then frozen (e.g., by snap freezing). In some embodiments, such storage in the frozen state is performed immediately after collection of the biological sample. In some embodiments, the biological sample may be kept at either room temperature or 4°C for some time (e.g., up to 1 hour, up to 8 hours, or up to 1 day, or for several days) in a preservative or in a buffer without a preservative before being frozen.

[0097] Non-limiting examples of preservatives include formalin solution, formaldehyde solution, RNALater or other equivalent solutions, TriZol or other equivalent solutions, DNA / RNA Shield or equivalent solutions, EDTA (e.g., Buffer AE (10 mM Tris-Cl, 0.5 mM EDTA, pH 9.0)) and other coagulants, and Acids Citrate Dextronse (e.g., for blood samples).

[0098] In some embodiments, special containers may be used to collect and / or store the biological sample. For example, a vacutainer may be used to store blood. In some embodiments, the vacutainer may contain a preservative (e.g., a coagulant, or an anticoagulant). In some embodiments, the container in which the biological sample is stored may be housed in a secondary container for better preservation or for the purpose of avoiding contamination.

[0099] Any of the biological samples from a subject described herein can be stored under any conditions that maintain the stability of the biological sample. In some embodiments, the biological sample is stored at a temperature that maintains the stability of the biological sample. In some embodiments, the sample is stored at room temperature (e.g., 25°C). In some embodiments, the sample is stored under refrigeration (e.g., 4°C). In some embodiments, the sample is stored under freezing conditions (e.g., -20°C). In some embodiments, the sample is stored under ultra-low temperature conditions (e.g., -50°C to -800°C). In some embodiments, the sample is stored under liquid nitrogen (e.g., -1700°C). In some embodiments, the biological sample is stored at -60°C to -8°C (e.g., -70°C) for up to 5 years (e.g., up to 1 month, up to 2 months, up to 3 months, up to 4 months, up to 5 months, up to 6 months, up to 7 months, up to 8 months, up to 9 months, up to 10 months, up to 11 months, up to 1 year, up to 2 years, up to 3 years, up to 4 years, or up to 5 years). In some embodiments, the biological sample is stored for up to 20 years (e.g., up to 5 years, up to 10 years, up to 15 years, or up to 20 years) as described by any of the methods described herein.

[0100] The methods of the present disclosure include obtaining one or more biological samples from a subject for analysis. In some embodiments, one biological sample is taken from the subject for analysis. In some embodiments, a plurality (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more) of biological samples are taken from the subject for analysis. In some embodiments, one biological sample from the subject is analyzed. In some embodiments, a plurality (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more) of biological samples are analyzed. When multiple biological samples from the subject are analyzed, the biological samples can be obtained simultaneously (e.g., multiple biological samples can be taken with the same procedure) or the biological samples can be taken at different times (e.g., in different procedures including a procedure 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 days after the first procedure, a procedure 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 weeks after the first procedure, a procedure 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 months after the first procedure, a procedure 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 years after the first procedure, or a procedure 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 years after the first procedure).

[0101] The second or subsequent biological sample can be taken or obtained from the same region (e.g., from the same tumor or tissue region) or from different regions (e.g., including different tumors). The second or subsequent biological sample may be taken or obtained from the subject after one or more treatments, and may be taken from the same or different regions. As a non-limiting example, the second or subsequent biological sample can be useful in determining whether the cancer in each biological sample has different characteristics (e.g., in the case of biological samples taken from two physically separate tumors in a patient), or whether the cancer has responded to one or more treatments (e.g., in the case of two or more biological samples taken from the same or different tumors before and after treatment). In some embodiments, each of at least one biological sample is a body fluid sample, a cell sample, or a tissue biopsy sample.

[0102] In some embodiments, one or more biological specimens are combined (e.g., placed in the same container for storage) prior to further processing. For example, a first sample of a first tumor obtained from a subject may be combined with a second sample of a second tumor obtained from the subject, and the first and second tumors may or may not be the same tumor. In some embodiments, the first and second tumors are similar but not the same (e.g., two tumors in the brain of a subject). In some embodiments, the first and second biological samples of a subject are samples of different types of tumors (e.g., a tumor in muscle tissue and a tumor in brain tissue).

[0103] In some embodiments, the sample from which RNA and / or DNA is extracted (e.g., a tumor sample or a blood sample) is of a sufficient size such that at least 2 μg (e.g., at least 2 μg, at least 2.5 μg, at least 3 μg, at least 3.5 μg or more) of RNA can be extracted therefrom. In some embodiments, the sample from which RNA and / or DNA is extracted may be peripheral blood mononuclear cells (PBMCs). In some embodiments, the sample from which RNA and / or DNA is extracted may be any type of cell suspension. In some embodiments, the sample from which RNA and / or DNA is extracted (e.g., a tumor sample or a blood sample) is of a sufficient size such that at least 1.8 μg of RNA can be extracted therefrom. In some embodiments, at least 50 mg (e.g., at least 1 mg, at least 2 mg, at least 3 mg, at least 4 mg, at least 5 mg, at least 10 mg, at least 12 mg, at least 15 mg, at least 18 mg, at least 20 mg, at least 22 mg, at least 25 mg, at least 30 mg, at least 35 mg, at least 40 mg, at least 45 mg, or at least 50 mg) of tissue sample is taken and RNA and / or DNA is extracted therefrom. In some embodiments, at least 20 mg of tissue sample is taken and RNA and / or DNA is extracted therefrom. In some embodiments, at least 30 mg of tissue sample is taken. In some embodiments, at least 10 - 50 mg (e.g., 10 - 50 mg, 10 - 15 mg, 10 - 30 mg, 10 - 40 mg, 20 - 30 mg, 20 - 40 mg, 20 - 50 mg, or 30 - 50 mg) of tissue sample is taken and RNA and / or DNA is extracted therefrom. In some embodiments, at least 30 mg of tissue sample is taken. In some embodiments, at least 20 - 30 mg of tissue sample is taken and RNA and / or DNA is extracted therefrom.In some embodiments, the sample from which RNA and / or DNA is extracted (e.g., a tumor sample, or a blood sample) is of a sufficient size such that at least 0.2 μg (e.g., at least 200 ng, at least 300 ng, at least 400 ng, at least 500 ng, at least 600 ng, at least 700 ng, at least 800 ng, at least 900 ng, at least 1 μg, at least 1.1 μg, at least 1.2 μg, at least 1.3 μg, at least 1.4 μg, at least 1.5 μg, at least 1.6 μg, at least 1.7 μg, at least 1.8 μg, at least 1.9 μg, or at least 2 μg) of RNA can be extracted therefrom. In some embodiments, the sample from which RNA and / or DNA is extracted (e.g., a tumor sample, or a blood sample) is of a sufficient size such that at least 0.1 μg (e.g., at least 100 ng, at least 200 ng, at least 300 ng, at least 400 ng, at least 500 ng, at least 600 ng, at least 700 ng, at least 800 ng, at least 900 ng, at least 1 μg, at least 1.1 μg, at least 1.2 μg, at least 1.3 μg, at least 1.4 μg, at least 1.5 μg, at least 1.6 μg, at least 1.7 μg, at least 1.8 μg, at least 1.9 μg, or at least 2 μg) of RNA can be extracted therefrom.

[0104] Subject Aspects of the present disclosure relate to biological samples obtained from a subject. In some embodiments, the subject is a mammal (e.g., a human, mouse, cat, dog, horse, hamster, cow, pig, or other domestic animal). In some embodiments, the subject is a human. In some embodiments, the subject is an adult (e.g., 18 years of age or older). In some embodiments, the subject is a child (e.g., under 18 years of age). In some embodiments, the human subject has or has been diagnosed with at least one form of cancer. In some embodiments, the cancer the subject has is a carcinoma, sarcoma, myeloma, leukemia, lymphoma, or a mixed-type cancer including a plurality of carcinoma, sarcoma, myeloma, leukemia, and lymphoma. A carcinoma refers to a malignant neoplasm of epithelial origin or cancer of the inner or outer membranes of the body. A sarcoma refers to a cancer derived from supportive and connective tissues such as bone, tendon, cartilage, muscle, and fat. Myeloma is a cancer derived from plasma cells in the bone marrow. Leukemia (“liquid cancer” or “blood cancer”) is a cancer of the bone marrow (the site of blood cell production). Lymphoma occurs in lymphatic glands or nodules, which are part of the lymphatic system, a network of blood vessels, nodules, and organs (especially the spleen, tonsils, thymus) that purify body fluids and produce white blood cells, or lymphocytes. Non-limiting examples of mixed-type cancers include adenosquamous carcinoma, mixed mesodermal tumor, carcinosarcoma, and teratocarcinoma. In some embodiments, the subject has a tumor. The tumor can be benign or malignant. In some embodiments, the cancer is any one of skin cancer, lung cancer, breast cancer, prostate cancer, colon cancer, rectal cancer, cervical cancer, and uterine cancer. In some embodiments, the subject is at risk of developing cancer, for example, because the subject has one or more genetic risk factors, or has been exposed to one or more carcinogens (e.g., tobacco smoke or smokeless tobacco), or is being exposed to them.

[0105] Single cell suspension In some embodiments, methods for characterizing cancer that a subject has or is suspected of having (e.g., RNA sequencing, DNA sequencing, or multiplexed flow cytometry) are performed at the single cell level to capture the heterogeneity of a single tumor or cancer tissue, or multiple tumors or cancer tissues. That is, the measurement and evaluation of single cells in a tumor sample provides information that is not confounded by the genotypic or phenotypic heterogeneity of the bulk sample. In some embodiments, a single cell suspension is prepared from one or more biological samples obtained from a subject for use in methods such as single cell RNA or DNA sequencing, or mass cytometry.

[0106] Accordingly, some embodiments of any one of the methods described herein include forming a single cell suspension of cells from a sample of a tumor (e.g., a first sample of a tumor). In some embodiments, forming a single cell suspension of cells from a sample of a tumor includes dissecting the tumor sample to obtain tumor sample fragments. Curved scissors can be used to dissect the tumor tissue sample. In some embodiments, the tumor sample fragments are 0.5 - 3 mm 3 (e.g., 1 - 2 mm 3 ). In some embodiments, the tumor tissue sample or its fragments are kept moist during dissection.

[0107] Methods for preparing a single cell suspension from a tumor sample can include any one or more of the steps of mincing, enzymatic and / or non-enzymatic digestion, vigorous pipetting, passing through a cell strainer, washing, and counting, in any order. In some embodiments, one or more of these steps are repeated (e.g., 1, 2, 3, 4, or 5 or more times).

[0108] In some embodiments, a tumor sample or a tumor sample fragment is incubated with an enzyme cocktail. The enzymes can be used in any number and in any combination, for example, see BioFiles: For Life Science Research, Issue 2, 2006, www.sigmaaldrich.com / content / dam / sigma-aldrich / docs / Sigma / General_Information / 2 / biofiles_issue2.pdf, which is hereby incorporated by reference in its entirety. This incorporates herein by reference either the enzymes or other components (such as media) described therein specifically.

[0109] Quatromoni et al., "An optimized disaggregation method for human lung tumors that preserves the phenotype and function of the immune cell", J Leukoc Biol. 2015 Jan, 97(1):201-209, provides a comparison of different enzyme cocktails, which are hereby incorporated by reference in their entirety. In some embodiments, the enzyme cocktail includes, as components, a medium (e.g., L-15 medium), an antibacterial agent (e.g., penicillin and / or streptomycin), an antifungal agent (e.g., amphotericin), a collagenase (e.g., collagenase I, collagenase II, collagenase IV), a DNAse (e.g., DNAseI), an elastase, a hyaluronidase, one or more of proteases (e.g., protease XIV, trypsin, papain, thermolysin). Coll I has the original balance of collagenase activity, caseinase activity, chymotrypsin activity, and trypsin activity. Coll II contains a higher relative level of protease activity, particularly chymotrypsin. Coll IV is designed to have a particularly low tryptic activity (Quatromoni et al., J Leukoc Biol. 2015 Jan, 97(1):201-209). In some embodiments, only collagenase I, collagenase II, or collagenase IV is used. In some embodiments, a mixture of two collagenases is used (e.g., collagenase I and collagenase II, collagenase I and collagenase IV, or collagenase II and collagenase IV). In some embodiments, more than two collagenases are used (e.g., collagenase I, collagenase II, and collagenase IV).

[0110] In some embodiments, the enzyme cocktail comprises, as components, one or more of a medium (e.g., complete medium), penicillin, streptomycin, collagenase (e.g., collagenase I or collagenase IV). The concentration of the enzyme in the cocktail is adjustable. Non-limiting examples of the enzyme cocktail are collagenase I (0.2 mg / ml), collagenase IV (1 mg / ml), complete medium, penicillin (0.001%), and DNAse.

[0111] In some embodiments, at least 25 ml (e.g., at least 25 ml, at least 26 ml, at least 27 ml, at least 28 ml, at least 29 ml, or at least 30 ml) of the enzyme cocktail is added per 0.5 gm of tumor tissue. In some embodiments, a sample of the tumor or a fragment thereof is incubated in the enzyme cocktail while the sample is being shaken or stirred (e.g., rotated at 85 RPM and / or subjected to vigorous pipetting). In some embodiments, a sample of the tumor or a fragment thereof is incubated in the enzyme cocktail at a temperature between 20 and 50 °C (e.g., 20 - 50 °C, 20 - 25 °C, 25 - 30 °C, 25 - 35 °C, 30 - 40 °C, 35 - 45 °C, 40 - 50 °C, or 30 - 50 °C). In some embodiments, the method of preparing a single cell suspension comprises filtering the enzyme cocktail through, for example, a cell strainer (e.g., 50 μm, 70 μm, or 100 μm). In some embodiments, if the filter is too fine, a cell composition with a high concentration of fibroblasts may result. In some embodiments, if the filter is too coarse, cell clumps may result. In some embodiments, the cell clumps are broken up using mechanical force (e.g., vigorous pipetting, application of pressure using a syringe).

[0112] In some embodiments, the filtered cells are lysed using an RBC lysis buffer to lyse red blood cells. RBC lysis buffers are commercially available (see, for example, www.abcam.com / red-blood-cell-rbc-lysis-buffer-ab204733.html).

[0113] In some embodiments, methods of preparing a single cell suspension include enzymatic and mechanical dissociation. Examples of methods for dissociating cells from tissue are described in the publications, Quatromoni et al., "An optimized disaggregation method for human lung tumors that preserves the phenotype and function of the immune cell" J Leukoc Biol. January 2015, 97(1):201-209, Pennartz et al., "Generation of Single-Cell Suspensions from Mouse Neural Tissue" JOVE Issue 29, doi:10.3791 / 1267, Published:7 / 07 / 2009, and www.youtube.com / watch?v=N0jftyYqM38.

[0114] In some embodiments, an enzyme-free cell dissociation buffer is used. See, for example, catalog numbers 13151014 and 13150016 from ThermoFisher Scientific, or catalog number S-014-B from Millipore Sigma Aldrich. Heng et al., Biol Proced Online. 2009, 11:161-169 presents a comparison of enzymatic and non-enzymatic means of dissociating cells, which is hereby incorporated by reference in its entirety.

[0115] In some embodiments, the number of cells in the single cell suspension is counted and viability assays are performed. The following example presents an example of the overall process of forming a single cell suspension from a sample of tumor tissue.

[0116] In some embodiments, the method includes forming a single-cell suspension of cells from a sample of a tumor and dividing it into at least a first and a second portion. The first and second portions of the single-cell suspension can be of equal size or different sizes (e.g., containing different numbers of cells). In some embodiments, all portions of the single-cell suspension (e.g., the first portion, the second portion, etc.) are stored in separate containers and stored under the same or similar conditions (e.g., in liquid nitrogen, or at -80 °C). In some embodiments, different portions of the single-cell suspension are stored under different conditions, either before or after any further processing (e.g., labeling with an antibody for protein expression studies). In some embodiments, cells isolated from a biological sample are cultured, grown, and then stored. In some embodiments, cells isolated from a biological sample are cultured and grown after storage.

[0117] In some embodiments, any one of the methods described herein further includes forming a lysate from at least a portion (e.g., the first or second portion) of the single-cell suspension. In some embodiments, different portions of the single-cell suspension contain different types of cells. In some embodiments, the portion of the single-cell suspension from which the lysate is formed contains at least 1×10 6 cells (e.g., at least 1×10 6 cells, at least 2×10 6 cells, at least 3×10 6 cells, at least 4×10 6 cells, or at least 5×10 6 cells). In some embodiments, the portion of the single-cell suspension from which the lysate is formed contains at least 2×10 6contains individual cells. The lysate can be stored in a storage medium that prevents degradation of DNA and / or RNA (e.g., RNALater). In some embodiments, the method includes extracting RNA from a single cell suspension or lysate from portions of a single cell suspension and performing RNA sequencing on the extracted RNA to obtain RNA expression data. These RNA expression data can be used to determine tumor heterogeneity.

[0118] An overview of single cell RNA sequencing is described at hemberg-lab.github.io / scRNA.seq.course / introduction-to-single-cell-rna-seq.html, the figure 2.1 of which is incorporated herein by reference. In some embodiments, the method of performing RNA sequencing on a single cell suspension includes isolation of single cell RNA, pre-amplification of reverse transcribed cDNA, preparation of a cDNA library (e.g., using the Fluidigm C1 protocol), and sequencing of the sequenced ones using a platform such as the Illumina HiSeq 2500.

[0119] Methods for performing single-cell RNA sequencing include Bagnoli et al., "Studying Cancer Heterogeneity by Single-Cell RNA Sequencing", Methods Mol Biol. 2019, 1956:305 - 319, which is hereby incorporated by reference in its entirety; Sun et al., "Single-cell RNA sequencing reveals gene expression signatures of breast cancer-associated endothelial cells", Oncotarget. February 16, 2018, 9(13): 10945 - 10961; Kulkarni et al., "Beyond bulk: a review of single cell transcriptomics methodologies and applications", Curr Opin Biotechnol. April 9, 2019, 58:129 - 136; Huang et al., "High Throughput Single Cell RNA Sequencing, Bioinformatics Analysis and Applications", Adv Exp Med Biol. 2018, 1068:33 - 43; Zilionis et al., "Single-Cell Transcriptomics of Human and Mouse Lung Cancers Reveals Conserved Myeloid Populations across Individuals and Species", Immunity. April 5, 2019, pii:S1074-7613(19)30126 - 8; Kashima et al., "An Informative Approach to Single-Cell Sequencing Analysis", Adv Exp Med Biol. 2019, 1129:81 - 96, doi:10.1007 / 978-981-13-6037-4_6; Seki et al., "An Informative Approach to Single-Cell Sequencing Analysis", Adv Exp Med Biol. 2019, 1129 : pp. 81-96, "Single-Cell DNA-Seq and RNA-Seq in Cancer Using the C1 System", Adv Exp Med Biol. 2019, 1129:27-50, doi:10.1007 / 978-981-13-6037-4_3), See et al., "A Single-Cell Sequencing Guide for Immunologists", Front Immunol. 2018, 9:2425, which is incorporated herein by reference in its entirety.

[0120] Gan et al., "Identification of cancer subtypes from single-cell RNA-seq data using a consensus clustering method", BMC Med Genomics. 2018, 11(Suppl 6):117, which describes a clustering method for single-cell RNA sequencing data and is incorporated herein by reference in its entirety.

[0121] In some embodiments, any one of the Fluidigm C1 system (SMART-seq), Fluidigm C1 system (mRNA Seq HT), SMART-seq2, 10X Genomics Chromium system, and MARS-seq is used as the single-cell RNA sequencing method. See et al., Front Immunol. 2018, 9:2425, which provides a comparison of these methods and is incorporated herein by reference in its entirety.

[0122] In some embodiments, any one of the methods described herein further includes performing measurements on a single-cell suspension. In some embodiments, different measurements are performed in parallel on the same cells. Macaulay et al., Trends Genet. February 2017, 33(2):155-168, which describes methods for performing multiple measurements from single cells and is incorporated herein by reference in its entirety.

[0123] In some embodiments, any one of the methods described herein further comprises performing mass cytometry on at least a first portion of a single cell suspension. Mass cytometry is a mass spectrometry technique based on inductively coupled plasma mass spectrometry and time-of-flight mass spectrometry used to determine cell characteristics. In some embodiments, mass cytometry involves conjugating an antibody to an isotopically pure element and then using it to label cellular molecules (e.g., proteins). In some embodiments, the cells are nebulized and passed through an argon plasma to ionize the metal antibody. The metal signal is then analyzed by a time-of-flight mass spectrometer to identify and quantify cellular molecules within the cells. In some embodiments, the single cell suspension or a portion thereof on which mass cytometry is performed comprises at least 1×10 6 cells (e.g., at least 1×10 6 cells, at least 2×10 6 cells, at least 3×10 6 cells, at least 4×10 6 cells, at least 5×10 6 cells, at least 6×10 6 cells, at least 7×10 6 cells, at least 8×10 6 cells, at least 9×10 6 cells, or at least 10×10 6 cells). In some embodiments, the single cell suspension or a portion thereof on which mass cytometry is performed comprises at least 5×10 6 cells.

[0124] Methods for performing mass cytometry are described in Galli et al., "The end of omics? High dimensional single cell analysis in precision medicine," Eur J Immunol. February 2019, 49(2):212-220, Brodin, "The biology of the cell - insights from mass cytometry," FEBS J. November 3, 2018, doi:10.1111 / febs.14693, Olsen et al., "The anatomy of single cell mass cytometry data," Cytometry A. February 2019, 95(2):156-172, Behbehani, "Applications of Mass Cytometry in Clinical Medicine: The Promise and Perils of Clinical CyTOF," Clin Lab Med. December 2017, 37(4):945-964, Gondhalekar et al., "Alternatives to current flow cytometry data analysis for clinical and research studies," Methods. February 1, 2018, 134-135:113-129, and Soares et al., "Go with the flow: advances and trends in magnetic flow cytometry," Anal Bioanal Chem. March 2019, 411(9):1839-1862, doi: 10.1007 / s00216-019-01593-9. Epub February 19, 2019, each of which is incorporated herein by reference in its entirety.

[0125] Other assays Any of the biological samples described herein can be used to obtain expression data using a conventional assay or those described herein. Expression data, in some embodiments, includes gene expression levels. Gene expression levels can be detected by detecting products of gene expression such as mRNA and / or protein.

[0126] In some embodiments, the gene expression level is determined by detecting the level of protein in the sample and / or by detecting the activity level of the protein in the sample. As used herein, the phrases "determine" or "detect" mean to evaluate the presence, absence, quantity, and / or amount (which may be an effective amount) of a substance in a sample, including evaluating, or otherwise, deriving a qualitative or quantitative concentration level of such a substance, or alternatively, evaluating the value and / or classification of such a substance in a sample from a subject.

[0127] The level of protein can be measured using an immunoassay. Examples of immunoassays include any known assay (without limitation), immunoblot analysis (e.g., Western blot), immunohistochemical analysis, flow cytometry assay, immunofluorescence assay (IF), enzyme-linked immunosorbent assay (ELISA) (e.g., sandwich ELISA), radioimmunoassay, electrochemiluminescence-based detection assay, magnetic immunoassay, lateral flow assay, and any of the related techniques. Additional suitable immunoassays for detecting the level of protein provided herein will be apparent to those skilled in the art.

[0128] Such immunoassays can involve the use of an agent (e.g., an antibody) specific for the target protein. The phrase "agent," such as an antibody, that "specifically binds" to a target protein is well understood in the art, and methods for determining such specific binding are also well known in the art. An antibody is said to "specifically bind" when it reacts or binds to a particular target protein more frequently, more rapidly, for a longer period of time, and / or with a higher affinity as compared to alternative proteins. Also, by reading this definition, it is understood that, for example, an antibody that specifically binds to a first target peptide may or may not also specifically or selectively bind to a second target peptide. As such, "specific binding" or "selective binding" does not necessarily require exclusive binding (although it may include exclusive binding). Generally, but not always, reference to binding means selective binding. In some examples, an antibody that "specifically binds" to a target peptide or its epitope may not bind to other peptides or other epitopes in the same antigen. In some embodiments, a sample can be contacted with a plurality of binding agents that bind different proteins, simultaneously or sequentially (e.g., multiplex assays).

[0129] As used herein, the term "antibody" refers to a protein comprising at least one immunoglobulin variable domain or immunoglobulin variable domain sequence. For example, an antibody can comprise a heavy (H) chain variable region (abbreviated herein as VH), and a light (L) chain variable region (abbreviated herein as VL). In another example, an antibody comprises two heavy (H) chain variable regions and two light (L) chain variable regions. The term "antibody" encompasses antigen-binding fragments of antibodies (e.g., single-chain antibodies, Fab and sFab fragments, F(ab')2, Fd fragments, Fv fragments, scFv, and domain antibody (dAb) fragments (de Wildt et al., Eur J Immunol. 1996, 26(3):629-39), as well as full antibodies. Antibodies can have the structural characteristics of IgA, IgG, IgE, IgD, IgM (and further subtypes thereof). Antibodies can be from any source including, but not limited to, primates (human and non-human primates) as well as primatized (humanized, etc.) antibodies.

[0130] In some embodiments, an antibody as described herein can be conjugated to a detectable label, and the binding of a detection reagent to a peptide of interest can be determined based on the intensity of a signal emitted from the detectable label. Alternatively, a secondary antibody specific for the detection reagent can be used. One or more antibodies can be conjugated to a detectable label. Any suitable label known in the art can be used in the assay methods described herein. In some embodiments, the detectable label comprises a fluorophore. As used herein, the term "fluorophore" (also referred to as a "fluorescent label" or "fluorescent dye") refers to a moiety that absorbs light energy at a defined excitation wavelength and emits light energy at a different wavelength. In some embodiments, the detection moiety is an enzyme or comprises an enzyme. In some embodiments, the enzyme produces a colored product from a colorless substrate (e.g., β-galactosidase).

[0131] One of ordinary skill in the art will appreciate that the present disclosure is not limited to immunoassays. Detection assays that are not antibody-based, such as mass spectrometry, are also useful for detecting and / or quantifying the proteins and / or protein levels provided herein. Assays that rely on chromogenic substrates are also useful for detecting and / or quantifying the proteins and / or protein levels as provided herein.

[0132] Alternatively, the level of nucleic acid encoding a gene in a sample can be measured via conventional methods. In some embodiments, measuring the expression level of a nucleic acid encoding a gene includes measuring mRNA. In some embodiments, the expression level of mRNA encoding a gene can be measured using real-time reverse transcriptase (RT) Q-PCR or nucleic acid microarray. Methods for detecting nucleic acid sequences include, but are not limited to, polymerase chain reaction (PCR), reverse transcriptase PCR (RT-PCR), in situ PCR, quantitative PCR (Q-PCR), real-time quantitative PCR (RT Q-PCR), in situ hybridization, Southern blot, Northern blot, sequence analysis, microarray analysis, detection of reporter genes, or other DNA / RNA hybridization platforms.

[0133] In some embodiments, the level of nucleic acid encoding a gene in a sample can be measured via a hybridization assay. In some embodiments, the hybridization assay includes at least one binding partner. In some embodiments, the hybridization assay includes at least one oligonucleotide binding partner. In some embodiments, the hybridization assay includes at least one labeled oligonucleotide binding partner. In some embodiments, the hybridization assay includes at least one pair of oligonucleotide binding partners. In some embodiments, the hybridization assay includes at least one pair of labeled oligonucleotide binding partners.

[0134] Any binding agent that specifically binds to a desired nucleic acid or protein can be used in the methods and kits described herein to measure the expression level in a sample. In some embodiments, the binding agent is an antibody or aptamer that specifically binds to a desired protein. In other embodiments, the binding agent may be one or more oligonucleotides complementary to a nucleic acid or a portion thereof. In some embodiments, a sample can be contacted with multiple binding agents that bind different proteins or different nucleic acids, simultaneously or sequentially (e.g., multiplex analysis).

[0135] To measure the expression level of a protein or nucleic acid, the sample may be considered to be in contact with the binding agent under suitable conditions. Generally, the term "contacting" refers to exposing the binding agent to the sample or cells taken therefrom for a suitable period sufficient for a complex to form between the binding agent and, if any, the target protein or target nucleic acid in the sample. In some embodiments, contacting is performed by capillary action in which the sample is moved across the surface of a support membrane.

[0136] In some embodiments, the assay can be performed on a low-throughput platform, including a single assay format. In some embodiments, the assay can be performed on a high-throughput platform. Such high-throughput assays can include using a binding agent immobilized on a solid support (e.g., one or more chips). Methods for immobilizing the binding agent depend on factors such as the nature of the binding agent and the material of the solid support and may require specific buffers. Such methods will be apparent to those skilled in the art.

[0137] Extraction of DNA and / or RNA In one embodiment of any of the methods described herein, RNA is extracted from a biological sample to prevent degradation of the RNA and / or to prevent inhibition of enzymes in downstream processing, such as the preparation of DNA (i.e., cDNA library from RNA). In one embodiment of any of the methods described herein, DNA is extracted from a biological sample to prevent degradation of the DNA and / or to prevent inhibition of enzymes in downstream processing, such as the preparation of DNA. In some embodiments, the term "extraction" in the context of obtaining DNA or RNA from a biological sample is used interchangeably with the term "isolation".

[0138] The methods described herein involve the extraction of RNA and / or DNA from biological samples (e.g., tumor samples or blood samples). As described above, a biological sample can be composed of multiple samples from one or more tissues (e.g., one or more different tumors). In some embodiments, RNA and / or DNA are extracted from the combined samples. In some embodiments, RNA and / or DNA are extracted from multiple biological samples from a subject and then combined prior to further processing (e.g., storage or DNA library preparation). In some embodiments, multiple samples of the extracted RNA and / or DNA are combined with each other after being retrieved from a storage location. In some embodiments, at least tumor DNA is extracted from one or more tumor tissues. In some embodiments, at least tumor RNA is extracted from one or more tumor tissues. In some embodiments, at least normal DNA is extracted from one or more normal tissues and used as a control. In some embodiments, at least normal RNA is extracted from one or more normal tissues and used as a control. The protocol for DNA / RNA extraction is described at least in Example 2.

[0139] Methods for extracting DNA and / or RNA from biological samples are known in the art, and reagents and kits therefor are commercially available. Gomez-Akata et al., "Methods for extracting 'omes from microbialites", J Microbiol Methods. March 12, 2019, 160:1-10, describes methods for extraction applicable to the extraction of DNA and RNA from microorganisms, and also describes its advantages and disadvantages, the whole of which is incorporated herein by reference. The methods described by Gomez-Akata et al. are generally applicable to RNA and / or DNA extracted from tissues. Moore, Curr Protoc Immunol. May 2001, Chapter 10:Unit 10.1, describes the purification and concentration of DNA from aqueous solutions, the whole of which is also incorporated herein by reference.

[0140] In some embodiments, extracting DNA and / or RNA comprises lysing the cells of the biological sample and isolating the DNA and / or RNA from other cellular components. Examples of methods for lysing cells include, but are not limited to, mechanical lysis, liquid homogenization, sonication, freeze-thaw, chemical lysis, alkaline lysis, and manual grinding.

[0141] Methods for extracting DNA and / or RNA include, but are not limited to, solution-phase extraction methods and solid-phase extraction methods. In some embodiments, the solution-phase extraction method includes an organic extraction method, such as a phenol-chloroform extraction method. In some embodiments, the solution-phase extraction method includes a high-salt concentration extraction method, such as a guanidinium thiocyanate (GuTC) or guanidinium chloride (GuCl) extraction method. In some embodiments, the solution-phase extraction method includes an ethanol precipitation method. In some embodiments, the solution-phase extraction method includes an isopropanol precipitation method. In some embodiments, the solution-phase extraction method includes an ethidium bromide (EtBr)-cesium chloride (CsCl) gradient centrifugation method. In some embodiments, extracting DNA and / or RNA includes a non-ionic detergent extraction method, such as a cetyltrimethylammonium bromide (CTAB) extraction method.

[0142] In some embodiments, extracting DNA and / or RNA includes a solid-phase extraction method. Any solid phase that binds to DNA and / or RNA can be used in the methods and systems described herein to extract DNA and / or RNA. Examples of solid phases that bind to DNA and / or RNA include, but are not limited to, silica matrices, ion-exchange matrices, glass particles, magnetizable cellulose beads, polyamide matrices, and nitrocellulose membranes.

[0143] In some embodiments, the solid-phase extraction method includes a spin column-based extraction method. In some embodiments, the solid-phase extraction method includes a bead-based extraction method. In some embodiments, the solid-phase extraction method includes a cation exchange resin, such as a styrene divinylbenzene copolymer resin.

[0144] The systems and methods described herein include extracting DNA and / or RNA from a single biological sample or multiple biological samples. In some embodiments, extracting DNA includes extracting DNA from a single sample. In some embodiments, extracting DNA includes extracting DNA from multiple samples. In some embodiments, extracting DNA includes extracting DNA from a first sample and a second sample. In some embodiments, extracting DNA includes extracting DNA from one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more samples.

[0145] In some embodiments, extracting RNA includes extracting RNA from a single sample. In some embodiments, extracting RNA includes extracting RNA from multiple samples. In some embodiments, extracting RNA includes extracting RNA from a first sample and a second sample. In some embodiments, extracting RNA includes extracting RNA from one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more samples.

[0146] DNA and / or RNA extracted from a biological sample can be combined with DNA and / or RNA extracted from another biological sample. This can be achieved by combining one or more biological samples to extract nucleic acids or by combining nucleic acids extracted from one or more biological samples. In some embodiments, a first biological sample is combined with a second biological sample to form a combined sample, and DNA and / or RNA is extracted from the combined sample. In some embodiments, DNA and / or RNA extracted from a first biological sample can be combined with DNA and / or RNA extracted from a second biological sample.

[0147] The systems and methods described herein include extracting any type of DNA and / or RNA from a biological sample. In some embodiments, extracting DNA includes extracting genomic DNA (gDNA). In some embodiments, extracting DNA includes extracting mitochondrial DNA (mtDNA). In some embodiments, extracting RNA includes extracting messenger RNA (mRNA). In some embodiments, extracting RNA includes extracting precursor mRNA (pre-mRNA). In some embodiments, extracting RNA includes extracting ribosomal RNA (rRNA). In some embodiments, extracting RNA includes extracting transfer RNA (tRNA).

[0148] In some embodiments, a single kit is used to purify DNA and RNA from the same sample. A non-limiting example of a kit for doing so is the Qiagen AllPrep DNA / RNA kit. In some embodiments, a robot is employed to perform the extraction of DNA and / or RNA.

[0149] In some embodiments, if the extracted RNA sample does not have sufficient yield and / or quality, any of the following results may occur. First, there may be overrepresentation of common transcripts and underrepresentation of low-abundance transcripts in the RNA sequencing data. Second, if the quality of the RNA is low, the read length may be insufficient (i.e., the reads are short), and / or if the quality of the reads is inappropriate, misidentification of the RNA may occur.

[0150] For whole exome sequencing, if the amount and quality of the DNA are low, base pair misidentification may occur, resulting in false discovery of variants (e.g., false positives) or false non-discovery of variants (e.g., false negatives). Another problem that may result from low DNA amount and quality is insufficient exome coverage (e.g., missing sequences).

[0151] In some embodiments, the quality and / or quantity of the extracted RNA and / or DNA is checked before the extracted RNA and / or DNA is further processed for RNA sequencing or whole exome sequencing (WES). In some embodiments, the sample of extracted RNA has a total mass of at least 1000 - 6000 ng. In some embodiments, the sample of extracted RNA has a total mass of at least 100 - 60000 ng (e.g., 100 - 60000 ng, 500 - 30000 ng, 800 - 20000 ng, 1000 - 15000 ng, 1000 - 10000 ng, 1000 - 8000 ng, 1000 - 6000 ng, 10000 - 20000 ng, 20000 - 60000 ng). In some embodiments, the acceptable total RNA amount for further sequencing is at least 100 - 1,000 ng (e.g., 100 - 1,000 ng, 500 - 1,000 ng, or 300 - 900 ng). In some embodiments, the target total RNA amount for further sequencing is more than 200 - 1,000 ng (e.g., 200 - 1,000 ng, 500 - 1,000 ng, or 300 - 1,000 ng). In some embodiments, the purity of the sample of extracted RNA is a value corresponding to a ratio of the absorbance at 260 nm to the absorbance at 280 nm of at least 1 (e.g., at least 1, at least 1.2, at least 1.4, at least 1.6, at least 1.8, or at least 2). In some embodiments, the purity of the sample of extracted RNA is a value corresponding to a ratio of the absorbance at 260 nm to the absorbance at 280 nm of at least 2. The ratio of the absorbance at 260 nm to the absorbance at 280 nm is used to evaluate the purity of DNA and RNA. A ratio of ~1.8 is generally accepted as "pure" for DNA, and a ratio of ~2.0 is generally accepted as "pure" for RNA. If this ratio is significantly lower in either case, it may indicate the presence of proteins, phenol, or other contaminants that strongly absorb near 280 nm. The absorbance can be measured using a spectrophotometer.

[0152] In some embodiments, the purity or integrity of the extracted RNA or DNA (e.g., DNA fragment library) by any one of the methods described herein is a value corresponding to an RNA integrity number (RIN) of at least 4 (e.g., at least 4, at least 5, at least 6, at least 7, at least 8, or at least 9). In some embodiments, the purity of the nucleic acid (e.g., RNA or DNA) extracted by any one of the methods described herein is a value corresponding to an RNA integrity number (RIN) of at least 7. The RIN has been demonstrated to be robust and reproducible in studies compared to other RNA integrity calculation algorithms, and has solidified its position as a preferred method for determining the quality of the RNA being analyzed (Imbeaud et al., "Towards standardization of RNA quality assessment using user-independent classifiers of microcapillary electrophoresis traces", Nucleic Acids Research. 33 (6):e56).

[0153] In some embodiments, the sample of extracted DNA has a total mass of at least 100 - 20000 ng (e.g., 100 - 20000 ng, 500 - 15000 ng, 800 - 10000 ng, 1000 - 15000 ng, 1000 - 10000 ng, 1000 - 8000 ng, 1000 - 6000 ng, or 1000 - 2000 ng). In some embodiments, the sample of extracted DNA has a total mass of at least 1000 - 2000 ng. In some embodiments, the acceptable total DNA amount for further sequencing is at least 20 - 200 ng (e.g., 20 - 200 ng, 30 - 200 ng, or 50 - 150 ng). In some embodiments, the target total DNA amount for further sequencing is more than 30 - 200 ng (e.g., 30 - 200 ng, 50 - 200 ng, or 100 - 200 ng). In some embodiments, the target purity of the sample of extracted DNA is a value corresponding to a ratio range of the absorbance at 260 nm to the absorbance at 280 nm of at least 1.8 - 2 (e.g., at least 1.8 - 2, at least 1.8 - 1.9). In some embodiments, the purity of the sample of extracted DNA is a value corresponding to a ratio of the absorbance at 260 nm to the absorbance at 280 nm of at least 1 (e.g., at least 1, at least 1.2, at least 1.4, at least 1.6, at least 1.8, or at least 2). In some embodiments, the acceptable purity of the sample of extracted DNA is a value corresponding to a ratio of the absorbance at 260 nm to the absorbance at 280 nm of at least 1.5 (e.g., at least 1.5, at least 1.7, at least 2). In some embodiments, the target purity of the sample of extracted DNA is a value corresponding to a ratio range of the absorbance at 260 nm to the absorbance at 230 nm of at least 2 - 2.2 (e.g., at least 2 - 2.2, at least 2 - 2.1). In some embodiments, the acceptable purity of the sample of extracted DNA is a value corresponding to a ratio of the absorbance at 260 nm to the absorbance at 230 nm of at least 1.5 (e.g., at least 1.5, at least 1.7, at least 2). In some embodiments The purity of the extracted DNA sample as described herein is analyzed by a spectrophotometer, such as a small-volume full-spectrum ultraviolet-visible spectrophotometer (e.g., NanoDrop spectrophotometer available from ThermoFisher Scientific, www.thermofisher.com).

[0154] In some embodiments, the extracted DNA sample has a target concentration of at least 4.5 ng / μl (e.g., 4.5 ng / μl, 5.5 ng / μl, 6.5 ng / μl). In some embodiments, the extracted DNA sample has an acceptable concentration of at least 3 ng / μl (e.g., 3 ng / μl, 5 ng / μl, 10 ng / μl). In some embodiments, the determination of the concentration of the extracted DNA is performed by, for example, a fluorometer for the quantification of DNA or RNA (e.g., Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com).

[0155] In some embodiments, the extracted DNA sample has a target concentration of at least 4 ng / μl (e.g., 4 ng / μl, 6 ng / μl, 8 ng / μl). In some embodiments, the extracted DNA sample has an acceptable concentration of at least 2.5 ng / μl (e.g., 2.5 ng / μl, 4.5 ng / μl, 5.5 ng / μl). In some embodiments, the determination of the concentration of the extracted DNA is performed by Tapestation.

[0156] In some embodiments, the sample of extracted RNA has a target concentration of at least 2 ng / μl (e.g., 2 ng / μl, 4 ng / μl, 6 ng / μl). In some embodiments, the sample of extracted RNA has an acceptable concentration of at least 4 ng / μl (e.g., 4 ng / μl, 6 ng / μl, 10 ng / μl). In some embodiments, the determination of the concentration of the extracted DNA is performed, for example, by a fluorometer for the quantification of DNA or RNA (e.g., the Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com).

[0157] In some embodiments, the sample of extracted RNA has a target concentration of at least 4 ng / μl (e.g., 4 ng / μl, 6 ng / μl, 8 ng / μl). In some embodiments, the sample of extracted RNA has an acceptable concentration of at least 1.5 ng / μl (e.g., 1.5 ng / μl, 3.5 ng / μl, 5.5 ng / μl). In some embodiments, the determination of the concentration of the extracted RNA is performed by Tapestation. In some embodiments, the acceptable RNA Integrity Number (RIN) is at least 5 (e.g., 5, 6, 7). In some embodiments, the target RNA Integrity Number (RIN) is at least 8 (e.g., 8, 9, 10). In some embodiments, the RIN is performed by Tapestation.

[0158] In some embodiments, the target purity of the extracted RNA sample is a value corresponding to a ratio range of the absorbance at 260 nm to the absorbance at 280 nm of at least 1.8 to 2 (e.g., at least 1.8 to 2, at least 1.8 to 1.9). In some embodiments, the purity of the extracted RNA sample is a value corresponding to a ratio of the absorbance at 260 nm to the absorbance at 280 nm of at least 1.8. In some embodiments, the acceptable purity of the extracted RNA sample is a value corresponding to a ratio of the absorbance at 260 nm to the absorbance at 280 nm of at least 1.5 (e.g., at least 1.5, at least 1.7, at least 2). In some embodiments, the target purity of the extracted RNA sample is a value corresponding to a ratio range of the absorbance at 260 nm to the absorbance at 230 nm of at least 2 to 2.2 (e.g., at least 2 to 2.2, at least 2 to 2.1). In some embodiments, the acceptable purity of the extracted RNA sample is a value corresponding to a ratio of the absorbance at 260 nm to the absorbance at 230 nm of at least 1.5 (e.g., at least 1.5, at least 1.7, at least 2). In some embodiments, the purity of the extracted RNA sample as described herein is analyzed by a spectrophotometer, such as a small-volume full-spectrum ultraviolet-visible spectrophotometer (e.g., NanoDrop spectrophotometer available from ThermoFisher Scientific, www.thermofisher.com). In some embodiments, the concentration of the extracted DNA is at least 10 to 2000 ng / μl (e.g., 10 to 2000 ng / μl, 10 to 1000 ng / μl, 10 to 200 ng / μl, 1 to 200 ng / μl, 0.5 to 400 ng / μl, 0.5 to 200 ng / μl, 100 to 200 ng / μl, 100 to 400 ng / μl, 100 to 500 ng / μl, 50 to 500 ng / μl, or 50 to 250 ng / μl).

[0159] Protocols for quality control of the extracted RNA or DNA samples are described at least in Example 6. In some embodiments, the purity of the extracted DNA and / or RNA samples as described herein can be analyzed by any other suitable technique or tool. In some embodiments, if the extracted RNA or DNA samples do not meet certain quantity or purity criteria as described above, they are not further processed. In some embodiments, if the extracted RNA or DNA samples do not meet certain quantity or purity criteria, the samples are combined with another sample.

[0160] Library preparation for RNA sequencing Methods for preparing cDNA libraries from RNA samples are known in the art. For example, www.illumina.com / content / dam / illumina - marketing / documents / applications / ngs - library - prep / for - all - you - seq - rna.pdf provides illustrations of different methods for preparing cDNA libraries for RNA sequencing. Non - limiting examples of cDNA library preparation include ClickSeq, 3Seq, and cP - RNA - Seq. In some embodiments, preparing a cDNA library from RNA includes purifying mRNA from the RNA sample (RNA enrichment). In some embodiments, the enriched RNA is fragmented. In some embodiments, after the appropriate RNA fraction has been selected, the molecules are fragmented into smaller pieces of a size between 50 - 1000 bp (e.g., 50 - 100 bp, 100 - 800 bp, 100 - 500 bp, or 200 - 500 bp), depending on the sequencing platform being used. This fragmentation can be achieved either by fragmenting double - stranded (ds) cDNA or by fragmenting RNA. Both methods result in the same final product of a double - stranded cDNA library with adapters attached to each fragment.

[0161] In some embodiments, the library preparation method includes adding functional elements (e.g., sample index, molecular barcode or flow cell oligo binding site), enriching sequencing-competent DNA fragments, and / or one or more amplification steps to generate a sufficient amount of library DNA for downstream processing. In some embodiments, the enriched RNA (e.g., fragmented and enriched RNA) is amplified using random primers (e.g., random hexamers). In some embodiments, the enriched RNA (e.g., fragmented and enriched RNA) is amplified using oligo dTs. In some embodiments, the RNA is then removed from the formed cDNA. In some embodiments, the cDNA is amplified to include sequencing adapters and indexes (i.e., multiple indexes). The adapter is a DNA sequence of 10-100 bp (e.g., 10-20, 10-100, 20-80, 30-70, 40-60, 20-100, 40-100, 40-80, 30-60, or 45-65 bp) that can bind to the flow cell for sequencing. Also, the adapter enables PCR enrichment of the adapter-ligated DNA fragments. The adapter can enable sample indexing or barcoding such that multiple cDNA libraries can be mixed together in one sequencing sample (or lane), i.e., enable multiplexing. In some embodiments, the index or barcode is 4-20 bp in length (e.g., 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 4-20, 5-15, 6-12, or 4-12 bp). Tucf-genomics.tufts.edu / documents / protocols / TUCF_Understanding_Illumina_TruSeq_Adapters.pdf provides an exemplary protocol for preparing a cDNA library using adapters and indexing, which is hereby incorporated by reference in its entirety.Protocols for constructing DNA or RNA libraries are described at least in Examples 3 and 5.

[0162] RNA Enrichment Methods for RNA enrichment (also described herein as "RNA enrichment") for enriching mRNA during cDNA library preparation are known in the art. RNA enrichment may be either targeted or non-targeted. Targeted methods of RNA enrichment include the use of sequence-specific capture probes. Non-limiting examples of targeted mRNA enrichment include CaptureSeq (sapac.illumina.com / science / sequencing-method-explorer / kits-and-arrays / aptureseq.html), which utilizes capture probes specific to the sequence of interest. Other platforms and tools suitable for targeted mRNA enrichment may also be used.

[0163] Examples of methods for non-targeted mRNA enrichment include poly(A) capture using oligo(dT) (e.g., conjugated to beads) and depletion of rRNA. Petrova et al., Scientific Reports volume 7, Article number: 41114 (2017) provides a comparison of various rRNA depletion methods, which is hereby incorporated by reference in its entirety. In some embodiments, rRNA depletion may be performed using an enzymatic approach (e.g., using an exonuclease that does not process mRNA). In some embodiments, the rRNA depletion method includes subtractive hybridization, whereby rRNA is captured using sequence-specific probes (see, e.g., www.sciencedirect.com / topics / immunology-and-microbiology / subtractive-hybridization).

[0164] In some embodiments, poly(A) capture involves capturing mRNA having a poly(A) tail using a poly(A)-specific capture probe (oligo(dT)). In some embodiments, the capture probe is immobilized to facilitate purification. In some embodiments, the capture probe is immobilized on beads (e.g., magnetic beads). In some embodiments, a commercially available kit is used to prepare a DNA library from an RNA sample. In some embodiments, the Illumina TruSeq RNA Library Prep kit is used.

[0165] The choice of mRNA enrichment can have a significant impact on the selection of the sequenced transcripts. For example, in some embodiments, compared to the rRNA depletion method, cDNA libraries prepared using poly(A) enrichment result in libraries that contain a higher fraction (e.g., greater than 80%, greater than 90%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, or greater than 99.9%) of protein-coding transcripts when compared to non-coding transcripts (e.g., rRNA, miRNA, and IncRNA).

[0166] In some embodiments, the prepared cDNA library is examined for quality. In some embodiments, quantification of the library for use in sequencing is generally performed before the library is pooled for target enrichment or amplification, in order to ensure equal representation of the indexed libraries in a multiplexed application. In some embodiments, quantification is also used to confirm that individual libraries or library pools are optimally diluted prior to sequencing. Accurate and reproducible quantification of adapter-ligated library molecules contributes to obtaining consistent and reproducible results and to maximizing the yield of sequencing. If the DNA to be loaded is more than the recommended amount, the flow cell may become saturated or the cluster density may increase. If the DNA to be loaded is too little, the cluster density may decrease, and the coverage and depth may decrease.

[0167] Methods for quantifying DNA libraries include electrophoresis, fluorometry, spectrophotometry, digital PCR, droplet digital PCR, and qPCR. There are various instruments for measuring the amount and / or quality of DNA libraries, such as the Agilent High Sensitivity D1000 ScreenTape System.

[0168] Aspects of the present disclosure provide quality control of nucleic acids for sequencing analysis. Aspects of the present disclosure provide quality control of DNA for sequencing analysis. Aspects of the present disclosure provide quality control of RNA for sequencing analysis. In some embodiments, the nucleic acid can include any suitable type of DNA or RNA. In some embodiments, quality control of the nucleic acid includes confirmation of biopsy conditions and documentation. In some embodiments, confirmation of biopsy conditions and documentation can include, but is not limited to, cataloging and registering nucleic acid materials. In some embodiments, confirmation of biopsy conditions and documentation includes acceptance of nucleic acid materials. By way of example, a patient sample received from a healthcare provider is checked to see whether the patient tissue is in a fresh frozen state or a formalin-fixed paraffin-embedded state. The laboratory staff verifies compliance of the registered entity's biopsy. The laboratory staff verifies proper storage of the biopsy sample during transportation. The laboratory staff verifies the physical condition of the biopsy sample. If the laboratory staff identifies any error regarding the biopsy sample, the source of the biopsy sample (e.g., the healthcare provider) can be notified. In some embodiments, if the received biopsy sample is a patient tissue cell line, the sample is prepared for extraction. In some embodiments, if the received biopsy sample is extracted DNA or RNA, the sample is stored at -80°C for further sequencing. In some embodiments, the extracted DNA can be reference gDNA. In some embodiments, the extracted RNA can be reference RNA.

[0169] In some embodiments, the quality control procedure defines a target range. The target range may represent the most ideal quality for a given step (e.g., extraction). In some embodiments, the quality control procedure defines an acceptable range. The acceptable range may represent the ideal or acceptable quality for a given step. In some embodiments, the quality control of nucleic acids includes ensuring the quality in the process of constructing a DNA library. In some embodiments, the quality control of nucleic acids includes ensuring the quality in the process of constructing an RNA library. As shown in FIG. 7 and Example 6, the preparation of a DNA or RNA library includes extracting DNA or RNA from a patient tissue sample. In some embodiments, a spectrophotometer, e.g., a small-volume full-spectrum ultraviolet-visible spectrophotometer (e.g., NanoDrop spectrophotometer available from ThermoFisher Scientific, www.thermofisher.com), can be used to determine the quality of DNA or RNA extraction. As an example, extracted DNA >100 ng / μl indicates that the extracted DNA has passed the quality control test. Extracted RNA >500 ng / μl indicates that the extracted RNA has passed the quality control test. In another example, an absorbance ratio of 260 nm to 280 nm (260 / 280) of the extracted DNA being 1.8 - 2.0 indicates that the extracted DNA has passed the quality control test. An absorbance ratio of 260 nm to 280 nm (260 / 280) of the extracted RNA being 2.0 indicates that the extracted RNA has passed the quality control test. In another example, an absorbance ratio of 260 nm to 230 nm (260 / 230) of the extracted DNA being 2.0 - 2.2 indicates that the extracted DNA has passed the quality control test. An absorbance ratio of 260 nm to 230 nm (260 / 230) of the extracted RNA being 2.0 - 2.2 indicates that the extracted RNA has passed the quality control test.In some embodiments, a fluorometer (e.g., a Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com) can be used to determine the quality of DNA or RNA extraction, for example, for the quantification of DNA or RNA. In some embodiments, an electrophoresis device, e.g., an automated electrophoresis device (e.g., a TapeStation System available from Agilent, www.agilent.com), can be used to determine the quality of DNA or RNA extraction. In some embodiments, any suitable technique or tool can be used to determine the quality of DNA or RNA extraction.

[0170] In some embodiments, the acceptable total DNA amount for further DNA library construction is at least 200 - 1,000 ng (e.g., 200 - 1,000 ng, 300 - 1,000 ng, or 300 - 1,000 ng). In some embodiments, the target total DNA amount for further sequencing is more than 500 - 1,000 ng (e.g., 500 - 1,000 ng, 600 - 1,000 ng, or 800 - 1,000 ng). In some embodiments, the acceptable total RNA amount for further RNA library construction is at least 0.5 - 4 nmol / l (e.g., 200 - 1,000 ng, 300 - 1,000 ng, or 300 - 1,000 ng). In some embodiments, the target total RNA amount for further RNA library construction is at least 0.5 - 4 nmol / l (e.g., 500 - 1,000 ng, 600 - 1,000 ng, or 800 - 1,000 ng).

[0171] In some embodiments, the acceptable DNA concentration for further DNA library construction is at least 17 ng / μl (e.g., 17 ng / μl, 25 ng / μl, 35 ng / μl). In some embodiments, the target DNA concentration for further DNA library construction is at least 42 ng / μl (e.g., 42 ng / μl, 50 ng / μl, 80 ng / μl). In some embodiments, the acceptable RNA concentration for further RNA library construction is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 1 ng / μl, 3 ng / μl). In some embodiments, the target RNA concentration for further RNA library construction is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 1 ng / μl, 3 ng / μl). In some embodiments, the DNA and RNA concentrations are detected, for example, by a fluorometer (e.g., Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com) for quantification of DNA or RNA.

[0172] In some embodiments, the acceptable DNA concentration for further DNA library construction is at least 15 ng / μl (e.g., 15 ng / μl, 25 ng / μl, 35 ng / μl). In some embodiments, the target DNA concentration for further DNA library construction is at least 40 ng / μl (e.g., 40 ng / μl, 50 ng / μl, 80 ng / μl). In some embodiments, the acceptable RNA concentration for further RNA library construction is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 1 ng / μl, 3 ng / μl). In some embodiments, the target RNA concentration for further RNA library construction is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 1 ng / μl, 3 ng / μl). In some embodiments, the acceptable RNA concentration for further RNA library construction is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 1 nmol / l, 5 nmol / l). In some embodiments, the target RNA concentration for further RNA library construction is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 1 nmol / l, 5 nmol / l). In some embodiments, the DNA and RNA concentrations are detected by Tapestation.

[0173] In some embodiments, the acceptable RNA concentration for further RNA library construction is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 1 nmol / l, 5 nmol / l). In some embodiments, the target RNA concentration for further RNA library construction is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 1 nmol / l, 5 nmol / l). In some embodiments, the DNA and RNA concentrations are detected by a nucleic acid amplification device (e.g., a PCR system), such as a real-time PCR system (e.g., a LightCycler Instrument available from Roche, www.lifescience.roche.com). In some embodiments, the DNA and RNA concentrations can be detected by any suitable technique or tool.

[0174] In some embodiments, when RNA is extracted, reverse transcription can be performed. In some embodiments, after reverse transcription is performed, an RNA library can be constructed. In some embodiments, a fluorometer (e.g., a Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com) for quantifying, for example, DNA or RNA can be used to determine the quality of a DNA or RNA library. In some embodiments, any suitable method can be used to determine the quality of a DNA or RNA library. In some embodiments, an electrophoresis device, e.g., an automated electrophoresis device (e.g., a TapeStation System available from Agilent, www.agilent.com), can be used to determine the quality of a DNA or RNA library. In some embodiments, a fluorometer (e.g., a Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com) for quantifying, for example, DNA or RNA can be used to determine the quality of DNA or RNA extraction. In some embodiments, a nucleic acid amplification device (e.g., a PCR system), e.g., a real-time PCR system (e.g., a LightCycler Instrument available from Roche, www.lifescience.roche.com), can be used to determine the quality of an RNA library. In some embodiments, one or more RNA libraries can be pooled. In some embodiments, when DNA is extracted, the extracted DNA can be used for DNA library construction. In some embodiments, DNA fragments within a constructed DNA library can be hybridized and / or captured. In some embodiments, a fluorometer (e.g., a Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com) for quantifying, for example, DNA or RNA can be used to determine the quality of the DNA hybridization and capture steps It can be used to determine. In some embodiments, an electrophoresis device, e.g., an automated electrophoresis device (e.g., TapeStation System available from Agilent, www.agilent.com), can be used to determine the quality of the DNA hybridization and capture steps. In some embodiments, a nucleic acid amplification device (e.g., a PCR system), e.g., a real-time PCR system (e.g., LightCycler Instrument available from Roche, www.lifescience.roche.com), can be used to determine the quality of the DNA hybridization and capture steps. In some embodiments, any suitable method can be used to determine the quality of the DNA hybridization and capture steps. In some embodiments, one or more DNA libraries can be pooled. In some embodiments, an electrophoresis device, e.g., an automated electrophoresis device (e.g., TapeStation System available from Agilent, www.agilent.com), can be used to determine the quality of DNA or RNA library pooling. In some embodiments, any suitable method can be used to determine the quality of DNA or RNA library pooling.

[0175] In some embodiments, the acceptable and / or target final DNA concentration range for pooling is at least 0.5 to 4 nmol / l (e.g., 0.5 to 4 nmol / l, 0.5 to 3 nmol / l, 2 to 4 nmol / l). In some embodiments, the acceptable DNA concentration for pooling is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 0.8 ng / μl, 4 ng / μl) when a fluorometer for quantification of DNA or RNA (e.g., Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com) is used. In some embodiments, the target DNA concentration for pooling is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 0.8 ng / μl, 4 ng / μl) when a fluorometer for quantification of DNA or RNA (e.g., Qubit fluorometer available from ThermoFisher Scientific, www.thermofisher.com) is used.

[0176] In some embodiments, the acceptable DNA concentration for pooling is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 0.8 ng / μl, 4 ng / μl) when an electrophoresis device, such as an automated electrophoresis device (e.g., TapeStation System available from Agilent, www.agilent.com), is used. In some embodiments, the target DNA concentration for pooling is at least 0.1 ng / μl (e.g., 0.1 ng / μl, 0.8 ng / μl, 4 ng / μl) when an electrophoresis device, such as an automated electrophoresis device (e.g., TapeStation System available from Agilent, www.agilent.com), is used. In some embodiments, the acceptable DNA concentration for pooling is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 0.8 nmol / l, 3 nmol / l) when an electrophoresis device, such as an automated electrophoresis device (e.g., TapeStation System available from Agilent, www.agilent.com), is used. In some embodiments, the target DNA concentration for pooling is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 0.8 nmol / l, 3 nmol / l) when an electrophoresis device, such as an automated electrophoresis device (e.g., TapeStation System available from Agilent, www.agilent.com), is used. In some embodiments, the acceptable concentration and / or concentration of DNA is in the range of 380 - 440 ng (e.g., 380 - 440 ng, 400 - 440 ng, 420 - 440 ng) when an electrophoresis device, such as an automated electrophoresis device (e.g., TapeStation System available from Agilent, www.agilent.com), is used.In some embodiments, the acceptable DNA concentration for pooling is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 0.8 nmol / l, 3 nmol / l) when a nucleic acid amplification device (e.g., a PCR system), such as a real-time PCR system (e.g., the LightCycler Instrument available from Roche, www.lifescience.roche.com), is used. In some embodiments, the target DNA concentration for pooling is at least 0.5 nmol / l (e.g., 0.5 nmol / l, 0.8 nmol / l, 3 nmol / l) when the LightCycler is used.

[0177] In some embodiments, quality control of nucleic acids includes ensuring the quality after DNA or RNA library construction, such as during the sequencing process. In some embodiments, cluster density may be a parameter for quality control of sample runs (Example 6). Cluster density is an important factor in optimizing the data quality and yield of sequencing. Without wishing to be bound by any theory, an optimal cluster density indicates that at least the DNA or RNA library is balanced. In some embodiments, quality score and signal-to-noise ratio may be parameters for quality control of sample runs.

[0178] In some embodiments, nucleic acid quality control includes ensuring sequencing quality. In some embodiments, sequencing quality control includes bioinformatics quality control. In some embodiments, the sequencing can be DNA sequencing. In some embodiments, the sequencing can be RNA sequencing. In some embodiments, the sequencing can be any type of sequencing technology known in the art for determining the expression profile of DNA or RNA of a given biological sample. By way of example, the sequencing can be whole exome sequencing. The sequencing can be transcriptome sequencing. The sequencing can be Sanger sequencing.

[0179] In some embodiments, up to 1 ng (e.g., up to 0, up to 0.1 μl, up to 0.5 μl, up to 0.8 μl, up to 0.9 μl, up to 1 μl, up to 1.2 μl, up to 1.4 μl, up to 1.5 μl, up to 1.8 μl, or up to 2 μl) of a solution library up to 2 μl (e.g., up to 0.1, up to 0.2, up to 0.3, up to 0.4, up to 0.5, up to 0.6, up to 0.7, up to 0.8, up to 0.9, or up to 1 ng) is used for quality control testing. In some embodiments, the parameters examined include the size and size distribution of DNA molecules, as well as purity.

[0180] In some embodiments, standard methods for preparing a library of cDNA fragments from RNA cannot preserve information regarding which DNA strand was the original template during transcription and subsequent synthesis of mRNA transcripts. Since antisense transcripts are likely to have regulatory roles distinct from those of protein-coding complements, such loss of strand information results in an incomplete understanding of the transcriptome. Strand-specific RNA-Seq can be performed to retain this strandness. Methods for preserving strandness and preparing a cDNA fragment library therefor are known in the art (e.g., Mills et al., "Strand-Specific RNA-Seq Provides Greater Resolution of Transcriptome Profiling", Curr Genomics. May 2013, 14(3):173-181). In some embodiments, library preparation for strand RNA-seq utilizes strand-specific adapters of known orientation. In some embodiments, the strand is chemically modified to remember its origin.

[0181] In some embodiments, methods using adapters include strand-specific 3'-end RNA-seq. In some embodiments, strand-specific 3'-end RNA-Seq includes anchor oligo(dT) primers, which are first used to select mRNA, resulting in the production of double-stranded cDNA molecules. Next, adapters for paired-end sequencing are ligated to both ends of the cDNA molecules. The fragments are then sequenced to generate paired-end reads aligned to the reference genome. Aligned reads containing a stretch of adenines at the end of the transcript must be transcripts from the DNA antisense strand, while any reads that align with a stretch of thymines at the front must be transcripts from the DNA sense strand.

[0182] In some embodiments, the method of using an adapter utilizes four DNA ligases that enable ligation of single-stranded (ss) cDNA, Illumina adapters, and 3' and 5' adapters to the ssDNA. Since the second strand is never synthesized and sequencing does not proceed, strand information is retained.

[0183] In some embodiments, any suitable technique or tool may be used to preserve strandness. For example, flow cell reverse transcription sequencing (FRT-Seq) may be used to preserve strandness. In some embodiments, FRT-Seq or equivalent techniques include ligating an adapter to either end of fragmented and purified polyadenylated mRNA. In some embodiments, each adapter includes two regions: a region to which a sequencing primer anneals and a region complementary to an oligonucleotide present on the flow cell. The complementary region enables the mRNA fragment to hybridize to the flow cell. The mRNA fragment is then reverse transcribed on the flow cell surface.

[0184] Other non-limiting adapter-based methods of preserving strandness include direct strand-specific sequencing (DSSS) and the SOLiDR Total RNA-Seq Kit (tools.thermofisher.com / content / sfs / manuals / cms_078610.pdf) that preserves strand specificity through the addition of directional adapters.

[0185] In some embodiments, chemical modification of the strand to remember its origin includes marking the original RNA template using bisulfite treatment. In some embodiments, dUTP is incorporated into the reverse transcription reaction, resulting in a ds cDNA where the original strand has deoxythymidine residues and the complementary strand has deoxyuridine residues. Then, uracil-DNA-glycosylase (UDG) treatment may be used to degrade the complementary strand.

[0186] Preparation of a Library for WES An "exome" is the sum total of all regions within the genome that consist of exons. Exons are DNA regions that are transcribed into messenger RNA, as opposed to introns, which are removed by splicing to make proteins. Exome sequencing is a capture-based method developed to identify variants present in the coding regions of genes that affect protein function. Since the coding portion of the genome contains only 1-2% of the entire genome, this approach represents a cost-effective strategy for detecting DNA modifications that can alter protein function, as compared to whole-genome sequencing. In some embodiments, whole exome sequencing (WES) includes preparing a library of DNA fragments for sequencing from a DNA sample. In some embodiments, the DNA is first fragmented to an appropriate size (depending on the sequencing platform being used), and then sequencing platform-specific adapters are added. In some embodiments, the library is amplified prior to the next step of the process (target enrichment or sequencing).

[0187] Kits for preparing libraries are commercially available, and non-limiting examples thereof include KAPA HyperPrep Kits, Agilent HaloPlex, Agilent SureSelect QXT, IDT xGEN Exome, Illumina Nextera Rapid Capture Exome, Roche Nimblegen SeqCap, and MYcroarray MYbaits. In some embodiments, any kit capable of preparing a DNA library for WES can be used. For example, the Agilent Human All Exon V6 Capture Kit (www.agilent.com / cs / library / datasheets / public / SureSelect%20V6%20DataSheet%205991-5572EN.pdf) is used to prepare a DNA library for WES. In some embodiments, the Clinical Research Exome kit (www.agilent.com / en / promotions / clinical-research-exome-v2) is used. The amount of DNA required depends on the specific reagents used to prepare the library. For example, 100 ng of genomic DNA is sufficient for the Agilent SureSelect XT2 V6 Exome, while 500 ng of genomic DNA is required for the IDT xGEN Exome Panel. A comparison of various capture kits is provided at www.genohub.com / exome-sequencing-library-preparation / .

[0188] In some embodiments, the library preparation method includes one or more amplification steps to add functional elements (such as sample indexes, molecular barcodes, or flow cell oligo binding sites), enrich sequencing-competent DNA fragments, and / or generate a sufficient amount of library DNA for downstream processing. As an example, the library preparation method is shown in Examples 3 and 5.

[0189] In some embodiments, the prepared DNA library is inspected for quality. In some embodiments, quantification of the library for use in sequencing is generally performed before the library is pooled for target enrichment or amplification to ensure equal representation of the indexed libraries in multiplexed applications. In some embodiments, quantification is also used to confirm that individual libraries or library pools are optimally diluted prior to sequencing. Accurate and reproducible quantification of adapter-ligated library molecules contributes to obtaining consistent and reproducible results and to maximizing the yield of sequencing. If too much DNA is loaded, the flow cell may become saturated or the cluster density may be high, and if too little DNA is loaded, the cluster density may be low, resulting in reduced coverage and depth.

[0190] Methods for quantifying DNA libraries include electrophoresis, fluorometry, spectrophotometry, digital PCR, droplet digital PCR, and qPCR. There are various instruments for measuring the quantity and / or quality of DNA libraries, such as the Agilent High Sensitivity D1000 ScreenTape System.

[0191] In some embodiments, the prepared DNA library is inspected for quality. In some embodiments, up to 1 ng (e.g., up to 0, up to 0.1 μg, up to 0.5 μg, up to 0.8 μg, up to 0.9 μg, up to 1 μg, up to 1.2 μg, up to 1.4 μg, up to 1.5 μg, up to 1.8 μg, or up to 2 μg) of the library in a solution of up to 2 μl (e.g., up to 0.1, up to 0.2, up to 0.3, up to 0.4, up to 0.5, up to 0.6, up to 0.7, up to 0.8, up to 0.9, or up to 1 ng) is used for quality control testing. In some embodiments, the parameters inspected include the size and size distribution of DNA molecules, as well as purity.

[0192] RNA sequencing RNA sequencing is a tool for measuring the transcriptome. The transcriptome consists of a diverse population of RNA molecules, including mRNA, rRNA, tRNA, and other non-coding RNAs (such as microRNA, lncRNA, etc.). In some embodiments, RNA sequencing is used to profile the transcriptome (e.g., coding regions and / or non-coding regions). In some embodiments, this is used to identify genes that are expressed in different forms in different biological samples (e.g., cells, tissues, or body fluids). In some embodiments, RNA sequencing is used to determine the genetic impact of splicing events, identify novel transcripts, detect structural variations (e.g., gene fusions and isoforms), and / or detect single nucleotide variants.

[0193] In some embodiments, the term “RNA sequencing” can be used interchangeably with “RNA seq,” “RNA-seq,” or variants thereof as known in the art to refer to any technique, tool, or platform for interrogating a transcriptome. When “RNA sequencing,” “RNA seq,” “RNA-seq,” or variants thereof are referenced in the present disclosure, it should be noted that no specific technology or tool associated with a particular platform or company is being referred to, unless otherwise shown using non-limiting examples to demonstrate the processes or systems as described herein. In some embodiments, RNA sequencing can be performed by using any suitable sequencing platform and / or sequencing method. Non-limiting examples of high-throughput sequencing platforms include mRNA-seq, total RNA-seq, targeted RNA-seq, single-cell RNA-seq, RNA exome capture platforms, or small RNA-seq (e.g., Illumina, www.illumina.com), SMRT (single molecule, real-time) sequencing (e.g., Pacific Biosciences, https: / / www.pacb.com), and RNA sequencing (e.g., ThermoFisher, https: / / www.thermofisher.com).

[0194] As described above, RNA sequencing can be targeted or untargeted. Targeted approaches involve using sequence-specific probes or oligonucleotides to sequence one or more specific regions of the transcriptome. In some embodiments, targeted RNA sequencing includes methods such as mRNA enrichment (e.g., by polyA enrichment or rRNA depletion).

[0195] In some embodiments, RNA sequencing is whole transcriptome sequencing. Whole transcriptome sequencing involves measuring the complete complement of transcripts in a sample. In some embodiments, whole transcriptome sequencing is used to determine the global expression levels of each transcript (e.g., both coding and non-coding) and to identify exons, introns, and / or their junctions.

[0196] In some embodiments, RNA is sequenced directly without preparing cDNA from an RNA sample. In some embodiments, direct RNA sequencing includes single molecule RNA sequencing (DRSTM).

[0197] In some embodiments, RNA sequencing is mRNA sequencing. In some embodiments, mRNA sequencing is sequencing of only the coding transcripts that aim to exclude non-coding regions. In some embodiments, mRNA sequencing is independent of polyA enrichment. In some embodiments, mRNA sequencing is dependent on polyA enrichment.

[0198] In some embodiments, RNA is extracted from a biological sample, mRNA is enriched from the extracted RNA, and a cDNA library is constructed from the enriched mRNA. In some embodiments, single pieces of cDNA from the cDNA library are attached to a solid matrix. In some embodiments, single pieces of cDNA from the cDNA library are attached to the solid matrix by limiting dilution. Then, in some embodiments, the cDNA pieces attached to the matrix are sequenced (e.g., using Pacbio or Pacifbio technology). In some embodiments, the cDNA pieces attached to the matrix are amplified and sequenced (e.g., using dedicated emulsion PCR (emPCR) in a connector based on the SOLiD, 454 Pyrosequencing, Ion Torrent, or bridge reaction (Illumina) platform).

[0199] In some embodiments, cDNA transcripts can be sequenced in parallel by measuring incorporation of fluorescent nucleotides (e.g., Illumina), incorporation of fluorescent short linkers (e.g., SOLiD), release of by-products derived from incorporation of normal nucleotides (454), measuring fluorescence emission, or measuring pH changes (e.g., Ion Torrent). In some embodiments, cDNA transcripts can be sequenced using any known sequencing platform. Jazayeri et al., "RNA-seq: a glance at technologies and methodologies", Acta biol. Colomb. vol.20 no.2 Bogota May / Aug. 2015 provides a comparison of different RNA-seq platforms, including Tables 3 and 4, which are hereby incorporated by reference in their entirety. Mestan et al., "Genomic sequencing in clinical trials", Journal of Translational Medicine 2011, 9:222 performs a similar analysis of sequencing in clinical trials.

[0200] In some embodiments, RNA sequencing is strand or strand-specific. cDNA synthesis from RNA results in a loss of strandness. In some embodiments, strandness is preserved by chemically labeling either or both of the RNA and cDNA strands formed by reverse transcription or antisense transcription, as described above, or by using adapter-based techniques to distinguish the original RNA strand from the complementary DNA strand.

[0201] In some embodiments, non-stranded RNA sequencing is performed. In some embodiments, strand RNA-seq should be avoided for clinical samples. In some embodiments, non-stranded RNA-seq is used to compare data obtained from biological samples to RNA sequencing data of established datasets (e.g., The Cancer Genome Atlas (TCGA) and International Cancer Genome Consortium (ICGC)).

[0202] In some embodiments, paired-end reads are obtained by RNA sequencing. Paired-end reads are reads of the same nucleic acid fragment and are reads starting from either end of the fragment. In some embodiments, RNA sequencing is performed with at least 2×25 (2×25, 2×50, 2×75, 2×100, 2×125, 2×150, 2×175, 2×200, 2×225, 2×250, 2×275, 2×300, 2×325, or 2×350) paired-end reads. In some embodiments, RNA sequencing is performed with at least 2×75 paired-end reads. RNA sequencing with 2×75 paired-end reads means that, on average, each read that is paired reads 75 base pairs. In some embodiments, RNA sequencing is performed with a total of at least 20 million (e.g., at least 20 million, at least 30 million, at least 40 million, at least 50 million, at least 60 million, at least 70 million, at least 80 million, at least 90 million, at least 100 million, at least 120 million, at least 140 million, at least 150 million, at least 160 million, at least 180 million, at least 200 million, at least 250 million, at least 300 million, at least 350 million, or at least 400 million) paired-end reads. In some embodiments, RNA sequencing is performed with a total of at least 50 million paired-end reads. In some embodiments, RNA sequencing is performed with a total of at least 100 million paired-end reads.

[0203] In some embodiments, quality control is performed for RNA sequencing. In some embodiments, cluster density or cluster PF% are parameters for determining the quality of a sample run. In some embodiments, the target range of cluster density or cluster PF% is at least 170 - 220 (e.g., 170 - 220, 190 - 220, 210 - 220). In some embodiments, the acceptable range of cluster density or cluster PF% is at least 280 (e.g., 280, 300, 450).

[0204] In some embodiments, %≧Q30 is a parameter for determining the quality of a sample run. In some embodiments, the target %≧Q30 is at least 85% (e.g., 85%, 90%, 95%). In some embodiments, the acceptable %≧Q30 is at least 75% (e.g., 75%, 85%, 95%).

[0205] In some embodiments, the error rate % is a parameter for determining the quality of a sample run. In some embodiments, the target error rate % is less than at least 0.7% (e.g., 0.6%, 0.5%, 0.4%). In some embodiments, the acceptable error rate % is less than at least 1% (e.g., 0.9%, 0.8%, 0.7%).

[0206] Whole exome sequencing (WES) Whole exome sequencing (WES) is a genomic technology for sequencing all of the protein-coding regions of genes within a genome. In some embodiments, WES is performed to identify genetic mutations that alter protein sequences. In some embodiments, WES is performed to identify genetic mutations that alter protein sequences at a lower cost than whole genome sequencing.

[0207] In some embodiments, whole exome sequencing (WES) is performed on a sample of DNA extracted from a biological sample. In some embodiments, a library of DNA fragments is prepared from the extracted DNA sample. In some embodiments, any one of the methods described herein includes performing whole exome sequencing (WES) on a library of DNA fragments. Preparation of the DNA library from the DNA sample for WES is as described above.

[0208] In some embodiments, the DNA library is quantified prior to sequencing (e.g., using next-generation sequencing (NGS)). In some embodiments, the DNA libraries are pooled prior to sequencing. In some embodiments, the DNA library is amplified prior to sequencing. In some embodiments, the DNA library is indexed prior to sequencing to track the origin of the DNA fragments.

[0209] In some embodiments, WES includes target enrichment that enables selective capture of genomic regions of interest prior to sequencing. In some embodiments, array-based capture is used (e.g., using a microarray). In some embodiments, solution-phase capture is used.

[0210] Any high-throughput DNA sequencing platform and / or method can be used in any one of the methods described herein. In some embodiments, DNA sequencing can be performed by using any suitable platform and / or method. Non-limiting examples of high-throughput sequencing methods include single molecule real-time sequencing, ion semiconductor (Ion Torrent sequencing), pyrosequencing (i.e., 454), sequencing by synthesis (Illumina), Illumina (Solexa) sequencing, combinatorial probe anchor synthesis (cPAS - BGI / MGI), sequencing by ligation (SOLiD sequencing), nanopore sequencing (e.g., using devices from Oxford Nanopore Technologies), chain termination (Sanger sequencing), massively parallel signature sequencing (MPSS) polony sequencing, Heliscope single molecule sequencing, and single molecule real-time (SMRT) sequencing (e.g., using devices from Pacific Biosciences). Other non-limiting examples of high-throughput sequencing technologies include tunneling current DNA sequencing, sequencing by hybridization, sequencing using mass spectrometry, microfluidic Sanger sequencing, and RNAP sequencing.

[0211] In some embodiments, paired-end reads are obtained by DNA sequencing. Paired-end reads are reads from the same nucleic acid fragment and are reads starting from either end of the fragment. In some embodiments, the DNA sequencing is performed with at least 2×25 (2×25, 2×50, 2×75, 2×100, 2×125, 2×150, 2×175, 2×200, 2×225, 2×250, 2×275, 2×300, 2×325, or 2×350) paired-end reads. In some embodiments, the DNA sequencing is performed with at least 2×75 paired-end reads. DNA sequencing with 2×75 paired-end reads means that, on average, each read that is paired-end reads 75 base pairs. In some embodiments, the DNA sequencing is performed with a total of at least 20 million (e.g., at least 20 million, at least 30 million, at least 40 million, at least 50 million, at least 60 million, at least 70 million, at least 80 million, at least 90 million, at least 100 million, at least 120 million, at least 140 million, at least 150 million, at least 160 million, at least 180 million, at least 200 million, at least 250 million, at least 300 million, at least 350 million, or at least 400 million) paired-end reads. In some embodiments, the DNA sequencing is performed with a total of at least 50 million paired-end reads. In some embodiments, the DNA sequencing is performed with a total of at least 100 million paired-end reads. In some embodiments, the DNA sequencing is performed such that at least 20-fold (e.g., at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, at least 100-fold, at least 120-fold, at least 125-fold, at least 150-fold, at least 175-fold, at least 200-fold, at least 250-fold, at least 300-fold, or at least 400-fold) coverage is obtained.Coverage, also referred to as depth, is the number of times a single base pair in a nucleic acid sample is on average read or sequenced. In some embodiments, the portion of the genome targeted for capture and sequencing is at least 10 Mb (e.g., at least 10 Mb, at least 20 Mb, at least 30 Mb, at least 40 Mb, at least 50 Mb, at least 60 Mb, at least 70 Mb, at least 80 Mb, at least 90 Mb, at least 100 Mb, at least 120 Mb, at least 150 Mb, at least 200 Mb, at least 250 Mb, at least 300 Mb, or at least 350 Mb). In some embodiments, the portion of the genome targeted for capture and sequencing is at least 48 Mb (e.g., after using the Agilent Human All Exon V6 Capture system). In some embodiments, the portion of the genome targeted for capture and sequencing is at least 54 Mb (e.g., after using the Clinical Research Exome capture system (Agilent)).

[0212] In some embodiments, quality control is performed for whole exome sequencing. In some embodiments, cluster density or cluster PF% is a parameter for determining the quality of a sample run. In some embodiments, the target range of cluster density or cluster PF% is at least 170 - 220 (e.g., 170 - 220, 190 - 220, 210 - 220). In some embodiments, the acceptable range of cluster density or cluster PF% is at least 280 (e.g., 280, 300, 450).

[0213] In some embodiments, the actual yield is a parameter for determining the quality of a sample run. In some embodiments, the target actual yield is at least 15 Gbp (e.g., 15 Gbp, 20 Gbp, 30 Gbp).

[0214] In some embodiments, %≧Q30 is a parameter for determining the quality of a sample run. In some embodiments, the target %≧Q30 is at least 85% (e.g., 85%, 90%, 95%). In some embodiments, the acceptable %≧Q30 is at least 75% (e.g., 75%, 85%, 95%).

[0215] In some embodiments, the error rate % is a parameter for determining the quality of a sample run. In some embodiments, the target error rate % is less than at least 0.7% (e.g., 0.6%, 0.5%, 0.4%). In some embodiments, the acceptable error rate % is less than at least 1% (e.g., 0.9%, 0.8%, 0.7%).

[0216] Reagents and Kits Contemplated herein are reagents and kits comprising reagents for performing any one of the methods described herein. In some embodiments, a kit as provided herein includes reagents (e.g., buffers, preservatives, inhibitors, or enzymes) and / or laboratory instruments (e.g., pipettes, filters, tubes, storage containers such as vacutainers, or dissection tools) for storing a biological sample obtained from a subject.

[0217] In some embodiments, a kit as provided herein comprises reagents (e.g., buffers, preservatives, inhibitors or enzymes) and / or laboratory instruments (e.g., pipettes, filters, or tubes) for extracting RNA and / or DNA from a biological sample or a sample derived from a biological sample (e.g., a single cell solution). In some embodiments, a kit as provided herein comprises reagents (e.g., buffers, preservatives, inhibitors, enzymes or dyes) and / or laboratory instruments (e.g., pipettes, filters, tubes, storage containers, or electrophoresis paper) for measuring the quality and quantity of RNA and / or DNA extracted from a biological sample. In some embodiments, a kit as provided herein comprises reagents (e.g., buffers, preservatives, inhibitors, enzymes or dyes) and / or laboratory instruments (e.g., pipettes, filters, tubes, storage containers, or electrophoresis paper) for measuring the quality and quantity of a DNA library for sequencing (e.g., RNA-seq or WES).

[0218] In some embodiments, a kit as provided herein comprises reagents (e.g., buffers, preservatives, inhibitors or enzymes) and / or laboratory instruments (e.g., pipettes, filters, tubes, storage containers such as vacutainers, or dissection tools) for preparing a single cell solution from a biological sample.

[0219] In some embodiments, a kit as provided herein comprises reagents (e.g., buffers, inhibitors, or enzymes such as reverse transcriptase) and / or laboratory instruments (e.g., pipettes, filters, tubes, storage containers) for preparing a DNA library for sequencing.

[0220] In some embodiments, a kit as provided herein is for any combination of two or more of storing a biological sample, extracting RNA and / or DNA from a biological sample, assaying the quality and quantity of the extracted RNA and / or DNA sample and / or DNA library prepared therefrom, preparing a single cell solution from a biological sample, and preparing a DNA library from the extracted RNA and / or DNA, and comprises reagents (e.g., buffers, preservatives, inhibitors, or enzymes) and / or laboratory instruments (e.g., pipettes, filters, tubes, storage containers such as vacutainers, or dissection tools).

[0221] In some embodiments, any one of the kits described herein comprises components for making a cell dissociation cocktail. The cell dissociation cocktail may be enzymatic or non-enzymatic. In some embodiments, the kit comprises one or more enzyme cocktails. In some embodiments, the kit includes, as components, one or more of a medium (e.g., L-15 medium), an antibacterial agent (e.g., penicillin and / or streptomycin), an antifungal agent (e.g., amphotericin), collagenase (e.g., collagenase I, collagenase II, collagenase IV), DNAse (e.g., DNAseI), elastase, hyaluronidase, protease (e.g., protease XIV, trypsin, papain, thermolysin). In some embodiments, any one of the kits described herein comprises, as an enzyme, one or more of collagenase I and collagenase IV. In some embodiments, these enzymes are housed in separate containers. In some embodiments, these enzymes are housed in a single container.

[0222] In some embodiments, the kit comprises a smaller device such as a spectrophotometer. In some embodiments, the kit comprises instructions for performing any one, or any two or more combinations, of storing a biological sample, extracting RNA and / or DNA from a biological sample, assaying the quality and quantity of the extracted RNA and / or DNA sample and / or DNA library prepared therefrom, preparing a single cell solution from a biological sample, and preparing a DNA library from the extracted RNA and / or DNA. In some embodiments, the kit comprises instructions for performing any one of the methods described herein. In some embodiments, the kit is made or adapted for a particular tissue type, such as a biopsy of a solid tumor, a liquid biopsy, a blood sample, or urine.

[0223] Data processing Aspects of the present disclosure relate to processing data obtained from RNA sequencing. In some embodiments, a method for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) comprises aligning genes in the RNA expression data with known sequences of the human genome, annotating to obtain annotated RNA expression data, removing non-coding transcripts from the annotated RNA expression data, converting the annotated RNA expression data to gene expression data in transcripts per kilobase million (TPM) format, identifying at least one gene that introduces a bias into the gene expression data, and obtaining bias-corrected gene expression data by removing at least one gene from the gene expression data. In some embodiments, a method for processing RNA expression data comprises obtaining RNA expression data from a subject having or suspected of having cancer.

[0224] In some embodiments, the non-coding transcripts include genes selected from the group consisting of pseudogenes, polymorphic pseudogenes, processed pseudogenes, transcribed processed pseudogenes, unitary pseudogenes, non-processed pseudogenes, transcribed unitary pseudogenes, constant chain immunoglobulin (IG C) pseudogenes, joining chain immunoglobulin (IG J) pseudogenes, variable chain immunoglobulin (IG V) genes, transcribed non-processed genes, translated non-processed genes, joining chain T cell receptor (TR J) genes, variable chain T cell receptor (TR V) genes, small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), microRNA (miRNA), ribozymes, ribosomal RNA (rRNA), mitochondrial tRNA (Mt tRNA), mitochondrial rRNA (Mt rRNA), Cajal body-specific RNA (scaRNA), retained introns, sense intron RNAs, sense overlapping RNAs, nonsense-mediated decay RNAs, non-stop decay RNAs, antisense RNAs, long intergenic non-coding RNAs (lincRNAs), macro long non-coding RNAs (macro lncRNAs), processed transcripts, 3' overlapping non-coding RNAs (3' overlapping ncrnas), small RNAs (sRNAs), other RNAs (miscRNAs), vault RNAs, and TEC RNAs.

[0225] In some embodiments, information (e.g., sequence information) for one or more of these types of transcripts can be obtained in a nucleic acid database (e.g., the Gencode database, e.g., Gencode V23, the Genbank database, the EMBL database, or other databases).

[0226] In some embodiments, a method for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) involves identifying cancer treatment (also referred to herein as anti-cancer therapy) for a subject using bias-corrected gene expression data. In some embodiments, any one of the methods for processing RNA expression data is further combined with administering one or more anti-cancer therapies or cancer treatments to the subject. In some embodiments, any one of the methods for processing RNA expression data is further combined with instructing or recommending administering one or more anti-cancer therapies or cancer treatments to the subject.

[0227] Obtaining RNA Expression Data In some embodiments, a method for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) involves obtaining RNA expression data of a subject (e.g., a subject having cancer or diagnosed with cancer). In some embodiments, obtaining RNA expression data involves obtaining a biological sample, processing it, and performing RNA sequencing using any one of the RNA sequencing methods described herein. In some embodiments, the RNA expression data is obtained from a laboratory or center that has performed an experiment to obtain the RNA expression data (e.g., a laboratory or center that has performed RNA-seq). In some embodiments, the laboratory or center is a clinical laboratory or center.

[0228] In some embodiments, RNA expression data is obtained by obtaining a computer storage medium (e.g., a data storage drive) in which the data resides. In some embodiments, RNA expression data is obtained via a secure server (e.g., an SFTP server, or Illumina BaseSpace). In some embodiments, the data is obtained in the form of a text-based file (e.g., a FASTQ file). In some embodiments, the file in which the sequencing data is stored also includes quality scores for the sequencing data. In some embodiments, the file in which the sequencing data is stored also includes sequence identifier information.

[0229] Alignment and Annotation In some embodiments, a method for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) includes aligning genes within the RNA expression data with known sequences of the human genome and annotating to obtain annotated RNA expression data.

[0230] In some embodiments, alignment of RNA expression data involves aligning the data to a known assembled genome (e.g., the human genome) for a particular species of subject, or to a transcriptome database. A variety of sequence alignment software is available and can be used to align the data to an assembled genome or transcriptome database. Non-limiting examples of alignment software include short (unspliced) aligners (e.g., BLAT, BFAST, Bowtie, Burrows-Wheeler Aligner, Short Oligonucleotide Analysis package, or Mosaik), spliced aligners, aligners based on known splice junctions (e.g., Errange, IsoformEx, or Splice Seq), or de novo splice aligners (e.g., ABMapper, BBMap, CRAC, or HiSAT). In some embodiments, any suitable tool can be used for alignment and annotation of the data. For example, Kallisto (github.com / pachterlab / kallisto) is used for alignment and annotation of the data. In some embodiments, a known genome is referred to as a reference genome. A reference genome (also called a reference assembly) is a digital nucleic acid sequence database assembled as a representative example of a set of genes for a species. In some embodiments, the human and mouse reference genomes used in any one of the methods described herein are maintained and improved by the Genome Reference Consortium (GRC). Non-limiting examples of human reference releases are GRCh38, GRCh37, NCBI Build 36.1, NCBI Build 35, and NCBI Build 34. Non-limiting examples of transcriptome databases include transcriptome shotgun assembly (TSA).

[0231] In some embodiments, annotating RNA expression data involves identifying the placement of genes and / or coding regions in the data to be processed by comparing to an assembled genomic or transcriptome database. Non-limiting examples of data sources for annotation include GENCODE (www.gencodegenes.org), RefSeq (e.g., see www.ncbi.nlm.nih.gov / refseq / ), and Ensembl. In some embodiments, annotating genes in RNA expression data is based on the GENCODE database (e.g., GENCODE V23 annotation, www.gencodegenes.org).

[0232] Consea et al., "A survey of best practices for RNA-seq data analysis", Genome Biology 2016 17:13, provides best practices for analyzing RNA-seq data applicable to any one of the methods described herein, which is hereby incorporated by reference in its entirety. Also, Pereira and Rueda, bioinformatics-core-shared-training.github.io / cruk-bioinf-sschool / Day2 / rnaSeq_align.pdf, describes methods for analyzing RNA sequencing data, which is applicable to any one of the methods described herein and is hereby incorporated by reference in its entirety.

[0233] Removal of non-coding transcripts In some embodiments, a method for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) includes removing non-coding transcripts from the annotated RNA expression data. Aligning and annotating the RNA expression data enables the identification of coding and non-coding reads. In some embodiments, non-coding reads for a transcript are removed to focus analysis efforts on the expression of proteins (e.g., those that may be involved in cancer pathology). In some embodiments, removing reads for non-coding transcripts from the data reduces the dispersion of the data in replicates of the same or similar samples (e.g., nucleic acids from the same cell or cell type).In some embodiments, non-limiting examples of expression data to be removed include pseudogenes, polymorphic pseudogenes, processed pseudogenes, transcribed processed pseudogenes, unitary pseudogenes, non-processed pseudogenes, transcribed unitary pseudogenes, constant chain immunoglobulin (IG C) pseudogenes, joining chain immunoglobulin (IG J) pseudogenes, variable chain immunoglobulin (IG V) genes, transcribed non-processed genes, translated non-processed genes, joining chain T cell receptor (TR J) genes, variable chain T cell receptor (TR V) genes, small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), microRNA (miRNA), ribozymes, ribosomal RNA (rRNA), mitochondrial tRNA (Mt tRNA), mitochondrial rRNA (Mt rRNA), Cajal body-specific RNA (scaRNA), retained introns, sense intron RNA, sense overlapping RNA, nonsense-mediated decay RNA, non-stop decay RNA, antisense RNA, long intergenic non-coding RNA (lincRNA), macro long non-coding RNA (macro lncRNA), processed transcripts, 3' overlapping non-coding RNA (3' overlapping ncrna), small RNA (sRNA), other RNA (miscRNA), vault RNA, and TEC RNA, including one or more non-coding transcripts (e.g., 10 - 50, 50 - 100, 100 - 1,000, 1,000 - 2,500, 2,500 - 5,000 or more non-coding transcripts) selected from the group consisting of.

[0234] In some embodiments, information (e.g., sequence information) about one or more of these types of transcripts can be obtained in a nucleic acid database (e.g., the Gencode database, e.g., Gencode V23, the Genbank database, the EMBL database, or other databases). In some embodiments, a portion (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 99.5% or more) of the non-coding transcripts, histone-coding genes, mitochondrial genes, interleukin-coding genes, collagen-coding genes, and / or T cell receptor-coding genes described herein is removed from the aligned and annotated RNA expression data.

[0235] Conversion to TPM and gene aggregation In some embodiments, a method for processing RNA expression data (e.g., data obtained from RNA sequencing (also referred to herein as RNA-seq data)) includes normalizing the RNA expression data with respect to the length of the transcripts read (e.g., in transcripts per kilobase million (TPM) format). In some embodiments, the RNA expression data normalized with respect to the length of the transcripts is first aligned and annotated. The conversion of the data to TPM allows the expression to be represented in the form of concentration rather than count, and thus enables the comparison of samples with different read count totals and / or read lengths.

[0236] In some embodiments, RNA expression data normalized with respect to the length of reads of the transcript is analyzed to obtain gene expression data (expression data regarding a gene). This is also referred to as gene aggregation. Gene aggregation includes combining expression data in reads for transcripts of all isoforms of a gene to obtain expression data for that gene. In some embodiments, gene aggregation to obtain gene expression data is performed after TPM normalization but before identifying genes that introduce bias. In some embodiments, gene aggregation is performed before converting the data to TPM.

[0237] Wagner et al., Theory Biosci. (2012) 131:281-285 provides an explanation of how TPM can be calculated, which is hereby incorporated by reference in its entirety. In some embodiments, to calculate TPM, the formula

[0238]

Equation

[0239] is used.

[0240] Removal of Bias Conversion of RNA expression data to obtain expression in TPM format requires dividing the read count of a given transcript by the length of the reads of the transcript, and thus bias can be introduced into the data for various reasons (as explained below). Accordingly, some embodiments of any one of the methods described herein include identifying at least one gene that introduces bias into the gene expression data. Some embodiments of any one of the methods described herein include identifying at least one gene that introduces bias into the gene expression data and removing expression data for at least one gene from the gene expression data to obtain bias-corrected gene expression data.

[0241] In some embodiments, removing data from a dataset may involve deleting the data from the dataset, marking the data so that it is not used in some or all subsequent processing of the dataset, and / or performing any other suitable processing so that the data is not used in some or all subsequent processing of the dataset. For example, removing certain expression data (e.g., expression data for at least one gene that introduces bias) from gene expression data may involve deleting the certain expression data from the gene expression data, marking the certain expression data, and / or performing any other suitable processing so that the certain expression data is not used in some or all subsequent processing of the gene expression data. As another example, removing non-coding transcripts from RNA expression data (as described above) may involve deleting the non-coding transcripts, marking the non-coding transcripts, and / or performing any other suitable subsequent processing so that the non-coding transcripts are not used in some or all subsequent processing of the RNA expression data. As yet another example, removing sequence data that has been determined to fail one or more quality control checks during the execution of the quality control techniques described herein may involve deleting the sequence data, marking the sequence data, and / or performing any other suitable processing so that the sequence data that failed the quality control check is not used in some or all subsequent processing.

[0242] In some embodiments, the bias in the expression data converted to TPM format is due to transcripts of an average length that is at least a threshold amount higher or lower than the average length of transcripts as read across the entire expression data set. For example, for genes where one or more transcripts of one or more isoforms have a length that is lower than a threshold (e.g., at least 1 standard deviation, 2 standard deviations, 3 standard deviations, 4 standard deviations, 5 standard deviations, 6 standard deviations, 7 standard deviations, 8 standard deviations, 9 standard deviations, 10 standard deviations, 11 standard deviations, 12 standard deviations, 13 standard deviations, 13 standard deviations, or 15 standard deviations or more) from the mean or median value of transcript lengths across the entire expression data set, the expression of the gene in TPM format appears to be artificially high. Conversely, for genes where one or more reads of one or more isoforms have a length that is higher than a threshold (e.g., at least 1 standard deviation, 2 standard deviations, 3 standard deviations, 4 standard deviations, 5 standard deviations, 6 standard deviations, 7 standard deviations, 8 standard deviations, 9 standard deviations, 10 standard deviations, 11 standard deviations, 12 standard deviations, 13 standard deviations, 13 standard deviations, or 15 standard deviations or more) from the mean or median value of read lengths across the entire expression data set, the expression of the gene in TPM format appears to be artificially low. In some embodiments, the threshold is set with respect to the standard deviation (e.g., at least 1 standard deviation, 2 standard deviations, 3 standard deviations, 4 standard deviations, 5 standard deviations, 6 standard deviations, 7 standard deviations, 8 standard deviations, 9 standard deviations, 10 standard deviations, 11 standard deviations, 12 standard deviations, 13 standard deviations, 13 standard deviations, or 15 standard deviations or more). In some embodiments, the threshold is set based on the length of the transcript and / or the length of the read, e.g., less than 5 bp, less than 10 bp, less than 15 bp, less than 20 bp, less than 25 bp, less than 50 bp, less than 75 bp, less than 100 bp, or less than 150 bp or greater.

[0243] In some embodiments, the bias is due to the length of the polyA tail on the transcript. In some embodiments, RNA transcripts having a polyA tail that is on average smaller or larger than the average length of the polyA tails of the RNA transcripts in the sample have a concentration that is higher or lower than the average concentration of all RNA transcripts in the sample. Thus, a gene can be associated with a polyA tail having a length that is at least a threshold amount smaller compared to the average length of the polyA tail of the gene from the sample from which the RNA expression data was obtained. In some embodiments, such expression data for such genes is also removed from the gene expression data, thereby obtaining bias-corrected gene expression data. Removing expression data associated with one or more genes from a data set to reduce bias can be considered a type of data filtering. In some embodiments, "filtering" may refer to removing expression data for genes that appear to be artificially high or low (e.g., due to the length of the transcript or the length of the polyA tail associated with the transcript), and / or removing expression data for non-coding RNAs from the data.

[0244] In some embodiments, identifying at least one gene that introduces bias into the gene expression data includes analyzing the lengths of the transcripts within the data set being analyzed. In some embodiments, removing the expression data of at least one gene that introduces bias from the gene expression data reduces the variability and improves the overall accuracy of subsequent gene expression-based analyses.

[0245] In some embodiments, identifying at least one gene that introduces bias into gene expression data involves using knowledge obtained from analyzing data outside of the expression data set in question, e.g., using a reference data set. The inventors have recognized that by removing (expression data) of genes having poly-A tail lengths outside the average range of poly-A tails in an RNA expression data set, biases and / or outliers within the gene expression data can be effectively removed. For example, the knowledge that a particular gene family introduces bias can be had a priori (from previously performed experiments or previously performed processing of data) with respect to processing RNA expression data and can be used to filter data for that family of genes.

[0246] In some embodiments, genes that introduce bias into an expression data set may belong to a family of genes having poly-A tails that are small or high on average compared to the average length of the poly-A tails of the genes from the sample (or another reference sample) from which the RNA expression data was obtained. In some embodiments, "small or high" may refer to a numerical value that is small or high with respect to a known average threshold of one or more genes.

[0247] In some embodiments, genes that introduce bias into an expression data set belong to a gene family selected from the group consisting of histone-coding genes, mitochondrial genes, interleukin-coding genes, collagen-coding genes, B-cell receptor-coding genes, and T-cell receptor-coding genes. In some embodiments, genes that introduce bias into an expression data set may be any other gene having a poly-A tail that is small or high on average compared to the average length of the poly-A tails of the genes from the sample (or another reference sample) from which the RNA expression data was obtained.

[0248] In some embodiments, a histone-coding gene, a mitochondrial gene, an interleukin-coding gene, a collagen-coding gene, a B cell receptor-coding gene, and / or a T cell receptor-coding gene is a gene in a human sample that has a polyA tail that is small or high on average compared to the average length of the polyA tails of genes from the sample from which RNA expression data was obtained. For example, a histone-coding gene contains a polyA tail that is small on average relative to the average length of the polyA tails of genes from the sample from which RNA expression data was obtained. In some embodiments, a histone-coding gene does not contain a polyA tail. In some embodiments, a polyA tail is minimally detected or not detected in a histone-coding gene.

[0249] In some embodiments, abbreviations or acronyms of one or more genes or proteins are used in the present application, and the genes (or genes encoding proteins) are referred to using their recognized scientific names. Additional information regarding genes and / or encoded proteins is described in one or more gene sequence databases, such as the NIH Gene Sequence Database (GenBank, www.ncbi.nlm.nih.gov), the EMBL Database (the European Molecular Biology Laboratory Nucleotide Sequence Database, www.ebi.ac.uk / embl / index.html), the EMBL European Bioinformatics Institute Database (EMBL-EBI European Nucleotide Archive, www.ebi.ac.uk / ena), the GENCODE Database (www.gencodegenes.org), or other suitable databases, the contents of which are incorporated herein by reference with respect to the different types of genes and gene names described herein. In some embodiments, abbreviations or acronyms of genes or proteins refer to human genes (or human genes encoding proteins).

[0250] In some embodiments, the histone-coding genes are HIST1H1A, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H1T, HIST1H2AA, HIST1H2AB, HIST1H2AC, HIST1H2AD, HIST1H2AE, HIST1H2AG, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H2AK, HIST1H2AL, HIST1H2AM, HIST1H2BA, HIST1H2BB, HIST1H2BC, HIST1H2BD, HIST1H2BE, HIST1H2BF, HIST1H2BG, HIST1H2BH, HIST1H2BI, HIST1H2BJ, HIST1H2BK, HIST1H2BL, HIST1H2BM, HIST1H2BN, HIST1H2BO, HIST1H3A, HIST1H3B, HIST1H3C, HIST1H3D, HIST1H3E, HIST1H3F, HIST1H3G, HIST1H3H, HIST1H3I, HIST1H3J, HIST1H4A, HIST1H4B, HIST1H4C, HIST1H4D, HIST1H4E, HIST1H4F, HIST1H4G, HIST1H4H, HIST1H4I, HIST1H4J, HIST1H4K, HIST1H4L, HIST2H2AA3, HIST2H2AA4, HIST2H2AB, HIST2H2AC, HIST2H2BE, HIST2H2BF, HIST2H3A, HIST2H3C, HIST2H3D, HIST2H3PS2, HIST2H4A, HIST2H4B, HIST3H2A, HIST3H2BB, HIST3H3, or HIST4H4. In some embodiments, the mitochondrial genes are MT-ATP6, MT-ATP8, MT-CO1, MT-CO2, MT-CO3, MT-CYB, MT-ND1, MT-ND2, MT-ND3, MT-ND4, MT-ND4L, MT-ND5, MT-ND6, MT-RNR1, MT-RNR2, MT-TA, MT-TC, MT-TD, MT-TE, MT-TF, MT-TG, MT-TH, MT-TI, MT-TK, MT-TL1, MT-TL2, MT-TM, MT-TN, MT-TP, MT-TQ, MT-TR, MT-TS1, MT-TS2, MT-TT, MT-TV, MT-TW, MT-TY, MTRNR2L1, MTRNR2L10, M It is TRNR2L11, MTRNR2L12, MTRNR2L13, MTRNR2L3, MTRNR2L4, MTRNR2L5, MTRNR2L6, MTRNR2L7, or MTRNR2L8.

[0251] In some embodiments, removing expression data for at least one gene that introduces bias into gene expression data includes removing expression data for one or more (2, 3, 4, 5, or all) genes in each of one or more gene families, including histone-coding genes, mitochondrial genes, interleukin-coding genes, collagen-coding genes, B cell receptor-coding genes, and T cell receptor-coding genes. One or more (e.g., at least 2, at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, a number between 2 and 1000, or any suitable number within these ranges of genes). In some embodiments, removing expression data for at least one gene that introduces bias into gene expression data includes removing expression data for any of one or more genes having a poly-A tail that is small or high when averaged compared to the average length of the poly-A tails of genes from the sample (or reference sample) from which the RNA expression data was obtained.

[0252] In some embodiments, after expression data for at least one gene that introduces a bias is removed from the gene expression data, the remaining gene expression data can be re-normalized ( "re-normalized") so as not to be biased by the expression data of the bias gene from which the normalized expression value has been removed (e.g., adjusted to any other suitable unit such as TPM or Reads Per Kilobase Million (RPKM) or Fragments Per Kilobase Million (FPKM)). In some embodiments, the expression data of the remaining genes can have expression data for at least one thousand genes, at least five thousand genes, at least ten thousand genes, genes between five hundred and five thousand genes, genes between one thousand and ten thousand genes, genes between five thousand and fifteen thousand genes, or any suitable number of genes within these ranges.

[0253] Quality control of nucleic acid data after sequencing As presented in the present disclosure, quality control is performed regularly in the sample preparation process. For example, the purity of the extracted nucleic acid or the size distribution of the DNA library is detected. If one or more quality control problems occur and cannot be corrected in the laboratory, the provider of the biological sample (e.g., a medical service provider) is notified before proceeding to the subsequent steps. After the quality-related problems are resolved, the sample preparation process is completed and bioinformatics analysis (e.g., post-sequencing processing) is performed.

[0254] The aspects of the methods and systems described herein provide quality control to be performed on gene expression data to improve subsequent expression analysis (e.g., to determine the diagnosis, prognosis, and / or treatment of a patient or subject) and the accuracy and reliability of the resulting recommendations.

[0255] In some embodiments, the bioinformatics quality control of array data can be performed as a stand-alone process (e.g., based on nucleic acid data received from a healthcare provider) or in association with a prior sample preparation process (e.g., when patient samples are provided by a healthcare provider as opposed to nucleic acid sequence data). As illustrated in FIG. 7, activities 301 through 310 exemplify a non-limiting sample preparation process as described in the present disclosure, while activities 311 through 315 exemplify a non-limiting quality control process as described in the present disclosure. In some embodiments, one or more of activities 301 through 310 can be performed independently (e.g., without one or more of activities 311 through 315). In some cases, one or more of activities 301 through 310 may be skipped or delayed. Activities 311 through 315 may be performed independently (e.g., without activities 301 through 310). In some cases, one or more of activities 311 through 315 may be skipped or delayed. In some cases, both a sample preparation (activities 301 through 310) and a quality control (activities 311 through 315) process may be performed. In some cases, one or more of the sample preparation processes and one or more of the quality control processes may be performed.

[0256] In some embodiments, the process pipeline 300 includes obtaining a first tumor sample from a subject having cancer, suspected of having cancer, or at risk of having cancer at activity 301, extracting RNA from the first sample of the first tumor at activity 302, concentrating the extracted RNA against the coding RNA to obtain concentrated RNA at activity 303, preparing a first library of cDNA fragments from the concentrated RNA for non-coding RNA sequencing at activity 304, obtaining RNA expression data for a subject having cancer, suspected of having cancer, or at risk of having cancer at activity 305, aligning and annotating genes in the RNA expression data with known sequences of the human genome to obtain annotated RNA expression data at activity 306, removing non-coding transcripts from the annotated RNA expression data at activity 307, converting the annotated RNA expression data to gene expression data in Transcripts Per Kilobase Million (TPM) format at activity 308, identifying at least one gene that introduces a bias into the gene expression data at activity 309, removing at least one gene from the gene expression data to obtain bias-corrected gene expression data at activity 310, obtaining sequence information and asserted information at activity 311, determining one or more features from the sequence information at activity 312, determining whether the one or more features match the asserted information at activity 313, performing at least one additional determination of a feature at activity 314, and identifying cancer treatment for the subject using the bias-corrected gene expression data at activity 315.

[0257] In some embodiments, activity 305 may include obtaining RNA expression data by using a sequencing platform or by receiving it from a healthcare provider or a laboratory. In some embodiments, activity 306 may include converting the RNA expression data into gene expression data. As described herein, the "known sequence of the human genome" may refer to a reference. In some embodiments, activity 307 may include converting the RNA expression data into gene expression data. In some embodiments, activity 307 may include obtaining filtered RNA expression data. In some embodiments, activity 308 may include normalizing the filtered RNA expression data to obtain gene expression data in units of Transcripts Per Kilobase Million (TPM). In some embodiments, the asserted information of activity 311 may indicate the asserted source and / or the asserted completeness of the asserted sequence data. In some embodiments, activity 312 may include determining one or more disease characteristics. In some embodiments, activity 312 may include processing the sequence information or data to obtain determined information indicating the determined source and / or the determined completeness of the sequence information or data. In some embodiments, activity 313 may include determining whether the determined information matches the asserted information. In some embodiments, at least one additional determination of a feature in process 314 may include determining a disease characteristic or a feature not directly related to the disease.

[0258] Aspects of the methods and systems described herein provide an approach for verifying the validity of nucleic acid sequence data by obtaining both the claimed information related to the sequence data and one or more features of the sequence data (e.g., source, nucleic acid type, expected completeness, etc.), determining one or more features from the sequence data, and verifying that the one or more features determined from the sequence data match the claimed information regarding those features. In some embodiments, the claimed information can be information regarding a patient, tissue type, tumor type, nucleic acid type (RNA, DNA, WES, polyA, etc.), sequencing protocol used, etc., or combinations thereof. In some embodiments, the claimed information can be expected and / or acceptable levels of completeness of the sequence information, including, for example, GC content, contamination, coverage (e.g., genomic, exome, exon, protein coding, or other coverage), or other measures of completeness, that are expected and / or acceptable (e.g., acceptable for subsequent analysis of the sequence data).

[0259] Nucleic acid sequencing, particularly next-generation sequencing (NGS), enables the generation of a large amount of information for a given nucleic acid (DNA, RNA, genome, exome, transcriptome, etc.). However, there are a number of different sequencing platforms available, and the sample preparation and sequencing protocols and techniques used are diverse, with variations and inconsistencies between platforms and protocols, resulting in substantial variation in the content and coverage of the resulting nucleic acid sequencing information. Furthermore, when evaluating a large set of sequence information from different sources, such as from several sequencing runs, or from multiple sequencing runs (including historical data from different medical visits for one or more patients), or from different studies (such as studies for creating a prognosis or diagnostic assessment, or studies for evaluating the effect of a drug or treatment on the progression of a disease, etc.), it can be difficult to combine sequence information from different sources. In addition, it can also be difficult to detect incorrectly recognized sequence data when large amounts of information from different sources are combined.

[0260] Currently, for example, for diagnostic, prognostic, and / or clinical applications, there is no robust method for verifying the validity of the source and / or completeness (which may also be referred to as quality herein) of sequence information that can be the subject of further use (such as used for analysis beyond the initial sequencing step), for example, to enhance reliability, reduce uncertainty, correct or omit low-quality sequence information, or provide signals for verifying or reexamining suspicious sequence information or outliers.

[0261] In the present disclosure, it is recognized that next-generation sequencing technologies and platforms are proliferating in various fields of the scientific community. The present disclosure also recognizes the various protocols and methodologies associated with the different technologies and platforms being employed. Variations in platforms and in the protocols for using the various platforms result in variations in the data and sequence information realized from their use, which presents a significant hurdle when using the sequence information for substantial analysis, particularly when such sequence information is used for analysis beyond the initial data generated by the first user of the sample (e.g., a secondary user beyond the user who procured and first performed the sequencing, a third party with respect to the sequencing, etc.).

[0262] Accordingly, the present disclosure presents various methods and processes for evaluating the quality of sequence information (e.g., for proper identification of the sequence information, identification of the sample, identification of the subject, etc.) and further for evaluating the completeness of the sequence information (e.g., creating checkpoints for screening for various completeness issues, such as contamination or degradation). For example, in some embodiments, described herein is a method for evaluating sequence information by obtaining sequence information from nucleic acids of a subject's sample, obtaining the claimed information, determining the characteristics of the sequence information (e.g., source, identity, status, properties), and comparing the claimed information to the determined information. The sequence information can be obtained (e.g., acquired) from any source or through any means known in the art. Accordingly, the sequence information can be generated using any suitable sequencing technology. Alternatively, the sequence information may be obtained electronically from a third party that generated the sequence information. In some embodiments, the sequence information (e.g., reference sequence information) is obtained from an existing database of sequences. In some embodiments, the sequence information is obtained from a company, non-profit organization, academic institution, or medical institution.

[0263] In some embodiments, the sample may be any specimen, biopsy, or biological component obtained (e.g., procured, collected, received) from a subject. For example, in some embodiments, the sample may be a blood sample, hair sample, tissue sample, body fluid sample, cell sample, blood component sample, or any other cell or tissue sample from which nucleic acid can be obtained for sequencing.

[0264] In some embodiments, the subject may be any living being in need of treatment or diagnosis using the methods or systems of the present disclosure. For example, without limitation, the subject may include mammals and non-mammals. As used herein, "mammal" refers to any animal belonging to the class Mammalia (e.g., human, mouse, rat, cat, dog, sheep, rabbit, horse, cow, goat, pig, guinea pig, hamster, chicken, turkey, or non-human primate (e.g., marmoset, macaque)). In some embodiments, the mammal is a human. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human.

[0265] In some embodiments, the sample may be a biological sample obtained from a subject, e.g., a patient. In some embodiments, the sample may be blood, serum, sputum, urine, or a tissue biopsy (e.g., from any tissue including, without limitation, heart, liver, pancreas, CNS, gastrointestinal tract, mouth, large intestine, kidney, and skin). In some embodiments, the sample may be suspected of being a disease sample (e.g., a cancer sample). In some embodiments, the sample may be a healthy sample (e.g., used as a reference).

[0266] In some embodiments, the sequence information is obtained from next-generation sequencing platforms (e.g., Illumina™, Roche™, Ion Torrent™, etc.), or any high-throughput or massively parallel sequencing platform. In some embodiments, these methods may be automated, and in some embodiments, there may be manual intervention. In some embodiments, the sequence information may be the result of non-next-generation sequencing (e.g., Sanger sequencing). In some embodiments, sample preparation may be according to the manufacturer's protocol. In some embodiments, sample preparation may be a custom-made protocol, or other protocol for research, diagnostic, prognostic, and / or clinical purposes. In some embodiments, the protocol may be experimental. In some embodiments, the origin or method of preparation of the sequence information may be unknown.

[0267] In some embodiments, the size of the obtained RNA and / or DNA sequence data comprises at least 5 kilobases (kb). In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 10 kb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 100 kb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 500 kb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 1 megabase (Mb). In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 10 Mb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 100 Mb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 500 Mb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 1 gigabase (Gb). In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 10 Gb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 100 Gb. In some embodiments, the size of the obtained RNA and / or DNA sequence data is at least 500 Gb.

[0268] In some embodiments, the sequence information can be generated using nucleic acids from a sample of a subject. In some embodiments, the sequence information may be sequence data indicating the nucleotide sequences of DNA and / or RNA from a previously obtained biological sample of a subject having a disease, suspected of having a disease, or at risk of having a disease. In some embodiments, the nucleic acid is deoxyribonucleic acid (DNA). In some embodiments, the nucleic acid is prepared such that the entire genome is present in the nucleic acid. In some embodiments, the nucleic acid is processed such that only the genomic protein-coding regions remain (e.g., exome). When the nucleic acid is prepared such that only the exome is sequenced, this is referred to as whole exome sequencing (WES). Various methods known in the art for separating the exome for sequencing, e.g., in solution-based separation, a labeled probe is used to hybridize to the target region (e.g., exome), which can then be further separated from other regions (e.g., unbound oligonucleotides). These labeled fragments can then be prepared and sequenced.

[0269] In some embodiments, the nucleic acid is ribonucleic acid (RNA). In some embodiments, the sequenced RNA includes both coding and non-coding transcribed RNAs found in the sample. When such RNA is used for sequencing, the sequencing is said to be generated from "total RNA" and is sometimes also referred to as whole transcriptome sequencing. Alternatively, the nucleic acid can be prepared such that the coding RNA (e.g., mRNA) is isolated and used for sequencing. This can be done through any means known in the art, e.g., by isolating or screening for RNA against polyadenylated sequences. This is sometimes referred to as mRNA-Seq.

[0270] Array information can include sequence data generated by a nucleic acid sequencing protocol (e.g., a series of nucleotides in a nucleic acid molecule identified by next-generation sequencing, Sanger sequencing, etc.), as well as information contained therein (e.g., information indicating source, tissue type, etc.), which may also be information inferred or determined from the sequence data and considered. For example, in some embodiments, RNA sequence information can be analyzed to determine whether the nucleic acid is primarily polyadenylated.

[0271] The claimed information may refer to the sequence data and, by extension, information about the nucleic acid, sample, and / or subject from which the sequence data was obtained. In some embodiments, the claimed information is provided with the sequence data and can be verified by analyzing the sequence data as described herein. The claimed information may relate to characteristics of the nucleic acid, sample, or subject and can be used to evaluate the quality of the nucleic acid (e.g., source or integrity of the nucleic acid). The claimed information can refer to the sequence data or the claimed source and / or claimed integrity of the information.

[0272] In some embodiments, a third party may provide array data, and even related asserted information. In some embodiments, the asserted information is obtained from the same entity from which the sequencing data was obtained. In some embodiments, the asserted information and the sequencing data are obtained from different parties. In some embodiments, the asserted information is obtained from a database. In some embodiments, the asserted information is a reference value or property. In some embodiments, the asserted information may assert the identity of the array information, the identity of the nucleic acid of the array information, the identity of the sample from which the array information was generated, or the identity of the subject from which the sample was obtained. In some embodiments, the asserted information may identify the array data as being obtained from polyadenylated RNA, as being derived from whole transcriptome sequencing, or as being from WES. In some embodiments, the asserted information may identify the cell or tissue type for the sample from which the nucleic acid was obtained. In some embodiments, the asserted information may assert the tumor type for the sample from which the nucleic acid was obtained. In some embodiments, the asserted information may identify the MHC profile (e.g., the sequence of the alleles of the MHC of the subject from which the nucleic acid was obtained) for the subject from which the sample was obtained. In some embodiments, the asserted information may identify the expected protein subunit ratio for the sample. In some embodiments, the asserted information may provide an expected complexity value for the array information. In some embodiments, the asserted information may provide an expected contamination value for the array information. In some embodiments, the asserted information may provide an expected coverage value for the array information. In some embodiments, the asserted information may provide an expected exon coverage value for the array information. In some embodiments, the asserted information may provide an expected read composition value for the array information. In some embodiments, the asserted information may provide an expected Phred score for the array information.In some embodiments, the claimed information may provide an expected single nucleotide polymorphism (SNP) value for the sequence information. In some embodiments, the claimed information may relate to a GC content value for the sequence information. In some embodiments, the claimed information may include additional information. In some embodiments, the claimed information may include information related to a number of or more than one feature of the sequence information. In some embodiments, the claimed information is any combination of the foregoing features (e.g., determined values, properties, characteristics, etc.).

[0273] As used herein, a "feature" may be a property or characteristic determined from an analysis of sequence information that provides the user with information regarding the sequence information, the sample from which the information was derived, and / or the subject from which the sample was taken, beyond the nucleotide sequence of the sequence information. The sequence information may be related to gene expression data obtained from a healthcare provider or laboratory. For example, a feature may indicate a source (e.g., patient, subject, nucleic acid type), the identity of the patient or subject, tissue type, tumor type, polyadenylation status, MHC sequence, protein subunit, complexity, contamination, coverage (e.g., whole sequence, exon, etc.), read composition, quality and / or Phred score, the location of a single nucleotide polymorphism (SNP), and / or GC content. The features of the sequence information can indicate whether the sequence information potentially matches or does not match the claimed information from a healthcare provider or laboratory.

[0274] When considering identity or source, this term not only refers to identifying a particular subject or patient as a particular individual, but it is also important to recognize that one or more features of the sequence information for a sample can be identified as being the same as one or more features of the sequence information obtained from another sample. For example, sequence information A can be presented as being from the same nucleic acid, subject or patient, tissue, or tumor and compared to the claimed sequence information B. Identity can be supported or suspected by the methods herein without knowing the actual identity of the subject, but the finding that the identity is consistent with another given sequence information can be supported. In some embodiments, the identity of the sequence information is used to compare to the claimed information for a given sample, subject, tissue, or tumor. In some embodiments, the identity of the sequence information is used to compare to the claimed information to another nucleic acid or reference value.

[0275] Next, in some embodiments, these determined features of the sequence information are evaluated against the claimed information (e.g., determined, collated, aligned, measured, evaluated). This evaluation can be done to increase confidence that the sequence information is of a particular origin (e.g., source), is correctly identified, or has specific or distinctive features (e.g., is from a polyadenylated nucleic acid). In this regard, these methods can be used to provide checkpoints and means for highlighting potential problems (e.g., inconsistent values (e.g., for determined features and claimed information), or determined values outside an accepted or established range). Such problems can indicate or suggest that there are issues with the integrity of the sequence data (e.g., degraded, contaminated) or the source (e.g., misidentified, mislabeled, etc.). By using the methods and processes herein to compare the determined features against the claimed information for a given sequence information, the likelihood that incorrect or low-quality sequence data is used in the analysis is reduced, and confidence is increased that the sequence data has sufficient quality to be used in diagnosis, prognosis, and / or clinical analysis.

[0276] In some embodiments, evaluating whether the determined information matches the asserted information involves determining whether the determined information exactly matches the asserted information or is within a specified threshold range. More generally, in some embodiments, evaluating whether two values “match” can involve determining whether the two values exactly match or are within a specified threshold range. The threshold can be 0, which in some embodiments requires exact match. The threshold can be greater than 0 such that when numerical values are being compared, the numerical values are said to “match” when they are within the threshold of each other (e.g., when the absolute difference between the numerical values is less than or equal to the threshold). In some embodiments, the threshold can be set as a function of a standard deviation (or multiple thereof), quantile, percentile, or any other suitable statistic. In some embodiments, evaluating whether two values “match” can involve determining whether, when there is a difference between the two values, the difference is statistically significant. Such determination may be performed using statistical hypothesis testing, thresholds, or any other suitable statistical or mathematical technique, as the aspects of the techniques described herein are not limited in this regard.

[0277] In some embodiments, one or more quality control parameters are checked for bioinformatics data. In some embodiments, tumor purity can be checked. Tumor purity may refer to the percentage of cancer cells in an admixture, as described herein. In some embodiments, the target tumor purity for WES is ≧20% (e.g., 20%, 40%, 60%). In some embodiments, the target tumor purity for RNA-seq is ≧20% (e.g., 20%, 40%, 60%).

[0278] In some embodiments, the depth of coverage can be checked. In some embodiments, the depth of coverage for WES is an average coverage of ≧150-fold (e.g., 150-fold, 180-fold, 200-fold) of the tumor sample. In some embodiments, the target depth of coverage for RNA-seq is ≧100-fold (e.g., 100-fold, 150-fold, 200-fold).

[0279] In some embodiments, the alignment rate can be checked. In some embodiments, the target alignment rate for WES is greater than 90% (e.g., 91%, 95%, 99%). In some embodiments, the target alignment rate for RNA-seq is greater than 90% (e.g., 91%, 95%, 99%).

[0280] In some embodiments, base call quality scores such as Phred scores can be checked. In some embodiments, the target Phred score for WES is greater than 30 (e.g., 35, 40, 50). In some embodiments, the target Phred score for RNA-seq is greater than 30 (e.g., 35, 40, 50).

[0281] In some embodiments, the uniformity of coverage can be checked. In some embodiments, the target uniformity of coverage for WES is 85% of the base pairs in the target region covered at ≧20-fold for tumor tissue (e.g., 85%, 95%, 99%). In some embodiments, the target uniformity of coverage for WES is 85% of the base pairs in the target region covered at ≧20-fold for normal tissue (e.g., 85%, 95%, 99%). In some embodiments, the target region for determining the uniformity of coverage may be the ExonV7 target region using the coding region from the CCDS (Consensus Coded Sequence) gene.

[0282] In some embodiments, the GC bias can be checked. In some embodiments, the target GC bias for WES is at least 50 (e.g., 50, 60, 70). In some embodiments, the acceptable range of the GC bias for WES is at least 45 - 65 (e.g., 45 - 65, 50 - 65, 55 - 65). In some embodiments, the target GC bias for RNA-seq is at least 50 (e.g., 50, 60, 70). In some embodiments, the acceptable range of the GC bias for RAN-seq is at least 45 - 65 (e.g., 45 - 65, 50 - 65, 55 - 65).

[0283] In some embodiments, the mapping quality can be checked. In some embodiments, the mapping quality for WES is ≧10 (e.g., 10, 20, 30).

[0284] In some embodiments, the duplication rate can be checked. In some embodiments, the duplication rate for WES is less than 30% (e.g., 29.9%, 25%, 15%). In some embodiments, the duplication rate for RAN-seq is less than 85% (e.g., 84.99%, 80%, 70%).

[0285] In some embodiments, the insertion size can be checked. In some embodiments, the acceptable median insertion size of tumor tissue for WES is about 150 (e.g., 150, 280, 250). In some embodiments, the target median insertion size of tumor tissue for WES is about 200 (e.g., 200, 250, 350). In some embodiments, the acceptable median insertion size of normal tissue for WES is about 150 (e.g., 150, 280, 250). In some embodiments, the target median insertion size of normal tissue for WES is about 200 (e.g., 200, 250, 350). In some embodiments, the acceptable median insertion size of tumor tissue for RNA-seq is about 150 (e.g., 150, 280, 250). In some embodiments, the target median insertion size of tumor tissue for RNA-seq is about 200 (e.g., 200, 250, 350).

[0286] In some embodiments, contamination can be checked. In some embodiments, the acceptable contamination for WES is less than 0.05% (e.g., 0.04%, 0.03%, 0.01%). In some embodiments, the acceptable contamination for RNA-seq is less than 0.05% (e.g., 0.04%, 0.03%, 0.01%).

[0287] In some embodiments, the SNP concordance of the tumor-to-normal sample pair from the same patient can be checked. In some embodiments, the target SNP concordance for WES is greater than 90% (e.g., 91%, 95%, 98%). In some embodiments, the acceptable SNP concordance for WES is greater than 85% (e.g., 86%, 90%, 98%). In some embodiments, the target SNP concordance for RNA-seq is greater than 90% (e.g., 91%, 95%, 98%). In some embodiments, the acceptable SNP concordance for RNA-seq is greater than 85% (e.g., 86%, 90%, 98%).

[0288] In some embodiments, HLA allele matching of tumor-to-normal sample pairs from the same patient can be checked. In some embodiments, the threshold of normal tissue to tumor tissue for WES is less than 5 (e.g., 4.5, 3, 2.5). In some embodiments, the threshold of tumor RNA-seq to normal WES tissue for RNA-seq is less than 5 (e.g., 4.5, 3, 2.5).

[0289] In some embodiments, the sequence information can be evaluated for genomic contamination (e.g., non-human genomic contamination). In some embodiments, the sample or sequence information is evaluated to determine whether it is contaminated by determining whether it contains sequences from other species or reference genomes such as mouse, zebrafish, Drosophila, Caenorhabditis elegans (celegans), Saccharomyces, Arabidopsis thaliana, microbiome, Mycoplasma, adapter, UniVec, and phiX rRNA. In some embodiments, the target threshold for ADA genomic contamination for WES is greater than 60 (e.g., 65, 70, 80). In some embodiments, the acceptable threshold for ADA genomic contamination for WES is greater than 40 (e.g., 45, 60, 80). In some embodiments, the target threshold for ADA genomic contamination for RNA-seq is greater than 40 (e.g., 50, 60, 80). In some embodiments, the acceptable threshold for ADA genomic contamination for RNA-seq is greater than 20 (e.g., 30, 50, 70).

[0290] In some embodiments, only one feature is evaluated against the claimed information. In some embodiments, multiple features are evaluated against the claimed information. In some embodiments, at least two or more features are evaluated against the claimed information. In some embodiments, at least three or more features are evaluated against the claimed information. In some embodiments, at least four or more features are evaluated against the claimed information. In some embodiments, at least five or more features are evaluated against the claimed information. In some embodiments, at least six or more features are evaluated against the claimed information. In some embodiments, at least seven or more features are evaluated against the claimed information. In some embodiments, at least eight or more features are evaluated against the claimed information. In some embodiments, at least nine or more features are evaluated against the claimed information. In some embodiments, at least ten or more features are evaluated against the claimed information. In some embodiments, at least eleven or more features are evaluated against the claimed information. In some embodiments, at least twelve or more features are evaluated against the claimed information. In some embodiments, at least thirteen or more features are evaluated against the claimed information. In some embodiments, at least fourteen or more features are evaluated against the claimed information. In some embodiments, at least fifteen or more features are evaluated against the claimed information.

[0291] In some embodiments, if a feature or determined value is found not to conform to or match the claimed information, additional steps are performed. In some embodiments, if a feature or determined value is found not to conform to or match the claimed information, the array information is rejected (e.g., not used in subsequent analysis). In some embodiments, if a feature or determined value is found not to conform to or match the claimed information, the array information is reexamined, i.e., any evaluation of the feature or determination is performed at least one more time, or a second or subsequent time (e.g., a third, fourth, fifth, sixth, etc. time). In some embodiments, if a feature or determined value is found not to conform to or match the claimed information, another, or a second, or subsequent (e.g., a third, fourth, fifth, sixth, etc.) array information is obtained and then examined, i.e., any evaluation of the feature or determination is performed at least one time, or a second or subsequent time (e.g., a third, fourth, fifth, sixth, etc. time) regardless of the initial determination and evaluation made for the first array information. In some embodiments, if a feature or determined value is found not to conform to or match the claimed information, the array information is reported to the user as such. In some embodiments, any combination of these steps may be performed if a feature or determined value is found not to conform to or match the claimed information.

[0292] In some embodiments, if a feature or determined value does not conform to or match the claimed information, the array information can still be evaluated for properties related to a disease (e.g., cancer), but information regarding quality (e.g., the degree and nature of one or more features of the determined array information that do not match the claimed information) can be provided to the user (e.g., a physician or other healthcare provider). In some embodiments, the properties relate to the type of cancer, its environment, its stage, its location, its primary tissue, its statistical likelihood of responding to various treatments or therapies, or other properties that can assist the practitioner in treating the subject.

[0293] In some embodiments, if it is determined that a feature or determined value conforms to or matches the claimed information (e.g., matches, exceeds, or otherwise meets a condition with a reference value or threshold), additional steps can be performed. In some embodiments, if it is determined that a feature or determined value conforms to or matches the claimed information (e.g., matches, exceeds, or otherwise meets a condition with a reference or threshold), additional steps may be performed. In some embodiments, if it is determined that a feature or determined value conforms to or matches the claimed information (e.g., matches, exceeds, or otherwise meets a condition with a reference or threshold), the array information is evaluated for properties related to cancer. In some embodiments, the properties relate to the type of cancer, its environment, its stage, its location, its primary tissue, its statistical likelihood of responding to various treatments or therapies, or other properties that can assist the practitioner in treating the subject.

[0294] In some embodiments, after one or more quality control steps are performed, a report is generated for the user, along with the results of the quality control steps that were performed.

[0295] Accordingly, in one aspect, the present disclosure relates to a method of evaluating sequence information of at least one nucleic acid to determine at least one characteristic thereof. The at least one characteristic can be used to evaluate the quality or completeness of the sequence information, interrogate the source of the sequence information, or enable analysis of other sequence information, whether from the same sequencing platform or from the same or different sample preparation protocols. Further, the at least one characteristic may be used as a quality control measure to ensure that subsequent analysis of threshold quality and low-quality sequence information is omitted.

[0296] Accordingly, in one aspect, the present disclosure relates to a method of evaluating sequence information by: (a) obtaining sequence information comprising sequence data from (1) a first ribonucleic acid (RNA) or (2) a first whole exome sequence (WES); and (b) determining one or more characteristics of the sequence data, the sequence data being selected from the group consisting of: (i) the identity of the subject from whom the nucleic acid was obtained, (ii) the primary tissue from which the nucleic acid was obtained, (iii) the tumor type from which the nucleic acid was obtained, (iv) a quality metric of the first RNA sequence data, (v) whether the RNA sequence data was obtained from polyadenylated (polyA) RNA or total RNA, (vi) whether the first sequence data set is the first WES sequence data, (vii) the sequencing platform used to generate the first sequence data set, and (viii) a quality metric of the first sequence data set.

[0297] In some embodiments, the method further comprises obtaining additional sequence information if one or more characteristics of the sequence information fall below a quality control threshold suitable for further analysis.

[0298] In some embodiments, the feature to be evaluated is the identity of the subject. In some embodiments, the identity of the subject is determined by performing one or more of the evaluations from the group including major histocompatibility complex evaluation and SNP match evaluation, and the result of the evaluation is compared with the claimed value for the subject or a second sequence dataset from the subject.

[0299] In some embodiments, the feature to be evaluated is the primary tissue. In some embodiments, the primary tissue is determined by performing one or more of the evaluations from the group including protein expression and biomarker analysis. In another aspect, the present disclosure relates to a method of evaluating a feature including assigning a primary tissue of a sample that generated sequence information. In some embodiments, the method includes evaluating the sequence information for a marker or gene expression indicative of the tissue type that gave rise to the sequence information. In some embodiments, the method includes evaluating the marker or gene expression by matching it against a database that is the same for different tissue types. Different tissues throughout a subject's body express different proteins that profile such tissues. Thus, by evaluating the protein expression profile and matching it to a tissue type, it is possible to identify the sample and, thus, the tissue from which the sequence information was obtained. This can be done through various methods known in the art. For example, evaluating the number of a given messenger RNA (mRNA) transcript (e.g., used as an alternative to evaluating protein expression) can be evaluated by matching it against a database of known tissue markers (e.g., protein expression profiles), against a provided set of markers for the subject, or against a second set of sequence information or tissue markers obtained from the subject. In some embodiments, the primary tissue is determined by evaluating sequence information for a marker (e.g., protein expression) and matching that marker to a database of tissues. In some embodiments, the primary tissue is determined by evaluating sequence information for a marker (e.g., protein expression) and matching that marker to a set of markers from the subject's tissue. In some embodiments, the primary tissue is determined by evaluating sequence information for a marker (e.g., protein expression) and matching the marker to a second set of sequence information obtained from a subject in which the primary tissue is known.

[0300] In some embodiments, the feature being evaluated is a measure of the completeness of the sequence information. In some embodiments, the completeness measure of the first RNA sequence data is determined by determining the coverage of one or more genes in the RNA sequence data, determining the relative coverage of two or more exons for at least one gene in the RNA sequence data, determining the expression ratio of two known reference genes from the RNA sequence data, or performing one or more of the evaluations from the group including other features or combinations of two or more of them. In some embodiments, the completeness measure of the DNA sequence data is determined by performing one or more of the evaluations from the group including the total coverage and / or chromosomal coverage of the DNA sequence data, or other features or combinations of two or more of these.

[0301] In some embodiments, the RNA sequence data is analyzed to determine whether it was obtained from polyA RNA or total RNA. In some embodiments, the RNA sequence data is analyzed by evaluating the expression levels of one or more mitochondrial genes or histone genes from the RNA sequence data, and / or other features typical of polyA or total RNA.

[0302] In some embodiments, the feature being evaluated is the sequencing platform used to generate the sequence. In some embodiments, the sequencing platform used to generate the WES sequence data is determined by determining the % variance of one or more reference genes in the WES sequence data, or performing one or more of the evaluations from the group including other properties of the sequencing data typical of the sequencing platform used to generate the sequence data.

[0303] In some embodiments, the method includes evaluating at least one of the features described herein. In some embodiments, the method includes evaluating at least two of the features described herein. In some embodiments, the method includes evaluating at least three of the features described herein. In some embodiments, the method includes evaluating at least four of the features described herein. In some embodiments, the method includes evaluating at least five of the features described herein. In some embodiments, the method includes evaluating at least six of the features described herein. In some embodiments, the method includes evaluating at least seven of the features described herein.

[0304] In some embodiments, the quality (e.g., source or integrity) of sequence information from one or more nucleic acid samples (e.g., at least two nucleic acid samples) is evaluated by (a) determining the sequences of two or more (e.g., 2, 3, 4, 5, 6 or more) major histocompatibility complex (MHC) and (b) determining whether the MHCs from one or more samples match. In some embodiments, if the MHCs do not match (e.g., if the calculated match value is less than a statistically significant threshold), the sequence information from each of the nucleic acids is considered likely to be from different sources with insufficient quality, removed, discarded, retested, and / or reported to the user as such. In some embodiments, if the calculated match value (x) between WES normal / tumor / RNAseq is 0 < x ≤ 2 (e.g., 1, 1.5, 2), this represents acceptable and "warning". A warning means that the calculated match value is within the range considered acceptable but is thought to be close to unacceptable. In some embodiments, if the calculated match value (x) between WES normal / tumor / RNAseq is > 5, this represents unacceptable or poor quality. In some embodiments, if the calculated match value (x) between WES normal / tumor / RNAseq is 0, this represents good quality. In some embodiments, if the MHCs match (e.g., if the match value is above a statistically significant threshold), the sequence information from each of the nucleic acid samples is considered likely to be from the same source with sufficient quality, retained for further analysis, and / or reported to the user as such.

[0305] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is evaluated by determining a concordance value for single nucleotide polymorphisms (SNPs) within the sequence information. In some embodiments, this method further includes evaluating the concordance value. In some embodiments, if the concordance value is less than 85%, less than 80%, or less than 75%, the sequence information from each of the nucleic acid samples is considered likely to be from different sources and of insufficient quality, removed, discarded, retested, and / or reported to the user as such. In some embodiments, if the concordance value is less than 75%, the sequence information is considered unacceptable. In some embodiments, if the concordance value is greater than 80% and less than 95%, the sequence information is considered to be in a range approaching unacceptable. In some embodiments, if the concordance value is greater than 95%, the sequence information is considered acceptable. In some embodiments, if the concordance value is at least 75%, at least 80%, or at least 85%, the sequence information from each of the nucleic acid samples is considered likely to be from the same source and of sufficient quality, retained, and / or reported to the user as such. In some embodiments, at least 5,000 SNPs can be evaluated for the concordance value. In some embodiments, at least 6,000 SNPs can be evaluated for the concordance value. In some embodiments, at least 7,000 SNPs can be evaluated for the concordance value. In some embodiments, at least 8,000 SNPs can be evaluated for the concordance value.

[0306] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is evaluated by determining a contamination value for the sequence information. In some embodiments, if the contamination value exceeds a statistically significant threshold, the sequence information is removed, discarded, retested, and / or reported to the user as such. In some embodiments, if the contamination value is greater than 0.05% (e.g., 0.06%, 1%, 2%), the sequence information is considered to be near unacceptable (e.g., a warning). In some embodiments, if the contamination value is greater than 0.1% (e.g., 0.1%, 0.5%, 1%), the sequence information is considered unacceptable for blood samples and fresh frozen tissue. In some embodiments, if the contamination value is below the threshold, the sequence information is retained and / or reported to the user as such.

[0307] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is analyzed by comparing the sequence information from the one or more nucleic acid samples to a set of tumor types, determining the tumor type predicted from the sequence information, and determining whether the predicted tumor type matches the tumor type provided (e.g., claimed) for the one or more nucleic acid samples. In some embodiments, determining the predicted tumor type as a quality control step can be performed using a computerized system or process as described herein. In some embodiments, determining the predicted tumor type as a quality control step can be performed by using machine learning techniques to determine a cancer grade from sequence data, as described in U.S. Provisional Patent Application No. 62 / 943,976, filed Dec. 5, 2019, entitled “Machine Learning Techniques for Gene Expression Analysis,” which is hereby incorporated by reference in its entirety. In some embodiments, if there is a discrepancy between the tumor type (e.g., cancer grade) obtained from the sequence evaluation and the claimed information, the sequence information is identified as suspect or of insufficient quality, removed, discarded, retested, and / or reported to the user as such. In some embodiments, if there is a match between the predicted tumor type and the expected tumor type for the one or more nucleic acid samples, the sequence information is considered to have sufficient quality, retained, and / or reported to the user as such.

[0308] In some embodiments, collating the predicted tumor type with the provided tumor type involves using a set of reference genes from a training dataset that includes multiple signature genes that are upregulated or downregulated in a particular tumor type, for normal healthy samples. For example, if the predicted tumor type is prostate cancer (e.g., as claimed information), the sample is checked against known reference genes for prostate cancer. In some embodiments, the predicted tumor type is evaluated against its tumor grade, which can help determine the signature genes for the cancer grade claimed at different stages of the cancer.

[0309] As described above, in some embodiments, determining the predicted tumor type as a quality control step can be performed by using a machine learning approach that employs a statistical model trained using training data to determine the cancer grade from sequence information.

[0310] For example, in some embodiments, a statistical model can be used to predict the characteristics of a biological sample using gene expression data, based on the input ranking of genes, ranked according to their respective expression levels, for a sequencing platform. Using the input ranking instead of a specific value for the expression level allows the same or a similar data processing pipeline to be used across different expression data, regardless of the specific method by which the expression levels were obtained (e.g., regardless of the sequencing platform, sequencing conditions, sample preparation, data processing to obtain the expression levels, etc.). In some embodiments, the statistical model can be used to predict the cancer grade of a biological sample. In some embodiments, the statistical model may be used to predict the primary tissue of a biological sample, which may also be used to perform quality control as described herein.

[0311] For example, in some embodiments, ranking of genes based on gene expression levels (in a biological sample) as determined by a sequencing platform can be provided as an input to a statistical model trained to predict primary tissue for a biological sample. The predicted primary tissue can be compared against the primary tissue asserted as part of the quality control techniques described herein. As another example, in some embodiments, ranking of genes based on gene expression levels (in a biological sample) as determined by a sequencing platform can be provided as an input to a statistical model trained to predict cancer grade for a biological sample. The predicted cancer grade can be compared against the cancer grade asserted as part of the quality control techniques described herein.

[0312] In some embodiments, the set of genes being ranked depends on the particular biological characteristic of interest. For example, one set of genes can be used to determine primary tissue and another set of genes can be used to determine cancer grade.

[0313] In some embodiments, the expression data can be obtained for cells in a biological sample, and the subject has, is suspected of having, or is at risk of having cancer. In the context where the primary tissue is the characteristic being determined, the primary tissue is with respect to the cells in the biological sample. The primary tissue can refer to a particular tissue type from which the cells originate, such as lung, pancreas, stomach, large intestine, liver, bladder, kidney, thyroid, lymph node, adrenal gland, skin, breast, ovary, prostate, etc.

[0314] For example, in some embodiments, for diffuse large B-cell lymphoma (DLBCL), such as germinal center B cells (GCB) and activated B cells (ABC), it involves using a gene set to predict the primary tissue that may contain the origin cells. The genes within the gene set can be selected from the group consisting of ITPKB, MYBL1, LMO2, BATF, IRF4, LRMP, CCND2, SLA, SP140, PIM1, CSTB, BCL2, TCF4, P2RX5, SPINK2, VCL, PTPN1, REL, FUT8, RPL21, PRKCB1, CSNK1E, GPR18, IGHM, ACP1, SPIB, HLA-DQA1, KRT8, FAM3C, and HLA-DMB.

[0315] In the context of a property for which the cancer grade has been determined, the cancer grade pertains to the cells in a biological sample. The cancer grade may well refer to the proliferation and differentiation characteristics of the cells in a biological sample, generally referring to a numerical grade determined by visual observation of the cells using a microscope, such as grade 1, grade 2, grade 3, and grade 4.

[0316] For example, in some embodiments, it involves using a gene set to predict breast cancer grade. The genes within the gene set can be selected from the group consisting of UBE2C, MYBL2, PRAME, LMNB1, CXCL9, KPNA2, TPX2, PLCH1, CCL18, CDK1, MELK, CCNB2, RRM2, CCNB1, NUSAP1, SLC7A5, TYMS, GZMK, SQLE, C1orf106, CDC25B, ATAD2, QPRT, CCNA2, NEK2, IDO1, NDC80, ZWINT, ABCA12, TOP2A, TDO2, S100A8, LAMP3, MMP1, GZMB, BIRC5, TRIP13, RACGAP1, ASPM, ESRP1, MAD2L1, CENPF, CDC20, MCM4, MKI67, PBK, CKS2, KIF2C, MRPL13, TTK, BUB1, TK1, FOXM1, CEP55, EZH2, ECT2, PRC1, CENPU, CCNE2, AURKA, HMGB3, APOBEC3B, LAGE3, CDKN3, DTL, ATP6V1C1, KIAA0101, CD2, KIF11, KIF20A, CDCA8, NCAPG, CENPN, MTFR1, MCM2, DSCC1, WDR19, SEMA3G, KCND3, SETBP1, KIF13B, NR4A2, NAV3, PDZRN3, MAGI2, CACNA1D, STC2, CHAD, PDGFD, ARMCX2, FRY, AGTR1, MARCH8, ANG, ABAT, THBD, RAI2, HSPA2, ERBB4, ECHDC2, FST, EPHX2, FOSB, STARD13, ID4, FAM129A, FCGBP, LAMA2, FGFR2, PTGER3, NME5, LRRC17, OSBPL1A, ADRA2A, LRP2, C1orf115, COL4A5, DIXDC1, KIAA1324, HPN, KLF4, SCUBE2, FMO5, SORBS2, CARD10, CITED2, MUC1, BCL2, RGS5, CYBRD1, OMD, IGFBP4, LAMB2, DUSP4, PDLIM5, IRS2, and CX3CR1.

[0317] As another example, in some embodiments, it involves using a gene set to predict the renal clear cell carcinoma grade. The genes within the gene set can be selected from the group consisting of PLTP, C1S, LY96, TSKU, TPST2, SERPINF1, SRPX2, SAA1, CTHRC1, GFPT2, CKAP4, SERPINA3, CFH, PLAU, BASP1, PTTG1, MOCOS, LEF1, SLPI, PRAME, STEAP3, LGALS2, CD44, FLNC, UBE2C, CTSK, SULF2, TMEM45A, FCGR1A, PLOD2, C19orf80, PDGFRL, IGF2BP3, SLC7A5, PRRX1, RARRES1, LHFPL2, KDELR3, TRIB3, IL20RB, FBLN1, KMO, C1R, CYP1B1, KIF2A, PLAUR, CKS2, CDCP1, SFRP4, HAMP, MMP9, SLC3A1, NAT8, FRMD3, NPR3, NAT8B, BBOX1, SLC5A1, GBA3, EMCN, SLC47A1, AQP1, PCK1, UGT2A3, BHMT, FMO1, ACAA2, SLC5A8, SLC16A9, TSPAN18, SLC17A3, STK32B, MAP7, MYLIP, SLC22A12, LRP2, CD34, PODXL, ZBTB42, TEK, FBP1, and BCL2.

[0318] Aspects of using statistical models to predict primary tissue, cancer grade, and / or other characteristics of a biological sample are described in U.S. Provisional Patent Application No. 62 / 943,976, filed December 5, 2019, entitled "Machine Learning Techniques for Gene Expression Analysis", which is hereby incorporated by reference in its entirety.

[0319] Returning to aspects of evaluating the quality of array information, in some embodiments, the quality of array information from one or more (e.g., at least two) nucleic acid samples is evaluated by determining the presence or absence of polyadenylated RNA genes in order to predict whether the array information was obtained from poly(A) RNA. In some embodiments, if there is a discrepancy between the predicted poly(A) status of one or more samples and the expected (e.g., claimed) poly(A) status, the array information for those samples is considered of suspect, insufficient quality, removed, discarded, re-examined, and / or reported to the user as such. In some embodiments, if there is a match between the predicted poly(A) status and the expected poly(A) status for one or more nucleic acid samples, the array information is considered to have sufficient quality, retained, and / or reported to the user as such.

[0320] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is evaluated by determining a complexity value of the sequence information. In some embodiments, determining the complexity value includes determining the number of duplicates. In some embodiments, the % duplication rate can be determined for a DNA or RNA library. In some embodiments, if a large proportion of the library is duplicated, either a low complexity or over-amplified library of DNA or cDNA fragments is indicated. In some cases, differences between libraries in complexity or amplification indicate that some biases in the data are introduced (e.g., different % GC content). In some embodiments, if the complexity value is less than 75% or less than 80%, the sequence information is considered of insufficient quality, suspect, removed, discarded, reexamined, and / or reported to the user as such. In some embodiments, if the complexity value is at least 80% or at least 85%, the sequence information is considered of sufficient quality for further analysis, retained, and / or reported to the user as such.

[0321] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is evaluated by predicting the tissue source for the nucleic acid. In some embodiments, if there is a match between the predicted tissue source for the nucleic acid and the claimed tissue source, the sequence information is considered of insufficient quality, suspect, removed, discarded, reexamined, and / or reported as such. In some embodiments, if there is a match between the predicted tissue source and the claimed tissue source, the sequence information is considered to have sufficient quality for further analysis, retained, and / or reported to the user as such.

[0322] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is evaluated by (a) determining the gene expression levels for two different subunits of a known protein and (b) determining the expression ratio for the two different subunits. In some embodiments, if the determined expression ratio does not match the expected expression ratio for the protein subunit, the sequence information is identified as being of insufficient quality, suspect, removed, discarded, retested, and / or reported to the user as such. In some embodiments, if the determined expression ratio matches the expected expression ratio for the protein subunit, the sequence information is considered to have sufficient quality for further analysis, retained, and / or reported to the user as such.

[0323] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is evaluated by determining the Phred score for the sequence information. In some embodiments, if the Phred score is less than 27, the sequence information is considered to be of insufficient quality, suspect, removed, discarded, retested, and / or reported to the user as such. In some embodiments, if the Phred score is less than 20, the sequence information is removed, discarded, retested, and / or reported to the user as such. In some embodiments, if the Phred score is greater than 20 and less than 27, the sequence information is considered to be close to being something that should be removed, discarded, retested, and / or reported to the user as such. In some embodiments, if the Phred score is at least 27, the sequence information is considered to have sufficient quality for further analysis, retained, and / or reported to the user as such.

[0324] In some embodiments, the quality of sequence information from one or more (e.g., at least two) nucleic acid samples is evaluated by determining the GC content for the sequence information. In some embodiments, when the GC content is at least 30% and 55% or less, the sequence information is considered sufficient information for further analysis, retained, and / or reported to the user as such. In some embodiments, when the GC content is within the range of 45 - 65%, the sequence information is considered sufficient information for further analysis, retained, and / or reported to the user as such (i.e., acceptable). In some embodiments, a GC content of at least 50% (e.g., 50%, 51%, 60%) is a target value, at least for human samples.

[0325] In some embodiments, at least two (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or more) different methods for evaluating the quality (e.g., source and / or completeness) of sequence information are performed.

[0326] In some embodiments, the methods performed herein evaluate sequence information from mammals. In some embodiments, the mammal is a human.

[0327] In some embodiments, the sample from which the sequence information was generated is a subject having a disease, suspected of having a disease, or at risk of having a disease. In some embodiments, the disease is cancer.

[0328] In some embodiments, a report is generated that includes one or more features or results of the methods described herein. In some embodiments, the report further includes an analysis of the results of the methods described herein.

[0329] In some embodiments, the methods or processes of the present disclosure may be performed on a system or computer processor (e.g., a laptop, desktop, server, or other computerized machine). The components of the system may be located in remote locations and communicate over a network such as a local area network or wide area network, or by Internet protocol. The system may interact with a user through an interface via a web-enabled browser and a graphical user interface (GUI). In some embodiments, the system is under the control of a user in one location. In some embodiments, the system is composed of components that are not in one location and may not be under the direct control of a user. In some embodiments, the information of the system is stored locally.

[0330] As described herein, the terms "process", "activity", "step" or variations thereof used in a computerized process or a flowchart therein may be used interchangeably unless otherwise indicated.

[0331] As described herein, the terms "patient", "subject", "human subject" or variations thereof may be used interchangeably unless otherwise indicated.

[0332] FIG. 6A is a flowchart showing an exemplary computerized process 200 for performing non-stranded RNA sequencing with code RNA enrichment. Process 200 begins with activity 201, in which a first sample of a first tumor is obtained from a subject having cancer, suspected of having cancer, or at risk of having cancer. Further aspects related to obtaining a first sample of a first tumor from a subject having cancer, suspected of having cancer, or at risk of having cancer are presented in the "biological sample" section.

[0333] Next, process 200 proceeds to activity 202, and RNA is extracted from the first sample of the first tumor. Aspects related to extracting RNA from the first sample of the first tumor are described in the item "Extraction of DNA and / or RNA".

[0334] Next, process 200 proceeds to activity 203, and the extracted RNA is concentrated for coding RNA to obtain concentrated RNA. Aspects related to concentrating the extracted RNA for coding RNA to obtain concentrated RNA are described in the item "RNA Concentration".

[0335] Next, process 200 proceeds to activity 204, and a first library of cDNA fragments is prepared from the concentrated RNA for non-coding RNA sequencing. Aspects related to preparing a first library of DNA fragments from the concentrated RNA for non-coding RNA sequencing are described in the item "Library Preparation for RNA Sequencing".

[0336] Next, process 200 proceeds to activity 205, and non-coding RNA sequencing is performed on the first library of cDNA fragments prepared from the concentrated RNA. Aspects related to performing non-coding NDA sequencing on the first library of DNA fragments prepared from the concentrated RNA are described in the item "RNA Sequencing". It should be understood that one or more activities of process 200 may be optional.

[0337] FIG. 6B is a flowchart showing a computerized process 210 for identifying cancer treatment by obtaining bias-corrected gene expression data. Process 210 begins with activity 211, and RNA expression data is obtained for a subject having cancer, suspected of having cancer, or at risk of having cancer. Aspects related to obtaining RNA expression data are described in the item "Obtaining RNA Expression Data".

[0338] Next, process 210 proceeds to activity 212, where genes within the RNA expression data are aligned against a reference and the RNA expression data is annotated. Aspects related to aligning genes within the RNA expression data with known sequences of the human genome and annotating to obtain annotated RNA expression data are described in the item "Alignment and Annotation".

[0339] Next, process 210 proceeds to activity 213, where non-coding transcripts are removed from the annotated RNA expression data to obtain filtered RNA expression data. Aspects related to removing non-coding transcripts from the annotated RNA expression data are described in the item "Removal of Non-Coding Transcripts".

[0340] Next, process 210 proceeds to activity 214, where the filtered RNA expression data is normalized to obtain gene expression data. The gene expression data can also be in the form of Transcripts Per Kilobase Million (TPM). Aspects related to normalizing the filtered RNA expression data to gene expression data in the Transcripts Per Kilobase Million (TPM) format are described in the item "Conversion to TPM and Gene Aggregation".

[0341] Next, process 210 proceeds to activity 215, where at least one gene that introduces bias into the gene expression data is identified. Aspects related to identifying at least one gene that introduces bias into the gene expression data are described in the item "Removal of Bias".

[0342] Next, process 210 proceeds to activity 216, where expression data associated with at least one gene that introduces bias is removed from the gene expression data to obtain bias-corrected gene expression data. The manner of removing expression data associated with at least one gene that introduces bias into the gene expression data to obtain bias-corrected gene expression data is described in the item "Removal of Bias".

[0343] Next, process 210 proceeds to activity 217, where a cancer treatment for a subject using the bias-corrected gene expression data is identified. The manner related to identifying a cancer treatment for a subject using the bias-corrected gene expression data is described in the item "Identification of Cancer Treatment".

[0344] FIG. 6C is a flowchart showing a computerized process 220 for identifying a cancer treatment for a subject having cancer, suspected of having cancer, or at risk of having cancer using bias-corrected gene expression data. Process 220 begins at activity 221, where RNA is enriched for coding RNA in a sample of RNA extracted from a first tumor sample from a subject having cancer, suspected of having cancer, or at risk of having cancer. The manner related to enriching RNA for coding RNA in a sample of extracted RNA is described in the item "Extraction of DNA and / or RNA".

[0345] Next, process 220 proceeds to activity 222, where non-stranded RNA sequencing is performed on a first library of cDNA fragments prepared from the enriched RNA to obtain RNA expression data. The manner related to performing non-stranded NDA sequencing on a first library of cDNA fragments prepared from the enriched RNA to obtain RNA expression data is described in the item "RNA Sequencing".

[0346] Next, process 220 proceeds to activity 223, where the RNA expression data is converted into gene expression data. Next, process 220 proceeds to activity 224, where at least one gene that introduces bias into the gene expression data is identified. Next, process 220 proceeds to activity 225, where the expression data associated with at least one gene that introduces bias is removed from the gene expression data to obtain bias-corrected gene expression data. Aspects related to activities 223, 224, and 225 are described in the section entitled "Removal of Bias".

[0347] Next, process 220 proceeds to activity 226, where a cancer treatment for a subject using the bias-corrected gene expression data is identified. Aspects related to identifying a cancer treatment for a subject using the bias-corrected gene expression data are described in the section entitled "Identification of Cancer Treatment".

[0348] FIG. 7 is an exemplary flowchart showing a computerized process 300 for preparing a patient sample for sequencing analysis and performing bioinformatics quality control, whereby a cancer treatment suitable for a patient or subject from whom nucleic acids have been extracted for sequencing analysis can be obtained.

[0349] In the illustrated embodiment, process 300 includes obtaining a first sample of a first tumor from a subject having cancer, suspected of having cancer, or at risk of having cancer in activity 301; extracting RNA from the first sample of the first tumor in activity 302; enriching the RNA for coding RNA to obtain enriched RNA in activity 303; preparing a first library of cDNA fragments from the enriched RNA for non-coding RNA sequencing in activity 304; obtaining RNA expression data for the subject in activity 305; aligning and annotating genes in the RNA expression data with known sequences of the human genome to obtain annotated RNA expression data in activity 306; removing non-coding transcripts from the annotated RNA expression data in activity 307; converting the annotated RNA expression data to gene expression data (e.g., in Transcripts Per Kilobase Million (TPM) format) in activity 308; identifying at least one gene that introduces a bias into the gene expression data in activity 309; removing expression data for at least one gene that introduces a bias from the gene expression data to obtain bias-corrected gene expression data in activity 310; obtaining sequence information and asserted information in activity 311; determining one or more features from the sequence information in activity 312; determining whether the one or more features match the asserted information in activity 313; making at least one additional determination of a feature in activity 314; and identifying cancer treatment for the subject using the bias-corrected gene expression data in activity 315.

[0350] It should be understood that one or more activities of process 300 may be optional. For example, in some embodiments, activities 301 and 303 may be performed, and activity 303 is optional. In some embodiments, activities 301, 302, and 303 are all performed. In some embodiments, activities 301, 302, and 303 are all omitted, but the remaining activities are performed. This is useful when the enriched RNA extracted from the patient sample is already available prior to the start of process 300. In some embodiments, one or more features in activity 312 include one or more features of source, patient, tissue type, tumor type, polyA status, MHC sequence, protein subunit ratio, complexity, contamination, coverage, exon coverage, read composition, Phred score, SNP match, and GC content. In some embodiments, one or more features in activity 312 further include the strandness of the RNA sequence analysis. In some embodiments, any one or more of the features in activity 312 may be determined. In some embodiments, additional determinations of features in activity 314 can include, but are not limited to, SNP match values, contamination values, polyA status, complexity values, Phred scores, and GC content. In some embodiments, additional determinations of any one or more of the features may be performed in activity 314. In some embodiments, any one or more of activity 303, process 307, and process 314 may be omitted. In some embodiments, all activities of the computerized process 300 may be performed.

[0351] FIG. 8 illustrates a non-limiting process pipeline 800. FIG. 8 illustrates a non-limiting process pipeline 800 for processing array data and the claimed information associated with the array data to confirm validity for subsequent analysis (e.g., for diagnosis, prognosis, therapy, and / or other clinical applications). Activity 801 is performed by obtaining nucleic acid data that includes the array data and the claimed information indicating the claimed source for the array data. In some embodiments, the nucleic acid data is obtained from a previously processed biological sample. In some embodiments, the biological sample was previously obtained from a subject having cancer, suspected of having cancer, or at risk of having cancer. In some embodiments, activity 801 is performed by obtaining nucleic acid data that includes the claimed completeness of the array data. In some embodiments, activity 801 is performed by obtaining nucleic acid data that includes the array data and the claimed information indicating the claimed source and the claimed completeness of the array data. In some embodiments, the claimed information indicates the claimed completeness of the array data. In some embodiments, the claimed information indicates the subject from which the nucleic acid was obtained. For example, in some embodiments, the claimed information includes MHC allele information and / or SNP information for one or more loci of the subject. After activity 801, process 800 proceeds to activities 802 and 803, where the validity of the nucleic acid data obtained in activity 801 is confirmed. Validity confirmation includes obtaining the determined completeness and / or the determined source by processing the array data in activity 802, and determining in activity 803 whether the determined completeness and / or the determined source match the claimed completeness and / or the claimed source, respectively. The array data is processed in activity 802 to obtain the determined information indicating the determined source of the array data in activity 802a and / or the determined information indicating the determined completeness of the array data in activity 802b.In some embodiments, activity 802a may include determining information indicative of at least one, two, three of the subject's MHC genotype, whether the nucleic acid data is RNA data or DNA data, the tissue type of the biological sample, the tumor type of the biological sample, the sequencing platform used to generate the sequence data, SNP matches (e.g., determining whether one or more SNPs in the sequence data match one or more SNPs in a reference sequence), and / or whether the RNA sample is poly-A enriched. In some embodiments, activity 802b may include determining a first level of a first nucleic acid encoding a first subunit of a multimeric protein, determining a second level of a second nucleic acid encoding a second subunit of the multimeric protein, and determining whether the ratio between the first level and the second level matches an expected ratio. In some embodiments, the first subunit and the second subunit are the first and second CD3 subunits, the first and second CD8 subunits, or the first and second CD79 subunits. In some embodiments, the determined information indicative of the determined completeness is indicative of at least one, two, three of total sequence coverage, exon coverage, chromosomal coverage, the ratio of nucleic acids encoding two or more subunits of a multimeric protein, species contamination, complexity, and / or the percentage (%) of guanine (G) and cytosine (C) in the sequence data. In some embodiments, activity 803 includes determining one or more MHC allele sequences from the sequence data and determining whether the one or more MHC allele sequences match the claimed MHC allele information for the subject. In some embodiments, determining the MHC alleles includes determining sequences for six MHC loci from the sequence data.

[0352] In activity 803, the determined integrity and / or source is evaluated by determining whether the determined source of the array data matches the claimed source of the array data and / or whether the determined integrity of the array data matches the claimed integrity of the array data.

[0353] If the information claimed in activity 803 matches the determined information (i.e., in the case of yes), process 800 proceeds to activity 804, and the array data is further evaluated to determine whether the array data indicates a diagnosis, prognosis, therapy, or other clinical outcome. For example, in some embodiments, the array data is further processed in activity 804 to provide a recommendation for cancer treatment for a subject having cancer, suspected of having cancer, or at risk of having cancer. In some embodiments, activity 804 is performed by determining a therapy for the subject, and the therapy is then administered to the subject.

[0354] In some embodiments, the process may further include administering a therapy to the subject. In some embodiments, the therapy is a cancer therapy.

[0355] In some embodiments, determining a therapy for the subject may include determining a plurality of gene group expression levels including gene group expression levels for each gene group in a set of gene groups. In some embodiments, the set of gene groups includes at least one gene group related to cancer malignancy and at least one gene group related to the cancer microenvironment. The therapy for the subject is identified by using the determined gene group expression levels.

[0356] If the information asserted in activity 803 does not match the determined information (i.e., in the case of no), process 800 proceeds to 805 and one or more corrective actions are performed. In some embodiments, the corrective actions include generating an indication that the determined information does not match the asserted information, generating an indication not to process the array data in subsequent analysis, and / or generating an indication to obtain additional array data and / or biological samples and / or other information regarding the subject.

[0357] In some embodiments, the method includes all of the activities illustrated in FIG. 8. However, in some embodiments, a subset of these activities are performed, any one or more of those activities may be omitted, duplicated, and / or performed in an order different from that illustrated in FIG. 8. For example, either activity 802a or activity 802b is performed at activity 802. For example, activity 803 may be performed twice to confirm the determination. For example, one or more activities in process 800 may be performed after one or more corrective actions at activity 805. In some embodiments, one or more of the activities of FIG. 8 are implemented on a computer.

[0358] In some embodiments, the expression levels of one or more genes in a sample are analyzed to evaluate the origin and / or quality of the sample. For example, the expression of one or more genes known to be expressed in a particular cell, tissue, or tumor type is evaluated to determine whether the expression level is as expected based on the cell, tissue, or tumor being analyzed. Similarly, the expression of one or more genes known not to be expressed (or not highly expressed) in a particular cell, tissue, or tumor type is evaluated to determine whether the expression level is as expected based on the cell, tissue, or tumor being analyzed.

[0359] In some embodiments, the expression levels of one or more genes are analyzed for each of a plurality of samples (e.g., 2, 3, 4, 5, 4-10, 1-50, 50-500, or more samples). If the expression of one or more genes is lower or higher than expected, this may indicate that the quality of the data being analyzed and / or the source / origin is not as expected. In some embodiments, data from samples having unexpected levels of expression (e.g., levels lower or higher than expected) for one or more genes are excluded from further analysis. In some embodiments, new sequence information is obtained for samples having unexpected levels of expression for one or more genes, for example to confirm whether the initial data is correct. In some embodiments, samples having unexpected levels of expression for one or more genes may be further analyzed to determine, for example, whether the samples are from a different source than initially indicated.

[0360] In some embodiments, the expression levels for one or more genes are analyzed (e.g., using tSNE, PCA, or other techniques), thereby determining whether gene expression or patterns of gene expression are similar or different in distinct samples. In some embodiments, if a dataset containing the same cell type or the same tissue type does not cluster within a group, or if one or more datasets are identified as being statistically different from other datasets containing the same cells or tissue, the identified different datasets are excluded and may be further analyzed or flagged as potentially suspect. In some embodiments, additional sequence data may be obtained for samples identified as potentially suspect.

[0361] An exemplary embodiment of a computer system 500 that may be used in connection with any of the embodiments of the technology described herein is shown in FIG. 9. The computer system 500 comprises one or more processors 510 and one or more manufactured articles comprising non-transitory computer-readable storage media (e.g., memory 520 and one or more non-volatile storage media 530). The processor 510 may be configured to control writing and reading of data to and from the memory 520 and the non-volatile storage device 530 in any suitable manner, and aspects of the technology described herein are not limited in this regard. To execute any of the functions described herein, the processor 510 may execute one or more processor-executable instructions stored in one or more non-transitory computer-readable storage media (e.g., memory 520) that may serve as a non-transitory computer-readable storage media storing processor-execution instructions for execution by the processor 510.

[0362] The computing device 500 may also include a network input / output (I / O) interface 540 that may be used for communication of the computing device with other computing devices (e.g., over a network), and may include one or more user I / O interfaces 550 that may be used when the computing device provides output to and receives input from a user. The user I / O interface may include devices such as a keyboard, mouse, microphone, display device (e.g., monitor or touch screen), speaker, camera, and / or other various types of I / O devices.

[0363] The embodiments described above can be implemented in a number of ways. For example, these embodiments can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor (e.g., a microprocessor) or collection of processors, whether provided on a single computing device or distributed among a plurality of computing devices. It should be understood that any component or collection of components that performs the functions described above can generally be regarded as one or more controllers that control the functions described above. The one or more controllers can be implemented in various ways, such as dedicated hardware or general-purpose hardware (such as one or more processors) programmed using microcode or software to perform the functions described above.

[0364] In this regard, one implementation of the embodiments described herein includes at least one computer-readable storage medium (e.g., RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or other tangible, non-transitory computer-readable storage medium) encoded with a computer program (i.e., a plurality of executable instructions) that, when executed on one or more processors, performs the above-described functions of one or more embodiments. The computer-readable medium may be portable such that the program stored thereon can be loaded onto any computing device for implementing aspects of the techniques described herein. Additionally, it should be understood that references to a computer program that, when executed, performs any of the above-described functions are not limited to application programs executed on a host computer. Rather, the terms computer program and software are used herein in a general sense to refer to any type of computer code (e.g., application software, firmware, microcode, or other forms of computer instructions) that can be employed to program one or more processors for the purpose of implementing aspects of the techniques described herein.

[0365] The aspects of the techniques described herein provide a computer-implemented method for evaluating, generating, visualizing, and / or classifying biological properties of sequence information of a subject (e.g., a cancer patient) or a person having, suspected of having, or at risk of having a disease (e.g., cancer) (e.g., cancer grade, primary tissue).

[0366] In some embodiments, a software program may provide a visual representation of a subject's (e.g., a patient's) characteristics and / or other information related to the subject's (e.g., a patient's) cancer to a user using an interactive graphical user interface (GUI). Such a software program may be executed in any suitable computing environment, including but not limited to a cloud computing environment, a device located at the same location as the user (e.g., the user's laptop, desktop, smartphone, etc.), one or more devices remote from the user (e.g., one or more servers), and the like.

[0367] For example, in some embodiments, the techniques described herein may be implemented in the exemplary environment 600 shown in FIG. 10. As shown in FIG. 10, within the exemplary environment 600, one or more biological samples of a subject 680 may be provided to a laboratory 670. The laboratory 670 may process the biological sample to obtain expression data (e.g., DNA, RNA, and / or protein expression data) and / or sequence information and provide it to at least one database 660 that stores information about the subject (e.g., a patient) 680 via a network 610.

[0368] The network 610 may be a wide area network (e.g., the Internet), a local area network (e.g., a corporate intranet), and / or any other suitable type of network. Any of the devices shown in FIG. 10 may be connected to the network 610 using one or more wired links, one or more wireless links, and / or any suitable combination thereof.

[0369] In the illustrated embodiment of FIG. 10, at least one database 620 may store expression data and / or sequence information for a subject (e.g., a patient), medical history data for the subject (e.g., a patient), test result data for the subject (e.g., a patient), and / or any other suitable information regarding subject 680. Examples of stored test result data for a subject (e.g., a patient) include biopsy test results, imaging test results (e.g., MRI results), and blood test results. The information stored in at least one database 620 may be stored in any suitable format and / or using any suitable data structure, and aspects of the techniques described herein are not limited in this regard. At least one database 620 may store data in any suitable manner (e.g., in one or more databases, in one or more files). At least one database 620 may be a single database or multiple databases.

[0370] As shown in FIG. 10, an exemplary environment 600 includes one or more external databases 620, which may store information of patients other than patient 680. For example, external database 660 may store expression data and / or array information (of any suitable type) for one or more patients, medical history data of one or more patients, test result data of one or more patients (e.g., image results, biopsy results, blood test results), demographic and / or personal information of one or more patients, and / or any other suitable type of information. In some embodiments, external database 660 may store information available in one or more publicly accessible databases such as TCGA (The Cancer Genome Atlas), one or more databases of clinical trial information, and / or one or more databases maintained by commercial sequencing suppliers. External database 660 may store such information in any suitable manner using any suitable hardware, and the aspects of the techniques described herein are not limited in this regard.

[0371] In some embodiments, at least one of database 620 and external database 660 may be the same database, may be part of the same database system, or may be physically in the same location, and the aspects of the techniques described herein are not limited in this regard.

[0372] For example, in some embodiments, server 640 may access the information stored in database 620 and / or 660 and use this information to perform the processes described herein with reference to FIG. 10 to determine one or more characteristics of a biological sample and / or array information.

[0373] In some embodiments, server 640 may comprise one or more computing devices. When server 640 comprises multiple computing devices, the devices may be physically co-located (e.g., in a single room), or may be distributed across multiple physical locations. In some embodiments, server 640 may be part of a cloud computing infrastructure. In some embodiments, one or more servers 640 may be within a facility that is co-located and operated by an organization (e.g., a hospital, a research institution) to which physician 650 belongs. In such embodiments, it may be easier to enable server 640 to access the private medical data of patient 880.

[0374] As shown in FIG. 10, in some embodiments, the results of the analysis performed by server 640 may be provided to physician 650 via computing device 630, which may be a portable computing device such as a laptop or smartphone, or a fixed computing device such as a desktop computer. The results may be provided in the form of a written report, an email, a graphical user interface, and / or other suitable means. In the embodiment of FIG. 10, the results are provided to physician 650, but it should be understood that in other embodiments, the analysis results may be provided to patient 680, or a caregiver of patient 680, a healthcare provider such as a nurse, or a person involved in a clinical trial.

[0375] In some embodiments, the results may be part of a graphical user interface (GUI) presented to physician 650 via computing device 630. In some embodiments, the GUI may be presented to the user as part of a web page displayed by a web browser running on computing device 630. In some embodiments, the GUI may be presented to the user using an application program (different from a web browser) running on computing device 630. For example, in some embodiments, computing device 630 may be a mobile device (e.g., a smartphone), and the GUI may be presented to the user via an application program (e.g., an “app”) running on the mobile device.

[0376] The GUI presented on computing device 630 can provide a wide range of oncological data related to both the patient and the patient's cancer in a compact and information-rich new way. Previously, oncological data was obtained from multiple sources of data in several separate steps, and the process of obtaining such information was costly from both a time and monetary perspective. By using the techniques and graphical user interfaces illustrated herein, the user can access the same amount of information at once, reducing the user's requirements and the computing resources required to provide such information. Reducing the user's requirements helps reduce clinician errors related to searching various sources of information. Reducing the computing resource requirements helps reduce the processor power, network bandwidth, and memory required to provide extensive oncological data, improving computing technology. In some embodiments, the reports of the present disclosure are presented to the user by the system or using the GUI.

[0377] Thus, in one aspect, the present disclosure relates to a method of evaluating array information to determine at least one characteristic. This evaluation can be performed on a computer or other automated machine capable of executing programmable instructions, or can be performed manually by an evaluator. The characteristic can be used to generate a report informing the evaluator of at least one characteristic of the array information. In some embodiments, the characteristic is the sequence of MHC alleles of the array information.

[0378] The major histocompatibility complex (MHC) (referred to as human leukocyte antigen (HLA) in humans) is a mechanism by which the immune system can distinguish between self and non-self cells. It is an aggregate of glycoproteins (proteins containing carbohydrates) present on the plasma membranes of almost all somatic cells. The (MHC) is a highly polymorphic gene that is important in the immune system of living organisms. It is derived from 20 genes and has more than 50 mutations per gene among individuals, and also allows codominance between alleles. These glycoproteins are part of the pathway that allows the immune system to distinguish between self and non-self cells by the abnormalities of MHC presented on the plasma membrane.

[0379] Due to these properties, such as the high polymorphism of MHC, its codominance, and the large number of alleles that can exist in a given species, the MHC profile of a subject is highly specific and unique. Thus, it is highly unlikely that two humans, except for identical twins, will possess cells with the same set of MHC molecules. Thus, by evaluating the sequence of the MHC profile of the array information, this can be used to confirm or consider as ineligible to distinguish information between the array information, the claimed information, other array information, or combinations thereof.

[0380] In some embodiments, one MHC allele is used for the evaluation. In some embodiments, at least two MHC alleles are used for the evaluation. In some embodiments, at least three MHC alleles are used for the evaluation. In some embodiments, at least four MHC alleles are used for the evaluation. In some embodiments, at least five MHC alleles are used for the evaluation. In some embodiments, at least six MHC alleles are used for the evaluation.

[0381] In some embodiments, the feature being evaluated is the concordance value of a single nucleotide polymorphism (SNP). As used herein, "SNP" or "single nucleotide polymorphism" refers to a difference in a nucleic acid sequence (e.g., a genome, a sequence dataset) in a single nucleotide (e.g., adenine (A), thymine (T), cytosine (C), and / or guanine (G)) shared among subjects of one species, or within an individual subject on paired chromosomes. An SNP may represent a changed nucleotide called a substitution (e.g., A changed to T, G changed to A, etc.), a removed nucleotide where the nucleotide is completely absent from the sequence called a deletion, or an added nucleotide where an additional nucleotide is added to the sequence. An SNP may cause a change in the encoded protein (e.g., a non-synonymous SNP) or may not (e.g., synonymous). Further, when an SNP is non-synonymous, it may cause a change in the encoded amino acid (e.g., a missense) or may cause a premature stop codon (e.g., a nonsense). Synonymous SNPs can also modify the message of a nucleic acid sequence by affecting or altering splice sites, transcription factor binding, and / or messenger RNA (mRNA) binding. These mutations (e.g., changes to the protein-coding ability of a sequence) can cause many effects including differences in phenotype and even various disease types. Further, SNPs occur in large numbers within a subject's genome, and a typical genome is estimated to differ from the reference human genome at 4 to 5 million sites, more than 99.9% of which are estimated to be SNPs.

[0382] Since SNPs are encoded in nucleic acids that are part of the genome, they are inherited from parents to offspring (in the subject being tested and in the subject's body when the nucleic acids are replicated). Thus, because this inheritance is stable and because there are many of them, SNPs can be used as genetic markers of a subject's relatedness and as a measure of the identity of two nucleic acid sequences as being from the same subject. In some embodiments, the SNP match value is determined between the sequence information and a reference sequence. In some embodiments, the SNP match value is determined between the sequence information and the claimed value. In some embodiments, the SNP match value should be equal to or greater than a threshold that should be acceptable for further analysis (e.g., considered to have sufficient quality and completeness). In some embodiments, the threshold is 80%. In some embodiments, the SNP match value is determined between a sequence dataset and a subject, and the subject is considered to be highly likely to be from the subject and identified as being from the subject if the SNP match value is at least 70% (e.g., at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 95.5%, at least 96%, at least 96.5%, at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, at least 99.95%, at least 99.99%, at least 99.999%, or more). As described herein, in some embodiments, the determination of the match value is as described in this disclosure.The detection of SNP concordance can be performed by any available or known means in the art. For example, the detection of SNP concordance can be performed by various online tools such as Conpair (github.com / nygenome / Conpair) or GATK GenotypeConcordance (software.broadinstitute.org / gatk / documentation / tooldocs / 3.8-0 / org_broadinstitute_gatk_tools_walkers_variantutils_GenotypeConcordance.php), or can be calculated manually. In other examples, the detection of SNP concordance can be performed by tools as described on publicly available websites (genome.sph.umich.edu / wiki / VerifyBamID or software.broadinstitute.org / cancer / cga / contest).

[0383] In some embodiments, the feature being evaluated is a quality score, such as a Phred score. As used herein, a "Phred score" (also sometimes known as or referred to as a "Phred quality score" herein) refers to a measure of quality for the identification of nucleotides sequenced by a nucleic acid sequencing system or platform (e.g., NGS). Phred scores are known in the art and are often generated from a sequencing platform based on several parameters (e.g., peak shape, resolution, etc.), and a score (Q) is assigned to each nucleotide base call (for a detailed review of the calculations, see Ewing B, Hillier L, Wendl MC, Green P. "Base-calling of automated sequencer traces using phred. I. Accuracy assessment.", Genome Res. Mar 1998, 8(3):175-85 and Ewing B, Green P. "Base-calling of automated sequencer traces using phred. II. Error probabilities", Genome Res. Mar 1998, 8(3):186-94.). The Phred score for each base indicates the probability that the nucleotide base call is incorrect (base call error probability (P)), and is given by the formula Q = -10log 10It is determined by P. Thus, a score (e.g., Q) indicates base call accuracy. For example, a Phred score of 10 indicates 90% call accuracy for the base of interest, and a Phred score of 40 indicates 99.99% call accuracy for the same base. In some embodiments, the Phred score of the sequence information is determined and compared to a reference value. In some embodiments, the reference value is at least 27, at least 28, at least 29, at least 30, or greater than 30. In some embodiments, the Phred score is determined and compared to the Phred score of other sequence information. In some embodiments, the Phred score is determined and compared to the claimed score. In some embodiments, the Phred score is used as a quality base-level determination. In some embodiments, it is used to compare identity by comparing the sequence to the claimed information, but it is unlikely that they are the same sequence information or from the same sample or subject if the Phred scores are different.

[0384] In some embodiments, the feature to be evaluated is the tumor type.

[0385] In some embodiments, the feature to be evaluated is the tissue type.

[0386] In some embodiments, the feature being evaluated is the polyadenylation status of the sequence information. As used herein, "polyadenylation" or "polyA" refers to a series of multiple adenosine monophosphate nucleotides attached to the 3' end of messenger RNA (mRNA), which occurs after cleavage of the 3' end of the transcript and release of the hydroxyl. The "polyA tail", often referred to, is characteristic of fully processed mRNA and aids in various cellular processes. For example, the polyA tail is a binding site for proteins (polyA-binding proteins), facilitates transport from the cell nucleus so that translation can occur, and further affects mRNA translation and stability. When only the transcript of protein-coding mRNA is present, the sample from which the sequence information was generated is likely to have been generated using mRNA-Seq (e.g., as indicated). In some embodiments, the polyA status indicates that mRNA-Seq was not used (e.g., whole transcriptome was used). In some embodiments, the polyA status is evaluated against the claimed information. In some embodiments, the polyA status is evaluated against a reference sequence. In some embodiments, the probability that sequence information is generated using either mRNA-Seq or whole transcriptome must be higher than a threshold. In some embodiments, the threshold is a reference value. In some embodiments, the threshold level is 90%. In some embodiments, the threshold is the claimed information. In some embodiments, the sequence information is from a sample that mainly contains polyadenylated nucleic acids. In some embodiments, the sequence information is from a sample that contains both polyadenylated and non-polyadenylated nucleic acids.

[0387] In some embodiments, the feature being evaluated is the GC content of the sequence information. "G / C content" or "guanine (G)-cytosine (C) content", as used herein, refers to the percentage of nucleotides in a nucleic acid sample that are either G or C. This can be calculated by summing all of the G and C reads of a given sequence information and dividing by the total number of nucleotides sequenced. In some embodiments, the sequence information is evaluated and the GC content is calculated by summing the number of base calls that result in a G or C (e.g., G+C) within the sequence information and dividing by the total number of base calls within the sequence information (e.g., the number of nucleotides in the sequence dataset), i.e., (G+C) / (number of nucleotides in the sequence dataset).

[0388] The GC content can also be used as a quality metric for the sequence information. Many known genomes are being sequenced along with their respective exomes, transcriptomes, and various other parts (e.g., for a measure of a particular RNA component). Further, many of these sequences are being sequenced multiple times and the average values and ranges for their various components are being generated, for example, for the GC content of the human genome. The GC content of the human genome is known to vary in the range of about 35% to 60% and has an average value of about 41% (e.g., mean value). Thus, as a quality metric, if sequence information identified as being from the human genome were to be ...

Claims

**Claim 1** A method comprising: obtaining a first biological sample of a first tumor, wherein the first biological sample was previously obtained from a subject having cancer, suspected of having cancer, or at risk of having cancer; extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; enriching the extracted RNA for coding RNA to obtain enriched RNA; sequencing the enriched RNA using at least one sequencing platform to obtain RNA expression data comprising at least 5 kilobases (kb); using at least one hardware processor to: obtain the RNA expression data using the at least one sequencing platform; convert the RNA expression data to gene expression data; determine bias-corrected gene expression data by removing expression data for at least one gene that at least partially introduces bias into the gene expression data from the gene expression data; generate information indicative of cancer treatment for the subject using the bias-corrected gene expression data; wherein the at least one gene is At least one histone code gene selected from the group consisting of (a) HIST1H1A, HIST1H1B, HIST1H1C, HIST1H1D, HIST1H1E, HIST1H1T, HIST1H2AA, HIST1H2AB, HIST1H2AC, HIST1H2AD, HIST1H2AE, HIST1H2AG, HIST1H2AH, HIST1H2AI, HIST1H2AJ, HIST1H2AK, HIST1H2AL, HIST1H2AM, HIST1H2BA, HIST1H2BB, HIST1H2BC, HIST1H2BD, HIST1H2BE, HIST1H2BF, HIST1H2BG, HIST1H2BH, HIST1H2BI, HIST1H2BJ, HIST1H2BK, HIST1H2BL, HIST1H2BM, HIST1H2BN, HIST1H2BO, HIST1H3A, HIST1H3B, HIST1H3C, HIST1H3D, HIST1H3E, HIST1H3F, HIST1H3G, HIST1H3H, HIST1H3I, HIST1H3J, HIST1H4A, HIST1H4B, HIST1H4C, HIST1H4D, HIST1H4E, HIST1H4F, HIST1H4G, HIST1H4H, HIST1H4I, HIST1H4J, HIST1H4K, HIST1H4L, HIST2H2AA3, HIST2H2AA4, HIST2H2AB, HIST2H2AC, HIST2H2BE, HIST2H2BF, HIST2H3A, HIST2H3C, HIST2H3D, HIST2H3PS2, HIST2H4A, HIST2H4B, HIST3H2A, HIST3H2BB, HIST3H3, and HIST4H4, and (b) At least one mitochondrial gene selected from the group consisting of MT-ATP6, MT-ATP8, MT-CO1, MT-CO2, MT-CO3, MT-CYB, MT-ND1, MT-ND2, MT-ND3, MT-ND4, MT-ND4L, MT-ND5, MT-ND6, MT-RNR1, MT-RNR2, MT-TA, MT-TC, MT-TD, MT-TE, MT-TF, MT-TG, MT-TH, MT-TI, MT-TK, MT-TL1, MT-TL2, MT-TM, MT-TN, MT-TP, MT-TQ, MT-TR, MT-TS1, MT-TS2, MT-TT, MT-TV, MT-TW, MT-TY, MTRNR2L1, MTRNR2L10, MTRNR2L11, MTRNR2L12, MTRNR2L13, MTRNR2L3, MTRNR2L4, MTRNR2L5, MTRNR2L6, MTRNR2L7, and MTRNR2L8, comprising, wherein each of at least one gene of (a) and (b) brings a bias in the length of the poly-A tail or a bias in the average transcript length into gene expression data to generate bias-corrected gene expression data in transcripts per kilobase million (TPM) format, A method comprising the step of performing poly-A enrichment to enrich the RNA for the coding RNA. **Claim 2.** The method according to claim 1, wherein each of at least one gene of (a) and (b) is a gene having an average transcript length longer or shorter than the average length of transcripts in the gene expression data, is a gene having at least one threshold variation in the average transcript expression level based on the transcript expression level in a reference sample, and / or a gene having a poly-A tail length that is at least a threshold amount smaller compared to the average length of the poly-A tails of genes from the first biological sample and / or reference sample from which the RNA expression data was obtained. **Claim 3.** The step of determining the bias-corrected gene expression data further comprises the step of re-normalizing the gene expression data after removing the expression data for the at least one gene that brings bias into the gene expression data, according to the method of claim 1 or 2. **Claim 4.** The step of converting the RNA expression data into gene expression data A step of obtaining filtered RNA expression data by removing non-coding transcripts from the RNA expression data; After removing the non-coding transcripts, normalizing the filtered RNA expression data to obtain gene expression data of Transcripts Per Million (TPM), and The step of removing the non-coding transcripts from the RNA expression data includes pseudogenes, polymorphic pseudogenes, processed pseudogenes, transcribed processed pseudogenes, unitary pseudogenes, non-processed pseudogenes, transcribed unitary pseudogenes, constant chain immunoglobulin (IG C) pseudogenes, joining chain immunoglobulin (IG J) pseudogenes, variable chain immunoglobulin (IG V) genes, transcribed non-processed genes, translated non-processed genes, joining chain T cell receptor (TR J) genes, variable chain T cell receptor (TR V) genes, small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), microRNA (miRNA), ribozyme, ribosomal RNA (rRNA), mitochondrial tRNA (Mt tRNA), mitochondrial rRNA (Mt rRNA), Cajal body-specific RNA (scaRNA), retained introns, sense intron RNA, sense overlapping RNA, nonsense mutation-dependent degradation RNA, non-stop degradation RNA, antisense RNA, long intergenic non-coding RNA (lncRNA), macro long non-coding RNA (macro lncRNA), processed transcripts, 3' overlapping non-coding RNA (3' overlapping ncRNA), small RNA (sRNA), other RNA (miscRNA), vault RNA, and TEC RNA. The method according to any one of claims 1 to 3, including the step of removing non-coding transcripts belonging to the selected group.

5. Before removing the non-coding transcripts, Aligning the RNA expression data to a reference; and Annotating the RNA expression data. The method according to claim 4, further including these steps.

6. The RNA expression data includes at least 25 million paired-end reads, The RNA expression data includes at least 50 million paired-end reads, and the average read length is at least 100 bp. The method according to any one of claims 1 to 5.

7. The step of generating information indicating the cancer treatment for the subject using the bias-corrected gene expression data includes determining, using the bias-corrected gene expression data, a plurality of gene group expression levels, wherein the plurality of gene group expression levels includes the gene group expression level of each gene group in a set of gene groups, and the set of gene groups includes at least one gene group related to the malignancy of cancer and at least one gene group related to the microenvironment of cancer, the step of generating information indicating the cancer treatment using the determined gene group expression levels, according to any one of claims 1 to 6. **Claim 8** The cancer treatment according to any one of claims 1 to 7 is selected from the group consisting of radiotherapy, surgical therapy, chemotherapy, and immunotherapy. **Claim 9** The method further includes obtaining a second biological sample of a second tumor, the second biological sample having been previously obtained from the subject, The method according to any one of claims 1 to 8 further includes forming a combined tumor sample by combining the first biological sample and the second biological sample, The step of extracting the RNA includes extracting the RNA from the combined tumor sample and / or extracting RNA from the second biological sample, and further includes forming combined extracted RNA by combining the RNA extracted from the second biological sample with the RNA extracted from the first biological sample. The step of enriching the RNA for coding RNA includes enriching the combined extracted RNA for coding RNA. The method according to any one of claims 1 to 8. **Claim 10** The extracted RNA contains at least 1 μg of RNA after RNA extraction, The extracted RNA has a total mass of at least 1000-6000 ng and a purity corresponding to a ratio of absorbance at 260 nm to absorbance at 280 nm of at least 2.0, according to any one of claims 1 to 9. **Claim 11** The quality control evaluation for the RNA expression data is at least partly obtaining the claimed information indicating the claimed source and / or claimed completeness of the RNA expression data, Process the RNA expression data to obtain determined information indicating the determined source and / or determined completeness of the RNA expression data, further comprising the step of performing by determining whether the determined information matches the claimed information, The step of processing the RNA expression data includes the step of processing the RNA expression data to determine the tissue type of the first biological sample, the tumor type of the first biological sample, and / or the percentage (%) of guanine (G) and / or cytosine (C). The method according to any one of claims 1 to 10.

12. A system for performing the method according to any one of claims 1 to 11, wherein the system comprises: At least one sequencing platform configured to generate gene expression data from enriched RNA obtained from a first biological sample previously obtained from the subject, wherein the enriched RNA comprises: (i) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA, wherein the RNA expression data comprises at least 5 kilobases (kb). At least one sequencing platform obtained by steps; At least one computer hardware processor; At least one non-transitory computer-readable storage medium storing processor-executable instructions, wherein the processor-executable instructions, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to: Obtaining the RNA expression data using the at least one sequencing platform; Converting the RNA expression data into gene expression data; Determining bias-corrected gene expression data from the gene expression data by removing expression data for at least one gene that, at least in part, introduces bias into the gene expression data from the gene expression data; Causing at least one non-transitory computer-readable storage medium to perform a step of generating information indicating cancer treatment for the subject using the bias-corrected gene expression data.

13. A system for performing the method according to any one of claims 1 to 11, the system comprising at least one computer hardware processor; at least one non-transitory computer-readable storage medium storing processor-executable instructions, the processor-executable instructions, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to obtain RNA expression data from at least one sequencing platform, the RNA expression data including at least 5 kilobases (5 kb), the RNA expression data being obtained from a first biological sample of a first tumor previously obtained from the subject, at least in part by (i) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA; convert the RNA expression data into gene expression data; determine bias-corrected gene expression data from the gene expression data, at least in part by removing expression data for at least one gene that introduces a bias into the gene expression data from the gene expression data; Causing at least one non-transitory computer-readable storage medium to perform a step of generating information indicating cancer treatment for the subject using the bias-corrected gene expression data.

14. The system according to claim 13, further comprising the at least one sequencing platform.

15. at least one non-transitory computer-readable storage medium storing processor-executable instructions, the processor-executable instructions, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to Obtaining RNA expression data from at least one sequencing platform, wherein the RNA expression data includes at least 5 kilobases (5 kb), and the RNA expression data is from a first biological sample of a first tumor previously obtained from a subject having cancer, suspected of having cancer, or at risk of having cancer, and at least a portion is obtained by: (i) extracting RNA from the first biological sample of the first tumor to obtain extracted RNA; and (ii) enriching the extracted RNA for coding RNA to obtain enriched RNA; and converting the RNA expression data into gene expression data; Determining bias-corrected gene expression data from the gene expression data, at least in part by removing expression data for at least one gene that introduces bias into the gene expression data from the gene expression data; A non-transitory computer-readable storage medium that causes execution of the method according to any one of claims 1 to 11, the method including generating information indicative of cancer treatment for the subject using the bias-corrected gene expression data.

Citation Information

Patent Citations

  • High-throughput RNA sequencing data quality control method and high-throughput RNA sequencing data quality control apparatus

    CN105349617A

  • Expression of gag proteins from retroviruses in eucaryotic cells

    EP0345242A2

  • Heterovesicular liposomes

    EP0524968A1

  • GB2,200,651

  • Transmembrane proteins differentially expressed in cancer

    JP2005505267A