Systems and methods for automating RNA expression calls in cancer prediction pipeline
The method addresses the inadequacy of single-sample RNA quality measurements by performing automated batch quality control through cohort-matched reference batches and global quality control tests, enhancing the accuracy of RNA sequencing sample analysis.
Patent Information
- Application Number
- JP2025022377
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-12-04
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Current single-sample RNA quality measurements are inadequate for detecting batch effects in RNA sequencing samples, which can affect model performance and clinical interpretation.
A method for performing automated batch quality control of RNA sequencing samples by obtaining batch data sets, determining a cohort-matched reference batch, and performing global batch quality control tests to identify and address batch effects.
This method enables rapid and accurate quality control analysis of entire batches of RNA sequencing samples, improving diagnostic systems by identifying batch effects that conventional single-sample measurements miss.
Smart Images

Figure 2025085645000001_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 943,712, filed December 4, 2019, the contents of which are incorporated by reference herein in their entirety for all purposes.
[0002] The present disclosure relates generally to using RNA sequence information to perform quality control on batches of RNA sequencing samples. [Background technology]
[0003] Currently implemented single-sample RNA quality measurements in bioinformatics quality control processes were designed to detect overall sample quality and transcriptome integrity. However, there are many potential sources of error that can affect sample quality, particularly effects that can affect the results of an entire batch of sequencing samples. For example, changes in library creation protocols (e.g., reagents, capture probe lots, or instruments) or changes in bioinformatics pipelines (e.g., program versions) can result in subtle transcriptome changes, called batch effects, that can affect model performance and clinical interpretation of downstream processes (e.g., determining whether genes are over- or under-expressed compared to previously sequenced samples, or machine learning models to diagnose tumors of unknown etiology). These batch effects are only detectable by leveraging data patterns across samples, requiring new processes to be implemented between RNA mapping and downstream analysis.
[0004] What is needed in the art are improved methods for automatically performing quality control detection and analysis of batch effects of heterogeneous RNA sequencing samples in high throughput. Summary of the Invention
[0005] In view of the above background, there is a need for improved systems and methods for performing batch quality control (e.g., quality assessment) of RNA samples. Advantageously, the present disclosure provides a solution to these and other shortcomings in the art. For example, in some embodiments, the systems and methods described herein provide automated quality control of entire batches of RNA sequencing samples (e.g., thereby performing faster quality control analysis than currently available). Similarly, in some embodiments, the methods and systems described herein improve diagnostic systems and methods that use RNA expression data, for example, for precision oncology, by identifying batch effects that are otherwise not identified by conventional single-sample quality control measurements.
[0006] One aspect of the present disclosure provides a method for performing quality control. The method is implemented in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors. The method begins by (a) obtaining, for each respective sample in a batch of samples, a batch data set in electronic format, the batch data set including a corresponding plurality of sequence reads obtained from the respective sample by targeted or whole transcriptome RNA sequencing and corresponding metadata for the respective sample. The method further includes (b) determining a cohort-matched reference batch for the batch data set, the cohort-matched reference batch being a reference batch that is ... The method continues by determining that the batch data set is balanced with respect to cancer type, collection method, sequencer identity, and / or sequenced date. The method performs one or more global batch quality control tests on the batch data set using at least the cohort-matched reference batch. The method further comprises (d) removing each sample from the batch data set that fails any one of the one or more global batch quality control tests or flagging for manual inspection each sample that fails any one of the one or more global batch quality control tests.
[0007] In some embodiments, determining a cohort-matched reference dataset for the batch dataset comprises, for each sample in the batch of samples, i) extracting a respective plurality of sequence features from a respective plurality of sequence reads, thereby obtaining a plurality of sequence features of the batch, and ii) extracting a respective plurality of sample metadata features, thereby obtaining a plurality of metadata features of the batch. In some embodiments, determining the cohort-matched reference dataset further comprises selecting a cohort-matched reference dataset comprising a plurality of reference samples from the reference dataset based at least in part on the plurality of sample processing and sequence features of the batch or the plurality of metadata features of the batch.
[0008] In some embodiments, the method further includes, for each respective sample in the batch of samples, performing one or more single-sample quality control tests on the respective sample from the corresponding plurality of sequence reads, and removing each sample from the batch of samples that fails any one of the one or more single-sample quality control tests or flagging for manual inspection each sample that fails any one of the one or more single-sample quality control tests.
[0009] In some embodiments, the one or more global batch quality control tests include tests for one or more batch effects from a set that includes bioinformatics pipeline analysis and sequencing methods.
[0010] In some embodiments, the method further comprises determining a linear or non-linear combination of the plurality of sequence features of the batch and the plurality of metadata features of the batch by subjecting the plurality of sequence features of the batch and the plurality of metadata features of the batch to a dimensionality reduction technique.
[0011] In some embodiments, the method further comprises (c) adjusting each sample in the batch dataset for one or more confounding covariates using a cohort-matched reference batch prior to performing the one or more global batch quality control tests.
[0012] In some embodiments, the method further includes providing a respective sample report for each sample in the batch of samples, each respective sample report including at least one of a set of expression calls, one or more matched therapies, or one or more matched clinical trials.
[0013] In some embodiments, the method is implemented in a computer system comprising a cloud server. In some embodiments, the one or more global batch quality control tests comprise a first module and the one or more single sample quality control tests comprise a second module.
[0014] Other embodiments are directed to systems, portable consumer devices, and computer readable media associated with the methods described herein. , where applicable, may be applied to any aspect of the methods described herein.
[0015] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only exemplary embodiments of the present disclosure have been shown and described. As will be understood, the present disclosure is capable of other different embodiments, and its several details may be modified in various obvious respects, all without departing from the present disclosure. Thus, the drawings and description should be regarded as illustrative in nature, and not as restrictive. [Brief description of the drawings]
[0016] [Figure 1] 1 illustrates a block diagram of an exemplary computing device according to some embodiments of the present disclosure. [Figure 2A] Summary Provided herein is a flow chart of processes and features for performing quality control on batches of RNA-seq samples according to several embodiments of the present disclosure, with optional blocks indicated by dashed boxes. [Figure 2B] Summary Provided herein is a flow chart of processes and features for performing quality control on batches of RNA-seq samples according to several embodiments of the present disclosure, with optional blocks indicated by dashed boxes. [Figure 3A] According to some embodiments of the present disclosure, the evaluation of technical batch effect on RNA samples collected using PAX or EDTA tubes is summarized. Figure 3A shows the UMAP embedding of pooled cohorts and tissue-matched samples. Figures 3B and 3C show that the Mann-Whitney U test programmatically finds the difference between matched samples collected using PAX or EDTA tubes in both UMAP coordinates. UMAP coordinates are not ordered, and the coordinates to detect batch effect are arbitrary. [Figure 3B] According to some embodiments of the present disclosure, the evaluation of technical batch effect on RNA samples collected using PAX or EDTA tubes is summarized. Figure 3A shows the UMAP embedding of pooled cohorts and tissue-matched samples. Figures 3B and 3C show that the Mann-Whitney U test programmatically finds the difference between matched samples collected using PAX or EDTA tubes in both UMAP coordinates. UMAP coordinates are not ordered, and the coordinates to detect batch effect are arbitrary. [Figure 3C] According to some embodiments of the present disclosure, the evaluation of technical batch effect on RNA samples collected using PAX or EDTA tubes is summarized. Figure 3A shows the UMAP embedding of pooled cohorts and tissue-matched samples. Figures 3B and 3C show that the Mann-Whitney U test programmatically finds the difference between matched samples collected using PAX or EDTA tubes in both UMAP coordinates. UMAP coordinates are not ordered, and the coordinates to detect batch effect are arbitrary. [Figure 4A] According to some embodiments of the present disclosure, the evaluation of technical batch effect on aligned samples using Kallisto or STAR bioinformatics software pipeline is summarized. Figure 4 shows UMAP embedding of cohort and tissue-matched samples. Figure 4B and 4C show the results of Mann-Whitney U test on each of the UMAP coordinates. [Figure 4B] According to some embodiments of the present disclosure, the evaluation of technical batch effect on aligned samples using Kallisto or STAR bioinformatics software pipeline is summarized. Figure 4 shows UMAP embedding of cohort and tissue-matched samples. Figure 4B and 4C show the results of Mann-Whitney U test on each of the UMAP coordinates. [Figure 4C] According to some embodiments of the present disclosure, the evaluation of technical batch effect on aligned samples using Kallisto or STAR bioinformatics software pipeline is summarized. Figure 4 shows UMAP embedding of cohort and tissue-matched samples. Figure 4B and 4C show the results of Mann-Whitney U test on each of the UMAP coordinates. [Figure 5A] Technological batch effects. Distributions of false discovery rates (FDR) for common sources of effects are summarized. For each technology class, i.e., flow cell (Figure 5A), pipeline (Figure 5B), and sequencer (Figure 5C), technological batch effects were analyzed for 15 subsamples per feature and Benjamini Hochberg corrected FDRs were calculated across subsamples. The presented distributions represent the median FDR across subsamples. [Figure 5B]Technological batch effects. Distributions of false discovery rates (FDR) for common sources of effects are summarized. For each technology class, i.e., flow cell (Figure 5A), pipeline (Figure 5B), and sequencer (Figure 5C), technological batch effects were analyzed for 15 subsamples per feature and Benjamini Hochberg corrected FDRs were calculated across subsamples. The presented distributions represent the median FDR across subsamples. [Figure 5C] Technological batch effects. Distributions of false discovery rates (FDR) for common sources of effects are summarized. For each technology class, i.e., flow cell (Figure 5A), pipeline (Figure 5B), and sequencer (Figure 5C), technological batch effects were analyzed for 15 subsamples per feature and Benjamini Hochberg corrected FDRs were calculated across subsamples. The presented distributions represent the median FDR across subsamples. [Figure 6A] The results of a PCA dimensionality reduction analysis of approximately 100 paired cancer samples in which RNA expression was determined before and after a production change, as described in Example 2, are summarized below. The technological batch effect resulting from the production change is identified in the third principal component term (PC3) in Figure 6A. This technological batch effect can be removed by applying a correction factor to the RNA expression data obtained after the production change, as shown in Figure 6B. [Figure 6B] The results of a PCA dimensionality reduction analysis of approximately 100 paired cancer samples in which RNA expression was determined before and after a production change, as described in Example 2, are summarized below. The technological batch effect resulting from the production change is identified in the third principal component term (PC3) in Figure 6A. This technological batch effect can be removed by applying a correction factor to the RNA expression data obtained after the production change, as shown in Figure 6B. [Figure 7] Shown are UMAP embeddings of heme cancer transcriptome analysis from biological samples collected using the three methodologies. [Figure 8] 1 shows UMAP embedding of cancer transcriptome analysis from biological samples where RNA extraction was performed by the clinician or immediately prior to RNA sequencing. [Figure 9] PCA embedding (PC8) of transcriptome analysis of the same cancer sample using two batches of the same capture probes. Matched samples are indicated by lines connecting the respective PC terms. [Figure 10] UMAP embeddings of cancer transcriptome analyses performed on either single, 3-fold, or 6-fold sample pools, with matched sample groups indicated by circles. [Figure 11] UMAP embedding of cancer transcriptome analysis performed after 7–9 PCR amplification cycles following enrichment. Matched sample groups indicated by circles. [Figure 12] UMAP embedding of cancer transcriptome analysis performed with different sequencer loading molarities (0.7 uM, 1 uM, and 1.5 uM). Matched sample groups indicated by circles. [Figure 13] UMAP embedding of cancer transcriptome analyses performed with different sequencing reagent chemistries, with matched sample groups indicated by lines between points. [Figure 14] 1 shows an exemplary RNA expression profiling pipeline according to some embodiments of the present disclosure. [Figure 15] 1 illustrates an exemplary method for performing quality control on a batch of samples according to some embodiments of the present disclosure. [Figure 16] 1 shows an exemplary method for validating changes in a bioinformatics pipeline, e.g., an RNA expression pipeline, according to some embodiments of the present disclosure. [Figure 17] 1 illustrates an exemplary method for expanding a reference database according to some embodiments of the present disclosure.
[0017] In accordance with several embodiments of the present disclosure, like reference numerals refer to corresponding parts throughout the several views of the drawings. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0018] The disclosure herein provides an improved method of performing quality control analysis of a batch of RNA-seq samples (e.g., an entire flow cell equivalent). Quality control (QC) testing is performed on a batch of RNA-seq samples, including sequence reads and sample metadata. The methods described herein serve the purpose of ensuring sufficient data quality to perform automated RNA expression call reporting and detect batch effects that may affect clinical reporting. Such methods also help ensure consistency of data quality, which is important for comparing data over time. The quality control methods herein provide for automated review and analysis of an entire flow cell of samples.
[0019] advantage The present disclosure provides novel methods for evaluating technical batch effects in a set of transcriptome samples (e.g., flow cells) by pooling them with a set of validated reference samples (e.g., cohort-matched reference batches) that are matched for cancer type and tissue site. These methods improve upon the prior art in that they allow for global (e.g., performed on the entire set of samples in a batch) and single-sample quality control analysis simultaneously. These quality control methods benefit patients by providing rapid and accurate analysis of sample quality, thus providing improved and more timely patient diagnosis and treatment.
[0020] Laboratories performing high volume RNA sequencing must be very careful of technical batch effects when comparing samples over time, which is essential for the analysis of cancer transcriptomes and the determination of patient response to treatment. Changes in reagents, protocols, or techniques used in nucleic acid extraction, library preparation, and sequencing can alter the transcriptome in ways that invalidate or complicate the comparison of samples from different batches, necessitating continuous monitoring of sample quality and consistency. This monitoring can be particularly challenging when analyzing samples from different tissue sites, as tumor type is the primary biological determinant of transcriptomic variance in cancer. This means, for example, that brain and liver cancer samples are expected to be highly transcriptomically distinct, and their comparison would not be informative for the detection of batch effects. The fact that the methods herein provide cohort matching between the reference sample and the samples in each individual flow cell makes these quality control metrics more accurate than previous methods.
[0021] definition The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the invention. When used in the description of the invention and in the claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly dictates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated items listed. It will also be understood that as used herein, the terms "includes", "comprising", or any variation thereof, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. In addition, "including", "includes", "including ... To the extent that the terms "having," "has," "with," or variations thereof are used in either the detailed description and / or claims, such terms are intended to be inclusive in a similar manner as the term "comprising."
[0022] As used herein, the term "if" may be interpreted to mean "when" or "upon," or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (a stated condition or event) is detected" may be interpreted to mean "upon determining" or "in response to determining," or "upon detecting" or "in response to detecting," depending on the context.
[0023] It will also be understood that, although terms such as first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first subject can be referred to as a second subject, and similarly, a second subject can be referred to as a first subject, without departing from the scope of the present disclosure. A first subject and a second subject are both the same subject, but are not the same subject. Furthermore, the terms "subject", "user", and "patient" are used interchangeably herein.
[0024] As used herein, the term "subject" or "patient" refers to any living or non-living human (e.g., a male human, a female human, a fetus, a pregnant woman, a child, etc.). In some embodiments, the subject is male or female at any stage (e.g., a man, a woman, or a child).
[0025] As used herein, the terms "control", "control sample", "reference", "reference sample", "normal" and "normal sample" refer to a sample from a subject that does not have a particular condition or is otherwise healthy. In one example, the methods disclosed herein can be performed on a subject with a tumor, and the reference sample is a sample taken from the healthy tissue of the subject. In some embodiments, the reference sample can be obtained from the subject (e.g., to serve as a benchmark control for the subject from a particular time). In some embodiments, the reference sample can be obtained from a database. The reference can be, for example, a reference genome used to map sequence reads obtained from sequencing a sample from a subject. The reference genome can refer to a haploid or diploid genome to which sequence reads from a biological sample and a somatic sample can be aligned and compared. An example of a somatic sample can be the DNA of white blood cells obtained from a subject. For a haploid genome, there can be only one nucleotide at each locus. For a diploid genome, heterozygous loci can be identified, each heterozygous locus can have two alleles, and either allele can allow for a match for alignment to the locus.
[0026] As used herein, the term "locus" refers to a position (e.g., site) within a genome, e.g., on a particular chromosome. In some embodiments, a locus refers to a single nucleotide position within a genome, e.g., on a particular chromosome. In the context of morphology, a locus refers to a small group of nucleotide positions in a genome, for example, as defined by a mutation (e.g., a substitution, insertion, or deletion) of consecutive nucleotides in a cancer genome. Because normal mammalian cells have diploid genomes, a normal mammalian genome (e.g., a human genome) will generally have two copies of every locus in the genome, or at least two copies of every locus on an autosome, for example, one copy on a maternal autosome and one copy on a paternal autosome.
[0027] As used herein, the term "allele" refers to a particular sequence of one or more nucleotides at a chromosomal locus.
[0028] As used herein, the term "reference allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is either the predominant allele represented at that chromosomal locus within a population of a species (e.g., a "wild type" sequence), or an allele that is predefined within a reference genome for the species.
[0029] As used herein, the term "mutant allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is either not the predominant allele represented at that chromosomal locus within a population of a species (e.g., not a "wild type" sequence) or is not a predefined allele within a reference genome for the species.
[0030] As used herein, the term "single nucleotide variant", "SNV", "single nucleotide polymorphism", or "SNP" refers to the substitution of one nucleotide with a different nucleotide at a position (e.g., site) of a nucleotide sequence, e.g., a sequence read from an individual. A substitution of a first nucleobase X with a second nucleobase Y may be indicated as "X>Y". For example, a cytosine to thymine SNP may be indicated as "C>T". The term "het-SNP" refers to a heterozygous SNP, where the genome is at least diploid and at least one (but not all) of two or more homologous sequences exhibits a particular SNP. Similarly, a "hom-SNP" is a homologous SNP, where each homologous sequence of a polyploid genome has the same variant compared to a reference genome. As used herein, the term "structural variant" or "SV" refers to large (e.g., greater than 1 kb) regions of the genome that have undergone physical transformations such as inversions, insertions, deletions, or duplications (see, e.g., review of human genome SVs by Spielmann et al., 2018, Nat Rev Genetics 19:453-467).
[0031] As used herein, the term "indel" refers to an insertion and / or deletion event of a stretch of one or more nucleotides, either within a single locus or across multiple genes.
[0032] As used herein, the term "copy number variant", "CNV" or "copy number variation" refers to a region of genome that is repeated. These can be classified as short repeats or long repeats according to the number of nucleotides that are repeated in the genome region. Long repeats typically refer to cases where the whole gene or a large part of the gene is repeated one or more times.
[0033] As used herein, the term "mutation" refers to a detectable change in the genetic material of one or more cells. In certain instances, one or more mutations may be found in a cancer cell and may identify the cancer cell (e.g., driver and passenger mutations). Mutations may be transmitted from a parent cell to a daughter cell. One skilled in the art will appreciate that a genetic mutation in a parent cell (e.g., a driver mutation) may induce additional, different mutations (e.g., passenger mutations) in a daughter cell. It will be understood that the mutation generally occurs in nucleic acid. In certain instances, the mutation can be a detectable change in one or more deoxyribonucleic acids or fragments thereof. The mutation generally refers to a nucleotide that is added, deleted, substituted, inverted, or transposed to a new position in a nucleic acid. The mutation can be a naturally occurring mutation or an experimentally induced mutation. A mutation in a sequence of a specific tissue is an example of a "tissue-specific allele". For example, a tumor can have a mutation that results in an allele at a locus that does not occur in normal cells. Another example of a "tissue-specific allele" is a fetal-specific allele that occurs in fetal tissue but not in maternal tissue.
[0034] As used herein, the term "genomic variant" may refer to one or more mutations, copy number variants, indels, single nucleotide variants, or variant alleles. Genomic variants may also refer to a combination of one or more of the above.
[0035] As used herein, the terms "cancer," "cancerous tissue," or "tumor" refer to an abnormal mass of tissue in which the growth of the mass exceeds and is uncoordinated with that of normal tissue. In the case of hematological cancers, this includes a large volume of blood or other bodily fluids that contain cancer cells. Cancers or tumors can be defined as "benign" or "malignant" depending on the following characteristics: degree of cellular differentiation, including morphology and functionality, rate of growth, local invasion, and metastasis. "Benign" tumors can be well differentiated and are characterized by slower growth than malignant tumors, remaining localized at the primary site. In addition, in some cases, benign tumors do not have the ability to invade, invade, or metastasize to distant sites. "Malignant" tumors can be poorly differentiated (anaplastic) and have characteristically rapid growth with progressive invasion, infiltration, and destruction of surrounding tissue. In addition, malignant tumors can have the ability to metastasize to distant sites. Thus, cancer cells are cells found within an abnormal mass of tissue whose growth is uncoordinated with that of normal tissue. Thus, a "tumor sample" or "somatic biopsy," as described herein, refers to a biological sample obtained or derived from a subject's tumor.
[0036] As used herein, the term "somatic biopsy" refers to a biopsy of a subject. In some embodiments, the biopsy is of solid tissue. In some embodiments, it is a liquid biopsy.
[0037] As used herein, "sequencing," "sequence determination," and like terms used herein generally refer to any and all biochemical processes that can be used to determine the order of biological macromolecules, such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule, such as an mRNA transcript or a genomic locus.
[0038] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence produced by any sequencing process described herein or known in the art. A read can be generated from one end of a nucleic acid fragment (a "single-end read"), or in some cases, from both ends of a nucleic acid (e.g., paired-end reads, double-end reads). The length of a sequence read is often related to a particular sequencing technology. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, sequence reads are between about 15 bp and 900 bp in length (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). ) of the average, median, or average length of the sequence reads. In some embodiments, the sequence reads are of an average, median, or average length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. For example, nanopore sequencing may provide sequence reads that may vary in size from tens to hundreds to thousands of base pairs. Illumina parallel sequencing may provide sequence reads that do not vary much, e.g., most sequence reads may be less than 200 bp. A sequence read (or sequencing read) may refer to sequence information that corresponds to a nucleic acid molecule (e.g., a series of nucleotides). For example, a sequence read may correspond to a series of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, may correspond to a series of nucleotides at one or both ends of a nucleic acid fragment, or may correspond to the nucleotides of an entire nucleic acid fragment. Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or using probes, e.g., hybridization arrays or capture probes, or amplification techniques such as polymerase chain reaction (PCR), or linear amplification using a single primer, or isothermal amplification.
[0039] As used herein, the term "read segment" or "read" refers to any nucleotide sequence, including a sequence read obtained from an individual and / or a nucleotide sequence derived from an initial sequence read from a sample obtained from an individual. For example, a read segment can refer to an aligned sequence read, a folded sequence read, or a stitched read. Additionally, a read segment can refer to an individual nucleotide base, such as a single base mutation.
[0040] As used herein, the term "read depth", "sequencing depth", or "depth" refers to the total number of read segments from a sample obtained from an individual at a given location, region, or locus. A locus can be as small as a nucleotide, as large as a chromosome arm, or as large as an entire genome. Sequencing depth can be expressed as "Y-fold", e.g., 50-fold, 100-fold, etc., where "Y" refers to the number of times a locus is covered with sequence reads. In some embodiments, depth refers to the average sequencing depth across a genome, across an exome, across a transcriptome, or across a targeted sequencing panel. Sequencing depth can also be applied to multiple loci, the whole genome, where Y refers to the average number of times a locus or haploid genome, whole genome, whole transcriptome, or whole exome is sequenced, respectively. When an average depth is quoted, the actual depth for different loci included in the dataset can span over the range of values. Ultra-deep sequencing can refer to at least 100-fold in sequencing depth at a locus.
[0041] As used herein, the term "sequencing breadth" refers to how many percentages of a particular reference transcriptome (e.g., human reference exome), a particular reference genome (e.g., human reference genome), or a portion of a transcriptome or genome have been analyzed. The denominator of the percentage may be the repeat-masked genome, so 100% may correspond to the entire reference genome except for the masked portion. Repeat-masked transcriptome or genome may refer to a transcriptome or genome in which sequence repeats are masked (e.g., sequence reads align to unmasked portions of the transcriptome or genome). Any portion of the transcriptome or genome may be masked, and thus any particular portion of the reference exome or genome may be focused on. Broad sequencing may refer to sequencing and analyzing at least 0.1% of the reference transcriptome or genome.
[0042] As used herein, the term "reference transcriptome" refers to any set of sequences from any organism or pathogen that can be used to reference identified sequences from a subject. Reference transcriptome refers to any particular known, sequenced, or characterized transcriptome, whether partial or complete, of a tissue. Exemplary reference transcriptomes used for human subjects are provided in the online MiTranscriptome database described in Iyer et al. 2015 The landscape of long noncoding RNAs in the human transcriptome.Nat Genet 47, 199-208, the CHESS database described in Pertea et al.2018 CHESS:a new human gene catalog curated from thousands of large-scale RNA sequencing experiments reveals extensive transcriptional noise.Gen Biol 19:208, and the online ENCODE database hosted by the ENCODE project.
[0043] As used herein, the term "expression calling" refers to an RNA expression differential call (e.g., determining whether a particular sample from a subject exhibits higher or lower expression for a particular RNA compared to a reference transcriptome). In some embodiments, the expression calling is based at least in part on gene abundance counts.
[0044] As used herein, the term "reference exome" refers to any particular known, sequenced, or characterized exome, whether partial or complete, of any tissue from any organism or pathogen that can be used to reference an identified sequence from a subject. Exemplary reference exomes used for human subjects, as well as many other organisms, are provided in the online GENCODE database hosted by the GENCODE Consortium, e.g., Release 29 of the Human Exome Assembly (GRCh38.p12).
[0045] As used herein, the term "reference genome" refers to any particular known, sequenced, or characterized genome, whether partial or complete, of any organism or pathogen that can be used to reference identified sequences from a subject. Exemplary reference genomes used for human subjects, as well as many other organisms, include the National Institute of Health (NIH) ... The genome is provided in an online genome browser hosted by the Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz (UCSC). "Genome" refers to the complete genetic information of an organism or pathogen expressed in a nucleic acid sequence. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genome sequence from an individual or multiple individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. A reference genome may be considered a representative set of genes or gene sequences for a species. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).
[0046] As used herein, the term "assay" refers to a technique for determining a characteristic of a substance, e.g., a nucleic acid, a protein, a cell, a tissue, or an organ. An assay (e.g., a first assay or a second assay) can include a technique for determining the change in copy number of a nucleic acid in a sample, the methylation state of a nucleic acid in a sample, the fragment size distribution of a nucleic acid in a sample, the mutation state of a nucleic acid in a sample, or the fragmentation pattern of a nucleic acid in a sample. Any assay known to those skilled in the art can be used to detect any of the characteristics of nucleic acid described herein.Nucleic acid characteristics can include sequence, genome identity, copy number, methylation state at one or more nucleotide positions, nucleic acid size, the presence or absence of mutation in nucleic acid at one or more nucleotide positions, and nucleic acid fragmentation pattern (e.g., the nucleotide positions at which nucleic acid fragments).Assays or methods can have a certain sensitivity and / or specificity, and their relative usefulness as diagnostic tools can be measured using ROC-AUC statistics.
[0047] As used herein, the term "relative abundance" can refer to the ratio of a first amount of nucleic acid fragments having a particular characteristic (e.g., aligning to a particular region of the exome) to a second amount of nucleic acid fragments having a particular characteristic (e.g., aligning to a particular region of the exome).In one example, relative abundance can refer to the ratio of the number of mRNA transcripts that code for a particular gene in a sample (e.g., aligning to a particular region of the exome) to the total number of mRNA transcripts in a sample.
[0048] As used herein, two datasets are "balanced" with respect to features if the percentage of the number of samples with each type of feature in the two datasets is within a set percentage of each other. Unless otherwise specified, two datasets are balanced with respect to features if the percentage of the number of samples with each type of feature in the two datasets is within 10%. For example, if in a batch dataset, 15% of the samples are lung cancer samples, 25% of the samples are brain cancer samples, and 60% of the samples are colon cancer samples, and in a reference dataset, 5%-25% of the samples are lung cancer samples, 15%-35% of the samples are brain cancer samples, and 50%-70% of the samples are colon cancer samples, then the reference dataset is considered to be balanced with respect to the batch dataset. In some embodiments, two data sets are balanced in terms of features if the percentage of the number of samples with each type of feature in the two data sets is within 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, or 25% of each other. In general, the balance of the first feature and the balance of the second feature are considered to be independent of each other. However, in some embodiments, the balance of the first feature and the balance of the second feature are considered together. That is, in some embodiments, it is a combination of at least two features that are balanced. For example, as opposed to just balancing the percentages of brain cancer samples, samples taken from skin tissue, and samples taken from lung tissue, the percentage of brain cancer samples taken from skin tissue and the percentage of brain cancer samples taken from lung tissue in the batch dataset are balanced against the percentage of brain cancer samples taken from skin tissue and the percentage of brain cancer samples taken from lung tissue in the reference dataset.In some embodiments, rare features in the batch dataset will be unbalanced in the reference dataset, for example, because a sufficient number of reference samples that share the rare feature are not available.
[0049] Some aspects are described below with reference to illustrative applications. It should be understood that numerous specific details, relationships, and methods are shown to provide a complete understanding of the features described herein. However, those skilled in the art will readily recognize that the features described herein can be implemented without one or more of the specific details, or in other ways. The features described herein are not limited by the illustrated order of acts or events, since some acts may occur in different orders and / or simultaneously with other acts or events. Furthermore, not all illustrated acts or events are required to implement a methodology in accordance with the features described herein.
[0050] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0051] Example of System Embodiment Having provided an overview of some aspects of the disclosure and some definitions used in the disclosure, details of an example system will now be described in conjunction with Figure 1. Figure 1 is a block diagram illustrating a system 100 according to some implementations. The system 100 in some implementations includes one or more processing units CPU(s) 102 (also referred to as processors), one or more network interfaces 104, a user interface 106 including (optionally) a display 108 and an input system 110, a non-persistent memory 111, a persistent memory 112, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 optionally include circuitry (sometimes referred to as a chipset) that interconnects and controls communication between the system components. Non-persistent memory 111 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, while persistent memory 112 typically includes CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, magnetic disk storage device, optical disk storage device, flash memory device, or other non-volatile solid-state storage device. Persistent memory 112 optionally includes one or more storage devices located remotely from CPU 102. Persistent memory 112 and non-volatile memory devices in non-persistent memory 112 comprise non-transitory computer-readable storage media. In some implementations, non-persistent memory 111, or non-transitory computer-readable storage media, in some cases in combination with persistent memory 112, stores the following programs, modules, data structures, or a subset thereof: · an optional operating system 116 that handles various basic system services and contains instructions for performing hardware-dependent tasks; · an optional network communications module (or instructions) 118 for connecting the system 100 to other devices and / or the communications network 104; · a batch quality control module 120 for performing batch quality control of a batch of RNA sequencing samples; one or more batch datasets 122, each batch dataset including at least a corresponding plurality of sequence reads 126 and corresponding sample metadata 128 for each sample 124 in a plurality (e.g., batch) of samples, and also including, in each batch dataset, a corresponding cohort-matched reference batch 130; a reference sample dataset 140 storing one or more reference samples 142, each reference sample including at least a corresponding plurality of reference sample sequences 144 and corresponding reference sample metadata 146; · a global quality control dataset 150 for storing one or more batch quality control tests 152, the one or more batch quality control tests 152 being performed via the quality control module 120 on a plurality of samples (e.g., 124-1, ... 124-A) included in a batch dataset (e.g., batch dataset 122-1); a single-sample quality control module 121 for performing single-sample quality control of the RNA sequencing samples; and One or more single sample quality control tests 162 are performed on each individual sample 124 included in the batch data set 122 via the quality control module 120. A single sample quality control dataset 160 for storing a single sample quality control test 162 .
[0052] In various implementations, one or more of the above-identified elements are stored in one or more of the aforementioned memory devices and correspond to a set of instructions for implementing the functions described above. The above-identified modules, data, or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, data sets, or modules, and thus various subsets of these modules and data may be combined or otherwise reconfigured in various implementations. In some implementations, the non-persistent memory 111 optionally stores a subset of the above-identified modules and data structures. Additionally, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system other than the visualization system 100, which is addressable by the visualization system 100, such that the visualization system 100 may retrieve all or a portion of such data when needed.
[0053] While FIG. 1 illustrates a "system 100," this diagram is intended solely as a functional description of various features that may be present in a computer system, and not as a structural schematic of the implementations described herein. In practice, and as will be appreciated by those skilled in the art, items shown separately may be combined and some items may be separate. Additionally, while FIG. 1 illustrates certain data and modules in non-persistent memory 111, some or all of these data and modules may instead be stored in persistent memory 112 or more memories. For example, in some embodiments, at least one batch dataset 122 is stored in a remote storage device that may be part of a cloud-based infrastructure. In some embodiments, at least one of the batch datasets 122 is stored in a cloud-based infrastructure. In some embodiments, the batch dataset 122, the batch quality control module 120, the single sample quality control module 121, the reference dataset 140, the global QC dataset 150, and / or the single sample QC dataset 160 may also be stored in a remote storage device(s). In some embodiments, other configurations of data and module storage are utilized.
[0054] Batch and sample analysis Having disclosed details of system 100 according to the present disclosure, details regarding the processes and features of the systems according to various embodiments of the present disclosure are now disclosed below. In particular, exemplary processes are described below with reference to Figures 2A, 2B, 15, 16, and 17. In some embodiments, such processes and features of the systems are performed by modules 118, 120, and / or 121, as shown in Figure 1.
[0055] Block 202. Referring to block 202 of FIG. 2A, the method performs quality control. In some embodiments, quality control is performed on a single batch data set (e.g., including a single flow cell of an RNA sample). In some embodiments, quality control is performed simultaneously on two or more batch data sets (e.g., two or more flow cells of an RNA sample, where each flow cell is analyzed on the same or different days). In some embodiments, quality control is performed simultaneously on multiple batch data sets (e.g., multiple flow cells).
[0056] Referring to block 204, in some embodiments, the methods described herein are implemented in a computer system that includes a cloud server. That is, in some embodiments, the methods described herein are implemented entirely or entirely on a remote system. It may be partially implemented. For example, as described above, in some embodiments, one or more of the datasets are stored locally and at least one of the batch quality control module 120 and / or the single sample quality control module 121 are stored on a cloud server (e.g., in the cloud). In some embodiments, the reference dataset 140, the global quality control dataset 150, and / or the single sample quality control dataset 160 are also stored on the cloud server. In some embodiments, the required data (e.g., one or more batch datasets 122) may be transmitted between the local server and the cloud server.
[0057] Get initial batch dataset 122 information Block 206. Referring to block 206 of FIG. 2A, a batch dataset is obtained in electronic format (e.g., information for each sample in the batch dataset is stored in a .csv file). The batch dataset includes, for each respective sample in the plurality (e.g., batch) of samples, a corresponding plurality of sequence reads obtained from the respective sample by targeted panel or whole transcriptome sequencing. In some embodiments, each of the corresponding plurality of sequence reads is obtained from a plurality of RNA molecules or a derivative of the plurality of RNA molecules (e.g., a derivative such as cDNA). In some embodiments, each of the corresponding plurality of sequence reads is obtained by full transcriptome sequencing. In some embodiments, one or more of the corresponding plurality of sequence reads are derived from RNA isolated from a solid or hematological tumor (e.g., a solid biopsy). In some embodiments, one or more of the corresponding plurality of sequence reads are derived from a germline sample obtained from a respective subject.
[0058] In some embodiments, one or more corresponding multiple sequence reads are generated by next-generation sequencing. In some embodiments, one or more corresponding multiple sequence reads are generated from short-read paired end next-generation sequencing. In some embodiments, one or more corresponding multiple sequence reads are generated from short-read next-generation sequencing with one or more spike-in controls. In some embodiments, one or more spike-in controls calibrate the variation of sequence reads across a population of cells (e.g., the amount of RNA reads obtained from each cell can vary significantly, and the spikes help normalize the reads across a set of cells). In some embodiments, one or more corresponding multiple sequence reads are obtained by targeted panel sequencing using multiple probes.
[0059] Methods for mRNA sequencing are known in the art. In some embodiments, mRNA is reverse transcribed into cDNA prior to sequencing. For example, a method for RNA-seq for use according to block 210 is described in Nagalakshmi et al. al., 2008, Science 320, 1344-1349, and Finotell and Camillo, 2014, Briefings in Functional Genomics 14(2), 130-142, each of which is incorporated herein by reference. In some embodiments, mRNA sequencing is performed by whole exome sequencing (WES). In some embodiments, WES is performed by isolating RNA from a tissue sample, generating a cDNA library, optionally selecting desired sequences and / or depleting unwanted RNA molecules, and then sequencing the cDNA library, for example, using next-generation sequencing technology. For reviews of the use of whole exome sequencing technology in cancer diagnosis, see Serrati et al., 2016, Onco Targets Ther. 9, 7355-7365 and Cieslik, M. et al. 2015 Genome Res. 25, 1372-81, the contents of each of which are incorporated herein by reference for all purposes. The entirety of the present invention is incorporated herein by reference. In some embodiments, mRNA sequencing is performed by nanopore sequencing. A review of the use of nanopore sequencing technology in the human genome can be found in Jain et al., 2018, Nature 36(4), 338-345. This list is not exhaustive of the RNA sequencing methods that can be used according to the methods described herein. In some embodiments, RNA sequencing is performed according to one or more sequencing methods known in the art. See, for example, Kukurba et al. 2015 Cold Spring Harb Protoc. 11:951-969, a review of RNA sequencing methods.
[0060] Methods of next-generation sequencing for use in accordance with the methods described herein are disclosed in Shendure 2008 Nat.Biotechnology 26:1135-1145 and Fullwood et al. 2009 Genome Res.19:521-532, each of which is incorporated herein by reference.Next-generation sequencing methods well known in the art include synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single molecule real-time sequencing (Pacific Biosciences), sequencing by ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing.In some embodiments, massively parallel sequencing is performed using synthesis-by-sequencing with reversible dye terminators.
[0061] RNA-seq is a next-generation sequencing-based RNA profiling methodology that allows for the measurement and comparison of gene expression patterns across multiple subjects. In some embodiments, millions of short strings, called "sequence reads," are generated from sequencing random positions of cDNA prepared from input RNA obtained from a subject's tumor tissue. In some embodiments, RNA-seq gene expression data is generated from formalin-fixed paraffin-embedded tumor samples using an exome capture-based RNA-seq protocol. These reads can then be computationally mapped to a reference genome to reveal a "transcription map," and the number of sequence reads aligned to each gene provides a measure of its expression level (e.g., abundance). In some embodiments, RNA-seq expression levels (e.g., raw read counts) are normalized (e.g., to correct for GC content, sequencing depth, and / or gene length). In some embodiments, methods for mapping raw RNA-seq reads to the transcriptome and quantifying and normalizing gene counts are performed as described in U.S. patent application Ser. No. 62 / 735,349, filed Sep. 24, 2018, entitled "Methods of Normalizing and Correcting RNA Expression Data."
[0062] In some alternative embodiments, rather than using RNA-seq, microarrays are used to examine RNA profiling. Such microarrays are described in Wang et al., 2009, Nat Rev Genet 10, 57-63; Roy et al., 2011, Brief Funct Genomic 10:135-150; Shendure, 2008 Nat Methods 5, 585-587; Cloonan et al., 2008, "Stem cell transcriptome profiling via massive-scale mRNA sequencing," Nat. Methods 5, 613-619; Mortazavi et al., 2008, "Mapping and quantifying mammalian transcriptomes by RNA-Seq," Nat Methods 5, 621-628; and Bullard et al. al., 2010, “Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments” BMC Bioinformatics 11, p. 94, each of which is incorporated herein by reference.
[0063] The first computational step in the RNA-seq data analysis pipeline is read mapping, where reads are aligned to a reference genome or transcriptome by identifying genic regions that match the read sequence. Any of a variety of alignment tools can be used for this task. See, for example, Hatem et al., 2013 BMC Bioinformatics 14, 184, and Engstrom et al. 2013 Nat Methods 10, 1185-1191, each of which is incorporated herein by reference. In some embodiments, the mapping process begins by building an index of either the reference genome or the read, which is then used to search for a set of positions in the reference sequence where the read is likely to align. Once this subset of possible mapping positions is identified, alignment is performed on these candidate regions using slower, more sensitive algorithms. See, for example, Flicek and Birney, 2009, Nat Methods 6(Suppl.11), S6-S12, which is incorporated herein by reference. In some embodiments, the mapping tool is a methodology that utilizes pseudo-alignment (e.g., alignment of read sequences to transcripts rather than genomic locations). See, e.g., Bray et al. 2016 Near-optimal probabilistic RNA-seq quantification. Nat Biotech 34, 525-527, which is incorporated herein by reference.
[0064] After mapping, read counts are calculated using the reads aligned to each coding unit, such as exons, transcripts, or genes, to provide an estimate of its abundance (e.g., expression) level. In some embodiments, such counts take into account the total number of reads that overlap with the exons of a gene. However, in some instances, because some of the sequence reads map outside the boundaries of known exons, alternative embodiments take into account the full length of the gene and also count reads from introns. Furthermore, in some embodiments, spliced reads are used to model the abundance of different splicing isoforms of a gene. See, for example, Trapnell et al., 2010 Nat Biotechnol 28, 511-515, and Gatto et al., 2014 Nucleic Acids Res 42, p.e71, each of which is incorporated herein by reference.
[0065] As explained above, quantification of transcript abundance from RNA-seq data is typically implemented in an analysis pipeline through two computational steps: alignment of reads to a reference genome or transcriptome, and subsequent estimation of transcript and isoform abundance based on the aligned reads. Unfortunately, the reads generated by most used RNA-Seq techniques are generally much shorter than the transcripts from which they were sampled. As a result, it is not always possible to uniquely assign short sequence reads to a specific gene in the presence of transcripts with similar sequences. Such sequence reads are called "multiple reads" because they are homologous to two or more regions of the reference genome. In some embodiments, such multiple reads are discarded, i.e., they do not contribute to the gene abundance count. In some embodiments, to resolve ambiguities, programs such as MMSEQ or RSEM are used to See for examples of methodologies used to resolve multiple reads in Turro et al., 2011 Genome Biol 12, p.R13, and Nicolae et al., Algorithms Mol Biol 6, 9, each of which is incorporated herein by reference.
[0066] Another aspect of RNA-seq is the normalization of sequence read counts. In some embodiments, this includes normalization to take into account different sequencing depths. See, for example, Lin et al., 2011 Bioinformatics 27, 2031-2037; Robinson Oshlack, 2010 Genome Biol 11, R25; and Li et al., 2012 Biostatistics 13, 523-538, each of which is incorporated herein by reference. In some embodiments, sequence read counts are normalized to account for gene length bias. See, Finotell and Camillo, 2014 Briefings in Functional Genomics 14(2), 130-142, which is incorporated herein by reference.
[0067] In an embodiment in which one or more corresponding sequence reads are generated from targeted panel sequencing using a plurality of probes, each respective probe in the plurality of probes uniquely represents a different portion of the reference genome. In such an embodiment, each sequence read in the corresponding plurality of sequence reads corresponds to at least one probe in the plurality of probes.
[0068] Each respective probe in the plurality of probes uniquely targets a different (e.g., each) part of a reference transcriptome (e.g., a human reference transcriptome). Each sequence read in the second plurality of sequence reads and each sequence read in the third plurality of sequence reads correspond to at least one probe in the plurality of probes. In some embodiments, for example, whole genome sequencing is used instead of targeted panel sequencing.
[0069] In some embodiments, the second plurality of sequence reads has an average depth across the plurality of probes of at least 50 times. In some embodiments, the second plurality of sequence reads has an average depth across the plurality of probes of at least 400 times. In other embodiments, the second plurality of sequence reads has an average depth of at least 10 times, 15 times, 20 times, 25 times, 30 times, 40 times, 50 times, 75 times, 100 times, 150 times, 200 times, 250 times, 300 times, 400 times, 500 times, or more.
[0070] In some embodiments, the plurality of probes includes probes for at least 300 different genes. In some embodiments, the plurality of probes includes probes for at least 500 different genes. In still other embodiments, the plurality of probes includes at least 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 3000, 4000, 5000, or more different genes.
[0071] Block 208. Referring to block 208 of FIG. 2A, a cohort-matched reference batch is determined for the batch dataset. The cohort-matched reference batch is balanced for tissue site, tumor purity, cancer type, sequencer identity, or sequenced data. In some embodiments, the size of the cohort-matched reference batch (e.g., the number of samples therein) is the same as the size of the batch dataset. In some embodiments, the cohort-matched reference batch is of a different size than the batch dataset.
[0072] Determining a cohort-matched reference dataset 130 from the reference dataset 140 2A, block 210, in some embodiments, determining a cohort-matched reference dataset for the batch dataset includes, for each sample in the batch of samples, i) extracting a respective plurality of sequence features from a respective plurality of sequence reads, thereby obtaining a batch of sequence features, and ii) extracting a respective plurality of sample metadata features, thereby obtaining a batch of metadata features. In some embodiments, a cohort-matched reference dataset comprising a plurality of reference samples is selected from the reference dataset based at least in part on the batch of sequence features or the batch of metadata features.
[0073] In some embodiments, the cohort-matched reference dataset comprises a plurality of reference samples. In some embodiments, each reference sample in the plurality of reference samples comprises a corresponding plurality of sequence reads obtained from each reference sample by targeted or whole-transcriptome RNA sequencing, and corresponding metadata for each reference sample. In some embodiments, the reference sample for the cohort-matched reference dataset is selected based at least in part on sample features from sample metadata, such as RNA transcription profile (e.g., determined from a plurality of sequence reads), clinical data (e.g., patient diagnosis, treatment outcome, etc.), gender, biopsy type (e.g., heme vs. solid biopsy), and / or molecular data (e.g., genomic mutations).
[0074] In some embodiments, the patient diagnosis includes a cancer type and / or a cancer stage. In some embodiments, the respective cancer type for each sample in the batch dataset is selected from the set consisting of a tumor of a predefined stage of brain cancer, a predefined stage of glioblastoma, a predefined stage of prostate cancer, a predefined stage of pancreatic cancer, a predefined stage of renal cancer, a predefined stage of colorectal cancer, a predefined stage of ovarian cancer, a predefined stage of endometrial cancer, or a predefined stage of breast cancer.
[0075] In some embodiments, the type of biopsy includes a somatic biopsy. In some embodiments, the somatic biopsy includes a macro-dissected formalin-fixed paraffin-embedded (FFPE) tissue section, a surgical biopsy, a skin biopsy, a punch biopsy, a prostate biopsy, a bone biopsy, a bone marrow biopsy, a needle biopsy, a CT-guided biopsy, an ultrasound-guided biopsy, a fine needle aspiration, an aspiration biopsy, a fresh tissue or a blood sample. In some embodiments, the somatic biopsy is of a breast tumor, a glioblastoma, a prostate tumor, a pancreatic tumor, a kidney tumor, a colorectal tumor, an ovarian tumor, an endometrial tumor, a breast tumor, or a combination thereof. A biopsy is typically performed after one or more minimally invasive clinical tests suggest that a patient has or may have one or more tumors. The type of biopsy often depends on the location of the tumor. For example, a biopsy of a kidney tumor is frequently performed endoscopically, while a biopsy of an ovarian tumor frequently includes a tissue scraping.
[0076] In some embodiments, the genomic alteration comprises a copy number variant, a somatic mutation, a germline mutation, an indication of microsatellite instability, a tumor mutational burden, an indication of pathogen burden, or a tumor cellularity.
[0077] Examples of copy number variation are described in Shilien and Malkin 2009 Genome Med 1, 62. Signatures of microsatellite instability can be determined as described in Buhard et al. 2006 J Clinical Onco 24(2), 241. An example of determining tumor mutational burden is described in Chalmers et al 2017 Genome Med 9,34. Signs of pathogen load and / or signs of immune infiltration can be seen, for example, in Barber et al 201 5 PLoS Pathog 11(1):e1004558 and Pages et al 2010 Oncogene 29,1093-1102. In some cases, the indication of tumor cellularity is determined from a somatic biopsy by comparing some normal cells and some cancerous cells obtained in the somatic biopsy. In some cases, the indication of tumor cellularity is determined from one or more images of the somatic biopsy (e.g., by counting and identifying cancerous and non-cancerous cells).
[0078] In some embodiments, the cohort-matched reference dataset is balanced to correspond as closely as possible to the sample types present in the batch dataset (e.g., by selecting reference samples similar to each sample in the plurality of samples in the batch dataset). In some embodiments, the similarity between the reference sample and the batch dataset samples is determined based on at least one of the sample features described above. In some embodiments, the cohort-matched reference batch is selected from the reference database to include as many reference samples as possible (e.g., to maintain a reference batch that is balanced for tissue site, tumor purity, cancer type, sequencer identity, sequenced date, and / or sample features obtained from sample metadata).
[0079] In some embodiments, each sample in a first subset of samples in the batch dataset has a corresponding first biopsy type, and each sample in a second subset of samples in the batch dataset has a corresponding second biopsy type. In some embodiments, the first biopsy type or the second biopsy type comprises a somatic biopsy selected from the set comprising macro-dissected formalin-fixed paraffin-embedded (FFPE) tissue section, surgical biopsy, skin biopsy, punch biopsy, prostate biopsy, bone biopsy, bone marrow biopsy, needle biopsy, CT-guided biopsy, ultrasound-guided biopsy, fine needle aspiration, aspiration biopsy, fresh tissue, or blood sample.
[0080] In some embodiments, the first and second biopsy types are specified in the respective metadata for each sample in the batch dataset. To provide a balanced cohort-matched reference dataset, each reference sample in the first subset of reference samples in the cohort-matched reference dataset has a corresponding first biopsy type, and each reference sample in the second subset of reference samples in the cohort-matched reference dataset has a corresponding second biopsy type. For example, if the samples in the batch dataset include 50% of samples with breast cancer, 20% of samples with lung cancer, and 30% of samples with brain cancer, the cohort-matched reference dataset will incorporate the maximum number of reference samples from the reference dataset that match the percentages of these cancer types. In some embodiments, a similar method of balancing the cohort-matched reference dataset with the batch dataset is utilized for other sample features.
[0081] 2A, in some embodiments, a linear or non-linear combination of the sequence features of the batch and the metadata features of the batch is determined by subjecting the sequence features of the batch and the metadata features of the batch to a dimensionality reduction technique. In some embodiments, the dimensionality reduction technique includes Uniform Manifold Approximation and Projection (UMAP). In some embodiments, the dimensionality reduction technique includes Principal Component Analysis (PCA).
[0082] Global quality control testing performed on batch data set 122 Block 214. Referring to block 214 of FIG. 2B, one or more global quality control tests (e.g., test 152) are performed on the batch dataset using at least a cohort-matched reference dataset (e.g., batch). In some embodiments, the one or more global batch quality control tests include tests for one or more batch effects from a set that includes bioinformatics pipeline analysis and sequencing methods. nothing.
[0083] Referring to block 216 of FIG. 2B, in some embodiments, the one or more global batch quality control tests include tests for one or more batch effects from a set including bioinformatics pipeline analysis (e.g., date of sample analysis, sequencer identity, type of pipeline, etc.), DNA contamination, sample processing (e.g., sample collection method, reagent changes, etc.), and sequencing method (e.g., UMI vs. UDI sequence adapters).
[0084] In some embodiments, there are different bioinformatics pipelines used to analyze samples (e.g., based on type of biopsy, acellular nucleic acid sample), and the use of different pipelines can contribute to batch effects. For example, in some embodiments, even the type of test tube used for blood sample collection (e.g., PAX vs. EDTA) is considered for possible impact on batch effects. In some embodiments, changes to equipment (e.g., sequencer or flow cell) can also contribute to batch effects. In some embodiments, reagents with potential batch effects include probe lots, controls (e.g., Horizon controls), and buffers.
[0085] Referring to block 218 of FIG. 2B, in some embodiments, the cohort-matched reference batch is used to adjust each sample in the batch dataset for one or more confounding covariates before performing one or more global batch quality control tests. In some embodiments, this adjustment involves normalizing the expression levels for each gene in the reference genome for each respective plurality (e.g., a reference genome that each sample in the batch dataset shares with each reference sample in the cohort-matched reference batch) to the sequence reads for each sample in the batch dataset. In some embodiments, Mostafavi 2013 contains a summary of relevant normalization methods. See PLOS ONE, e68141, in the section titled "Unified Representation of Existing Normalization Methods."
[0086] In some embodiments, at least one sample in the batch dataset is a control sample (e.g., a Horizon control sample). In some embodiments, at least one control sample in the batch dataset is used to adjust each other sample in the batch dataset. A Horizon control sample is a commercially available control derived from a cell line containing a known fusion variant. In some embodiments, the expression of the fusion variant is expected to be constant regardless of experimental conditions (e.g., sequencer identity, sequencing method, sequencing date, etc.). These Horizon controls are useful for normalizing samples between and across batch datasets, and also for providing information about sequencing trends over time (e.g., by comparing Horizon controls evaluated at different time points to each other). In some embodiments, any commercially available control sample can be used in the methods described herein.
[0087] In some embodiments, each global batch quality control test includes: i) determining the average number of sequence reads per sample across the batch dataset, ii) obtaining a reference average number of sequence reads per sample from a reference dataset (or, for example, from a cohort-matched reference batch), and iii) comparing the average number of sequence reads across the batch dataset to the reference average number of sequence reads per sample. In some embodiments, if the average number of sequence reads is less than the reference average number of sequence reads per sample, the batch dataset fails the respective global batch quality control test.
[0088] In some embodiments, each global batch quality control test: i) measures the average percentage of mapped sequence reads per sample across the batch dataset; determining a reference average percentage of mapped sequence reads per sample from a reference dataset (or, for example, from a cohort-matched reference batch); and iii) comparing the average percentage of mapped sequence reads across the batch dataset to the reference average percentage of mapped sequence reads per sample. In some embodiments, if the average percentage of mapped sequence reads falls below the reference average percent of mapped sequence reads per sample, the batch dataset fails the respective global batch quality control test.
[0089] In some embodiments, the respective metadata for each sample in the batch dataset includes a respective cancer type. In some such embodiments, the respective global batch quality control test includes applying, for each respective sample in the batch dataset, the corresponding plurality of sequence reads and the corresponding metadata to a second trained classification model, whereby the second trained classification model provides a respective predicted cancer type for each sample. In some embodiments, the respective global batch quality control test further includes comparing, for each sample, the respective cancer type from the respective metadata to the respective predicted cancer type. In some embodiments, one or more samples having a respective predicted cancer type that does not match a respective known cancer type fail the global batch quality control test. In some embodiments, if one or more samples in the batch dataset fail the global batch quality control test, the entire batch dataset fails the respective global batch quality control test. In some embodiments, the second trained classification method includes any of the classification methods described in U.S. Provisional Patent Application No. 62 / 855,750, entitled "Systems and Methods for Multi-Label Cancer Classification," filed May 31, 2019.
[0090] In some embodiments, the respective metadata for each sample in the batch dataset includes a respective cancer type, and in some embodiments, the respective global batch quality control test includes determining a respective tumor purity percentage for each respective sample in the batch dataset. In some embodiments, the tumor purity is determined based at least in part on a variant allele fraction, and in some embodiments, the variant allele fraction is determined based at least in part on a variant allele fraction as described in Shin et al. 2017 “Prevalence and detection of low-allele-fraction variants in clinical cancer samples” Nat Comm 8, 1377. In some embodiments, a respective sample fails a respective global batch quality control test if the respective sample has a corresponding tumor purity of less than 20%, less than 30%, less than 40%, or less than 50%. In some embodiments, a batch data set fails a global batch quality control test if at least 30%, at least 40%, at least 50%, or at least 60% of the samples in the batch data set fail a respective global batch quality control test.
[0091] Block 220. Referring to block 220 of FIG. 2B, each sample from the batch data set that fails any one of the one or more global quality control tests 152 is removed from the batch data set or flagged for manual inspection. In some embodiments, the removing step further includes providing an updated batch data set lacking each of the respective samples that failed any one of the one or more global quality control tests.
[0092] Sample Report Referring to block 222, in some embodiments, a respective sample report is provided for each sample in the batch of samples (even those that failed the global batch quality control tests in some embodiments). In alternative embodiments, a sample report is provided only for samples that did not fail any one of the global batch quality control tests (e.g., samples included in the updated batch data set). In some embodiments, each respective sample report includes at least one of a set of expression calls, one or more matched therapies, or one or more matched clinical trials. In some embodiments, the appropriate matched therapy is determined based on the expression calls and the cancer type information. In some embodiments, the appropriate matched therapy is determined based on the organoid test. Examples of organoid tests and correlations between organoid test results and therapy sensitivity are provided in U.S. Provisional Patent Application No. 62 / 924,621, filed October 22, 2019, entitled "Systems and Methods for Predicting Therapeutic Sensitivity." In some embodiments, the appropriate matched clinical trial is determined based at least in part on the corresponding expression calls for each sample.
[0093] In some embodiments, the sample report may further include a summary that provides the patient and / or healthcare provider with a concise summary of the most important findings from the full sample report. In some embodiments, a sample (e.g., patient) report is provided as described in U.S. Provisional Patent Application No. 62 / 855,750, entitled "Systems and Methods for Multi-Label Cancer Classification," filed May 31, 2019.
[0094] In some embodiments, each sample in a batch of samples is further associated with corresponding clinical data after quality control analysis is performed. In some embodiments, the association between the RNA-seq samples and clinical data is used to verify or refine expression calls. In some embodiments, the clinical data includes DNA mutations, patient response to therapy, organoid experimental results (e.g., organoids obtained from a patient can be tested to determine whether the organoids are sensitive to a matched therapy), and / or histopathological images. Examples of histopathological images include H&E (hematoxylin and eosin) and IHC (immunohistochemistry) staining images.
[0095] Conducting single sample quality control testing Block 230. Referring to block 230 of FIG. 2B, in some embodiments, for each respective sample in a batch of samples, one or more single-sample quality control tests are performed on the respective sample from the corresponding plurality of sequence reads. Each sample from the batch of samples that fails any one of the one or more single-sample quality control tests is removed from the batch data set or flagged for manual inspection. In some embodiments, any single-sample quality control test can be applied to the entire batch data set. In some embodiments, a single-sample QC test is performed before a batch QC test. In some embodiments, a single-sample QC test is performed after a batch QC test.
[0096] In some embodiments, each single sample quality control test compares the total number of sequence reads per sample in the batch of samples to the average number of sequence reads per reference sample in the entire reference dataset 140. In other words, in some embodiments, each single sample quality control test includes i) determining the average number of sequence reads per sample across the batch dataset, and ii) obtaining a reference average number of sequence reads per sample from the reference dataset (or, for example, from a cohort-matched reference batch, or from a subset of the reference dataset). Each single sample dataset is a set of samples that are generated from the batch dataset. The average number of sequence reads across the datum is compared to a reference average number of sequence reads per sample. In some embodiments, if the respective total number of sequence reads for each sample is below the reference average number of sequence reads per sample, then the respective sample fails the respective quality control test.
[0097] In some embodiments, the Wilcoxon test is used to evaluate whether the batch data set fails each single sample quality control test. In some embodiments, the two-sample Wilcoxon test compares paired groups (e.g., compares the most similar paired samples between the two groups). For example, this is useful for directly comparing samples from the same set of subjects sequenced on both HiSeq1 and HiSeq2 systems. In some embodiments, an unpaired Wilcoxon test is used (e.g., when there are some differences in the batches being compared). In some embodiments, a modified p-value threshold is used to determine whether there is a significant difference (e.g., batch effect). In some embodiments, the modified p-value threshold is determined based on at least a plurality of high-quality reference samples from the reference data set 140. In some embodiments, a reference sample is determined to be of high quality if the corresponding read count of the reference sample is at least 5 million sequence reads, at least 10 million sequence reads, at least 20 million sequence reads, at least 30 million sequence reads, at least 40 million sequence reads, at least 50 million sequence reads, at least 100 million sequence reads, or at least 200 million sequence reads.
[0098] In some embodiments, each single sample quality control test includes applying the corresponding plurality of sequence reads and corresponding sample metadata for each respective sample in the batch dataset to a first trained classification model, whereby the first trained classification model provides a set of predicted gender assignments, including a respective predicted gender assignment for each sample. In some embodiments, each single sample quality control test further includes comparing the set of predicted gender assignments to an expected set of gender assignments (e.g., to detect unintended sample exchanges). In some embodiments, one or more samples having a respective predicted gender assignment that does not match the expected set of gender assignments (e.g., the proportion of one gender is too high) fail the respective single sample quality control test. In some embodiments, if the set of predicted gender assignments does not match the expected set of gender assignments (e.g., if there appears to be too much sample exchange), the entire batch dataset fails the respective single sample quality control test.
[0099] In some embodiments, by way of non-limiting example, the first classification model comprises a decision tree. Decision tree algorithms suitable for use as the classifier in block 244 are described, for example, in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc. & Sons, Inc., New York, pp. 395-396, which is incorporated herein by reference. In some embodiments, the decision tree is a random forest regression. One particular algorithm that may be used as the classifier of block 244 is a classification and regression tree (CART). Other examples of particular decision tree algorithms that may be used as the classifier of block 244 include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York. pp. 396-408 and pp. 411-412, which is incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference. Random forests are described in Breiman, 1999, "Random Forests--Random Features," Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is incorporated by reference in its entirety.
[0100] In some embodiments, each single-sample quality control test in the one or more single-sample quality control tests includes: i) determining a respective number of non-redundant mapped sequence reads (e.g., reads that are not the result of PCR duplication) in the multiple sequence reads for each respective sample in the batch of samples; and ii) comparing the respective number of non-redundant mapped sequence reads with an expected number of non-redundant mapped sequence reads. Each non-redundant mapped sequence read is mapped to a corresponding portion of a reference genome (e.g., has a unique start and end site in the reference genome). In some embodiments, if the respective number of non-redundant mapped sequence reads is below a predetermined number of non-redundant mapped reads, the respective sample fails the respective single-sample quality control test.
[0101] In some embodiments, the expected number of non-overlapping mapped sequence reads is predicted based on whether the sample was obtained from a solid or liquid biopsy of the subject (e.g., from a solid tumor or from a blood sample). In some embodiments, the respective number of overlapping reads is determined for each sample in the batch dataset (e.g., by identifying sequence reads with identical start and end sites). In some embodiments, the method further provides a graphical display of the respective number of overlapping reads as part of the sample report or global batch report.
[0102] In some embodiments, each single-sample quality control test includes determining a respective quality score for each base pair read position in the corresponding plurality of sequence reads of each respective sample in the batch data set. In some embodiments, if the respective quality scores of one or more of the one or more respective base pair read positions are below a threshold quality score, the respective sample fails the respective single-sample quality control test. In some embodiments, the threshold quality score includes 20.0 (e.g., as calculated by FastQC). In some embodiments, a respective quality score is determined for each sequence read in the plurality of sequence reads for each respective sample in the batch data set, and in some such embodiments, one or more sequence reads having a corresponding quality score below the threshold read quality score are discarded.
[0103] In some embodiments, each single-sample quality control test comprises determining an average quality score for each sequence read in the corresponding plurality of sequence reads of each respective sample in the batch data set. In some embodiments, if the average for the average quality score across the corresponding plurality of sequence reads is below a threshold quality score, each sample fails each single-sample quality control test. In some embodiments, the threshold quality score comprises 20.0 (e.g., as calculated by FastQC).
[0104] In some embodiments, the respective single-sample quality control test includes determining, for each respective sample in the batch of samples, a respective percentage of properly paired sequence reads (e.g., a percentage of properly paired sequence reads that are paired-end reads), where if the percentage of properly paired sequence reads is below a predetermined paired read threshold, the respective sample fails the respective single-sample quality control test. In some embodiments, the predetermined paired read threshold Threshold values include at least 90%, at least 95%, or at least 99%.
[0105] In some embodiments, each single sample quality control test includes determining a respective number of expressed genes (e.g., the number of genes with non-zero supporting sequence reads) for each respective sample in the batch of samples. In some embodiments, each sample fails each single sample quality control test if the corresponding expression read score is below the predetermined number of expressed reads. In some embodiments, when each sample is obtained from a solid biopsy, the predetermined number of expressed genes is at least 18,000, at least 19,000, or at least 20,000 genes. In some embodiments, when each sample is obtained from a liquid (e.g., hematological) biopsy, the predetermined number of expressed genes is at least 15,000, at least 16,500, or at least 17,000 genes. Some cancer types include different sets of expressed genes (e.g., some cancer types are transcriptionally distinct). For example, Li See, e.g., Wang et al. 2017 "Transcriptional landscape of human cancers" Oncotarget 8(21), 34534-34551. In some embodiments, the predetermined number of expressed genes is determined based at least in part on the cancer type of each sample (e.g., from corresponding metadata).
[0106] In some embodiments, each single-sample quality control test includes determining the GC content of each of the corresponding plurality of sequence reads for each sample in the batch of samples. In some embodiments, each sample fails each single-sample quality control test if the GC content is outside a predetermined GC content threshold. In some embodiments, the predetermined GC content threshold includes 35-60%, 40-60%, 45-60%, 50-60%, or 55-60%. GC content varies widely among genes in the human genome. See, for example, Versteeg et al. 2003 "The Human Transcriptome Map Reveals Extremes in Gene Density, Intron Length, GC content, and Repeat Patterns for Domains of Highly and Weekly Expressed Genes" Genome Res 13(9),1998-2004. GC content can affect how well a nucleic acid molecule is amplified during PCR. See, e.g., Mammedov et al. 2009 "A Fundamental Study of the PCR Amplification of GC-Rich DNA Templates" Comput Biol Chem 32(6), 452-457.
[0107] In some embodiments, each single-sample quality control test comprises determining, for each respective sample in the batch of samples, a content analysis for each respective base sequence across the corresponding plurality of sequence reads of each respective sample. In some embodiments, if the distribution of A, T, C or G content drifts by a percentage greater than a threshold across the base positions collectively represented by the corresponding plurality of sequence reads for each respective sample, each sample fails each single-sample quality control test. In some embodiments, the threshold percentage comprises at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, or at least 10% drift as determined from FastQC.
[0108] In some embodiments, each single sample quality control test includes determining, for each respective sample in the batch of samples, a respective per-base GC content profile across the corresponding plurality of sequence reads. Each sample fails each single sample quality control test if there is a higher percentage drift across base positions collectively represented by the corresponding sequence reads. In some embodiments, the threshold percentage includes a drift of more than 5%, more than 6%, more than 7%, more than 8%, more than 9%, or more than 10%.
[0109] In some embodiments, each single-sample quality control test includes determining, for each respective sample in the batch of samples, a corresponding distribution of read GC content per sequence across the corresponding plurality of sequence reads. In some embodiments, if the goodness-of-fit test determines that the respective distribution of read GC content per sequence deviates from a normal distribution at a threshold significance level (e.g., by analysis with a chi-squared goodness-of-fit test at a significance level of 0.05), then the respective sample fails the respective single-sample quality control test.
[0110] In some embodiments, each single-sample quality control test includes determining, for each respective sample in the batch of samples, a content analysis for each corresponding missing base across the corresponding plurality of sequence reads that determines a percentage of missing calls for each base position represented by the corresponding plurality of sequence reads. In some embodiments, if the corresponding percentage of missing calls for the base positions represented by the corresponding plurality of sequence reads exceeds a threshold percentage, the respective sample fails the respective single-sample quality control test. In some embodiments, the threshold percentage includes more than 10%, more than 15%, more than 20%, or more than 25%.
[0111] In some embodiments, each single-sample quality control test includes a sequence read length distribution analysis that determines, for each respective sample in the batch dataset, a respective range of sequence read lengths across the corresponding plurality of sequence reads. In some embodiments, if the respective range of sequence read lengths deviates from the sequence read expectation (e.g., fixed length sequence reads are flagged / removed, a distribution of sequence read lengths is observed, a range of observed sequence reads is flagged / removed, if the distribution does not meet the expectation for that distribution), then each sample fails the respective single-sample quality control test.
[0112] In some embodiments, each single-sample quality control test includes an over-represented sequence analysis that determines whether any sequence reads in the corresponding plurality of sequence reads are over-represented for each respective sample in the batch data set. In some embodiments, if the over-represented sequence analysis identifies one or sequence read sequences that are represented by a higher percentage than a threshold value of the corresponding plurality of sequence reads, each sample fails each single-sample quality control test. In some such embodiments, the threshold percentage includes at least 0.05%, at least 0.10%, at least 0.15%, or at least 0.2%.
[0113] Segmenting modules into containers Block 240. Referring to block 240 of FIG. 2B, in some embodiments, one or more batch quality control tests (e.g., detection of whole batch outliers) include a first module (e.g., module 120) and one or more single sample quality control tests (e.g., detection of single sample outliers) include a second module (e.g., module 121). In some embodiments, each of the first module and the second module includes a respective docker (e.g., a computational container that enables the performance of the batch quality control tests and the single sample quality control tests regardless of the operating system 116). In some embodiments, the first module and the second module are implemented on the same computer system. In some embodiments, the first module and the second module are implemented on different computer systems.
[0114] Examples of Docker (also written as "containers" or "Docker containers") are provided by Boettiger 2015 "An introduction to Docker for reproducible research,with examples from the R environment" arXiv:1410.0846v1, and Felter et al. 2014 "An updated performance comparison of virtual machines and linux containers" at IEEE International Sympo. Docker containers are often useful to facilitate workflows and may allow multiple applications to be used in concert. See, for example, Di Tommaso et al 2015 "The impact of Docker containers on the performance of genomic pipelines" PeerJ 3:e1273. In some embodiments, the use of more than one Docker (e.g., modules or containers) provides flexibility (e.g., ease of application regardless of the type of operating system available) to the implementation of the methods described herein.
[0115] In some embodiments, the first module 120 (e.g., a batch quality control module) tests the global transcriptomic quality of a batch of RNA samples (e.g., multiple samples from an entire flow cell or set of flow cells). In some embodiments, the first module 120 evaluates the global transcriptomic quality of the batch of RNA samples against a balanced set of samples from a reference (e.g., a cohort-matched reference batch).
[0116] In some embodiments, the input to the first module 120 includes, for each sample in the batch of samples, i) a corresponding plurality of sequence reads, and ii) corresponding metadata (e.g., including at least one or more bioinformatics values). In some embodiments, each of the plurality of sequence reads is normalized (e.g., as described above with respect to block 218). In some embodiments, the first module 120 includes or has access to reference data (e.g., a reference dataset 140). In some embodiments, the reference dataset 140 includes a plurality of reference samples 142, including a corresponding plurality of sequence reads 144 and corresponding reference metadata 146 for each reference sample in the plurality of reference samples.
[0117] In some embodiments, the corresponding plurality of sequence reads 144 comprises a .csv file or a .parquet file. In some embodiments, the corresponding plurality of sequence reads 144 comprises any file format known in the art. In some embodiments, the bioinformatics values included in the sample metadata comprise LIMS (e.g., laboratory information management system) values.
[0118] In some embodiments, the global batch quality control tests evaluated by the first module include statistical tests and / or dimensionality reduction. In some embodiments, these statistical tests include any method for distinguishing between the batch dataset 122 and the corresponding cohort-matched reference batch 130. In some embodiments, the statistical batch quality control tests are performed on a subset of the batch dataset and the corresponding cohort-matched reference batch (e.g., only samples of a particular cancer type are compared).
[0119] In some embodiments, the evaluation of one or more batch quality tests over time includes a third module, which in some embodiments includes a respective third docker. In some embodiments, the third module includes a third docker over time (e.g., This is useful both to assess trends across batches of RNA-seq (e.g., at multiple time points) and to ensure the stability of the sequencing method (e.g., by assessing whether control samples are similar across multiple time points).
[0120] In some embodiments, the method further provides a global batch quality control report (e.g., following application of module 120). In some embodiments, the global batch report includes at least i) a list of one or more samples from the batch dataset that failed any one of the one or more bath quality control tests, and ii) a list of one or more reference samples from the reference dataset 140 (e.g., identified from corresponding reference sample metadata) that were evaluated within a predefined period of time.
[0121] In some embodiments, the predefined period includes at least 1 day, at least 2 days, at least 3 days, at least 4 days, at least 5 days, at least 6 days, at least 7 days, at least 10 days, at least 14 days, at least 21 days, at least 28 days, or at least 30 days. In some embodiments, a global batch report is provided by a third module. In some embodiments, each global batch report is provided at least daily, at least weekly, at least biweekly, at least monthly, at least three months, or at least annually.
[0122] In some embodiments, the quality metrics assessed by the third module include one or more metrics from the set of at least GC content, contamination level (e.g., DNA contamination due to improper DNase application, among others, to the biological sample), number of reads, percentage of mapped reads, gene duplication rate, number of genes expressed as a function of number of reads, number of transcript completeness, accuracy of tumor of unknown cause determination, or accuracy of gender prediction.
[0123] In some embodiments, the third module further provides a graphical representation of the results of the evaluation of the one or more metrics compared to one or more of a set including time, pipeline version, sequencer type, flow cell, or cancer type. In some embodiments, these graphical representations include any graphical representations described elsewhere herein or known in the art.
[0124] In some embodiments, the method further provides one or more graphical representations of the overall characteristics of the batch dataset. In some embodiments, each graphical representation includes detailed information about the corresponding batch dataset characteristic. In some embodiments, the method provides one or more graphical representations of the results of one or more global batch quality control tests performed according to embodiments described herein. In some embodiments, the method provides one or more graphical representations of the results of one or more single sample quality control tests performed according to embodiments described herein. In some embodiments, the batch dataset characteristic includes a combination of respective metadata features for each respective sample in a batch of samples of the batch dataset. For example, in some embodiments, the metadata features of each sample in the batch dataset are combined to provide an overall metric (e.g., characteristic) for the batch dataset.
[0125] Identifying technical batch effects in an RNA expression pipeline Another aspect of the disclosure provides a method of performing quality control in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, the method comprising: determining, for each respective test sample in a batch of test samples, a pairwise comparison for each respective gene in a first set of genes; The method includes obtaining a batch dataset in electronic format, the batch dataset including a corresponding expression profile including corresponding gene expression values, and a corresponding set of metadata including a value for each respective feature in the first set of features for the sample.
[0126] The method includes determining a cohort-matched reference dataset for the batch dataset, the cohort-matched reference dataset including a corresponding expression profile for each respective reference sample in the plurality of reference samples, the corresponding expression profile including a corresponding gene expression value for each respective gene in the first set of genes. Each respective reference sample in the plurality of reference samples is associated with a corresponding set of metadata including a corresponding value for each respective feature in the second set of features for the respective reference sample. The aggregated values for each respective feature in the third set of one or more features present in both the first set of features and the second set of features are balanced between the batch dataset and the cohort-matched reference dataset.
[0127] The dimensionality reduction is performed on the combined data set consisting of the corresponding expression profile for each respective test sample in the plurality of test samples and the corresponding expression profile for each respective reference sample in the plurality of reference samples. Thus, for each respective test sample and each respective reference sample, a corresponding set of coordinates is obtained that is embedded in a dimensional space lower than the dimensionality of the corresponding expression profile.
[0128] The method further includes determining a statistical measure of similarity between the set of coordinates acquired for the test sample and the set of coordinates acquired for the reference sample. The statistical measure of similarity is compared to a threshold, and the batch data set is validated for reporting if the statistical measure of similarity meets the threshold, or is not validated if the statistical measure of similarity does not meet the threshold.
[0129] For example, FIG. 15 illustrates a method of performing quality control (e.g., for a batch of samples compared to a reference dataset) according to some embodiments of the present disclosure. The batch dataset is obtained from a batch of test samples (e.g., "batch i") that includes a plurality of RNA samples from a plurality of N patients 1502 (e.g., 1502-1, 1502-2, ..., 1502-N). In some embodiments, with reference to block 1504, the batch dataset is obtained using sequencing analysis (e.g., RNAseq exome analysis) of the test samples. The batch dataset includes an expression profile for each test sample, where each expression profile includes expression data for each respective patient 1506 in the plurality of N patients (e.g., 1506-1, 1506-2, ..., 1506-N). In some embodiments, the expression data includes corresponding gene expression values for each respective gene in the plurality of genes in the test samples (e.g., sequenced by RNAseq). In some embodiments, the expression data further includes metadata indicating values for multiple features associated with the test sample, such as an RNA transcription profile (e.g., determined from multiple sequence reads), clinical data (e.g., patient diagnosis, treatment outcome, etc.), gender, biopsy type (e.g., heme vs. solid biopsy), molecular data (e.g., genomic mutations), and / or other features (e.g., tissue site, tumor purity, cancer type, collection method, sequencer identity, and / or date sequenced).
[0130] In some embodiments, the expression data in each respective expression profile is normalized, referring to block 1508. In some embodiments, the normalization of the expression data in each expression profile generates a plurality of normalized data sets 1510 (e.g., 1510-1, 1510-2, ..., 1510-N).
[0131] In some embodiments, the method further includes obtaining a cohort-matched reference dataset 1512 for the plurality of reference samples. In some embodiments, the cohort-matched reference dataset is identified by matching a proportion of one or more features 1509 of the samples in the batch dataset with a reference sample having the same proportion of those one or more features (e.g., tissue site, tumor purity, cancer type, collection method, sequencer identity, sequenced date, clinical data (e.g., patient diagnosis, treatment outcome, etc.), gender, biopsy type (e.g., heme vs. solid biopsy), molecular data (e.g., genomic mutations), and / or other features) in, for example, the reference database 1511. The cohort-matched reference dataset includes an expression profile for each reference sample in the plurality of reference samples. Each respective expression profile includes a corresponding gene expression value for each respective gene in the plurality of genes (e.g., where the plurality of genes included in each reference sample expression profile is the same as the plurality of genes included in each test sample expression profile).
[0132] In some embodiments, the plurality of reference samples includes the same number of reference samples as the number of test samples in the batch of test samples (e.g., N reference samples). In some embodiments, the plurality of reference samples are selected for the batch of test samples based on one or more similarities between features (e.g., tissue site, tumor purity, cancer type, collection method, sequencer identity, and / or sequenced date) associated with the reference sample and the test sample. In some embodiments, the metadata (e.g., values for features) between the batch of test samples and the plurality of reference samples are balanced such that the distribution (e.g., percentage) of the features of the test samples among the plurality of test samples in the batch of test samples is similar to the distribution (e.g., percentage) of the features of the reference samples among the plurality of reference samples.
[0133] In some embodiments, the cohort-matched reference dataset is normalized.
[0134] According to this method, referring to block 1514, a dimensionality reduction (e.g., PCA, latent component analysis, partial least squares regression, etc.) is performed on a combined dataset consisting of a corresponding expression profile for each respective test sample in the plurality of test samples and a corresponding expression profile for each respective reference sample in the plurality of reference samples. For each respective test sample P and each respective reference sample M, the corresponding set of coordinates 1516-1518 are embedded into a lower dimensional space (e.g., m-space) than the dimensionality of the corresponding expression profile (e.g., 1516-1, 1516-2, ..., 1516-N, and 1518-1, 1518-2, ..., 1518-N).
[0135] Referring to block 1520, the method further includes evaluating the component values using the combined dataset after dimensionality reduction. The method includes determining a statistical measure of similarity between the set of coordinates acquired for the test sample and the set of coordinates acquired for the reference sample, and comparing the statistical measure of similarity to a threshold. Referring to block 1522, the batch dataset is verified for reporting if the statistical measure of similarity meets the threshold, and is not verified if the statistical measure of similarity does not meet the threshold. In some embodiments, if the statistical measure of similarity does not meet the threshold, the batch dataset is flagged for rejection and / or further evaluation of batch effects. In some embodiments, further analysis identifies individual samples or groups of individual samples in the batch dataset that drive dissimilarity with the reference sample. In some embodiments, one or more of these samples that drive the statistical difference are removed from the batch dataset, and global quality control tests are re-run on the modified batch dataset (e.g., with the individual samples that contribute to the statistical dissimilarity removed), and the modified dataset is verified if it passes the batch quality control tests. In some embodiments, once one or more samples are identified as contributing to the identified batch effect, one or more A correction factor is determined and applied to normalize the samples to the reference data set. In this way, these samples can be validated and used for downstream analysis. In some embodiments, if one or more samples are identified as contributing to the identified batch effect, one or more samples are rejected (e.g., after manual inspection or automatically). In some embodiments, the rejected samples are rerun through the RNA expression pipeline.
[0136] It should be noted that other process details described herein with respect to other methods described herein (e.g., the methods shown in Figures 2, 16, and 17) are also applicable in a similar manner to the method described above with respect to Figure 15. For example, details relating to data collection, data processing, cohort matching, dimensionality reduction analysis, etc., described above with reference to the method outlined in Figure 15, optionally have one or more of the features of data collection, data processing, cohort matching, dimensionality reduction analysis, etc., described herein with reference to other methods described herein (e.g., the methods outlined in Figures 2, 16, and 17). For the sake of brevity, these details will not be repeated here.
[0137] Validating changes to the bioinformatics pipeline Another aspect of the disclosure provides a method for validating changes in an RNA expression pipeline in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors. The method includes obtaining, in electronic format, a batch dataset for each respective test sample in a batch of test samples, the batch dataset including a corresponding expression profile prepared using a first RNA expression pipeline. The corresponding expression profile includes a corresponding gene expression value for each respective gene in a first set of genes. The batch dataset further includes a corresponding set of metadata including a value for each respective feature in the first set of features for each respective test sample.
[0138] The method includes determining a cohort-matched reference dataset for the batch dataset, the cohort-matched reference dataset including a corresponding expression profile prepared using a second RNA expression pipeline (e.g., the pipeline existing prior to the modification) for each respective reference sample in the plurality of reference samples, the corresponding expression profile including a corresponding gene expression value for each respective gene in the first set of genes. Each respective reference sample in the plurality of reference samples is associated with a corresponding set of metadata including a corresponding value for each respective feature in the second set of features for the respective reference sample. The aggregated values for each respective feature in the third set of one or more features present in both the first set of features and the second set of features are balanced between the batch dataset and the cohort-matched reference dataset.
[0139] The dimensionality reduction is performed on the combined data set consisting of the corresponding expression profile for each respective test sample in the plurality of test samples and the corresponding expression profile for each respective reference sample in the plurality of reference samples. Thus, for each respective test sample and each respective reference sample, a corresponding set of coordinates is obtained that is embedded in a dimensional space lower than the dimensionality of the corresponding expression profile.
[0140] The method further includes determining a statistical measure of similarity between the set of coordinates obtained for the test sample and the set of coordinates obtained for the reference sample. The statistical measure of similarity is compared to a threshold, and if the statistical measure of similarity meets the threshold, the change in the RNA expression pipeline is validated, or if the statistical measure of similarity does not meet the threshold, the change in the RNA expression pipeline is not validated.
[0141] For example, FIG. 16 illustrates some embodiments of the present disclosure (e.g., the utility of the new process). 10 illustrates a method for validating changes in an RNA expression pipeline by determining a gene expression profile for each respective test sample in a batch of test samples and / or determining a correction factor for a new process. A batch dataset is obtained from a batch of test samples, the batch dataset including a corresponding expression profile for each respective test sample in the batch of test samples. In some embodiments, each respective expression profile includes a corresponding gene expression value for each respective gene in a plurality of genes in the test sample (e.g., as determined by the RNA expression pipeline). In some embodiments, each respective expression profile further includes metadata indicating values for a plurality of features associated with the test sample, such as an RNA transcription profile (e.g., determined from a plurality of sequence reads), clinical data (e.g., patient diagnosis, treatment outcome, etc.), gender, biopsy type (e.g., heme vs. solid biopsy), molecular data (e.g., genomic mutations), and / or other features (e.g., tissue site, tumor purity, cancer type, collection method, sequencer identity, and / or sequenced date).
[0142] In some embodiments, each expression profile is prepared using a sequencing analysis (e.g., an RNA expression pipeline) of each respective test sample in the batch of test samples. In some embodiments, each respective test sample in the batch of test samples is subjected to a first sequencing analysis (e.g., a first RNA expression pipeline, "vX"). Referring to block 1602, in some embodiments, the process change in the RNA expression pipeline includes changing the sequencing analysis from a first process (e.g., vX) to a second process (e.g., vY). Referring to block 1604, in some embodiments, each respective test sample in the batch of test samples is further subjected to a second sequencing analysis (e.g., a second RNA expression pipeline, "vY").
[0143] Thus, in some embodiments, the method includes obtaining a first batch dataset 1608 (e.g., a vX batch dataset) that includes, for each respective test sample in the batch of test samples, a corresponding first expression profile obtained using a first process (e.g., vX), and a second batch dataset (e.g., a vY batch dataset) that includes, for each respective test sample in the batch of test samples, a corresponding second expression profile obtained using a second process (e.g., vY).
[0144] In some embodiments, the expression data in each respective first expression profile in the first batch dataset and the expression data in each respective second expression profile in the second batch dataset are normalized.
[0145] In some embodiments, the method further includes obtaining a cohort-matched reference dataset 1606 for the plurality of reference samples. In some embodiments, the cohort-matched reference dataset is identified by matching a proportion of one or more features 1603 of the samples in the batch dataset with a reference sample having the same proportion of those one or more features (e.g., tissue site, tumor purity, cancer type, collection method, sequencer identity, sequenced date, clinical data (e.g., patient diagnosis, treatment outcome, etc.), gender, biopsy type (e.g., heme vs. solid biopsy), molecular data (e.g., genomic mutations), and / or other features) in, for example, the reference database 1605. The cohort-matched reference dataset includes an expression profile for each reference sample in the plurality of reference samples. Each respective expression profile includes a corresponding gene expression value for each respective gene in the plurality of genes (e.g., where the plurality of genes included in each reference sample expression profile is the same as the plurality of genes included in each test sample expression profile).
[0146] In some embodiments, the cohort-matched reference dataset is a sample-matched dataset. That is, in some embodiments, the same samples are run through both versions of the RNA expression pipeline and compared to each other as a batch dataset (e.g., generated from samples run through a new version of the RNA expression pipeline) and a reference dataset (e.g., generated from samples run through a previous version of the RNA expression pipeline).
[0147] In some embodiments, the plurality of reference samples are selected for a batch of test samples based on one or more similarities between features (e.g., tissue site, tumor purity, cancer type, collection method, sequencer identity, and / or sequenced date) associated with the reference sample and the test sample processed using process vY. In some embodiments, the metadata (e.g., values for features) between the batch of test samples and the plurality of reference samples are balanced such that the distribution (e.g., proportions) of the features of the test samples among the plurality of test samples processed using process vY is similar to the distribution (e.g., proportions) of the features of the reference samples among the plurality of reference samples.
[0148] In some embodiments, the cohort-matched reference dataset is normalized.
[0149] According to this method, referring to block 1610, dimensionality reduction (e.g., PCA, latent component analysis, partial least squares regression, etc.) is performed on a combined dataset including a corresponding second expression profile for each respective test sample in the plurality of test samples processed using process vY, and a corresponding expression profile for each respective reference sample in the cohort-matched reference dataset processed using process vX. In some embodiments, the combined dataset further includes a corresponding first expression profile for each respective test sample in the plurality of test samples processed using process vX.
[0150] Thus, for each respective test sample and each respective reference sample processed using process vX, and for each respective test sample processed using process vY, the corresponding sets of coordinates 1612-1614 are embedded in a lower dimensional space (e.g., m-space) than the dimensionality of the corresponding expression profile (e.g., 1612-1, 1612-2, ..., 1612-N, and 1614-1, 1614-2, ..., 1614-N).
[0151] Referring to block 1616, the method further includes evaluating the component values using the combined data set after dimensionality reduction. The method includes determining a statistical measure of similarity between sets of coordinates obtained for the test sample and the reference sample processed using process vX and the test sample processed using process vY. Referring to block 1618, the statistical measure of similarity is compared to a threshold value to determine whether there is significant variance between process vX and process vY. Referring to block 1620, if the statistical measure of similarity meets the threshold value, the change in the RNA expression pipeline is verified. Referring to block 1622, if the statistical measure of similarity does not meet the threshold value, the change in the RNA expression pipeline is rejected and / or flagged for further evaluation of the process change and / or for determination of a correction factor for process vY.
[0152] It should be noted that other process details described herein with respect to other methods described herein (e.g., the methods shown in Figures 2, 15, and 17) are also applicable in a similar manner to the method described above with respect to Figure 16. For example, details relating to data collection, data processing, cohort matching, dimensionality reduction analysis, etc., described above with reference to the method outlined in Figure 16, may optionally be applied to other methods described herein (e.g., the methods shown in Figures 2, 15, and 17). The method may include one or more of the features of data collection, data processing, cohort matching, dimensionality reduction analysis, etc., described herein with reference to the methods outlined in (e.g., and 17). For the sake of brevity, these details will not be repeated here.
[0153] Reference database expansion Another aspect of the disclosure provides a method of adding RNA expression data to a reference database in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors. The method includes obtaining a new expression dataset in electronic format. The new expression dataset includes, for each respective test sample in the plurality of test samples, a corresponding expression profile prepared using a first RNA expression pipeline, the corresponding expression profile including a corresponding gene expression value for each respective gene in the first set of genes. The new expression dataset further includes a corresponding set of metadata including a value for each respective feature in the first set of features for each respective test sample.
[0154] The method includes determining a cohort-matched reference dataset for the new expression dataset, the cohort-matched reference dataset including a corresponding expression profile for each respective reference sample in the plurality of reference samples, the corresponding expression profile including a corresponding gene expression value for each respective gene in the first set of genes. Each respective reference sample in the plurality of reference samples is associated with a corresponding set of metadata including a corresponding value for each respective feature in the second set of features for the respective reference sample. Each expression profile corresponding to a reference sample in the plurality of reference samples is from a reference database. The aggregated values for each respective feature in the third set of one or more features present in both the first set of features and the second set of features are balanced between the batch dataset and the cohort-matched reference dataset.
[0155] The dimensionality reduction is performed on the combined data set consisting of the corresponding expression profile for each respective test sample in the plurality of test samples and the corresponding expression profile for each respective reference sample in the plurality of reference samples. Thus, for each respective test sample and each respective reference sample, a corresponding set of coordinates is obtained that is embedded in a dimensional space lower than the dimensionality of the corresponding expression profile.
[0156] The method further includes determining a statistical measure of similarity between the set of coordinates obtained for the test sample and the set of coordinates obtained for the reference sample, where the statistical measure of similarity is compared to a threshold. The method includes adding a new expression dataset to the reference database if the statistical measure of similarity meets the threshold or if the statistical measure of similarity does not meet the threshold, determining a set of conversion factors for normalizing expression profiles in the new expression dataset to expression profiles in the reference database, normalizing the expression profiles in the new expression dataset using the set of conversion factors, thereby obtaining a normalized new expression dataset, and adding the normalized new expression dataset to the reference database.
[0157] For example, FIG. 17 illustrates a method of adding RNA expression data to a reference database (eg, validating newly acquired expression data used to update the reference database) according to some embodiments of the present disclosure.
[0158] Referring to block 1702, a new expression dataset (e.g., of RNA expression data) is obtained. In some embodiments, the new expression dataset is prepared for each respective test sample in the plurality of test samples using the first RNA expression pipeline and includes an expression profile including gene expression values for a plurality of genes. The current dataset further includes a corresponding set of metadata that includes values for features associated with the test samples.
[0159] Referring to block 1704, in some embodiments, the new expression dataset is normalized.
[0160] In some embodiments, the method further includes obtaining a cohort-matched reference dataset 1706 for the new expression dataset. In some embodiments, the cohort-matched reference dataset is identified by matching a proportion of one or more features 1703 of the samples in the batch dataset with a reference sample having the same proportion of those one or more features (e.g., tissue site, tumor purity, cancer type, collection method, sequencer identity, sequenced date, clinical data (e.g., patient diagnosis, treatment outcome, etc.), gender, biopsy type (e.g., heme vs. solid biopsy), molecular data (e.g., genomic mutations), and / or other features) in, for example, a reference database 1705. The cohort-matched reference dataset includes an expression profile for each reference sample in the plurality of reference samples. Each respective expression profile includes a corresponding gene expression value for each respective gene in the plurality of genes (e.g., where the plurality of genes included in each reference sample expression profile is the same as the plurality of genes included in each test sample expression profile).
[0161] In some embodiments, the cohort-matched reference dataset comprises an expression profile for each reference sample in the plurality of reference samples. In some embodiments, each respective expression profile is from a reference database (e.g., an existing database) and comprises a corresponding gene expression value for each respective gene in the plurality of genes (e.g., included in the new expression dataset).
[0162] Each respective reference sample in the plurality of reference samples is associated with a corresponding set of metadata that includes values for the features associated with the reference sample. In some embodiments, the cohort-matched reference dataset is matched (e.g., balanced) according to any of the methods described in Figures 15 and / or 16.
[0163] In some embodiments, the cohort-matched reference dataset is normalized.
[0164] According to this method, referring to block 1708, a dimensionality reduction (e.g., PCA, latent component analysis, partial least squares regression, etc.) is performed on a combined dataset consisting of a corresponding expression profile for each respective test sample in the plurality of test samples and a corresponding expression profile for each respective reference sample in the plurality of reference samples. For each respective test sample N and each respective reference sample C, the corresponding set of coordinates 1710-1712 are embedded into a lower dimensional space (e.g., m-space) than the dimensionality of the corresponding expression profiles (e.g., 1710-1, 1710-2, ..., 1710-N, and 1712-1, 1712-2, ..., 1712-N).
[0165] Referring to block 1714, the method further includes evaluating component values using the combined dataset after dimensionality reduction. The method includes determining a statistical measure of similarity between the set of coordinates obtained for the test sample and the set of coordinates obtained for the reference sample. Referring to block 1716, the statistical measure of similarity is compared to a threshold to determine whether there is significant variance between the new expression dataset and data from the existing database.
[0166] Referring to block 1718, if the statistical measure of similarity meets a threshold, the new The new expression dataset is added to the reference database. Referring to block 1720, if the statistical measure of similarity does not meet a threshold, a set of conversion coefficients for normalizing the expression profiles in the new expression dataset to the expression profiles in the reference database is determined, the expression profiles in the new expression dataset are normalized using the set of conversion coefficients to obtain a normalized new expression dataset, and the normalized new expression dataset is added to the reference database.
[0167] It should be noted that other process details described herein with respect to other methods described herein (e.g., the methods shown in Figures 2, 15, and 16) are also applicable in a similar manner to the method described above with respect to Figure 17. For example, details relating to data collection, data processing, cohort matching, dimensionality reduction analysis, etc., described above with reference to the method outlined in Figure 17, optionally have one or more of the features of data collection, data processing, cohort matching, dimensionality reduction analysis, etc., described herein with reference to other methods described herein (e.g., the methods outlined in Figures 2, 15, and 16). For the sake of brevity, these details will not be repeated here.
[0168] Example of an embodiment In some embodiments of the systems and methods described herein (e.g., the methods outlined in Figures 2, 15, 16, and 17, as described above), obtaining the batch data set includes obtaining, for each respective sample in the batch of samples, a corresponding plurality of sequence reads obtained from the respective sample by targeted or whole-transcriptome RNA sequencing in electronic format, and determining a corresponding gene expression value for each respective gene in the first set of genes from the corresponding plurality of sequence reads. In some embodiments, the methods described herein also include a step of generating sequencing data. However, in other aspects, the methods described herein start after sequencing has already been performed. For example, in some embodiments, the methods described herein start by obtaining, for each sample in the batch of samples, a sequence read in electronic format, determining an expression profile for each sample in the batch of samples based on the sequence reads, and then performing one or more quality control methods, as described herein. Similarly, in some embodiments, the methods described herein start by obtaining, for each sample in the batch of samples, an expression profile in electronic format, and then performing one or more quality control methods, as described herein.
[0169] In some embodiments, for each respective sample in the batch of samples, the corresponding plurality of sequence reads is at least 10,000 sequence reads. In some embodiments, the corresponding plurality of sequence reads is at least 100,000 sequence reads. In some embodiments, the corresponding plurality of sequence reads is at least 1,000,000 sequence reads. In some embodiments, the corresponding plurality of sequence reads is at least 10,000,000 sequence reads. In some embodiments, the corresponding plurality of sequence reads is between 10,000 and 100,000,000 sequence reads. In some embodiments, the corresponding plurality of sequence reads is between 100,000 and 50,000,000 sequence reads. In some embodiments, the corresponding plurality of sequence reads is between 1,000,000 and 50,000,000 sequence reads.
[0170] In some embodiments, the batch of test samples includes at least 10 test samples. In some embodiments, the batch of test samples includes at least 25 test samples. In some embodiments, the batch of test samples includes at least 100 test samples. In some embodiments, the batch of test samples includes at least 1000 test samples. In some embodiments, the batch of test samples includes at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 25, at least 50, at least 100, at least 250, at least 500, at least 750, at least 1000, at least 2500, at least 5000, at least 10,000, at least 100,000, at least 1,000,000, or more samples. In some embodiments, the batch of test samples includes between 5 and 100 test samples. In some embodiments, the batch of test samples includes between 50 and 500 test samples. In some embodiments, the batch of test samples includes between 100 and 1000 test samples. In some embodiments, the batch of test samples includes between 1000 and 100,000 test samples.
[0171] In some embodiments, the first set of genes comprises at least 10 genes. In some embodiments, the first set of genes comprises at least 100 genes. In some embodiments, the first set of genes comprises at least 1000 genes. In some embodiments, the first set of genes comprises at least 10,000 genes. In some embodiments, the batch of test samples comprises at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 25, at least 50, at least 100, at least 250, at least 500, at least 750, at least 1000, at least 2500, at least 5000, at least 10,000, at least 20,000, at least 30,000, or more genes.
[0172] In some embodiments, the set of features for which the batch and reference datasets are balanced includes at least one feature selected from tissue site (the site where the biological sample was obtained), tumor purity, cancer type, sequencer identity, and sequencing date. In some embodiments, the set of features for which the batch and reference datasets are balanced includes at least two features selected from tissue site, tumor purity, cancer type, sequencer identity, and sequencing date. In some embodiments, the set of features for which the batch and reference datasets are balanced includes at least three features selected from tissue site, tumor purity, cancer type, sequencer identity, and sequencing date. In some embodiments, the set of features for which the batch and reference datasets are balanced includes at least tissue site and cancer type. In some embodiments, the set of features for which the batch and reference datasets are balanced is tissue site and cancer type.
[0173] In some embodiments, the set of features by which the batch and reference datasets are balanced comprises at least one feature selected from the nucleic acid extraction method, the cDNA library preparation method, the RNA sequencing method, the type of reagents used, and the type of instrument used.
[0174] In some embodiments, the plurality of reference samples (reference datasets) comprises at least 50 reference samples. In some embodiments, the plurality of reference samples (reference datasets) comprises at least 100 reference samples. In some embodiments, the plurality of reference samples (reference datasets) comprises at least 500 reference samples. In some embodiments, the plurality of reference samples (reference datasets) comprises at least 1000 reference samples. In some embodiments, the plurality of reference samples (reference datasets) comprises at least 5000 reference samples. In some embodiments, the plurality of reference samples (reference datasets) comprises at least 10,000 reference samples. In some embodiments, the plurality of reference samples (reference datasets) comprises at least 100,000 reference samples. In some embodiments, the plurality of reference samples (reference datasets) comprises at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 16, at least 18, at least 20, at least 24, at least 28, at least 26, at least 28 ... The plurality of reference samples may comprise at least 25, at least 50, at least 100, at least 250, at least 500, at least 750, at least 1000, at least 2500, at least 5000, at least 10,000, at least 100,000, at least 1,000,000, or more samples. In some embodiments, the plurality of reference samples comprises 5 to 100 reference samples. In some embodiments, the plurality of reference samples comprises 50 to 500 reference samples. In some embodiments, the plurality of reference samples comprises 100 to 1000 reference samples. In some embodiments, the plurality of reference samples comprises 1000 to 100,000 reference samples.
[0175] In some embodiments, the set of reference samples balanced to the batch dataset includes at least the same number of samples as present in the batch dataset. In some embodiments, the set of reference samples balanced to the batch dataset has the same number of samples as present in the batch dataset. In some embodiments, the set of reference samples balanced to the batch dataset includes at least 25% more samples than present in the batch dataset. In some embodiments, the set of reference samples balanced to the batch dataset includes at least 50% more samples than present in the batch dataset. In some embodiments, the set of reference samples balanced to the batch dataset includes at least 100% more samples than present in the batch dataset. In some embodiments, the set of reference samples balanced to the batch dataset includes at least 5 times more samples than present in the batch dataset. In some embodiments, the set of reference samples balanced to the batch dataset includes at least 10 times more samples than present in the batch dataset.
[0176] The method of any one of claims 27 to 36, wherein the plurality of reference samples comprises at least 1000 reference samples.
[0177] In some embodiments, the aggregated values for each feature are balanced between the batch dataset and the cohort-matched reference set if the percentage of each test sample in the batch of test samples that has the respective value for each feature is within 2.5% of the percentage of each reference sample in the multiple reference samples that has the same respective value for each feature.For example, in some embodiments where cancer type is a feature that is balanced between the batch dataset and the cohort-matched dataset, if the batch dataset is composed of 20% brain cancer samples, 30% lung cancer samples, and 50% colon cancer samples, the reference dataset includes 17.5%-22.5% brain cancer samples, 27.5%-32.5% lung cancer samples, and 47.5%-52.5% colon cancer samples.
[0178] In some embodiments, the aggregate values for each feature are balanced between the batch dataset and the cohort-matched reference set if the percentage of each test sample, within a batch of test samples, having a respective value for each feature is within 1%, within 2%, within 3%, within 4%, within 5%, within 6%, within 7%, within 8%, within 9%, within 10%, within 11%, within 12%, within 13%, within 14%, within 15%, within 16%, within 17%, within 18%, within 19%, within 20%, within 21%, within 22%, within 23%, within 24%, or within 25% of the percentage of each reference sample, within the multiple reference samples, having the same respective value for each feature.
[0179] In some embodiments, the dimensionality reduction comprises embedding the corresponding expression profiles for each respective test sample and each respective reference sample into a two-dimensional representation. In some embodiments, the dimensionality reduction includes embedding into two coordinates using uniform manifold approximation and projection (UMAP). In some embodiments, the dimensionality reduction includes embedding the corresponding expression profile for each respective test sample and each respective reference sample into two, three, four, five, six, seven, eight, nine, ten, or more coordinates. In some embodiments, the dimensionality reduction includes embedding into a smaller coordinate system using principal component analysis (PCA). EXAMPLES
[0180] Example 1 - Robust detection of sequencing batch effects in RNA by low-dimensional embedding with subtype-matched reference samples Technical batch effects, such as changes in protocols, reagents, or sequencing techniques, can invalidate large-scale transcriptomic studies. Laboratories that analyze the transcriptome across tumor types, time, or multiple sites need to have a systematic method for validating the compatibility of data between batches.
[0181] Batch effects can manifest as large changes in a few genes or small changes in many genes. A robust batch effect detection method would identify either. Furthermore, results from bulk RNAseq are driven by cancer type and tissue site. This complicates the detection of batch effects in studies across multiple cancer types. To overcome these challenges, we developed a method to assess technical batch effects in a heterogeneous set of transcriptomic samples.
[0182] Briefly, samples were selected from validated reference data to match the transcriptome set based on cancer type and tissue site. Gene expression profiles of the transcriptome set and the matched reference data were embedded into two coordinates using uniform manifold approximation and projection (UMAP). The clustering property of UMAP makes it ideal for detecting batch effects. A Mann-Whitney U test was then performed on the x and y UMAP coordinates. If either test returns a p-value below a threshold, e.g., 0.01, there is a possibility of a batch effect.
[0183] As a first example, we applied this method to determine whether batch effects arise when using different blood collection methodologies. Briefly, RNAseq data were generated for paired cohort and tissue-matched blood samples collected using either PAX or EDTA collection tubes. All other sample preparation, data collection, and data processing steps were performed identically for all samples. RNAseq data were then embedded into two coordinates using UMAP (Figure 3A; 302 = PAX collection tube; 304 = EDTA collection tube). Mann-Whitney U tests were then applied separately to the x and y coordinates of the UMAP embedding. As shown in Figures 3B and 3C, both Mann-Whitney U tests identified statistically significant differences (p = 7.08E-10) between RNAseq data generated from blood collected in PAX and EDTA collection tubes, proving that batch effects arose from the use of different blood collection methodologies.
[0184] As a second example, we applied this method to determine whether batch effects arise when using different bioinformatics pipelines for the analysis of RNAseq data. Briefly, RNAseq data were generated for paired cohorts and tissue-matched samples that were processed using either the STAR or kallisto pipelines (Dobin A. et al., Bioinformatics, 29(1):15-21(2013) (describing STAR) and Bray NL et al., Nature Biotechnology, 34:525- 27 (2016) (depicting Kallisto). All sample preparation, data collection, and data processing steps were performed identically for all samples, except for the differences in the bioinformatics pipelines used to align the RNAseq data and quantify transcript abundance. The RNAseq data were then embedded into two coordinates using UMAP (Figure 4A; 402 = Star alignment; 404 = Kallisto alignment). The Mann-Whitney U test was then applied separately to the x and y coordinates of the UMAP embedding. As shown in Figures 4B and 4C, both Mann-Whitney U tests identified statistically significant differences between the RNAseq data aligned using the STAR algorithm and the RNAseq data aligned using the Kallisto algorithm (p = 2.80E-9 and p = 1.69E-19), proving a batch effect resulting from the use of different RNAseq alignment algorithms.
[0185] Finally, we used this method to test common sources of technological batch effects, namely flow cells, bioinformatics pipeline updates, and sequencers. For each effect, for each technology class (flow cells, pipelines, and sequencers), we ran the method on 15 subsamples per feature and calculated the Benjamini-Hochberg corrected false discovery rate across subsamples. The distributions presented in Figures 5A-5C represent the median FDR calculated across subsamples.
[0186] Thus, the above method is an effective and easy-to-implement method for automatically testing technical and software batch effects across multiple cancer types and tissue sites.
[0187] Example 2 - Batch correction applied after capture probe redesign Capture RNA-Seq methods have many advantages over total mRNA capture methods, especially when analyzing gene expression in FFPE samples. For example, polyA selection in total mRNA capture methods does not work optimally in FFPE samples because RNA molecules are fragmented during the fixation process. Thus, many mRNA fragments are not captured because they are no longer associated with a polyA tail. In contrast, capture RNA-Seq methods isolate mRNA fragments using probes designed against the coding sequence of the target mRNA, so these methods are significantly less affected by fragmentation. Furthermore, when using hematological samples, capture RNA-Seq methodologies are less affected by ribosome depletion and hemoglobin depletion.
[0188] However, we observed that changing the design of the exome capture probes used for RNA-Seq capture resulted in a small technical batch effect. To correct for these batch effects, 450 samples representing various cancer types that had previously been exome sequenced using a set of first-generation exome capture probes were selected for exome resequencing using a set of second-generation exome capture probes. We then determined a linear correction factor for each gene by comparing the original exome sequencing results generated using the set of first-generation exome capture probes to the new exome sequencing results generated using the set of second-generation exome capture probes.
[0189] A linear gene-by-gene linear correction was found to be sufficient to remove all systematic differences between the two datasets. For each gene i, the expression value E ci was calculated as follows: E ci =(E i *m i )+b i In the formula, E i is the uncorrected expression level determined for gene i (log2 TPM) using a set of second-generation exome capture probes, and m i is the gene i is the gradient correction coefficient for b i is the intercept correction coefficient for gene i. These slope and intercept correction coefficients for each gene are trained to match the distribution of v1 and corrected v2 in the matched dataset. This was optimized by minimizing a weighted loss function that can take into account the context of the paired samples. By utilizing this side information, the same linear correction can be applied to all samples processed using the set of second-generation exome probes, so the resulting correction coefficients will work robustly for sequenced samples of any cancer type.
[0190] To analyze the effect of the correction factor, principal component analysis (PCA) was performed on expression values for 100 samples, each representing 37 different cancer types, that were treated twice, once using the first generation set of exome capture probes and once using the second generation set of exome capture probes. These 100 samples were not part of the training cohort used to generate the correction factor. PCA was first performed on uncorrected expression values determined using the second generation set of exome capture probes, and then on expression values corrected using the correction factor described above.
[0191] As shown in Figure 6A, when uncorrected expression values were used for PCA analysis, the technical batch effect was clearly observable by the association of the third principal component with assay type (+=first generation exome capture probe set, O=second generation exome capture probe set, lines connect paired samples). However, as shown in Figure 6B, when uncorrected expression values were used for PCA analysis, none of the principal components were associated with assay type, and all samples clustered together across samples and cancer types.
[0192] Example 3 - Identification of technological batch effects resulting from differences in collection site and storage method of biological samples Heme cancers can be collected from whole blood, bone marrow, and sometimes preserved in formalin-fixed paraffin-embedded (FFPE). However, these differences in sample collection methodologies result in the introduction of technical batch effects. Briefly, RNA expression data from cohort-matched cancer samples collected by either blood sampling, bone marrow sampling, or preserved in FFPE were analyzed by dimensionality reduction analysis, embedded into two coordinates using UMAP, according to some embodiments of the present disclosure. The results presented in Figure 7 show that transcriptomic samples are clustered and separated by FFPE vs. EDTA blood / bone marrow tubes on the y-axis and by whole bone marrow vs. on the x-axis. This separation is driven by both biological and technical differences and requires appropriate reference matching.
[0193] Example 4 - Identification of technical batch effects resulting from differences in RNA extraction methodology Different extraction methods and chemicals used during transcriptome analysis may introduce batch effects. For example, batch effects may occur when RNA samples are extracted, for example, before the samples are sent for sequencing by the clinician (external extraction) as opposed to immediately before RNAseq analysis (internal extraction). Briefly, RNA expression data from cohort-matched cancer samples acquired before or after RNA isolation were analyzed by dimensionality reduction analysis, embedded into two coordinates using UMAP, according to some embodiments of the present disclosure. The results presented in Figure 8 show that heme extracted samples by external sources cluster separately from internally extracted samples. A second batch effect is observed by separating internally extracted FFPE samples from blood and bone marrow as described in Example 3.
[0194] Example 5 - Identification of technical batch effects resulting from different reagent lots The capture RNASeq methodology involves a step in which a cDNA fragment library is enriched using capture probes. In this example, two samples of cDNA fragment libraries prepared for several cancer samples were hybridized to two batches of the same set of biotinylated oligonucleotide probes complementary to the genomic regions of interest. These capture probe libraries can sometimes themselves be pools of different genomic capture designs. The manufacture of probe lots and the pooling of capture libraries can introduce batch effects. The RNA expression data generated using both lots of capture probes were then analyzed by PCA dimensionality reduction analysis. The results presented in Figure 9 show the batch effects introduced by the different probe lots, as detected in PC8 (x-axis).
[0195] Example 6 - Analysis of technical batch effects resulting from different hybrid capture plexities Hybridization plexity refers to the number of cDNA samples pooled together during targeted capture. Assays can vary from a single sample in a pool to more than 12 samples. In this experiment, nine tumor samples and two cell controls were sequenced under three plexity conditions (single, 3x, and 6x sample pools). Following RNA sequencing, expression data generated using samples prepared under different plexity conditions was analyzed by dimensionality reduction analysis, embedded into two coordinates using UMAP, according to some embodiments of the present disclosure. The results presented in Figure 10 show that matched samples were clustered regardless of the plexity condition used, indicating that plexity does not introduce batch effects on transcriptome analysis.
[0196] Example 7 - Analysis of technical batch effects resulting from different numbers of PCR amplification cycles In some RNAseq methodologies, post-capture PCR is an amplification step after amplicon fragments are captured by the probe and before sequencing. Unbound fragments are washed away and remaining fragments are amplified for a set number of cycles (more cycles means more amplification). Too many cycles can cause unbalanced duplication rates based on sequence features. In this experiment, the effect of the number of amplification cycles (7-9) on six tumor and one control sample was determined. Following RNA sequencing, expression data generated using samples prepared under different amplification conditions (7-9 cycles) were analyzed by dimensionality reduction analysis, embedded into two coordinates using UMAP, according to some embodiments of the present disclosure. The results presented in Figure 11 show that matched samples were clustered regardless of the number of amplification cycles used, indicating that variation in the number of PCR amplification cycles from 7 to 9 does not introduce batch effects on transcriptome analysis.
[0197] Example 8 - Analysis of technical batch effects resulting from different sequencer loading molarities Loading molarities refer to the amount of sample loaded onto the sequencer. Typically, too low a molar concentration can lead to high replication rates and high noise in the data. In this batch effect experiment, 11 tumor samples and 3 control samples were sequenced under three molar conditions (0.7, 1, and 1.5 uM). Following RNA sequencing, expression data generated using samples prepared under different loading molarities were analyzed by dimensionality reduction analysis, embedded into 2 coordinates using UMAP, according to some embodiments of the present disclosure. The results presented in Figure 12 show that matched samples were clustered regardless of the loading molarities used, indicating that variation in loading molarities from 0.7 to 1.5 uM does not introduce batch effects on transcriptome analysis.
[0198] Example 9 - Analysis of technical batch effects resulting from changes in sequencing reagent chemistry Modifications to Illumina's sequencing reagents are general and proprietary changes to the chemistry used to better accommodate additional features for their technology. The example was adapted to allow additional reads in the Universal Molecular Index (UMI). In a batch effect control experiment, 28 samples were sequenced under two reagent versions to confirm that no batch effect was detected between the previous and current versions of the reagent. Following RNA sequencing, expression data generated using samples prepared with different versions of the reagent were analyzed by dimensionality reduction analysis, embedded into two coordinates using UMAP, according to some embodiments of the present disclosure. Overall, samples clustered by sample (Figure 13, connecting lines), with the reagents being slightly, but within acceptable limits for transcriptome variance.
[0199] conclusion The method described herein provides an improved quality control method for evaluating batches of RNA sequencing samples.With improved accuracy and higher resolution than previous methods, the prediction algorithm provided herein can be used to identify single samples and entire batches that meet quality control criteria.This enhanced quality control results in more accurate information used to provide diagnosis and determine appropriate treatment for patients, leading to improved diagnosis and more informed treatment recommendations for patients.
[0200] References and Alternative Embodiments All references cited in this specification are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety for all purposes.
[0201] The present invention can be implemented as a computer program product that includes a computer program mechanism embedded in a non-transitory computer-readable storage medium. For example, the computer program product can include program modules such as those shown in Figure 1 and / or described in Figures 2A and 2B. These program modules can be stored on a CD-ROM, DVD, magnetic disk storage product, USB key, or any other non-transitory computer-readable data or program storage product.
[0202] As will be apparent to those skilled in the art, many modifications and variations of this application can be made without departing from the spirit and scope of the present application. The specific embodiments described herein are provided by way of example only. The embodiments have been selected and described in order to best explain the principles of the invention and its practical use, so as to enable those skilled in the art to best utilize the invention and various embodiments with various modifications suited to the particular applications contemplated. The present invention should be limited only by the terms of the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
1. 1. A method for performing quality control, the method comprising:
1. A computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, a) obtaining, for each respective sample in a batch of samples, a batch dataset in electronic format, the batch dataset including a corresponding plurality of sequence reads obtained from said respective sample by targeted or whole-transcriptome RNA sequencing and corresponding metadata for said respective sample; b) determining a cohort-matched reference batch for the batch dataset, the cohort-matched reference batch being balanced for tissue site, tumor purity, cancer type, collection method, sequencer identity, and / or sequencing date; c) performing one or more global batch quality control tests on said batch data set using at least said cohort-matched reference batch; and d) after conducting said one or more global batch quality control tests; validating the batch data set if each of the one or more global batch quality control tests is satisfied; or not validating the batch data set if one or more of the global batch quality control tests are not satisfied.
2. 10. The method of claim 1, further comprising, if one or more of the global batch quality control tests are not satisfied, removing each sample from the batch data set that fails any one of the one or more global batch quality control tests or flagging each sample that fails any one of the one or more global batch quality control tests for manual inspection.
3. The method of claim 1 or 2, wherein the targeted panel RNA sequencing uses multiple probes.
4. each probe in said plurality of probes uniquely targets a respective portion of a reference transcriptome; 4. The method of claim 3, wherein each sequence read in the corresponding plurality of sequence reads corresponds to at least one probe in the plurality of probes.
5. The method of any one of claims 1 to 4, wherein whole transcriptome sequencing comprises next generation sequencing.
6. The method of any one of claims 2 to 5, wherein the removing further comprises providing an updated batch data set.
7. Determining a cohort-matched reference dataset for the batch dataset, For each sample in said batch of samples, i) extracting a respective plurality of sequence features from each of the plurality of sequence reads, thereby obtaining a batch of sequence features; and ii) extracting a respective plurality of sample metadata features, thereby obtaining a batch plurality of metadata features; and selecting the cohort-matched reference dataset comprising a plurality of reference samples from a reference dataset based at least in part on a plurality of sequence features of the batch or a plurality of metadata features of the batch.
8. the cohort-matched reference dataset comprises a plurality of reference samples; 8. The method of claim 7, wherein each reference sample in the plurality of reference samples comprises a corresponding plurality of sequence reads obtained from a respective reference sample by targeted or whole-transcriptome RNA sequencing and corresponding metadata for the respective reference sample.
9. each sample in a first subset of samples in the batch dataset has a corresponding first biopsy type; 8. The method of claim 7, wherein each sample in the second subset of samples in the batch data set has a corresponding second biopsy type.
10. 10. The method of claim 9, wherein the first biopsy type or the second biopsy type comprises a somatic biopsy selected from the set comprising: macro-dissected formalin-fixed paraffin-embedded (FFPE) tissue section, surgical biopsy, skin biopsy, punch biopsy, prostate biopsy, bone biopsy, bone marrow biopsy, needle biopsy, CT-guided biopsy, ultrasound-guided biopsy, fine needle aspiration, aspiration biopsy, fresh tissue or blood sample.
11. The method further comprising: for each respective sample in the batch of samples, performing one or more single sample quality control tests on the respective sample from the corresponding plurality of sequence reads; 11. The method of any one of claims 1 to 10, further comprising removing each sample from the batch of samples that fails any one of the one or more single-sample quality control tests or flagging for manual inspection each sample that fails any one of the one or more single-sample quality control tests.
12. each single sample quality control test in the one or more single sample quality control tests i) determining, for each sample in the batch of samples, a respective number of non-redundant mapped sequence reads in the plurality of sequence reads, where each non-redundant mapped sequence read maps to a corresponding portion of a reference genome; ii) comparing the respective number of non-redundant mapped sequence reads to an expected number of non-redundant mapped sequence reads, wherein if the respective number of non-redundant mapped sequence reads is below a predetermined number of non-redundant mapped reads, the respective sample fails the respective single sample quality control test.
13. 13. The method of claim 11 or 12, wherein each single sample quality control test in the one or more single sample quality control tests comprises determining, for each sample in the batch of samples, a respective percentage of properly matched sequence reads, wherein if the percentage of properly matched sequence reads is below a predetermined matched read threshold, the respective sample fails the respective single sample quality control test.
14. 14. The method of any one of claims 11-13, wherein each single sample quality control test in the one or more single sample quality control tests comprises determining, for each sample in the batch of samples, a GC content of each of the corresponding plurality of sequence reads, wherein if the respective GC content is outside a predetermined GC content threshold, the respective sample fails the respective single sample quality control test.
15. Each single sample quality control test in the one or more single sample quality control tests 15. The method of claim 11, further comprising determining for each sample in the batch of samples a respective number of expressed genes, wherein if a corresponding expression read score is below a predetermined number of expression reads, then the respective sample fails the respective single sample quality control test.
16. 16. The method of any one of claims 1 to 15, wherein the one or more global batch quality control tests comprise tests for one or more batch effects from a set including bioinformatics pipeline analysis, DNA contamination, sample processing, and sequencing methods.
17. Each global batch quality control test is i) determining the average number of sequence reads per sample across the batch data set; ii) obtaining a reference average number of sequence reads per sample from a reference dataset; and 17. The method of claim 16, comprising: iii) comparing the average number of sequence reads across the batch data set to a reference average number of sequence reads per sample, wherein if the average number of sequence reads is below the reference average number of sequence reads per sample, the batch data set fails the respective global batch quality control test.
18. Each global batch quality control test is for each respective sample in the batch dataset, applying the corresponding plurality of sequence reads and corresponding metadata to a first trained classification model, whereby the first trained classification model provides a set of predicted gender assignments including a respective predicted gender assignment for each sample; and comparing the set of predicted gender assignments to an expected set of gender assignments, wherein each sample having a respective predicted gender assignment that does not match the expected set of gender assignments fails the respective global batch quality test.
19. 19. The method of any one of claims 1 to 18, further comprising determining a linear or non-linear combination of the plurality of sequence features of the batch and the plurality of metadata features of the batch by subjecting the plurality of sequence features of the batch and the plurality of metadata features of the batch to a dimensionality reduction technique.
20. 20. The method of any one of claims 1-19, wherein the method further comprises: (c) adjusting each sample in the batch dataset for one or more confounding covariates using the cohort-matched reference batch prior to performing the one or more global batch quality control tests.
21. 21. The method of claim 20, wherein at least one sample in the batch data set is a control sample.
22. 22. The method of claim 21, wherein the at least one control sample in the batch data set is used to adjust each other sample in the batch data set.
23. 23. The method of any one of claims 1-22, wherein the method further comprises providing, for each sample in the batch of samples, a respective sample report, each respective sample report comprising at least one of a set of expression calls, one or more matched therapies, or one or more matched clinical trials.
24. The method according to any one of claims 1 to 23, wherein the method is implemented in a computer system comprising a cloud server.
25. the one or more global batch quality control tests include a first module; The method of any one of claims 11 to 24, wherein the one or more single sample quality control tests comprises a second module.
26. 1. A method for performing quality control, the method comprising:
1. A computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, a) for each respective test sample within the batch of test samples: a corresponding expression profile comprising a corresponding gene expression value for each respective gene in the first set of genes; and and a corresponding set of metadata including a value for each respective feature in the first set of features for each of the test samples; b) determining for said batch dataset a cohort-matched reference dataset comprising, for each respective reference sample in a plurality of reference samples, a corresponding expression profile comprising a corresponding gene expression value for each respective gene in said first set of genes, each respective reference sample in the plurality of reference samples is associated with a corresponding set of metadata including a corresponding value for each respective feature in a second set of features for the respective reference sample; determining that a summary value for each respective feature in a third set of one or more features that are present in both the first set of features and the second set of features is balanced between the batch dataset and the cohort-matched reference dataset; c) performing a dimensionality reduction on a combined dataset consisting of the corresponding expression profile for each respective test sample in the plurality of test samples and the corresponding expression profile for each respective reference sample in the plurality of reference samples, thereby obtaining, for each respective test sample and each respective reference sample, a corresponding set of coordinates embedded in a dimensional space lower than the dimensionality of the corresponding expression profile; d) determining a statistical measure of similarity between said set of coordinates obtained for said test sample and said set of coordinates obtained for said reference sample; e) comparing said statistical measure of similarity to a threshold; validating said batch data set for reporting if said statistical measure of similarity meets said threshold; or and not validating the batch data set for reporting if the statistical measure of similarity does not meet the threshold.
27. 27. The method of claim 26, further comprising flagging the batch data set for further review if the statistical measure of similarity does not meet the threshold.
28. 1. A method for validating changes in an RNA expression pipeline, the method comprising:
1. A computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, a) obtaining a batch dataset in electronic form, the batch dataset comprising, for each respective test sample in a batch of test samples: a corresponding expression profile prepared using the first RNA expression pipeline, the corresponding expression profile including a corresponding gene expression value for each respective gene in the first set of genes; a corresponding set of metadata including a value for each respective feature in the first set of features for each test sample; b) determining for said batch dataset a cohort-matched reference dataset comprising, for each respective reference sample in the plurality of reference samples, a corresponding expression profile comprising corresponding gene expression values for each respective gene in said first set of genes prepared using a second RNA expression pipeline; each respective reference sample in the plurality of reference samples is associated with a corresponding set of metadata including a corresponding value for each respective feature in a second set of features for the respective reference sample; determining that a summary value for each respective feature in a third set of one or more features that are present in both the first set of features and the second set of features is balanced between the batch dataset and the cohort-matched reference dataset; c) performing a dimensionality reduction on a combined dataset consisting of the corresponding expression profile for each respective test sample in the plurality of test samples and the corresponding expression profile for each respective reference sample in the plurality of reference samples, thereby obtaining, for each respective test sample and each respective reference sample, a corresponding set of coordinates embedded in a dimensional space lower than the dimensionality of the corresponding expression profile; d) determining a statistical measure of similarity between said set of coordinates obtained for said test sample and said set of coordinates obtained for said reference sample; e) comparing said statistical measure of similarity to a threshold; if the statistical measure of similarity meets the threshold, validating a change in the RNA expression pipeline; or if the statistical measure of similarity meets the threshold, then not validating the change in the RNA expression pipeline.
29. 29. The method of claim 28, further comprising: if the statistical measure of similarity does not meet the threshold, flagging the change in the RNA expression pipeline for further evaluation if the statistical measure of similarity does not meet the threshold.
30. 30. The method of claim 28 or 29, wherein the cohort-matched reference dataset is a sample-paired reference dataset.
31. 1. A method for adding RNA expression data to a reference database, the method comprising:
1. A computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, a) obtaining a new expression dataset in electronic form, said new expression dataset comprising, for each respective test sample in a plurality of test samples: a corresponding expression profile prepared using the first RNA expression pipeline, the corresponding expression profile including a corresponding gene expression value for each respective gene in the first set of genes; a corresponding set of metadata including a value for each respective feature in the first set of features for each test sample; b) for each respective reference sample in the plurality of reference samples, a cohort-matched reference dataset comprising a corresponding expression profile comprising a corresponding gene expression value for each respective gene in said first set of genes is calculated for said new expression dataset. To determine, each respective reference sample in the plurality of reference samples is associated with a corresponding set of metadata including a corresponding value for each respective feature in a second set of features for the respective reference sample; each expression profile corresponding to a reference sample in said plurality of reference samples is from said reference database; determining that a summary value for each respective feature in a third set of one or more features that are present in both the first set of features and the second set of features is balanced between the batch dataset and the cohort-matched reference dataset; c) performing a dimensionality reduction on a combined dataset consisting of the corresponding expression profile for each respective test sample in the plurality of test samples and the corresponding expression profile for each respective reference sample in the plurality of reference samples, thereby obtaining, for each respective test sample and each respective reference sample, a corresponding set of coordinates embedded in a dimensional space lower than the dimensionality of the corresponding expression profile; d) determining a statistical measure of similarity between said set of coordinates obtained for said test sample and said set of coordinates obtained for said reference sample; e) comparing said statistical measure of similarity to a threshold; adding the new expression dataset to the reference database if the statistical measure of similarity meets the threshold; or if the statistical measure of similarity does not meet the threshold, determining a set of conversion factors for standardizing the expression profiles in the new expression dataset to expression profiles in the reference database; standardizing the expression profiles in the new expression dataset using the set of transformation coefficients, thereby obtaining a standardized new expression dataset; and adding said new standardized expression dataset to said reference database.
32. Obtaining the batch data set comprises, for each respective sample in the batch of samples: obtaining, in electronic form, a corresponding plurality of sequence reads obtained from each of said samples by targeted or whole-transcriptome RNA sequencing; and determining the corresponding gene expression value for each respective gene in the first set of genes from the corresponding plurality of sequence reads.
33. 33. The method of Claim 32, wherein for each respective sample in the batch of samples, the corresponding plurality of sequence reads comprises at least 10,000 sequence reads.
34. The method of any one of claims 26 to 33, wherein the batch of test samples comprises at least 10 test samples.
35. 35. The method of any one of claims 26 to 34, wherein said first set of genes comprises at least 10 genes.
36. 35. The method of any one of claims 26 to 34, wherein said first set of genes comprises at least 10,000 genes.
37. The third set of features may include tissue site, tumor purity, cancer type, sequencer identity, 37. The method of any one of claims 26 to 36, comprising features selected from the group consisting of: and sequencing date.
38. 38. The method of any one of claims 26 to 37, wherein the third set of features comprises features selected from a nucleic acid extraction method, a cDNA library preparation method, an RNA sequencing method, types of reagents used, and types of equipment used.
39. The method of any one of claims 26 to 38, wherein the plurality of reference samples comprises at least 100 reference samples.
40. The method of any one of claims 26 to 38, wherein the plurality of reference samples comprises at least 1000 reference samples.
41. 41. The method of any one of claims 26-40, wherein the aggregated values for each feature are balanced between the batch data set and the cohort-matched reference set if the percentage of respective test samples, within the batch of test samples, having the respective value for the respective feature is within 2.5% of the percentage of respective reference samples, in the plurality of reference samples, having the same respective value for the respective feature.
42. 42. The method of any one of claims 26 to 41, wherein dimensionality reduction comprises embedding, for each respective test sample and each respective reference sample, the corresponding expression profiles into a two-dimensional representation.
43. 43. The method of any one of claims 26-42, wherein for each respective test sample in a batch of test samples, the corresponding expression profile is determined from sequence reads obtained from said respective sample by targeted or whole-transcriptome RNA sequencing.
44. Targeted panel RNA sequencing uses multiple probes, each probe in said plurality of probes uniquely targets a respective portion of a reference transcriptome; 44. The method of Claim 43, wherein each sequence read in the corresponding plurality of sequence reads corresponds to at least one probe in the plurality of probes.
45. 44. The method of claim 43, wherein the whole transcriptome sequencing comprises next generation sequencing.
46. For each respective sample within said batch of samples, performing one or more single sample quality control tests on said respective sample; 46. The method of any one of claims 26-45, further comprising removing each sample from the batch of samples that fails any one of the one or more single-sample quality control tests or flagging for manual inspection each sample that fails any one of the one or more single-sample quality control tests.
47. each single sample quality control test in the one or more single sample quality control tests i) determining, for each sample in the batch of samples, a respective number of non-redundant mapped sequence reads in the plurality of sequence reads, where each non-redundant mapped sequence read maps to a corresponding portion of a reference genome; ii) dividing each number of non-redundant mapped sequence reads into a predicted and comparing the respective number of non-redundant mapped sequence reads to a predetermined number of non-redundant mapped sequence reads, wherein if the respective number of non-redundant mapped sequence reads is below a predetermined number of non-redundant mapped reads, the respective sample fails the respective single sample quality control test.
48. 48. The method of claim 46 or 47, wherein each single sample quality control test in the one or more single sample quality control tests comprises determining, for each sample in the batch of samples, a respective percentage of properly matched sequence reads, wherein if the percentage of properly matched sequence reads is below a predetermined matched read threshold, the respective sample fails the respective single sample quality control test.
49. 49. The method of any one of claims 46-48, wherein each single sample quality control test in the one or more single sample quality control tests comprises determining, for each sample in the batch of samples, a GC content of each of the corresponding plurality of sequence reads, wherein if the respective GC content is outside a predetermined GC content threshold, the respective sample fails the respective single sample quality control test.
50. 50. The method of any one of claims 46-49, wherein each single sample quality control test in the one or more single sample quality control tests comprises determining, for each sample in the batch of samples, a respective number of expressed genes, wherein if the corresponding expression read score is below a predetermined number of expression reads, then the respective sample fails the respective single sample quality control test.
51. 51. The method of any one of claims 26-50, wherein the method further comprises providing, for each sample in the batch of samples, a respective sample report, each respective sample report comprising at least one of a set of expression calls, one or more matched therapies, or one or more matched clinical trials.
52. The method of any one of claims 26 to 51, wherein the method is implemented in a computer system comprising a cloud server.
53. 1. A non-transitory computer-readable storage medium storing at least one program for performing quality control, the at least one program being configured to be executed by a computer, the at least one program comprising: a) obtaining, for each respective sample in a batch of samples, a batch dataset in electronic format, the batch dataset including a corresponding plurality of sequence reads obtained from said respective sample by targeted or whole-transcriptome RNA sequencing and corresponding metadata for said respective sample; b) determining a cohort-matched reference batch for the batch data set, the cohort-matched reference batch being balanced for tissue site, tumor purity, cancer type, sequencer identity, or sequencing date; c) performing one or more global batch quality control tests on said batch data set using at least said cohort-matched reference batch; d) removing each sample from the batch data set that fails any one of the one or more global batch quality control tests or flagging for manual inspection each sample that fails any one of the one or more global batch quality control tests.
54. 1. A computer system comprising: at least one processor; a memory for storing at least one program executed by said at least one processor; The at least one program a) obtaining, for each respective sample in a batch of samples, a batch dataset in electronic format, the batch dataset including a corresponding plurality of sequence reads obtained from said respective sample by targeted or whole-transcriptome RNA sequencing and corresponding metadata for said respective sample; b) determining a cohort-matched reference batch for the batch data set, the cohort-matched reference batch being balanced for tissue site, tumor purity, cancer type, sequencer identity, or sequencing date; c) performing one or more global batch quality control tests on said batch data set using at least said cohort-matched reference batch; d) removing each sample from the batch data set that fails any one of the one or more global batch quality control tests or flagging for manual inspection each sample that fails any one of the one or more global batch quality control tests.
55. A non-transitory computer readable storage medium storing at least one program configured for execution by a computer, said at least one program comprising instructions for performing a method according to any one of claims 1 to 52.
56. 1. A computer system comprising: at least one processor; a memory for storing at least one program executed by said at least one processor; A computer system, wherein said at least one program comprises instructions for carrying out the method according to any one of claims 1 to 52.
Citation Information
Patent Citations
Methods and systems for identifying a physiological state of a target cell
US20160026754A1