Systems and methods for cell-free RNA sequencing
Targeted sequencing of rare gene transcripts in cell-free RNA using a panel of nucleic acid molecules addresses the low signal strength issue, improving diagnostic sensitivity for medical disorders and cancer detection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
- Filing Date
- 2024-05-15
- Publication Date
- 2026-05-29
AI Technical Summary
Existing methods for sequencing cell-free RNA in liquid biopsies face challenges due to low signal strength from hematopoietic cells, making it difficult to detect rare transcripts associated with medical disorders and cancer, necessitating improved methods to enhance detection sensitivity.
Targeted sequencing of rare gene transcripts (RAGs) using a panel of nucleic acid molecules, such as capture probes or primers, to enrich for low-abundance cfRNA molecules, excluding those from hematopoietic cells, and performing differential transcript analysis to improve diagnostic accuracy.
Enhances the detection of rare cfRNA molecules, improving diagnostic sensitivity for medical disorders and cancer by reducing interference from hematopoietic cell-derived signals, thereby enhancing the detection of biological phenomena and medical conditions.
Smart Images

Figure 2026517402000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 502,368, filed May 15, 2023, entitled "Systems and Methods for Sequencing of Cell - Free RNA", the disclosure of which is hereby incorporated by reference in its entirety.
[0002] Sequence Listing This application includes a sequence listing submitted electronically in XML format, which is hereby incorporated by reference in its entirety. A copy of the XML created on May 14, 2024, is named 08574 Seq Listing.xml and is 4,382 bytes in size.
[0003] Statement Regarding Federally Sponsored Research or Development This invention was made with government support under Contract No. CA254179 awarded by the National Institutes of Health. The government has certain rights in this invention.
[0004] Technical Field This disclosure provides an improved method for performing sequencing on cell - free RNA.
Background Art
[0005] Background Liquid biopsies based on blood enable non - invasive medical characterization, including the detection of biological phenomena such as pregnancy and cancer. Liquid biopsies offer many advantages over tissue biopsies because they are minimally invasive, easily repeatable over time, and more accurately reflect the geographic heterogeneity in cell sources. In patients with progressive cancer diseases, analysis of circulating tumor DNA (ctDNA) is clinically used for non - invasive genotyping. However, for a sufficient clinical assessment, for example, to identify tumor types or subtypes, expression - based analysis is usually required. [Overview of the Initiative] [Means for solving the problem]
[0006] overview In some implementations, methods for sequencing cell-free RNA include a step of performing targeted sequencing. This sequencing targets cell-free RNA molecules that are rarely expressed in liquid biopsies of control organisms.
[0007] In some implementations, the nucleic acid panel is designed to target transcripts that are rare as cell-free RNA molecules.
[0008] In some implementations, the nucleic acid panel includes nucleic acid molecules derived from or having sequences complementary to gene transcripts that are rare as cell-free RNA molecules in a control liquid biopsy.
[0009] In some implementations, transcripts that are rarely found as cell-free RNA molecules in a control fluid biopsy are defined as transcripts expressed in less than 50% of the control fluid biopsy population.
[0010] In some implementations, transcripts that are rarely found as cell-free RNA molecules in control fluid biopsies are defined as transcripts expressed in less than 5% of the control fluid biopsy population.
[0011] In some implementations, transcripts that are rare as cell-free RNA molecules within a control fluid biopsy are defined as transcripts that fall within the bottom 60% of gene expression relative to normalized expression across the control fluid biopsy population.
[0012] In some implementations, transcripts that are rare as cell-free RNA molecules in a control fluid biopsy are defined as transcripts that fall into the bottom 30% of gene expression relative to the control fluid biopsy population when normalized.
[0013] In some implementations, transcripts that are rare as cell-free RNA molecules in a control fluid biopsy are defined as transcripts with logarithmically transformed and normalized expression values less than zero across the control fluid biopsy population.
[0014] In some implementations, the control fluid biopsy population includes at least five different fluid biopsies.
[0015] In some implementations, the control fluid biopsy population includes at least 50 different fluid biopsies.
[0016] In some implementations, control fluid biopsies are collected from individuals who, at the time the biopsy is collected, do not have one or more of the following: observed pathogenic infection, diagnosed cancer, diagnosed metabolic disorder, diagnosed neurological disorder, diagnosed immunodeficiency disorder, diagnosed autoimmune disorder, diagnosed inflammatory disorder, diagnosed cardiovascular disorder, diagnosed renal disorder, diagnosed hepatic disorder, active pregnancy, diagnosed pregnancy complications, diagnosed fetal complications, organ transplant, active organ transplant rejection, obesity, malnourishment, cachexia, or abnormalities in clinical trials.
[0017] In some implementations, transcripts that are rarely found as cell-free RNA molecules in control fluid biopsies are defined as transcripts expressed in less than 5% of the control fluid biopsy population.
[0018] In some implementations, transcripts that are rare as cell-free RNA molecules within a control fluid biopsy are defined as transcripts that fall within the bottom 60% of gene expression relative to normalized expression across the control fluid biopsy population.
[0019] In some implementations, transcripts that are rare as cell-free RNA molecules in a control fluid biopsy are defined as transcripts that fall into the bottom 30% of gene expression relative to the control fluid biopsy population when normalized.
[0020] In some embodiments, transcripts that are rare as cell-free RNA molecules in a control liquid biopsy are defined as transcripts with log-transformed and normalized expression values less than zero across the population of control liquid biopsies.
[0021] In some embodiments, the population of control liquid biopsies includes at least five liquid biopsies.
[0022] In some embodiments, the population of control liquid biopsies includes at least fifty liquid biopsies.
[0023] In some embodiments, a control liquid biopsy is collected from an individual who does not have one or more of the following abnormalities at the time the biopsy is collected: observed pathogenic infection, diagnosed cancer, diagnosed metabolic disorder, diagnosed neuropathy, diagnosed immunodeficiency disorder, diagnosed autoimmune disorder, diagnosed inflammatory disorder, diagnosed cardiovascular disorder, diagnosed kidney disorder, diagnosed liver disorder, active pregnancy, diagnosed pregnancy complication, diagnosed fetal complication, organ transplantation, active rejection of an organ transplantation, obesity, malnourishment, cachexia, and an abnormality during a clinical trial.
[0024] In some embodiments, transcripts that are rare as cell-free RNA molecules in a control liquid biopsy include 50% of the transcripts in Table 3.
[0025] In some embodiments, transcripts that are rare as cell-free RNA molecules in a control liquid biopsy include 90% of the transcripts in Table 3.
[0026] In some embodiments, transcripts that are rare as cell-free RNA molecules in a control liquid biopsy include 100% of the transcripts in Table 3.
[0027] In some embodiments, a panel of nucleic acid molecules excludes at least 50% of all exome gene transcripts that are not transcripts that are rare as cell-free RNA molecules.
[0028] In some implementations, the nucleic acid molecule panel excludes at least 90% of all exome gene transcripts that are not transcripts, as they are rare in existence as cell-free RNA molecules.
[0029] In some implementations, the panel of nucleic acid molecules consists of 5,000 or fewer gene transcripts, in addition to transcripts that are rare as cell-free RNA molecules.
[0030] In some implementations, the panel of nucleic acid molecules consists of 500 or fewer gene transcripts, in addition to transcripts that are rare as cell-free RNA molecules.
[0031] In some implementations, the panel of nucleic acid molecules further includes tissue-specific transcripts, cell type-specific transcripts, clinically relevant transcripts, B cell receptor and T cell receptor transcripts, biomarkers, and generally mutagenic transcripts.
[0032] In some implementations, biomarkers are associated with one of the following biological characteristics: medical disorders, pregnancy, fetal complications, pregnancy complications, neoplasm growth, cancer, specific cancer types, pathogenic infections, immune activation, organ transplant rejection, neurodegeneration, originating tissue, originating cell type, or activation of biochemical pathways.
[0033] In some implementations, the panel of nucleic acid molecules further includes a set of control transcripts for normalization between samples.
[0034] In some implementations, a panel of nucleic acid molecules is a set of probes for targeted capture hybridization.
[0035] In some implementations, a panel of nucleic acid molecules is a set of primers for targeted amplification.
[0036] In some implementations, a method for preparing cell-free RNA sequencing includes the step of providing a sample containing nucleic acids for sequencing.
[0037] In some implementations, the nucleic acid used for sequencing is cell-free RNA, or a representative nucleic acid derived from cell-free RNA.
[0038] In some implementations, a method for preparing cell-free RNA sequencing involves contacting a sample with a panel of nucleic acid molecules containing sequences derived from or complementary to gene transcripts that are rare as cell-free RNA molecules in a control liquid biopsy, thereby annealing the panel of nucleic acid molecules with a subset of nucleic acids for sequencing.
[0039] In some implementations, cfRNA samples are derived from blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, or feces.
[0040] In some implementations, methods for preparing cell-free RNA sequencing include the step of generating a sequencing library derived from the sample.
[0041] In some implementations, the method for preparing cell-free RNA sequencing includes the step of performing targeted sequencing of a sequencing library to obtain cell-free RNA sequencing results. The sequencing is targeted against a panel of nucleic acid molecules.
[0042] In some implementations, methods for preparing cell-free RNA sequencing include a step of removing platelet expression from in silico sequencing results.
[0043] In some implementations, the method for preparing cell-free RNA sequencing includes the step of performing differential transcript analysis using the sequencing results and the second sequencing results.
[0044] In some implementations, a method for preparing cell-free RNA sequencing includes the step of detecting enrichment of at least one expression signature within the sequencing results.
[0045] In some implementations, methods for preparing cell-free RNA sequencing include a step of detecting sequence mutagenicity within the sequencing results.
[0046] In some implementations, methods for preparing cell-free RNA sequencing include a step of inferring the copy number state of one or more genes from the sequencing results.
[0047] In some implementations, a method for preparing cell-free RNA sequencing includes a step of using the sequencing results together with several other sequencing results to train a computational model to predict the categorical state or likelihood of the biological properties. The cell-free RNA sample has a known categorical state of biological properties.
[0048] In some implementations, the method for preparing cell-free RNA sequencing includes a step of using the sequencing results as input to a trained computational model to predict the categorical state or likelihood of a biological trait. The computational model is trained using a cohort of RNA sequencing results that have known categorical states of biological traits.
[0049] In some implementations, a method for preparing cell-free RNA sequencing includes a step of deriving one or more features from the sequencing results. These features may include enrichment of one or more gene signatures, enrichment of biochemical pathways, a collection of sequence variants, and copy number states.
[0050] In some implementations, methods for preparing cell-free RNA sequencing include a step of using one or more derived features as input to a trained computational model to predict the categorical state or likelihood of a biological trait. The computational model is trained using a cohort of RNA sequencing results that have known categorical states for the biological traits.
[0051] In some implementations, a method for extracting RNA from a cell-free source includes the step of adding glycogen to a sample containing cell-free nucleic acids.
[0052] In some implementations, a method for extracting RNA from a cell-free source includes the step of (b) contacting a silica column with a sample containing cell-free nucleic acids.
[0053] In some execution modes, step (a) is performed before step (b).
[0054] In some implementations, a method for extracting RNA from a cell-free source includes the step of eluting the cell-free nucleic acid from a silica column to obtain a solution of the extracted cell-free nucleic acid.
[0055] In some implementations, a method for extracting RNA from a cell-free source includes the step of contacting a solution of the extracted cell-free nucleic acid with DNase.
[0056] In some implementations, a method for quantifying cell-free RNA for downstream molecular application includes the step of providing a sample containing cell-free RNA.
[0057] In some implementations, methods for quantifying cell-free RNA for downstream molecular applications include the step of reverse transcribing the cell-free RNA to obtain cDNA.
[0058] In some implementations, methods for quantifying cell-free RNA for downstream molecular applications include a step of quantifying the concentration of cell-free RNA in solution using quantitative real-time polymerase chain reactions and cDNA.
[0059] In some implementations, a method for quantifying cell-free RNA for downstream molecular application includes the step of using a predetermined amount of substance and performing one or more downstream steps of a molecular protocol based on the quantification of cell-free RNA.
[0060] In some implementations, the step of quantifying the concentration of cell-free RNA further includes a step of generating a standard curve based on a set of control standards with known concentrations. The control standards are also evaluated using quantitative real-time polymerase chain reactions.
[0061] In some implementations, the sample further includes cell-free DNA. Methods for quantifying cell-free RNA for downstream molecular applications further include the step of quantifying the concentration of cell-free DNA in the sample using a quantitative real-time polymerase chain reaction. Cell-free RNA is quantified by using primers that span gene introns that are relatively stable across cell-free RNA samples, and cell-free DNA is quantified by using primers that anneal to transcriptionally silent regions of the genome that are relatively stable across cell-free DNA samples.
[0062] In some implementations, primers for quantifying cell-free RNA span the introns of GAPDH, while primers for quantifying cell-free DNA targets cover a 78 bp transcriptionally silent region on chromosome 12.
[0063] In some implementations, methods for sequencing cell-free RNA include a step of providing a library of nucleic acid molecules. This library of nucleic acid molecules is derived from cell-free RNA, which is obtained from a liquid biopsy.
[0064] In some implementations, methods for sequencing cell-free RNA include the step of sequencing a library of nucleic acid molecules to obtain sequencing results.
[0065] In some implementations, methods for sequencing cell-free RNA include a step of removing variations caused by the expression of platelets and associated transcripts.
[0066] In some implementations, libraries of nucleic acid molecules were generated by capturing or amplifying nucleic acid molecules.
[0067] In some execution models, the nucleic acid library is a whole-exome library.
[0068] In some implementations, nucleic acid libraries are libraries that target genes that are rare in existence.
[0069] In some implementations, a method for generating a targeted sequencing panel for cell-free RNA sequencing includes the step of collecting a population of control liquid biopsies, each containing cell-free RNA.
[0070] In some implementations, a method for generating a targeted sequencing panel for cell-free RNA sequencing includes the step of performing sequencing on cell-free RNA from a control liquid biopsy.
[0071] In some implementations, a method for generating a targeted sequencing panel for cell-free RNA sequencing includes the step of identifying a set of genes that are rare in a control fluid biopsy population, defined by at least one or more of their expression levels in a certain percentage of the control fluid biopsy population or across the control fluid biopsy population.
[0072] In some implementations, a method for generating a targeted sequencing panel for cell-free RNA sequencing includes the step of synthesizing a set of nucleic acid molecules that are for capturing or amplifying genes that are rare in presence, in order to obtain a targeted sequencing panel for cell-free RNA sequencing.
[0073] In some implementations, a set of rare genes is defined at least by its expression within a certain percentage of the control fluid biopsy population and its expression level across the control fluid biopsy population.
[0074] This specification and the claims shall be understood more fully with reference to the following drawings and data, which are presented as examples of the disclosure and should not be construed as a complete enumeration of the scope of the disclosure. [Brief explanation of the drawing]
[0075] [Figure 1-1] Figures 1A-1F present schematic diagrams and data charts for the optimization of blood collection and cfRNA extraction. Figure 1A: Representative bioanalyzer traces of cell-free RNA (yellow) and leukocyte RNA (red). Figure 1B: cfRNA concentration per 1 mL of plasma in healthy controls measured by quantitative PCR (see methods; n=117). Plasma was collected for technical experiments using 2,500 G centrifugation for 10 minutes. Figure 1C: Relationship between hemolysis and cfRNA plasma concentration. Hemolysis was measured using optical density (OD) at 414 nm. Pearson and Spearman correlations are shown. Figure 1D: Relationship between time in a -80°C freezer and cfRNA plasma concentration. Pearson and Spearman correlations are shown. Figure 1E: Analysis of plasma cfRNA concentration using various blood collection tubes (BCTs). All samples were rotated at 2500 G during plasma isolation. Ro, Roche Cell-free DNA BCT. D, Streck Cell-free DNA BCT. R, Streck RNA Complete BCT. E, EDTA. Figure 1F, Analysis of plasma cfRNA concentration using various extraction methods. TR, TRIzol-LS + Qiagen RNeasy Kit. V, QIAamp Viral RNA Kit. Ro, Roche High Pure Vial RNA Kit. M, Qiagen miRNeasy Kit. MPS, Qiagen miRNeasy Serum / Plasma Kit. TM, TRIzol-LS + Qiagen miRNeasy Kit. P, mirVana PARIS Kit. CCF, QIAamp ccfDNA / RNA Kit. TV, TRIzol-LS + QIAamp Viral RNA Kit. CNA, QIAamp Circulating Nucleic Acid Kit. [Figure 1-2] Same as above.
[0076] [Figure 2-1]Figures 2A–2K present data charts for optimizing RARE-Seq library preparation and capture. Figure 2A: Correlation between cfRNA expressions in blood samples collected in EDTA tubes or Streck RNA Complete tubes (mean values for n=3 pairs). Figure 2B: Expression correlation of cfRNAs extracted using the TRIzol LS + QIAamp Viral RNA method (TV) and the QIAamp Circulating Nucleic Acid Kit (CNA) (mean values for n=3 pairs). Figure 2C: Expression correlation between cfRNA libraries generated using the stranded method and the non-stranded method (mean values for n=3 pairs). Figure 2D: Dilution analysis showing the relationship between total sequencing depth and unique sequencing depth in cfRNA libraries. Inserted plots show the percentage increase in unique sequencing depth for non-stranded libraries compared to their paired stranded libraries. Sequencing depth was downsampled so that pairs had equivalent depths. Figure 2E, Expression correlation between cfRNA libraries generated with and without S1 nuclease end repair (mean of n=3 pairs). Figure 2F, Dilution analysis showing the relationship between total sequencing depth and unique sequencing depth of cfRNA libraries. Log2NX, log2 normalized expression. Figure 2G, Estimated DNA contamination when DNase I digestion is performed either on-Qiagen column or off-column. DNA contamination was measured as the percentage of exon boundary aligned reads containing adjacent intron sequences. Conditions were compared using paired t-tests. Figure 2H, Expression correlation between cfRNA samples digested on-column and off-column (mean of n=5 pairs). Figure 2I, Distribution of RNA biotypes in cfRNA using the SMART-Seq whole transcriptome method (n=3). Figure 2J, Expression correlation between cfRNA libraries generated using SMART-Seq and whole coding transcriptome RARE-Seq (mean of n=3 pairs). The histograms represent the log2nx distribution for each method, and the Venn diagrams represent the number of coding genes with high (or equivalent) log2NX for each method.The sequencing depth was downsampled so that pairs had equivalent depths. Blue = coding genes, gray = non-coding genes. Figure 2K, heatmap showing the Pearson correlation between pairs of all technical and biological replicas sequenced from a single individual (n=8 replicas). [Figure 2-2] Same as above. [Figure 2-3] Same as above. [Figure 2-4] Same as above.
[0077] [Figure 3-1]Figures 3A-3L present schematic diagrams and data showing that platelet and non-hematopoietic transcripts are enriched with cfRNA. Figure 3A: Schematic diagram of the RARE-Seq method. Figure 3B: Differential expression analysis (n=10 pairs) comparing cfRNA from healthy donors with matched leukocyte RNA using whole-code transcriptome RARE-Seq. Figure 3C: Pre-ranked gene set enrichment analysis (n=170 cell types) using cell type-specific signatures from PanglaoDB. The top 15 positively enriched signatures and the top 15 negatively enriched signatures are shown, with dotted lines indicating signatures with significant enrichment (P<0.05). Figure 3D: Relationship between centrifugation rate during plasma isolation and cfRNA concentration per 1 mL of plasma (n=10 1200G, n=6 1800G, n=16 2500G). Comparisons were performed using the Kruskal-Wallis test. Figure 3E: Relationship between centrifugation speed and platelet-specific gene expression in cfRNA. Mean log2NX, normalized expression mean. Comparison was performed using the Kruskal-Wallis test. Figure 3F: Principal component analysis of cfRNA expression, clustering cfRNA samples according to rotation speed (PC1) and sex (PC2). Figure 3G: Relationship between mean platelet-specific gene expression in Figure 3E and the first principal component in Figure 3F. Pearson and Spearman correlations are shown. Figure 3H: Relationship between mean platelet-specific gene expression and the first principal component when only samples processed using the same centrifugation speed (2500G) were considered. Figure 3I: Schematic diagram of the platelet transcript correction method. See Methods for details. Figure 3J: Relationship between platelet-specific gene expression and the first principal component after platelet transcript correction of cfRNA expression. Figure 3K: Hierarchical clustering of the top 1,000 most variably expressed genes in cfRNA. Clustering was performed using Euclidean distance. Figure 3L shows hierarchical clustering of the top 1,000 genes that exhibited the most variability after platelet-specific correction. [Figure 3-2] Same as above. [Figure 3-3] Same as above. [Figure 3-4] Same as above. [Figure 3-5] Same as above.
[0078] [Figure 4-1] Figures 4A-4I present schematic diagrams and data demonstrating improved analytical sensitivity for lung cancer detection by targeting transcripts not present in healthy cfRNAs. Figure 4A: Schematic diagram representing the selection of rare genes (RAGs) in plasma. Log2NX, log2 normalized expression. Figure 4B: Percentage of tissue-enriched genes and Figure 4C: Percentage of cancer-enriched genes in RAGs. Tissue-enriched and cancer-enriched genes were selected as defined in the methods. Comparison was performed using Fisher's exact test. Figure 4D: Unique depth in libraries captured using either the whole-code transcriptome capture panel or the RAG capture panel (n=5 fitted libraries). Libraries were downsampled to have equivalent sequencing depths and filtered to include regions captured by both panels. Comparison was performed using a paired t-test. Figure 4E: Number of co-genes detected in the libraries of Figure 4D. The detection threshold was ≥1 unique read. Comparison was performed again using a paired t-test. Histograms showing differences in unique depth per gene for the libraries in Figures 4F and 4D. Figure 4G, schematic diagram of the enrichment score (ES) analysis framework (see Methods). Figure 4H, association between ES detection rate and NCI-H1975 proportion for spikes captured using the RAG panel (n=3 replicas per spike). The presence of cancer RNA in each spike was detected using NCI-H1975 ES. LOD95 was calculated from in silico spikes using logistic regression (shown as gray circles). In vitro spike results are shown as red triangles. Figure 4I, empirical LOD of NCI-H1975 spikes (n=1 replica per spike) captured using the full code transcriptome panel, evaluated using the same method as in Figure 4F. [Figure 4-2] Same as above. [Figure 4-3] Same as above. [Figure 4-4] Same as above.
[0079] [Figure 5-1] Figures 5A-5I present data charts describing and validating the RARE-Seq panel targeting rare genes. Figure 5A shows the association (n=50) between the mean gene expression in healthy cfRNA and the overall percentage of healthy cfRNA samples with low or absent gene expression. Rare genes (RAGs) are shown in red, and other genes in gray. Log2NX, log2 normalized expression. Figure 5B shows the association (n=50) between the mean gene expression in healthy cfRNA and the expression stability in healthy cfRNA as measured by the Gini coefficient. Housekeeping genes added to the RAG-focused capture panel are shown in yellow, and other genes in gray. Figure 5C shows the cumulative frequency (n=122) of non-small cell lung cancer (NSCLC) patients with one or more gene variants, added to the RAG-focused capture panel because they are repeatedly modified or frequently abnormally expressed in NSCLC. The top 25 genes are shown. Modified data from the TCGA lung adenocarcinoma and lung squamous cell carcinoma project were downloaded from cBioPortal. Figure 5D, Venn diagram representing the gene set included in the RAG-focused capture panel (n=5,546 total genes). Figures 5E and 5F, expression correlations between whole-coding transcriptome capture and RAG capture for all covalent genes (Figure 5E) and housekeeping genes (Figure 5F) (n=5 pairs). Figure 5G, schematic diagram of a platelet transcript correction method using a meta-reference control with variable platelet expression. This method was developed for samples captured using a RAG-focused panel that excludes platelet genes. Figure 5H, association between platelet-specific gene expression and the first major component after platelet transcript correction of cfRNA expression shown in Figure 5G. The same control cohort as in Figures 3D-3K was used (n=32). Figure 5I, association between NCI-H1975 spike detection and total sequencing depth. Logistic regression was used to estimate the LOD95 of spikes in silico. [Figure 5-2] Same as above. [Figure 5-3] Same as above. [Figure 5-4] Same as above.
[0080] [Figure 6-1] Figures 6A–6L present data charts for the detection of lung adenocarcinoma ctRNA. Figures 6A, 6B, 6C, and 6D show the relationship between sample library metrics for healthy controls (n=24), low-dose computed tomography controls (LDCT, n=26), and lung adenocarcinoma patients (LUAD, n=50). Box plots show Figure 6A, cfRNA concentration (ng per 1 mL of plasma), Figure 6B, cfRNA mass used for library preparation, Figure 6C, total sequencing depth, and Figure 6D, unique sequencing depth. Each comparison was performed using the Kruskal-Wallis test. Figure 6E shows the cfRNA detection rate for unique sequencing depth per cfRNA sample. Figure 6F shows the cfRNA detection sensitivity summarized by driver oncogene subtype using LUAD Sig ES detection. A detection threshold achieving a specificity of 95% or higher was used. Figures 6G and 6H show representative Integrated Genomics Viewer (IGV) images of sequencing reads including EGFR exon 19 deletion (Figure 6G) and alternative splicing of MET exon 14 (Figure 6H). Figure 6I shows read counts mapped to standard splice junctions and fusion breakpoints of the ROS1 and CD74 genes in a representative sample with tissue-adjudicated ROS1 / CD74 gene fusions. Junctions / breakpoints with five or more aligned reads are shown. Figures 6J and 6K show the top 30 LUAD Elastic Net (EN) model coefficients (n=202 total features) selected by negative coefficients (Figure 6J) and positive coefficients (Figure 6K). Figure 6L shows the relationship between model training cohort size and 10-fold cross-validation model performance. The dotted line represents the 10-fold cross-validation (10CV) AUC of the final LUAD EN model, and the error bars represent the standard deviation of the 10CV AUC for each subsampling level (n=10 samples per level). [Figure 6-2] Same as above. [Figure 6-3] Same as above. [Figure 6-4] Same as above.
[0081] [Figure 7-1]Figures 7A-7F present data charts showing the detection of ctRNA in plasma derived from NSCLC patients. Figure 7A: Differential expression analysis of lung adenocarcinoma (LUAD) cfRNA (n=50) compared with non-cancer cfRNA (n=50). Figure 7B: Mean expression of enriched genes in LUAD cfRNA (n=94 genes) and meta-reference cfRNA (n=15) from LUAD tumor tissue (n=385). Figure 7C: Pre-ranked gene set enrichment analysis using cell type-specific signatures from PanglaoDB11. The top 15 positively enriched signatures and the top 15 negatively enriched signatures are shown, with dotted lines indicating signatures with significant enrichment (P<0.05). Figure 7D. Enrichment scores (ES) for LUAD-specific gene signatures (LUAD Sig; n=72 genes) in cfRNAs from LUAD patients (n=50), risk-matched controls (n=26), and healthy controls (n=24). LDCT, low-dose computed tomography. Figure 7E. cfRNA detection sensitivity, summarized by stage, using LUAD Sig ES detection. Detection thresholds were used to achieve specificity of ≥95%. Figure 7F. Association between LUAD Sig ES in cfRNA and mean variant allele frequencies (VAF) in matched cfDNA (n=15). Comparisons were made using Pearson and Spearman correlations. ND, not detected. [Figure 7-2] Same as above.
[0082] [Figure 8-1]Figures 8A–8G present schematic diagrams and data charts for the identification of somatically modified NSCLC ctRNAs. Figure 8A: Schematic diagram representing the ctRNA variant calling method (see Methods). Figure 8B: Oncoprints of single nucleotide variants (SNVs), insertions / deletions (indels), gene fusions, and splice variants found in cfRNA. Stage IV NSCLC samples (n=55) and risk-matched LDCT controls (n=36) with one or more tissue or ctDNA-determined variants were considered. Using healthy controls, variant-specific error rates were defined as described in Methods. Bar plots represent the percentage of samples with each mutation in each gene. LDCT, low-dose computed tomography. Figure 8C: Percentage of samples with one or more variants detected in cfRNA. Figure 8D: Percentage of tissue or ctDNA-determined variants detected in cfRNA and summarized by variant type. Figure 8E, association between EGFR expression levels in cfRNA and EGFR variant detection in samples with tissue or ctDNA-determined EGFR variants. Enrichment analysis was performed using fgsea. Figure 8F, predicted probability of lung cancer, calculated for each sample using a weighted elastic net (EN) model and summarized by stage. Results are shown for the training cohort (LUAD n=50, control n=50) and the pending validation cohort (LUAD n=28 time points with 14 individuals, control n=10). Figure 8G, receiver operating characteristic (ROC) curves summarizing cfRNA detection using LUAD Sig ES and EN probabilities. AUC was compared using DeLong's test. Val, validation. AUC, area under the ROC curve. ns, not significant. [Figure 8-2] Same as above. [Figure 8-3] Same as above.
[0083] [Figure 9-1]Figures 9A-9G present schematic diagrams and data charts regarding the detection mechanism of resistance to EGFR tyrosine kinase inhibitors in cfRNA. Figure 9A, Summary of the EGFR TKI cohort (n=10 patients, n=24 time points). Resistance mechanisms were determined via tissue biopsy. METamp, MET amplification. Histological transformation for C797S, EGFR C797S SNV, tSCLC, and small cell lung cancer. Figure 9B, Enrichment scores (ES) of small cell lung cancer (SCLC) gene signatures (SCLC Sig; n=73 genes) in patients with tSCLC diagnosed by biopsy. The dotted line indicates the threshold representing specificity of 95% or higher in the control. Figure 9C, Vignette of patients with tSCLC after EGFR TKI treatment. The left axis represents the ES or EN score, and the right axis represents the tumor-identified variant AF detected in cfRNA. Figure 9D, Enrichment scores (ES) of MET amplification-specific gene signatures (METamp Sig; n=9 genes) in patients with MET amplification diagnosed by biopsy. Figure 9E, Schematic diagram of a patient who developed MET amplification after EGFR TKI treatment. Figure 9F, EGFR C797S AF in cfRNA from a patient with C797S diagnosed by biopsy. Figure 9G, Schematic diagram of a patient who developed the EGFR C797S mutation after EGFR TKI treatment. [Figure 9-2] Same as above. [Figure 9-3] Same as above. [Figure 9-4] Same as above.
[0084] [Figure 10-1]Figures 10A-10H present data charts regarding the application of RARE-Seq. Figure 10A shows the sensitivity of cfRNA detection for LUAD Sig ES (n=50) in LUAD cfRNA, hepatocellular carcinoma (LIHC) Sig ES (n=10) in LIHC cfRNA, pancreatic adenocarcinoma (PAAD) Sig ES (n=10) in PAAD cfRNA, and prostate adenocarcinoma (PRAD) Sig ES (n=9) in PRAD cfRNA. A detection threshold was used to achieve a specificity of 90% or higher. Figure 10B shows the accuracy of origin tissue (TOO) determination using cancer TOO gene signatures. Shading in the bar plot represents the ES rank of the true cancer type. Samples were considered for TOO analysis when detected by any cancer-specific signature. Figure 10C shows the confusion matrix comparing predicted cancer types with clinically diagnosed cancer types. Figure 10D, Enrichment scores (ES) of normal lung tissue signatures (n=5 genes) in cfRNAs from patients with benign lung conditions (n=74 time points from 70 individuals). COPD, chronic obstructive pulmonary disease. COVID, COVID-19 infection. ARDS, acute respiratory distress syndrome. Figure 10E, Association between normal lung ES and patient smoking history. All samples from Figure 10D without known lung conditions were considered (n=42). Figure 10F, Association between normal lung ES and ventilator status. All samples from Figure 10D with known lung conditions were considered (n=32 time points from 28 individuals). Figure 10G, Time course of unique read counts in cfRNAs aligned to COVID-19 mRNA vaccine sequences, collected at various time points before vaccination and after two vaccine doses (n=9 time points). Log2NX, log2 normalized expression. Figure 10H, Gene ontology (GO) enrichment analysis of significantly differentially expressed genes at the post-vaccination time point. This shows the top 10 genes in the MSigDb Hallmark gene set and C5 collection. [Figure 10-2] Same as above. [Figure 10-3] Same as above.
[0085] [Figure 11]Figure 11 presents summaries of cancer cfRNA cohorts and non-cancer cfRNA cohorts. LDCT (Low-Dose Computed Tomography). ARDS (Acute Respiratory Distress Syndrome). COVID (COVID-19 Infection). VACC (Post-COVID-19 mRNA Vaccination). LUAD (Lung Adenocarcinoma). Treated with EGFR TKI (Epidermal Growth Factor Receptor Tyrosine Kinase Inhibitor). PAAD (Pancreatic Adenocarcinoma). PRAD (Prostate Adenocarcinoma). LIHC (Hepatocellular Carcinoma of the Liver). [Modes for carrying out the invention]
[0086] Detailed explanation Hereinafter, following the drawings and data, systems and methods for performing cell-free RNA (cfRNA) molecule sequencing are described. The systems and methods may include means for performing targeted sequencing of cfRNA. Furthermore, the systems and methods may include various steps and / or components for improving cfRNA processing and input.
[0087] When sequencing cfRNA from cell-free sources, analysis can be difficult because the nucleic acids in these sources lack high-quality nucleic acid molecules. Historically, the analysis of cfRNA from cell-free sources has focused on microRNAs (miRNAs). However, expressed nucleic acid molecules of other RNA types (e.g., messenger RNAs) can greatly improve the detection of biological phenomena and / or medical disorders, thus offering significant advantages in the field of diagnostics. Unfortunately, liquid biopsies derived from plasma primarily contain cfRNA molecules from hematopoietic cells, masking signals from other potential sources. Therefore, performing cfRNA analysis for diagnostic purposes, such as cancer detection, is difficult due to the low signal strength. Consequently, new systems and methods are needed to improve the detection of cfRNA from cell-free sources.
[0088] Here, the systems and methods are directed toward sequencing cfRNAs derived from cell-free sources (also referred to as liquid biopsy or excrement samples). Throughout this disclosure, the term liquid biopsy is used to refer to cfRNA sources (including excrement) and includes, but is not limited to, blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, and feces. The systems and methods can perform targeted sequencing to improve the detection of certain cfRNA molecules in liquid biopsies. In some embodiments, targeted sequencing includes a step of targeting cfRNA molecules that are rarely found in liquid biopsies of control individuals (e.g., healthy individuals). In some embodiments, a panel of capture probes or a panel of primer sets is used to target the desired cfRNA molecule for sequencing.
[0089] Various systems and methods are directed toward transcript panels and their use in methods for targeted sequencing of cfRNA molecules. A transcript panel may refer to a capture-based panel (e.g., ssDNA molecules for hybridization capture) or an amplification-based panel (e.g., a set of primers for amplification). Therefore, transcript panels can be utilized during the preparation of sequencing libraries for targeted sequencing.
[0090] In some implementations, the transcript panel includes rare genes (RAGs), which are transcripts (including non-coding transcripts) that rarely contain cfRNA molecules expressed in fluid biopsies of healthy individuals. Targeting RAGs may offer several advantages, including overcoming the drowning of signals mediated by hematopoietic cell-derived cfRNAs and improving the detection of various biological characteristics, including medical disorders, pregnancy, fetal complications, pregnancy complications, neoplasm growth, cancer, specific cancer types, pathogen infections, immune activation, organ transplant rejection, neurodegeneration, originating tissue, originating cell type, or activation of biochemical pathways.
[0091] In some implementations, RAGs are defined as genes expressed below a threshold in a control fluid biopsy. A population of individuals may be used to obtain control fluid biopsies for analysis. In some implementations, a control fluid biopsy is a biopsy from an individual that is generally in good health (a control fluid biopsy may also be called a healthy fluid biopsy). Generally good health may mean an individual that, at the time the biopsy is collected, does not have one or more of the following: observed pathogenic infections, diagnosed cancers, diagnosed metabolic disorders, diagnosed neurological disorders, diagnosed immunodeficiency disorders, diagnosed autoimmune disorders, diagnosed inflammatory disorders, diagnosed cardiovascular disorders, diagnosed renal disorders, diagnosed hepatic disorders, active pregnancy, diagnosed pregnancy complications, diagnosed fetal complications, organ transplants, active rejection of organ transplants, obesity, poor nutritional status, cachexia, and abnormalities in clinical trials. In some implementations, cfRNAs are collected to diagnose specific disorders. In some practices for diagnosing specific disorders (or traits), the control fluid biopsy is a biopsy taken from an individual that has not been diagnosed with the specific disorder.
[0092] To identify RAGs, any number of appropriate control fluid biopsies can be used to establish genes that are expressed below the threshold of the control fluid biopsies. In various cases, the number of control fluid biopsies used to identify RAGs is 5 or more, 10 or more, 15 or more, 20 or more, 50 or more, 100 or more, 200 or more, 500 or more, 1000 or more, 2000 or more, 500 or more, 1000 or more, 2000 or more, 5000 or more, 10,000 or more, 20,000 or more, 5000 or more, 10,000 or more, 20,000 or more, 50,000 or more, or 100,000 or more.
[0093] Various definitions of RAG can be used based on gene expression in a control fluid biopsy. In various implementations, RAG is defined as a transcript expressed in less than 50% of the control fluid biopsy, or in less than 40%, or in less than 30%, or in less than 20%, or in less than 10%, or in less than 5%, or in less than 1%. Some implementations utilize clustering techniques to categorize transcripts expressed in the control fluid biopsy from those not expressed in the control fluid biopsy.
[0094] In some implementations, RAG is defined by having an expression level below a threshold in a control fluid biopsy. In some implementations, expression values are normalized for comparison. In some implementations, expression values are logarithmically transformed for comparison (e.g., Log2NX). In some examples, RAG is defined as a transcript with an expression value Log2NX below a threshold (e.g., Log2NX<0). In some implementations, RAG is defined as a transcript with the minimum expression value relative to normalized expression in a control fluid biopsy. In various implementations, RAGs are defined as transcripts that fall in the bottom 60% of genes with respect to normalized expression, or as transcripts that fall in the bottom 50% of genes with respect to normalized expression, or as transcripts that fall in the bottom 40% of genes with respect to normalized expression, or as transcripts that fall in the bottom 30% of genes with respect to normalized expression, or as transcripts that fall in the bottom 20% of genes with respect to normalized expression, or as transcripts that fall in the bottom 10% of genes with respect to normalized expression, or as transcripts that fall in the bottom 5% of genes with respect to normalized expression.
[0095] In some implementations, a RAG is defined as being detected in a control fluid biopsy below the threshold of the control fluid biopsy and / or having an expression level below the threshold. Thus, any threshold defined herein for presence in a control biopsy can be combined with any threshold defined herein for expression level. For example, in the following Examples and Data section, a RAG is defined as a transcript expressed in less than 5% of control fluid biopsies and falling in the bottom 30% of genes with respect to normalized expression. Other definitions can be combined with RAGs for various applications, e.g., tissue-specific transcripts, cell-type-specific transcripts, transcripts with known clinical relevance, transcripts with known biological relevance (e.g., biomarkers), and transcripts with known and common mutagenicity profiles, e.g., fusion events and variants.
[0096] Table 3 presents an example list of RAGs defined as those identified in less than 5% of control fluid biopsies, with the mean log2NX falling in the bottom 30% of total genes, based on 307 samples of whole blood gene expression data from 50 replication and genotype-tissue expression (GTEx) projects from 28 healthy controls using whole exome sequencing (WES) of cell-free RNA (cfRNA). In some implementations, the transcript panel includes nucleic acid molecules to detect multiple RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 1% of the RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 5% of the RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 10% of the RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 20% of the RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 30% of the RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 40% of the RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 50% of the RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 60% of the RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 70% of the RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 80% of the RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 90% of the RAGs listed in Table 3. In some implementations, the transcript panel includes nucleic acid molecules to detect 100% of the RAGs listed in Table 3.
[0097] Transcript panels can target nucleic acids using capture or amplification techniques. To target cfRNA, some implementations first convert the RNA to cDNA before targeting. Other implementations target the cfRNA before cDNA conversion. Thus, transcript panels may contain nucleic acid molecules complementary to either the RNA strand or the cDNA strand so that they can anneal to and / or hybridize to RNA or cDNA. Some implementations utilize transcript panels to specifically target specific molecules of RNA and / or double-stranded cDNA. In some implementations, capture-based panels include a set of single-stranded nucleic acid probes for hybridization capture of specific molecules of RNA and / or specific molecules of double-stranded cDNA. In some implementations, amplification-based panels include a set of primers for specific reverse transcription of specific RNA and / or specific amplification of specific molecules of double-stranded cDNA.
[0098] The sequences of the capture probe and / or primer are complementary to the target transcript. While complete complementarity is not required, the capture probe and / or primer must be sufficiently complementary to the target to capture via hybridization or prime for amplification through annealing. Probe design can be based on any suitable target sequence. For example, a reference database, such as hg19 or hg38, can be used to design probes or primers for human transcript targets.
[0099] Specific targeting is based on sequence complementarity and selected genes within the panel. In some implementations, the transcript panel specifically targets a set of rare genes (RAGs). In some implementations, the transcript panel excludes non-RAG genes (defined by criteria or listed in Table 3). Excluding non-RAG genes can facilitate streamlined sequencing protocols and improved sequencing results (e.g., greater sequencing depth at the targeted sequence compared to the depth provided by the whole exome panel). Better sequencing results lead to greater sensitivity and better outcomes in various applications, such as cfRNA-based diagnostics.
[0100] In some implementations, the transcript panel excludes at least 50% of all exome genes that are not RAG. In some implementations, the transcript panel excludes at least 60% of all exome genes that are not RAG. In some implementations, the transcript panel excludes at least 70% of all exome genes that are not RAG. In some implementations, the transcript panel excludes at least 80% of all exome genes that are not RAG. In some implementations, the transcript panel excludes at least 90% of all exome genes that are not RAG. In some implementations, the transcript panel excludes at least 95% of all exome genes that are not RAG. In some implementations, the transcript panel excludes at least 99% of all exome genes that are not RAG.
[0101] Sequencing protocols and their applications can be enhanced by including a set of control genes, which provides a positive guarantee of sequencing results and facilitates normalization between samples. In some implementations, the transcript panel includes nucleic acid molecules for detecting a set of control transcripts to perform normalization between samples. Generally, the set of control transcripts can be any set of transcripts that are commonly detected and have relatively stable expression levels in control fluid biopsies. In some embodiments, the set of control transcripts includes one or more housekeeping transcripts (e.g., GAPDH, actin, ubiquitin).
[0102] In some implementations, the transcript panel consists of one control gene. In some implementations, the transcript panel consists of five or fewer control genes. In some implementations, the transcript panel consists of ten or fewer control genes. In some implementations, the transcript panel consists of twenty or fewer control genes. In some implementations, the transcript panel consists of fifty or fewer control genes. In some implementations, the transcript panel consists of 100 or fewer control genes. In some implementations, the transcript panel consists of 500 or fewer control genes. In some implementations, the transcript panel consists of 1000 or fewer control genes.
[0103] Sequencing protocols and their applications can be improved by evaluating genes that provide further insights. For example, in cancer diagnosis, certain transcripts, such as genes that are repeatedly mutagenic in cancer, can provide additional diagnostic insights. Therefore, transcript panels may include additional genes (defined by criteria or listed in Table 3) that are not RAGs.
[0104] In some execution models, the transcript panel consists of RAG plus 10 or fewer genes. In some execution models, the transcript panel consists of RAG plus 20 or fewer genes. In some execution models, the transcript panel consists of RAG plus 50 or fewer genes. In some execution models, the transcript panel consists of RAG plus 100 or fewer genes. In some execution models, the transcript panel consists of RAG plus 200 or fewer genes. In some execution models, the transcript panel consists of RAG plus 500 or fewer genes. In some execution models, the transcript panel consists of RAG plus 1000 or fewer genes. In some execution models, the transcript panel consists of RAG plus 2000 or fewer genes. In some execution models, the transcript panel consists of RAG plus 5000 or fewer genes. In some execution models, the transcript panel consists of RAG plus 10,000 or fewer genes.
[0105] In addition to RAGs, the transcript panel may include various types of genes, such as tissue-specific transcripts, cell-type-specific transcripts, clinically relevant transcripts, biomarker transcripts, generally mutagenic transcripts, and B-cell clone transcripts and T-cell clone transcripts. In some implementations, the transcript panel targets a set of tissue-specific transcripts. In some implementations, the transcript panel targets a set of cell-type-specific transcripts. In some implementations, the transcript panel targets a set of clinically relevant transcripts. In some implementations, the transcript panel targets a set of transcripts that are biomarker transcripts. In some implementations, the transcript panel targets a set of mutagenic events. In some implementations, the transcript panel targets a set of B-cell clones and T-cell clones.
[0106] In some implementations, the transcript panel includes nucleic acid molecules for detecting tissue-specific transcripts. Tissue-specific transcripts are transcripts that are uniquely highly expressed in a particular tissue compared to expression in all other tissues. In various implementations, tissue-specific transcripts are expressed at least twice as much in a particular tissue compared to all other tissues, at least three times as much in a particular tissue compared to all other tissues, at least four times as much in a particular tissue compared to all other tissues, at least five times as much in a particular tissue compared to all other tissues, at least six times as much in a particular tissue compared to all other tissues, at least seven times as much in a particular tissue compared to all other tissues, at least eight times as much in a particular tissue compared to all other tissues, at least nine times as much in a particular tissue compared to all other tissues, or at least ten times as much in a particular tissue compared to all other tissues. Unique expression can be determined empirically or obtained from databases, such as the Genotype-Tissue Expression (GTEx) project or the Human Protein Atlas (HPA).
[0107] In some implementations, the transcript panel includes nucleic acid molecules for detecting cell type-specific transcripts. Cell type-specific transcripts are transcripts that are uniquely and highly expressed within a particular cell type compared to their expression in other cell types. In various implementations, cell type-specific transcripts are expressed at least twice as much in a particular cell type compared to other cell types, at least three times as much in a particular cell type compared to other cell types, at least four times as much in a particular cell type compared to other cell types, at least five times as much in a particular cell type compared to other cell types, at least six times as much in a particular cell type compared to other cell types, at least seven times as much in a particular cell type compared to other cell types, at least eight times as much in a particular cell type compared to other cell types, at least nine times as much in a particular cell type compared to other cell types, or at least ten times as much in a particular cell type compared to other cell types. Unique expression can be determined empirically or may be obtained from database data, for example, from PanglaoDb.
[0108] In some implementations, the transcript panel includes nucleic acid molecules for detecting clinically relevant transcripts (e.g., in terms of diagnostic relevance). Clinically relevant transcripts may include transcripts for the diagnosis of medical disorders. For example, transcripts of EGFR, KRAS, MET, ALK, RET, and ROS1 are useful in the diagnosis of non-small cell lung cancer (NSCLC).
[0109] In some implementations, the transcript panel includes nucleic acid molecules for detecting transcripts associated with biological characteristics (e.g., biomarker transcripts). Biological characteristics include, but are not limited to, medical disorders, pregnancy, fetal complications, pregnancy complications, neoplasm growth, cancer, specific cancer types, pathogen infections, immune activation, organ transplant rejection, neurodegeneration, originating tissue, originating cell type, and activation of biochemical pathways.
[0110] In some implementations, the transcript panel includes nucleic acid molecules for detecting transcripts associated with mutagenic events, which may be useful for identifying de novo mutagenicity and / or oncogenes in cancer. Mutagenic events include, but are not limited to, gene fusions, insertions, deletions, transpositions, single nucleotide variants, and splice modifications.
[0111] In some implementations, transcript panels include nucleic acid molecules for detecting B cell clones and T cell clones and associated transcripts, which can be useful for tracking immunological activity against specific antigens. Because transcript panels target V(D)J recombination of B cell receptors and T cell receptors, the sequences of B cell clones and T cell clones can be identified. In some implementations, clones are detected at a single time point to identify the current activity of an immunogen (e.g., an immunogen for pathogenic infection, cancer, vaccine, or autoimmune disorder). In some implementations, clones are detected across multiple time points to detect changes in immunogen activity (e.g., minimal residual disease, success in cancer treatment, decreased immunogenicity from vaccination or pathogenic infection, detection of autoimmune flares).
[0112] Several components can be used, and sequencing of cfRNA molecules obtained from cell-free samples can be performed by carrying out several steps. Liquid biopsies can be derived from any suitable biological source, including (but not limited to) blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, and feces. Any suitable kit (e.g., the QIAamp Circulating Nucleic Acid kit) can be used to extract RNA. In some implementations, glycogen is added to the cell-free sample containing nucleic acids before adding it to a silica column. Typically, the sample also includes cell-free DNA, which can be used in other evaluations.
[0113] In some implementations, cell-free nucleic acids are quantified using quantitative real-time polymerase chain reaction (qRT-PCR). In some implementations, both RNA and DNA are quantified simultaneously in cell-free samples. qRT-PCR can be performed to simultaneously quantify RNA and DNA, allowing for primer-assisted RNA detection and quantification using a set of primers spanning introns, and primer-assisted DNA detection and quantification using a set of primers targeting non-transcribed regions of the genome. For RNA quantification, introns of commonly expressed genes (e.g., housekeeping genes) can be utilized. A standard curve for nucleic acid quantification can be generated using a control standard. After quantification, the cell-free sample can be aliquoted for evaluation of DNA and RNA. In some implementations, DNA is digested from the cell-free sample (e.g., DNase digestion) to evaluate RNA. In some implementations, DNA digestion is performed after elution from a silica column to evaluate RNA. In some implementations, one or more downstream steps in the sequencing preparation protocol use predetermined amounts of substances based on the quantification of cell-free RNA.
[0114] In some implementations, cell-free RNA is converted to cDNA. In some implementations, double-stranded cDNA is generated. In some implementations, the double-stranded cDNA is treated with a nuclease (e.g., S1 endonuclease) that removes single-stranded nucleic acid molecules.
[0115] After capturing and / or amplifying the target set, the library may be further processed (e.g., amplified) and sequenced using high-throughput sequencing. Some executions may involve sequencing to a depth of less than 10,000 reads, more than 10,000 reads, more than 100,000 reads, more than 1,000,000 reads, more than 10,000,000 reads, or more than 100,000,000 reads.
[0116] In some execution modes, sequencing is performed to a depth of less than 1x genome, to a depth greater than 1x genome, to a depth greater than 5x genome, to a depth greater than 10x genome, to a depth greater than 20x genome, to a depth greater than 30x genome, to a depth greater than 40x genome, to a depth greater than 50x genome, to a depth greater than 100x genome, to a depth greater than 150x genome, or to a depth greater than 200x genome.
[0117] After sequencing, various analyses of the sequencing results can be performed to improve the detection of targeted cell-free nucleic acid molecules. In some implementations, only transcripts contained within the transcript panel are included in the downstream analysis. In some implementations, transcript counts are converted to logarithmically transformed counts per million (CPM) to account for library size and transcriptome complexity. In some implementations, transcript counts are normalized to account for transcript size. In some implementations, counts are normalized to account for inter-sample variability. In some implementations, unwanted variability is removed, for example, variability due to the expression of platelet-related transcripts.
[0118] Platelet-derived RNA has been found to confound cfRNA analysis. Platelets are found in almost all types of fluid biopsies, particularly when cytotoxicity is present. Therefore, platelets can confound the evaluation of fluid biopsies (e.g.) derived from blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, and stool for assessment of several biological characteristics, such as medical impairment, pregnancy, fetal complications, pregnancy complications, neoplasm growth, cancer, specific cancer types, pathogen infection, immune activation, organ transplant rejection, neurodegeneration, originating tissue, originating cell type, and activation of biochemical pathways. In some implementations, cfRNA samples are centrifuged to remove platelets. In some implementations, platelet expression is removed in silico after sequencing. To remove platelet expression in silico, programs that remove unwanted variations (e.g., the RUVseq R package) are used. In some implementations, transcript panels for targeted sequencing specifically exclude the expression of platelet-associated genes. Even if a transcript panel for targeted sequencing specifically excludes platelet expression, the coefficients of unwanted variation can be estimated using control cfRNAs to establish a correction factor that correlates with platelet expression. For example, platelet correction can be performed by generalized least squares regression of log2NX on selected coefficients of unwanted platelet expression.
[0119] In some implementations, differential transcript expression analysis is performed using sequencing results. In some implementations, differential transcript expression analysis can be used for comparing samples (e.g., medical impairment versus control; e.g., comparing from one time point to another). Any suitable method for performing differential transcript expression analysis (e.g., the DESeq2 R package) can be used. For example, a generalized linear model can be constructed using sample type as a covariate. By using appropriate cutoffs and significance levels (e.g., greater than 1 / log ratio, corrected p-value < 0.05), significantly differentially expressed genes can be identified. Further downstream analyses can be performed on differentially expressed genes to gain insights into biological phenomena associated with the sample. For example, gene set enrichment analysis can be performed and used to identify molecular signatures associated with various phenomena.
[0120] In some implementations, sequencing results are used to detect enrichment of expression signatures associated with biological characteristics. Generally, many biological characteristics can be identified through expression signatures within the sequencing results. Examples of biological characteristics include (but are not limited to) medical disorders, pregnancy, fetal complications, pregnancy complications, neoplasm growth, cancer, specific cancer types, pathogen infections, immune activation, organ transplant rejection, neurodegeneration, originating tissue, originating cell type, and activation of biochemical pathways. In some implementations, an enrichment score of expression signatures is calculated to provide a scaled assessment of whether the biological characteristics are present in the sample. Based on the various biological characteristics present in the sequencing results, a diagnosis can be established. For example, the health of a pregnant woman or fetus can be assessed by using a liquid biopsy to evaluate cfRNA for expression signatures associated with various health conditions and / or complications.
[0121] In some implementations, sequencing results are used to detect mutagenicity, including somatic or de novo mutagenic events. Examples of detectable mutagenic events include, but are not limited to, gene fusions, insertions, deletions, transpositions, single nucleotide variants, and splice modifications. In some implementations, sequencing results are used to infer the copy number state of one or more genes. For example, expression signatures can be used to infer amplification of the MET gene associated with targeted therapy resistance in non-small cell lung cancer. As mentioned above, assessment of mutagenicity is diagnostic, and this can provide information on treatment options, particularly for various cancer types.
[0122] In some implementations, a computational model is trained to predict the categorical state or likelihood of biological characteristics present in a cfRNA sample based on sequencing results. In other implementations, sequencing results are directly used as input to train the computational model. In some implementations, the computational model utilizes acquired features, such as (for example) normalized expression of individual genes, enrichment of one or more gene signatures, enrichment of biochemical pathways, collection of sequence variants, and copy number state. The trained computational model can then be used to evaluate the patient's cfRNA sequencing results for use, for example, as a diagnostic method in a clinical setting.
[0123] To predict biological traits, any appropriate computational model can be used. Examples of computational models include (but are not limited to) logistic regression, elastic networks, LASSO, random forests, XGBoost, and neural networks. Training can be performed using expression data from samples with the biological traits and control samples. A cohort of individuals with known categorical states of biological phenomena has cfRNAs that have been sequenced and processed to train the computational model. Alternatively, since cfRNA samples are based on expression within a specific cell type, samples may be from solid tissues or any other source representative of cfRNAs. During training, a set of features, body weight, and hyperparameters can be selected to provide robust predictiveness.
[0124] Sequencing methods and diagnostic methods Various methods are being developed for performing cfRNA sequencing. Furthermore, cfRNA sequencing methods can be used in several diagnostic evaluations. Therefore, an individual may have a liquid biopsy extracted or collected for cfRNA sequencing, which can be prepared using a RAG transcript panel. Various biomedical characteristics can be screened for and / or diagnosed through cfRNA sequencing. Based on the diagnostic evaluation, the individual can be further evaluated through clinical evaluation. Also, since the diagnostic evaluation can provide information on treatment options, in some cases, treatment can be implemented by healthcare professionals, such as physicians, nurses, dietitians, or similar specialists. Additionally, cfRNA sequencing can be used to evaluate treatment response, treatment outcomes, and / or the presence of minimal residual disease.
[0125] Several sequencing methods can be used to evaluate cfRNAs. Generally, samples containing cfRNAs are collected using sequencing methods, and these are then prepared for targeted sequencing using a RAG panel.
[0126] Examples of methods for sequencing cfRNA may include: - Steps to extract or collect cfRNA samples - Step of enriching cfRNA samples using a targeted RAG panel - Steps to sequence the enriched cfRNA sample and obtain sequencing results.
[0127] In some implementations, the cfRNA sample is (or derived from) a liquid biopsy. In some implementations, the cfRNA sample is (or derived from) a fecal sample. In some implementations, the cfRNA sample is (or derived from) blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, or feces.
[0128] In some implementations, the RAG panel includes a set of genes expressed at less than a certain percentage of the control fluid biopsy population. In some implementations, the RAG panel includes a set of genes with subthreshold expression levels within the control fluid biopsy population. In some implementations, the RAG panel includes a set of genes expressed at less than a certain percentage of the control fluid biopsy population and with subthreshold expression levels within the control fluid biopsy population. In some implementations, the RAG panel includes the set of genes in Table 3. In some implementations, the RAG panel includes nucleic acid molecules for RAG amplification. In some implementations, the RAG panel includes nucleic acid molecules for RAG capture via hybridization.
[0129] In some implementations, the method for sequencing cfRNA further includes the following: - Step to collect a control cfRNA sample population. - Steps to perform sequencing on each cfRNA sample and obtain sequencing results. - Steps to identify RAGs from a set of sequencing results
[0130] In some implementations, control cfRNA samples are extracted or collected from healthy individuals. In some implementations, RAG is identified by being expressed in less than a certain percentage of the control fluid biopsy population. In some implementations, RAG is identified by having a subthreshold expression level within the control fluid biopsy population. In some implementations, RAG is identified by being expressed in less than a certain percentage of the control fluid biopsy population and having a subthreshold expression level within the control fluid biopsy population.
[0131] Sequencing of cfRNA samples can provide information about the biological characteristics of the individual from which the cfRNA sample was extracted or collected. Diagnostic methods can be used for a variety of purposes, such as screening individuals for biological characteristics or biomedical complications, diagnosing specific medical disorders, evaluating treatment responses, assessing treatment outcomes, and screening for minimal residual disease. Diagnostic methods can be used to evaluate several biomedical conditions, such as medical disorders, pregnancy, fetal complications, pregnancy complications, neoplasm growth, cancer, specific cancer types, pathogenic infections, immune activation, organ transplant rejection, neurodegeneration, originating tissue, originating cell type, and activation of biochemical pathways.
[0132] In some implementations, the screening or diagnostic method includes the following: - Steps to extract or collect cfRNA samples - Step of enriching cfRNA samples using a targeted RAG panel - Steps to sequence the enriched cfRNA sample and obtain sequencing results. - Steps to identify one or more biomedical conditions within the sequencing results.
[0133] In some implementations, cfRNA samples are (or derived from) blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, or feces. In some implementations, screening methods are generalized to assess multiple biomedical conditions. In some implementations, diagnostic methods are specific to particular biomedical conditions. In some implementations, biomedical conditions are identified by gene expression signatures. In some implementations, biomedical conditions are identified by sequencing results or computational models trained to predict biomedical conditions using features obtained from sequencing results.
[0134] In some practices, the biomedical condition is pregnancy. In some practices, the biomedical condition is fetal complications. In some practices, the biomedical condition is pregnancy complications. When evaluating fetal or pregnancy complications, in some practices, cfRNA samples are derived from amniotic fluid. When evaluating fetal or pregnancy complications, in some practices, screening is performed at various points throughout pregnancy. In some practices, when identifying fetal or pregnancy complications, further diagnostic procedures are performed, e.g., evaluation for gestational diabetes, fetal genetic evaluation, fetal ultrasound, and maternal blood tests. In some practices, when identifying fetal or pregnancy complications, treatment is performed, e.g., induction of labor, administration of uterine contraction inhibitors, and cesarean section.
[0135] In some implementations, the biomedical condition is a pathogenic infection. In some implementations, the biomedical condition is an immunological condition, e.g., vaccination status or previous pathogenic infection. In some implementations, when evaluating pathogenic infection, cfRNA samples are also enriched with pathogen sequences. In some implementations, when identifying a pathogenic infection, treatment, e.g., administration of antipathogenic medication (e.g., antibiotics, antivirals, antiparasitic agents, etc.), is performed. In some implementations, the screening method monitors pathogenic infection, antipathogenic treatment response, immunological response, health status, or any combination thereof.
[0136] In some practices, the biomedical condition is immune activation, e.g., activation in response to a pathogen, activation in response to immunization, or activation of an autoimmune disorder. In some practices, the biomedical condition is inflammation. In some practices, if the biomedical condition is an autoimmune disorder or inflammation, treatment is performed, e.g., administration of immunosuppressants and anti-inflammatory agents.
[0137] In some implementations, the biomedical condition is organ transplant rejection. In some implementations, organ transplant rejection is identified by gene expression signatures associated with cytotoxicity, gene expression signatures associated with originating tissue, gene expression signatures associated with originating cell type, gene sequences that differentiate the donor from the host, or a combination thereof. In some implementations, screening methods are performed periodically after the host receives the graft. In some implementations, if organ transplant rejection is identified, further diagnostic procedures are performed, such as (for example) organ tissue biopsy and medical imaging of the organ. In some implementations, if organ transplant rejection is identified, treatment is performed, such as (for example) administration of increased doses of immunosuppressants and administration of more potent immunosuppressants.
[0138] In some practices, the biomedical condition is neurodegeneration. When assessing neurodegeneration, in some practices, the cfRNA sample is (or derived from) cerebrospinal fluid. In some practices, neurodegeneration is identified by gene signatures associated with neurodegenerative disorders, gene signatures associated with originating nerve tissue, gene signatures associated with originating neuronal cell type, gene signatures associated with inflammation, or a combination thereof. In some practices, when neurodegeneration is identified, further diagnostic procedures are performed, e.g., medical screening, assessment of motor activity or speech, and assessment of cognition. In some practices, when neurodegeneration is identified, treatment is performed, e.g., drug therapy to reduce neurodegenerative symptoms.
[0139] In some implementations, the biomedical condition is cancer. In some implementations, screening methods are performed as part of a cancer surveillance effort (e.g., before cancer symptoms are present or recognized). In some implementations, screening methods are performed during treatment to assess the treatment response. In some implementations, screening methods are performed after treatment to assess whether residual cancer (e.g., MRD) is present after treatment, and this may be done periodically. In some implementations, the diagnostic method provides information on the cancer subtype, cancer stage, and / or treatment strategy.
[0140] Screening can be performed for several types of neoplasms, including acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), anal cancer, astrocytoma, basal cell carcinoma, cholangiocarcinoma, bladder cancer, breast cancer, Burkitt lymphoma, cervical cancer, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myeloproliferative neoplasms, colorectal cancer, diffuse large B-cell lymphoma, endometrial cancer, ependymoma, esophageal cancer, sensory neuroblastoma, Ewing's sarcoma, fallopian tube cancer, follicular lymphoma, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumors, hairy cell leukemia, hepatocellular carcinoma, and more. This includes, but is not limited to, digkin's lymphoma, hypopharyngeal cancer, Kaposi's sarcoma, renal cancer, Langerhans cell histiocytosis, laryngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, melanoma, Merkel cell carcinoma, mesothelioma, oral cancer, neuroblastoma, non-Hodgkin's lymphoma, non-small cell lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic neuroendocrine tumor, pharyngeal cancer, pituitary tumor, prostate cancer, rectal cancer, renal cell carcinoma, retinoblastoma, skin cancer, small cell lung cancer, small intestine cancer, cervical squamous cell carcinoma, T-cell lymphoma, testicular cancer, thymoma, thyroid cancer, uterine cancer, upper urothelial carcinoma, vaginal cancer, and hemangiomas. In some practices, if the cancer being evaluated is colorectal or gastric cancer, the cfRNA sample is (or derived from) a stool sample. In some implementations, if the cancer being evaluated is bladder, kidney, prostate, or upper urothelial carcinoma, the cfRNA sample is (or derived from) a urine sample.
[0141] In some forms of practice, when cancer is indicated, several follow-up clinical evaluations may be performed, including but not limited to physical examinations, medical imaging, mammography, endoscopy, stool tests, PAP tests, alpha-fetoprotein blood tests, CA-125 tests, prostate-specific antigen (PSA) tests, biopsy extractions, bone marrow aspirations, and tumor marker detection tests. Medical imaging includes but is not limited to X-ray, magnetic resonance imaging (MRI), computed tomography (CT), ultrasound, and positron emission tomography (PET). Endoscopy includes but is not limited to bronchoscopy, colonoscopy, vaginal and cervical endoscopy, cystoscopy, esophagoscopy, gastroscopy, laparoscopy, neuroendoscopy, proctoscopy, and sigmoidoscopy.
[0142] In some forms of practice, when cancer is indicated, several treatments may be administered, including (but not limited to) surgery, chemotherapy, radiotherapy, immunotherapy, targeted therapy, hormone therapy, stem cell transplantation, and blood transfusion. In some forms of practice, anticancer drugs and / or chemotherapeutic agents may be administered, including (but not limited to) alkylating agents, platinum agents, taxanes, vinca agents, anti-estrogens, aromatase inhibitors, ovarian suppressants, endocrine / hormone agents, bisphosphonate therapy agents, and targeted biological therapy agents. The drugs include cyclophosphamide, fluorouracil (or 5-fluorouracil or 5-FU), methotrexate, thiotepa, carboplatin, cisplatin, taxane, paclitaxel, protein-bound paclitaxel, docetaxel, vinorelbine, tamoxifen, raloxifene, toremifene, fulvestrant, gemcitabine, irinotecan, ixabépirone, temozolmide, topotecan, vincristine, vinblastine, eribulin, mutamycin, capecitabine, anastrozole, exemestane, letrozole, leuprolide, and abarelix. Buserelin, goserelin, megestrol acetate, risedronate, pamidronate, ibandronate, alendronate, zoledronate, tykerb, daunorubicin, doxorubicin, epirubicin, idarubicin, barurubicin, mitoxantrone, bevacizumab, cetuximab, ipilimumab, ad-trastuzumab emtansine, afatinib, aldesleukin, alectinib, alemtuzumab, atezolizumab, avelumab, axitinib, belimumab, belinostat, bevacizumab, blinatumomab, bortezomib, bosutinib, brentuximab vedotinVedoitn), briguchinib, cabozantinib, canakinumab, carfilzomib, ceritinib, cetuximab, cobimetinib, crizotinib, dabrafenib, daratumumab, dasatinib, denosumab, dinutuximab, durvalumab, elotuzumab, enasidenib, Erlotinib, everolimus, gefitinib, ibritumomab tiuxetan, ibrutnib, idelalisib, imatinib, ipilimumab, ixazomib, lapatinib, lenvatinib, midostaurin, nectiumumab, neratinib, nilotinib, niraparib, nivolumab, obinutuzumab, ofatum This includes, but is not limited to, mab, olaparib, loaratumab, osimertinib, palbocicilib, panitumumab, panobinostat, pembrolizumab, pertuzumab, ponatinib, ramucirumab, reorafenib, ribociclib, rituximab, romidepsin, rucaparib, ruxolitinib, siltuximab, ciproisel T, sonidebib, sorafenib, temsirolimus, tocilizumab, tofacitinib, tocitumomab, trametinib, trastuzumab, vandetanib, vemurafenib, venetoclax, bismodegib, vorinostat, and ziv-aflibercept. Individuals may be treated with a single drug or a combination of drugs described herein. A common combination of drugs is cyclophosphamide, methotrexate, and 5-fluorouracil (CMF). [Examples]
[0143] Examples and supporting data The systems and methods of this disclosure are better understood through several examples and supporting evidence provided. The validation results demonstrate that the use of a targeted sequencing panel based on rare genes significantly improves cancer detection using cfRNA samples. Using the RAG panel, the detection of expression signatures was 50-fold improved compared to whole transcriptome sequencing. Furthermore, this method provides better detection of resistance to targeted therapy compared to cfDNA-based methods. The method also allows for evaluation of origin tissue, which may be useful in identifying the primary pathological site. These results can be extrapolated to other cell-free diagnostic techniques, including the diagnosis of medical disorders, pregnancy, fetal complications, pregnancy complications, neoplasm growth, cancer, specific cancer types, pathogen infections, immune activation, organ transplant rejection, neurodegeneration, origin tissue, origin cell type, and activation of biochemical pathways.
[0144] Cell-free RNA analysis for non-invasive cancer detection and characterization RARE-Seq (Random priming and affinity capture of cell-free RNA fragments for enrichment analysis by sequencing) R andom priming & A ffinity capture of cell-free R NA fragments for E This document describes a system and method for cfRNA sequencing analysis, known as RARE-Seq (refined analysis by sequencing). This method improves the detection limit of circulating tumor RNA (ctRNA) compared to conventional methods while maintaining high specificity in the detection of tumor-naive cancers. RARE-Seq is highly useful in various clinical applications, including non-invasive ctRNA genotyping, monitoring of treatment resistance, histological analysis, and molecular characterization of non-malignant states.
[0145] result Pre-analysis factors affecting cfRNA recovery and analysis Cell-free RNA was highly fragmented, lacked detectable 18S and 28S rRNA peaks (Figure 1A), and showed high levels of degradation. In blood samples from healthy donors without a history of cancer, the median cfRNA concentration was 220 pg per mL of plasma, representing approximately 20 cell equivalents (n=117, Figure 1B). To improve recovery, factors affecting cfRNA recovery and downstream expression analysis were evaluated (Figures 1C-1F). In detail, analysis and RNA extraction protocols involving multiple blood collections, enabling optimal RNA recovery from plasma, led to the finding that factors such as hemolysis and freezer storage time had little to no correlation with cfRNA yield (Figures 1C and 1D). Multiple steps in the assay workflow, including contaminating DNA removal, complementary DNA (cDNA) synthesis, and end repair, were optimized. This optimization improved the efficiency of library preparation from low cfRNA input and minimized the impact of contaminating cfDNA (Figures 2A-2K). Further confirmation was that the optimization resulted in reproducible cfRNA expression profiles in healthy controls. The cfRNA-based method optimization can be utilized in the RARE-Seq (Random Priming and Affinity Capture of Cell-Free RNA Fragments for Sequencing-Based Enrichment Analysis) protocol (Figure 3A).
[0146] Platelet and non-hematopoietic cell types are enriched in cfRNA. Expression differences between cellular RNA and cell-free RNA were characterized by applying RARE-Seq to cfRNA and matched leukocyte RNA from 10 healthy adults. Differential expression analysis identified 8,336 significant genes, including 5,271 genes overexpressed in cfRNA and 3,065 genes overexpressed in leukocyte RNA (Figure 3B). Using gene set enrichment analysis (GSEA) and cell type-specific gene sets from single-cell RNA sequencing atlases, cfRNA was found to be significantly enriched for non-hematopoietic cell types, including endothelial cells, hepatocytes, fibroblasts, and several neuronal cell types (n=71 total) (Figure 3C). In contrast, leukocyte RNA was significantly enriched for 12 major hematopoietic cell types from myeloid and lymphoid lineages, including T cells, NK cells, neutrophils, and monocytes. Only two hematopoietic cell types, erythrocyte precursors and platelets, were enriched in cfRNA, suggesting that these significantly contribute to the cfRNA pool in erythrocyte formation and megakaryocyte formation.
[0147] Given the significant enrichment of platelet transcripts in cfRNA (Figure 3C), it was suspected that residual platelets not removed during plasma isolation could contribute RNA to the cfRNA sample. To test this idea, cfRNA concentrations were compared between protocols by separating cell-free plasma components from cellular plasma components using three different centrifugation rates (n=32). The median concentration was approximately 7.5 times higher at 1200×G than at 2500×G (1.72 ng / mL vs. 0.23 ng / mL; P=0.000052; Figure 3D). Furthermore, mean expression levels of platelet-specific genes were also compared (n=130). 11The expression was significantly associated with the centrifugation rate (Figure 3E). Principal component analysis (PCA) revealed that the first principal component (PC1) captured 70% of the expression changes and was strongly associated with platelet expression (R=-0.99, P=<2.2e-16) (Figures 3F and 3G). Similar results were found between samples processed at fixed centrifugation rates, suggesting that differences in platelet counts may also contribute to the variation in platelet transcripts in cfRNA samples (R=0.98, P=<1.3e-11) (Figure 3H). In summary, these results suggest that platelets are the largest source of cfRNA expression changes, and that platelet contamination is an important pre-analysis variable for regulating cfRNA analysis.
[0148] Individual platelets have a lower RNA content than white blood cells, but the number of platelets in the blood is approximately two orders of magnitude higher than the number of white blood cells. In addition, physiological variability in platelet counts ranges from approximately threefold in healthy adults, and further variability is observed in cancer patients. Therefore, even with efficient methods to minimize plasma contamination by platelets, it may be difficult to effectively reduce the level of platelet-related RNA. Separately, from a practical standpoint, standardizing blood storage protocols is often impractical, especially when profiling previously collected samples. Therefore, we investigated whether the contribution of expression from platelet contamination can be algorithmically removed. To this end, we developed an in silico method to remove unwanted gene expression driven by altering the platelet contamination level (see also Figures 3I and 5G). As shown by the lack of correlation between corrected platelet expression and PC1 (R=0.07, P=0.69) (see also Figures 3J and 5H), this method successfully removed unwanted variability contributed by platelets. Hierarchical clustering further confirmed this effect in samples clustered according to platelet gene expression before correction (Figure 3K), but it was not observed in samples clustered according to inter-individual expression differences after correction (Figure 3L). Therefore, addressing platelet contamination in cfRNA analysis is important to reliably clarify the contribution of gene expression from non-platelet sources.
[0149] Maximizing sensitivity for ctRNA detection cfRNAs are highly enriched for gene expression from non-hematopoietic sources, but the absolute expression levels of such genes are relatively low compared to hematopoietic transcripts (e.g., globin) and highly expressed housekeeping genes (e.g., beta-actin). Therefore, we hypothesized that selective capture of genes absent or low-expressed in healthy control cfRNAs could improve the detection of non-hematopoietic tissues or disease signatures, extending the usefulness of RARE-Seq to a wide range of clinical applications. To validate this, we identified “rare genes” (RAGs) using whole-coding transcriptome RARE-Seq data from healthy cfRNA samples (n=50) and RNA-seq data from healthy whole blood samples (n=307) derived from the GTEx project (Figure 4A). RAGs were defined based on consistently low or absent expression in both cfRNA and whole blood (Figure 5A; Table 3). Notably, RAGs substantially enriched both non-hematopoietic tissue-specific genes and tumor-specific genes from various cancers compared to all protein-coding transcripts (Figures 4B and 4C).
[0150] Considering the enrichment of genes within RAGs that may reflect the presence of tissue damage or cancer, we investigated whether targeted sequencing of RAGs could improve the sensitivity of detecting such transcripts. Therefore, we designed a capture panel targeting 4,323 RAGs. This panel was supplemented with 50 housekeeping genes to prioritize genes with relatively uniform expression across healthy cfRNA and whole blood samples, while covering a broad dynamic range of expression (Figure 5B). We also added genes known to repeatedly mutate in lung cancer (n=123 genes), including genes such as TP53, EGFR, KRAS, ALK, RET, and ROS1 (Figure 5C). The complete collection of genes targeted by our RAG capture panel is summarized in Figure 5D.
[0151] Using both the full coding transcriptome and the RAG capture panel, five healthy cfRNA samples were profiled using RARE-Seq. Gene expression correlated highly across the fitted libraries (R=0.98, P=<2.2e-16) (Figure 5E), and also correlated for the housekeeping gene set (R=0.97, P=<2.2e-16) (Figure 5F). Within the subset of genes shared across both capture panels, RAG capture significantly increased the unique sequencing depth (Figure 4D), detected more genes (Figure 4E), and increased the unique sequencing depth per detected gene (Figure 4F). Therefore, targeted RAG capture effectively improves the sensitivity for detecting genes that are low in expression in healthy cfRNA.
[0152] The 95% limit of detection (LOD95) was determined for the target transcriptional signature using either the whole coding transcriptome or a RAG-targeted panel. An analytical framework was developed to score the presence of the target gene signature in cfRNA samples. A gene signature enrichment score (ES) was calculated by comparing the expression of a given signature gene in the target sample to a "meta-reference" consisting of cfRNA samples from healthy controls, and estimating significance using bootstrapping (Figure 4G). A key advantage of this method is that, unlike machine learning-based methods, it does not require a large training cohort of patient cfRNA samples. This is because the ES strategy relies on existing gene expression datasets for defining gene signatures and weights for relevant genes based on expression levels in the target tissue or condition.
[0153] To establish LOD95 in RARE-Seq, serial dilutions of RNA were performed in silico from the NCI-H1975 NSCLC cell line to generate healthy control cfRNAs from various cancer fractions. To enable the use of the ES framework, an NCI-H1975-specific gene signature was generated. This included top differentially expressed genes that identified NCI-H1975 from a meta-reference healthy cfRNA profile (H1975 Sig; n=122 genes). Using this signature, LOD95 in RAG RARE-Seq was 0.05% (Figure 4H), demonstrating over 50-fold sensitivity compared to whole-coding transcriptome RARE-Seq (2.8%; Figure 4I), thus proving its value in targeting RAG. To experimentally confirm these results, in vitro dilution experiments were performed by spiking NCI-H1975 RNA into healthy donor cfRNA at specified concentrations. These in vitro spikes demonstrated a strong agreement with the in silico results (Figures 4H and 4I). The effect of sequencing coverage depth (ranging from 10 million to 50 million read pairs) on LOD95 of RAG-targeted RARE-Seq was evaluated, and it was found that LOD95 increased with sequencing depth (Figure 5I). Therefore, 50 million read pairs were targeted for subsequent experiments using RARE-Seq with the RAG panel.
[0154] Detection of non-small cell lung cancer With the development of RARE-Seq, the remainder of the study focused on generating proof-of-concept data for the potential utility of cfRNA analysis across several clinical applications. Given the observed high analytical sensitivity by targeting RAG, a RAG capture panel was used for such experiments. Initially, detection performance was evaluated for non-invasive and tumor-naive lung adenocarcinoma (LUAD), the most common type of NSCLC. cfRNA samples were analyzed from 50 LUAD patients (stages I-IV) and 50 non-cancer controls, including 26 risk-matched controls with a history of substantial tobacco exposure, collected at the time of low-dose computed tomography (LDCT) screening for lung cancer. cfRNA concentrations, inputs to library preparations, and sequencing depths were kept similar across groups (Figures 6A-6C). Platelet-regulated differential expression analysis between LUAD and control cells identified 94 genes overexpressed in LUAD, including key genes known to be highly expressed in LUAD, such as SFTA2, SLC34A2, and NKX2-1 (also known as TTF1) (Figure 7A). LUAD DEG was found to be highly expressed in LUAD tumors mediated by TCGA but depleted in meta-reference cfRNA controls (Figure 7B). Furthermore, using cell type-specific gene signatures, GSEA revealed that LUAD cfRNA was most enriched in epithelial cells, including type I and type II alveolar cells, and most depleted in naive B lymphocytes and memory B lymphocytes (Figure 7C). Therefore, cfRNA derived from LUAD patients is enriched with transcripts highly expressed in LUAD tumors.
[0155] To further quantify the enrichment of lung adenocarcinoma-derived transcripts, the ES framework developed above was applied using LUAD-specific gene signatures (LUAD Sig; n=72 genes) identified using RNA-seq data from LUAD tumors within TCGA, LUAD cell lines derived from CCLE, whole blood (GTEx), and meta-reference cfRNA controls (Methods). LUAD Sig was detected in 39 out of 50 LUAD targets (78% sensitivity with 95% specificity in non-cancer controls; Figure 7D) and was significantly associated with stage (sensitivity by stage: I=40%, II=60%, III=80%, IV=87%; P=4.7e-14; Figure 7E). In addition, in a subset of plasma samples where circulating tumor DNA was also examined, LUAD ES was significantly correlated with the ctDNA variant allele frequency (VAF) (R=0.75, P=0.03) (Figure 7F), suggesting that RARE-Seq can quantify cancer burden in cfRNA. However, 40% of samples with both ctRNA and ctDNA data were detected by RARE-Seq alone, suggesting that ctRNA and ctDNA analysis may be complementary for cancer detection. Group-to-group differences in unique sequencing depth were not associated with the likelihood of detection (Figures 6D and 6E), indicating that this variable does not facilitate detection. Furthermore, detection success did not appear to vary significantly depending on specific oncogenic driver mutations in LUAD (Figure 6F).
[0156] Genotyping of somatic mutations in ctRNA We evaluated the ability of RARE-Seq to genotype cancer-derived somatic mutations. While non-invasive tumor genotyping is already routinely performed clinically using ctDNA, ctRNA-based variant detection may be useful in addition to cfRNA analysis. Therefore, we developed a custom ctRNA genotyping method to detect repetitive, clinically active single nucleotide variants (SNVs), insertions / deletions (indels), splice variants, and gene fusions in NSCLC (Figure 8A). To ensure specificity, putative variants were censored if the VAF did not significantly exceed the background error rate observed in healthy cfRNA controls. The performance of this method was evaluated in 55 stage IV LUAD patients and 36 LDCT controls. One or more somatic driver mutations were detected in 44% of LUAD patients and 5.6% of LDCT controls, including key NSCLC oncogenic driver mutations such as EGFR L858R and exon 19 deletion, KRAS G12C, ROS1 and RET fusions, and MET exon skipping events (Figures 8B and 8C; Figures 6G and 6I). Of the 74 variants identified by tumor tissue genomic DNA or ctDNA genotyping in the same patient, 22 (30%) were also detectable in ctRNA, with similar detection rates across variant types (Figure 8D). A significant association was observed between variant detection rates and corresponding expression levels. For example, EGFR variants were detected significantly more frequently in cfRNA samples and were associated with high EGFR expression (P=0.007; Figure 8E). Therefore, RARE-Seq analysis enables the simultaneous identification of potentially active somatic mutations in cancer patients.
[0157] Next, considering that ctRNA expression signatures and ctRNA somatic mutations can be used individually to distinguish cancer patients from non-cancer controls, we investigated whether machine learning techniques could combine these features into a single detection algorithm. A weighted elastic network (EN) classifier was trained and LUAD cfRNA and control cfRNA were identified using 10-fold nested cross-validation (CV) in a training cohort of 50 patients and 50 controls. To reduce the risk of model overfitting, gene weights based on LUAD tumor expression data were utilized during model training. The model demonstrated high classification performance (AUROC = 0.9; Figures 8F and 8G), with 279 features selected, including gene expression features (e.g., key LUAD markers, e.g., SLC34A2, SFTPB, ROS1) and mutation-based features (e.g., detection of at least one ctRNA SNV; Figures 6J and 6K). To verify whether the training cohort was large enough to estimate classification accuracy, we iterated classifier training and cross-validation using subsampling and found stable performance when using more than 70% of the training cohort (Figure 6L). Importantly, the EN model performed similarly well in a pending validation cohort collected at an independent institution (AUROC=0.92) (Figures 8F and 8G). The performance of the EN classifier was statistically similar to the ES framework in the training cohort (P=0.60), but improved in the validation cohort (P=0.04). Therefore, if sufficient samples are available for training, a robust classifier can be trained using RARE-Seq data as input, even using machine learning-based methods.
[0158] Identification of resistance mechanisms to EGFR-targeted therapy We hypothesized that ctRNA analysis could enable the detection of non-gene therapy resistance mechanisms, such as histological transformation. To investigate this question, our experiment focused on EGFR-mutant LUAD patients treated with EGFR tyrosine kinase inhibitors (TKIs). Acquired resistance to EGFR TKIs arises from heterologous mechanisms, including both EGFR-dependent mechanisms (e.g., secondary point mutations in the EGFR gene) and EGFR-independent mechanisms (e.g., bypass pathway activation or histological transformation). We profiled plasma samples from 10 patients with stage IV EGFR-mutant LUAD who received osimertinib after an exacerbation during at least one prior EGFR TKI administration and subsequently developed resistance to osimertinib. The resistance mechanisms present in each patient were determined by tissue biopsy at the time of exacerbation, revealing small cell histological transformation in four patients, MET gene amplification in three patients, and emergent EGFR C797S point mutations in three other patients (Figure 9A). Each patient received blood samples, initiated osimertinib, and had one or more blood samples taken after disease progression (n=24 total time points).
[0159] Histological transformation in small cell lung cancer (SCLC) was evaluated using the SCLC signature (SCLC Sig), defined using publicly available tumor gene expression data. The signature included 73 genes, including canonical SCLC markers such as ASCL1, NEUROD1, and INSM1. At the time of radiological progression, SCLC Sig was detected in 3 out of 4 patients (75%) with histological transformation as diagnosed by biopsy (Figure 9B). In one of the three patients, a decrease in SCLC Sig ES was observed in a plasma sample collected 274 days after the initiation of SCLC-targeted chemotherapy (e.g., carboplatin / etoposide), suggesting a therapeutic response to the small cell component (Figure 9C).
[0160] To identify the genetic mechanisms of resistance, RARE-Seq was evaluated. To detect MET pathway activation caused by MET amplification using ctRNA, a MET amplification signature (METamp Sig) was developed, containing nine genes differentially expressed in three EGFR mutant NSCLC cell lines with MET amplification-acquired resistance. METamp Sig was detected in 2 out of 3 patients (67%) with MET amplification diagnosed by biopsy (Figure 9D). Notably, METamp Sig ES decreased in one patient after adding the MET inhibitor savolitinib to osimertinib, and a partial response was observed on imaging (Figure 9E). EGFR C797S mutant reads were identified in 2 out of 3 patients (67%) who were known to have acquired this mutation based on tumor DNA or cfDNA analysis (Figures 9F and 9G). Importantly, these resistance mechanisms were not detected at any pre-resistance time point, confirming the specificity of this method. In summary, these results demonstrate a proof of concept that RARE-Seq can identify both non-genetic and genetic mechanisms underlying EGFR TKI resistance, including histological transformations that are undetectable by mutation-based ctDNA methods.
[0161] Use of ctRNA analysis for origin tissue determination Because the transcriptional profile of cancer can reflect the tissue of origin (TOO), ctRNA analysis can be used to identify cancer type. The ability to non-invasively identify the tissue of origin of malignant disease may be useful in several clinical settings, including patients with metastatic cancer of unknown primary origin. To investigate this potential application, plasma samples from patients with four advanced-stage carcinoma subtypes, including lung adenocarcinoma (LUAD, n=50), pancreatic adenocarcinoma (PAAD, n=10), prostate adenocarcinoma (PRAD, n=9), and hepatocellular carcinoma (LIHC, n=10), were analyzed. Corresponding tumor-specific gene signatures were generated for each cancer type using TCGA data and evaluated using the ES framework. The sensitivity of cancer-specific ES detection was 78% for LUAD, 80% for LIHC, 90% for PAAD, and 100% for PRAD (Figure 10A). To identify cancer types, cancer signatures were generated, including genes uniquely expressed in each cancer, and compared to cfRNAs from all other cancers, as well as from patients without cancer. Using these signatures, the top predicted TOOs were accurate with 84% probability, and the top two predicted TOOs identified the tumor type accurately with approximately 89% probability (Figures 10B and 10C). Therefore, RARE-Seq may potentially enable non-invasive identification of TOOs.
[0162] Application of RARE-Seq to non-malignant conditions RARE-Seq has several applications that surpass oncological evaluation. We investigated whether cfRNA expression could reveal lung injury resulting from benign lung conditions. We profiled cfRNAs from 6 patients with active COVID-19 infection (10 time points), 19 patients with acute respiratory distress syndrome (ARDS), and 3 patients with chronic obstructive pulmonary disease (COPD). Using gene expression signatures from normal lungs derived from GTEx, lung-derived cfRNAs were detectable in 47% of patients across this diverse lung condition (Figure 10D). In contrast, in adult subjects without known lung conditions (n=42), lung cfRNAs were detected in 14%, and detection was significantly more frequent in current smokers than in former smokers or non-smokers (45.5% vs. 3.6%, P=0.004; Figure 10E). This result suggests that lung cfRNAs are induced by ongoing injury to lung epithelium resulting from recent exposure to tobacco smoke. Conversely, in patients with a known lung condition, detection was significantly higher in patients using mechanical ventilation at the time of blood collection (Figure 10F), potentially reflecting the severity of the lung condition and damage to the lung epithelium due to mechanical ventilation. Therefore, cfRNA analysis may be useful for evaluating benign conditions, including acute and chronic patterns of tissue damage.
[0163] Separately, given the rapidly growing interest in RNA-based vaccines and therapeutics, we investigated whether RARE-Seq has utility in simultaneously tracking RNA-based therapies and their effects on the host. To this end, we profiled plasma from individuals that had received the first two doses of mRNA-based COVID-19 vaccination over the long term. RARE-Seq was performed on these samples using a RAG capture panel supplemented with bait targeting mRNA vaccine sequences. RNA aligned by the vaccine was detectable after both vaccinations and persisted at high levels for at least 10 days post-injection (Figure 10G). Interestingly, genes differentially expressed after vaccination compared to pre-vaccination were enriched for interferon-gamma response, antiviral response, chemokine activity, and leukocyte migration pathways, suggesting the detection of a host response to vaccination (Figure 10H). These results highlight the potential of RARE-Seq for the simultaneous measurement of the pharmacokinetic and pharmacodynamic aspects of RNA-based therapies.
[0164] Exam summary and interpretation A novel framework for cfRNA profiling offers diverse potential clinical applications. RARE-Seq was found to achieve a level of analytical sensitivity comparable to tumor-naive and tumor-information-based ctDNA-based SNV tracking methods, demonstrating its potential usefulness for non-invasive profiling of tumor gene expression in patients with low disease burden. A key innovation of this method is the specific capture of transcripts that are absent or expressed at very low levels in the plasma of healthy controls. While analysis of such rare transcripts is possible after sequencing the entire transcriptome, this method is significantly less efficient because it primarily measures the expression of leukocyte-derived genes, thus missing these rare transcripts. These latter transcripts, rare in healthy plasma cfRNA, are highly enriched for cancer and solid organ-specific genes, making them crucial for detecting pathophysiology occurring in extra-blood tissues.
[0165] One important finding was the high proportion of platelet-derived transcripts present in plasma cfRNA. In contrast to cfDNA evicted by megakaryocytes, such transcripts are at least partially derived from intact platelets that separate from plasma during blood sample processing and appear to be potential confounding factors for cfRNA analysis. Indeed, expression differences in cfRNA between cancer patients and controls, reported in several conventional studies, appear to be primarily due to differences in platelet-derived transcripts. While pre-analysis techniques, such as increasing the centrifugation rate, can reduce platelet contamination, platelet-poor plasma preparations are rarely completely platelet-free. Furthermore, such procedures are often impractical when analyzing historical samples, including those from completed clinical trials. Therefore, the techniques described herein for removing the contribution of platelets to cfRNA gene expression profiles are broadly useful in future cfRNA liquid biopsy studies.
[0166] Investigations into potential applications suggest that ultra-sensitive cfRNA analysis may be useful in a variety of clinical settings. For example, because this method enables tumor-naive detection of ctRNA, it could potentially be used for cancer screening, either alone or in combination with other diagnostic modalities. The fact that RARE-Seq detected lung cancer RNA in some samples that were negative with ctDNA analysis suggests that the two methods may be complementary. A second promising application is non-invasive identification of cancer types. Such methods could be useful for non-invasive identification of TOO in patients with cancer of unknown primary origin, or for distinguishing between cancer subtypes.
[0167] When considering EGFR-mutant NSCLC patients treated with EGFR TKIs, these results suggest that RARE-Seq can provide a more comprehensive non-invasive analysis of treatment resistance mechanisms, including mechanisms such as histological transformation not facilitated by recurrent emergent somatic alterations. Histological transformation is a common mechanism of resistance occurring in various cancer types, and its diagnosis currently requires invasive tissue biopsy; therefore, non-invasive detection using cfRNA can significantly improve diagnostic methods. The ability of RARE-Seq to simultaneously examine the presence of somatic mutations (e.g., EGFR C797S) and the transcriptional activity of pathway activation mediated by somatic alterations (e.g., MET amplification) or non-genetic mechanisms (e.g., histological transformation) enables profiling of a broader and more diverse range of resistance mechanisms than currently possible.
[0168] Ultimately, RARE-Seq is useful in detecting non-cancerous cfRNA signatures in patients with both malignant and benign conditions. For example, data demonstrating the presence of normal lung RNA in the plasma of patients with acute lung injury caused by conditions such as COVID-19 infection or ARDS suggest that cfRNA analysis can also enable blood-based monitoring of tissue injury. Separately, analysis of cfRNA in plasma samples after COVID-19 mRNA vaccination demonstrates that RARE-Seq can enable simultaneous tracking of mRNA vaccines or therapeutics and measurement of host responses.
[0169] In summary, we developed a versatile method for cfRNA analysis. This method is useful in a variety of clinical settings. Furthermore, the results include several important biological and technical insights that will enable more sensitive cfRNA analysis and be broadly useful for the application of cfRNA-based liquid biopsies. The method will enable novel non-invasive diagnostic applications and is expected to advance precision medicine and improve patient care.
[0170] method Human participants and cohorts All samples analyzed in this study were collected with informed consent and using protocols approved by the respective institutional review boards. Collection sites included Stanford University, Memorial Sloan Kettering Cancer Center (MSKCC), Massachusetts General Hospital (MGH), and University Hospital Zurich, as detailed below. A total of 269 blood samples were collected from 201 individuals (Figure 11). Clinical and demographic characteristics are presented in Table 1 for the non-cancer cohort and in Table 2 for the cancer cohort.
[0171] Cancer cohort. Blood samples were collected from non-small cell lung cancer (NSCLC) patients enrolled at Stanford University (n=33), MSKCC (n=17), and MGH (n=14). Of these, 64 (98%) were diagnosed with lung adenocarcinoma (LUAD), and one patient (2%) had a mixed adenosquamous histological diagnosis. Stage information was determined using the American Joint Committee on Cancer (AJCC) 8th edition. In 10 EGFR-mutant NSCLC patients enrolled at MGH, blood samples were collected at multiple time points before and after progression during treatment with the EGFR tyrosine kinase inhibitor (TKI) osimertinib, where available (n=24 plasma samples). Tissue biopsies were collected at the time of progression to identify resistance mechanisms. Further exploratory analyses involved collecting blood samples from patients diagnosed with advanced pancreatic adenocarcinoma (PAAD, n=10) and advanced prostate adenocarcinoma (PRAD, n=9) at Stanford University, and from patients with hepatocellular carcinoma (LIHC, n=10) at Zurich University Hospital.
[0172] Non-cancer cohort. Blood samples from individuals without known cancer or with benign lung conditions were collected at Stanford University for technical experimentation and methods development control (n=45). A subset of these samples included a set of "meta-references" used to define expected expression in healthy cfRNAs for various downstream analyses (n=15). Blood samples were collected at Stanford University (n=27) and MGH (n=10) from individuals eligible for low-dose computed tomography (LDCT) screening based on smoking history (≥30 pack years) and age (55–80 years). Three individuals were found to have chronic obstructive pulmonary disease (COPD) at the time of LDCT screening. In addition, blood samples were collected from patients treated for acute respiratory distress syndrome (ARDS; n=20) and coronavirus disease 2019 (COVID-19; n=6) at Stanford Hospital. Additionally, plasma was collected from one individual at Stanford University after they received two doses of Pfizer-BioNTech COVID-19 mRNA vaccine (n=8 time points).
[0173] Blood collection and plasma processing Peripheral blood samples were collected and processed according to the protocols of each facility. In the technical experiment conducted at Stanford, whole blood was collected in EDTA tubes, and plasma was isolated using centrifugation at 2,500 G for 10 minutes at 4°C. The optical density at 414 nm was measured using a Nanodrop spectrophotometer to quantify the level of hemolysis in the plasma. After centrifugation, all plasma was stored at -80°C until cell-free nucleic acid isolation.
[0174] cfRNA extraction Cell-free nucleic acids were extracted from plasma using the miRNA protocol of the QIAamp Circulating Nucleic Acid kit (Qiagen, range 0.2–8.0 mL). Extraction was performed according to the manufacturer's instructions, with some modifications. The resulting eluate was incubated with 14U DNase I (RNase-Free DNase Set, Qiagen) at room temperature for 30 minutes to digest the DNA. The digested eluate was purified using the Zymo RNA Clean & Concentrator kit and stored at -80°C.
[0175] Plasma processed before February 2021 was extracted using phenol / chloroform phase separation and purified using the QIAamp Viral RNA Kit (Qiagen). Briefly, the plasma was first incubated with 3 volumes of TRI Reagent LS (Molecular Research Center), and then incubated with 0.4 volumes of chloroform (relative to the plasma input). Phase separation was performed using a Maxtract 50 mL conical tube (Qiagen) rotated at 1500 G for 5 minutes at room temperature. The RNA-containing aqueous phase was carefully collected, and RNA was extracted according to the QIAamp Viral RNA Kit protocol. DNA digestion and cleanup were performed as described above. For a subset of samples digested using the "on-column" protocol, 28U DNase I was added directly to the sample conjugated to a silica column and incubated for 15 minutes as recommended by the manufacturer. Careful analysis was performed to compare cfRNA yield and gene expression distribution between the various extraction methods (Figures 1E and 2B). Since no significant difference was found, the cfRNA samples extracted using both methods were mixed and the analysis presented in this manuscript was performed.
[0176] cfRNA quantification cfRNA was quantified using quantitative real-time polymerase chain reaction (qRT-PCR). RNA-specific primers were designed to encompass a 97 bp amplicon spanning two exon boundaries within the housekeeping gene GAPDH (forward 5'-GATCATCAGCAATGCCTCCT-3' (SEQ ID NO: 1), reverse 5'-TGTGGTCATGAGTCCTTCCA-3' (SEQ ID NO: 2)). DNA-specific primers were designed to cover a 78 bp transcriptionally silent region on chromosome 12 (forward 5'-TACGGTTGGTCCTTTCTTCG-3' (SEQ ID NO: 3), reverse 5'-TTTCCTTTGGGTCTGAATGC-3' (SEQ ID NO: 4)). First, reverse transcription was performed using a High-Capacity cDNA Reverse Transcription Kit (Applied Biosystems). Next, quantitative PCR (qPCR) was performed using 2× Power SYBR Green PCR Master Mix (Thermo Fisher Scientific) on an Applied Biosystems 7500 Fast Real-Time PCR or QuantStudio 7 Pro instrument. In parallel, Universal Human Reference RNA (Thermo Fisher Scientific) was run to generate a standard curve, and cfRNA concentration was calculated by comparing the RNA-specific Ct values of the samples to the standard curve. DNA was quantified using Human Genomic DNA (Promega) using a similar method. When detecting DNA using DNA-specific primers, DNA digestion, cleanup, and quantification were repeated. The cfRNA size distribution was evaluated using an Agilent Bioanalyzer RNA 6000 Pico tip.
[0177] cfRNA library preparation and sequencing RARE-Seq. The input mass was 424 pg of cfRNA, which represents the 25th percentile of cfRNA yield in 4 mL of healthy control plasma (Figure 1B). For samples with less than 424 pg, all extracted cfRNA was used for library preparation (range 8–424 pg). Double-stranded complementary DNA (cDNA) was synthesized from cfRNA using the NEBNext Ultra® II RNA First-Strand Synthesis Module and Non-Directional Second Strand Synthesis Module (New England Biolabs). Double-stranded cDNA was treated with 100 U S1 nuclease (Thermo Fisher) at room temperature for 30 minutes to hydrolyze incomplete (i.e., single-stranded) regions. Libraries for sequencing were prepared using the KAPA Hyper Prep kit (Kapa Biosystems), mainly following the manufacturer's instructions. Full coding transcriptome capture was performed using the Roche Nimblegen SeqCap EZ MedExome Target Enrichment Kit and / or the Twist Biosciences Comprehensive Exome Hybridization Kit, according to the respective manufacturers' instructions. Singleplex capture using a custom gene panel targeting rare genes (see "RAG Capture Panel Design" below) was performed using the Twist Biosciences Hybridization Kit. Captured libraries were sequenced using 2 × 150 bp paired-end reads on an Illumina HiSeq4000 or NovaSeq6000 instrument.
[0178] SMART-Seq. cfRNAs from three control subjects were input into the library preparation using the SMART-Seq Stranded kit (TaKaRa) according to the manufacturer's instructions, without further fragmentation. The library was amplified and sequenced on an Illumina NovaSeq6000 using 150 bp paired-end runs.
[0179] Processing of leukocyte RNA After plasma separation, leukocytes were isolated from the plasma-depleted whole blood sample using SepMate PBMC isolation tubes (Stem Cell Technologies) according to the manufacturer's instructions. The leukocyte cell pellet was stored at -80°C. RNA was extracted from the leukocyte cell pellet by mixing with 800 ul of TRIzol (Invitrogen) and 200 ul of chloroform, and then centrifuged at 13,000 G for 15 minutes at 4°C. RNA was then purified from the aqueous supernatant using the Qiagen RNeasy kit according to the manufacturer's instructions. Total RNA was quantified using a NanoDrop instrument and an Agilent Bioanalyzer RNA 6000 Pico tip. RARE-Seq library preparation and whole-coding transcriptome capture were performed using 5-10 ng of RNA as previously described. After fragmentation, first-strand synthesis was performed according to the manufacturer's recommendations.
[0180] Treatment of NCI-H1975 cell line RNA NCI-H1975 lung adenocarcinoma cells were obtained from the American Cell Culture and Cell Lineage Preservation Center (ATCC). Cellular RNA was extracted from the cell pellet, and leukocyte RNA was quantified as described. To generate an in vitro NCI-H1975 cell line spike, NCI-H1975 RNA was serially diluted in cfRNA from a single healthy organism to produce samples with cancer fractions of 10%, 1%, 0.1%, 0.01%, and 0.001% by mass. Triple RARE-seq libraries were generated from each mixture, as well as from NCI-H1975 RNA (100% cancer fraction) and cfRNA alone (0% cancer fraction), and captured by a whole-coding transcriptome panel. Further mixtures using cfRNA from different healthy organisms were generated and captured by a RAG panel (see "RAG Capture Panel Design" below).
[0181] cfDNA library preparation and sequencing Cell-free DNA (cfDNA) was extracted from plasma using the standard cfDNA protocol of the QIAamp Circulating Nucleic Acid Kit (Qiagen). Following extraction, cfDNA was quantified using the Qubit Double-Stranded DNA High Sensitivity Kit (Thermo Fisher Scientific) and a High Sensitivity NGS Fragment Analyzer (Agilent). A sequencing library was generated from 32 ng of cfDNA using deep sequencing-based cancer individualized profiling (CAPP-Seq). Subsequently, genes that repeatedly mutate in lung cancer were targeted using hybridization-based capture (Roche NimbleGen). The library was sequenced on an Illumina HiSeq4000 instrument using paired-end reads of 2 × 150 bp.
[0182] Mapping, deduplication, and quality control for RARE-Seq First, we deduplication of FASTQ files using a custom pipeline. Fastp (v0.20.0) trimmed the first 10 bases from the 5' end of read 1 and the 3' end of read 2, while simultaneously removing low-quality or too-short read pairs from each sample. Using STAR 2-pass (v2.7.0), we aligned the remaining high-quality reads against a reference transcriptome (GENCODE v27) and the human genome (hg19). Using a custom barcoding method, we removed PCR duplication from both the transcriptome and genome-aligned files. Using the deduplication-free reads, we estimated gene-level expression using RSEM (v1.2.28).
[0183] Quality control was evaluated using metrics calculated by the RNASeQC package (v2.3.5), such as read mapping quality and mapping rate, which included exon rate, intron rate, and inter-gene ratios, as well as ribosomal RNA rate. In addition, DNA contamination was estimated by calculating the percentage of reads that map to intron sequences out of the total number of reads that straddle exon boundaries. Samples with fewer than 20 million reads or estimated DNA contamination exceeding 20% were removed from downstream analysis. In summary, three samples were removed based on these QC thresholds.
[0184] Gene expression normalization and platelet correction Expression analysis was performed using RSEM predicted counts of the captured genes (tximport R package v1.22). First, counts were normalized using the trimmed mean (TMM) method of M values to account for inter-sample variability in library size and transcriptome complexity (edgeR R package v3.36). The logarithmically transformed and normalized expression values are referred to herein as "log2NX". Next, expression variations caused by platelets were corrected using a modified remove unwanted variation (RUV) method (RUVseq R package v1.28; D. Risso, et al., Nat. Biotechnol. 32, 896-902 (2014), its disclosure incorporated herein by reference). In the full code transcriptome library, the coefficient of unwanted platelet variability was estimated using RUVg, with 130 platelet cell type marker genes from PanglaoDB used as a set of negative control genes (Figure 3I) (O Franzen, et al., Database (Oxford). 2019 Jan 1;2019:baz046, its disclosure is incorporated herein by reference). In the RAG capture library, since most platelet genes were intentionally excluded from the panel (see "RAG Capture Panel Design" below), the coefficient of unwanted variability was estimated using healthy control meta-reference cfRNA (n=15) as a negative control sample within RUVs. The coefficient of unwanted platelet variability was selected to be the coefficient that best correlated with platelet expression and was defined as the mean log2NX across 21 captured platelet cell type marker genes (e.g., PPBP) (Figure 8G). In both RUVg and RUVs, platelet correction was performed by generalized least-squares regression of log2NX on the selected coefficient of unwanted platelet expression. RUVg and RUVs were performed similarly on all coding transcriptome samples (Figure 3J and Figure 6H). After normalization and correction, quality control was performed by calculating the Pearson correlation (i.e., "intra-group correlation") between the gene expression profile of each sample and the mean value of all other samples in the same sample group.In controls with two or more biological copies, the copy with the highest in-group correlation was selected for downstream analysis.
[0185] Differential gene expression analysis Differential gene expression analysis was performed using DESeq2 (DESeq2 R package v1.34; M. Love, et al., Genome Biol. 15, 550 (2014), its disclosure incorporated herein by reference). For differential expression analysis of cfRNA and cellular RNA, a generalized linear model (GLM) was constructed using sample type as a covariate. In the differential expression analysis of cancer versus control, the GLM used the state (i.e., cancer or control) and the coefficient of unwanted platelet variation determined by RUV as covariates. After determining the significance of the log-fold change estimate using Wald's test, Benjamini-Hochberg's multiple hypothesis correction was performed. In the COVID-19 vaccination time series analysis, the significance of differential expression was assessed using likelihood ratio tests comparing the full GLM, including vaccination time and coefficients of unwanted variation, to a reduced model with the vaccination period removed. In all analyses, significantly differentially expressed genes were defined by an absolute log-fold change greater than 1 and a corrected p-value < 0.05. For functional analysis of differentially expressed genes, gene set enrichment analysis (GSEA) (fgsea R package v1.20) or gene ontology (GO) analysis (goseq R package v1.46) of pre-ranked log-fold change estimates was used. GSEA and GO analyses were performed using cell marker genes from either the PanglaoDB single-cell RNA sequencing database and / or the molecular signature database (MSigDb) (O. Franzen, et al., Database (Oxford). 2019 Jan 1;2019:baz046; A. Subramanian, et al., Proc. Natl. Acad. Sci. USA 102, 15545-15550 (2005); A. Liberzon, et al., Cell Syst 1, 417-425 (2015); and A. Liberzon, et al., Bioinformatics 27, 1739-1740 (2011); these disclosures are incorporated herein by reference, respectively).
[0186] Design of a capture panel focused on RAG To identify rare genes (RAGs) in cfRNA, gene expression data from controls generated by the whole-coding transcriptome RARE-Seq method were analyzed (n=50 biological copies from n=28 individuals). In addition, gene expression data from whole blood samples generated by the genotype-tissue expression (GTEx) project were obtained via the UCSC Xena repository (n=307). Expression was TMM normalized, and k-mean clustering (k=2) was used to categorize genes as expressed or not expressed in each sample. RAGs were defined as genes expressed in less than 5 percent of all samples for both cfRNA and whole blood, with their mean log2NX falling in the bottom 30 percent of all genes. RAGs are listed in Table 3. Next, expression homogeneity in cfRNA and whole blood was calculated using the Gini coefficient, and endogenous control genes were selected from circulating housekeeping genes with a Gini coefficient of <0.2. In addition, lung cancer-related genes, including those that recurrently mutate or rearrange in lung cancer or are abnormally expressed in lung cancer histological types, were manually selected and included in the panel. In total, the RAG-focused capture panel included 7,766,820 bp and 80,866 probes targeting 5,546 unique genes. The probes were also designed to target sequences from Pfizer-BioNTech and Moderna's COVID-19 mRNA vaccines and were used to supplement the RAG capture panel with samples collected after vaccination.
[0187] Genetic signatures of cell types, tissues, and cancers For cell type-specific signatures, 7,481 marker genes derived from 170 human cell types were downloaded from the PanglaoDB single-cell RNA sequencing database (O Franzen, et al., Database (Oxford). 2019 Jan 1;2019:baz046, its disclosure is incorporated herein by reference).
[0188] Tissue and cancer-enriched signatures identified using gene expression data from the Genotype-Tissue Expression (GTEx) and Cancer Genome Atlas (TCGA) projects were analyzed. Gene-level read counts generated by the UCSC Toil RNA sequencing bioinformatics pipeline were obtained in the UCSC Xena repository. Gene expression was TMM normalized, and outlier samples were removed using the intra-group correlation metric. The TCGA cohort was further filtered to exclude tumors with tumor purity less than 60 percent using the Consensus Purity Estimate (CPE). Normal tissue types evaluated by GTEx included bladder (n=9), brain (n=1,107), breast (n=168), colon (n=261), esophagus (n=598), kidney (n=24), liver (n=97), lung (n=241), ovary (n=79), pancreas (n=145), prostate (n=86), skin (n=497), stomach (n=167), and whole blood (n=290). Genes were defined as tissue-enriched if their expression in each tissue was 5 times higher than in all other tissues evaluated, and if the mean tissue log2NX was greater than zero. Similarly, cancer histological types evaluated by TCGA included bladder (BLCA, n=344), brain (GBM, n=144; LGG, n=183), breast (BRCA, n=914), colon (COAD, n=270), esophagus (ESCA, n=181), kidney (KICH, n=65; KIRC, n=350; KIRP, n=262), liver (LIHC, n=164), lung (LUAD, n=309; LUSC, n=365), ovary (OV, n=412), pancreas (PAAD, n=177), prostate (PRAD, n=473), melanoma (SKCM, n=94), and stomach (STAD, n=413). A gene was defined as cancer-enriched if its expression in each cancer tissue type was five times higher than in all other cancer types evaluated, and if the mean cancer tissue log2NX was greater than zero.
[0189] Cancer-specific gene signatures were generated to include the genes most differentially expressed in each cancer tissue compared to whole blood and meta-reference cfRNA, and were therefore most likely to be detected in cfRNA mixtures with a high hematopoietic background. For this purpose, gene expression data from the Cancer Cell Line Encyclopedia (CCLE) (LIHC, n=25; LUAD, n=76; PAAD, n=40; PRAD, n=7; SCLC, n=50) obtained via Xena and from further SCLC studies (n=79) were added to the TCGA cancer cohort. DESeq2 differential gene expression analysis quantified the gene-by-gene ratio changes between cancer tissue expression and whole blood and meta-reference cfRNA expression. Each gene was also categorized using K-mean clustering (k=2) of the mean gene expression in each group, indicating whether it was expressed or not in cancer tissue, whole blood, and meta-reference cfRNA. Genes were selected for cancer-specific signatures when: 1. When found in the top 1% of genes overexpressed in the target cancer tissue or whole blood and meta-reference cfRNA, and 2. When categorized as being expressed in the target cancer tissue or in whole blood and meta-reference cfRNA.
[0190] Instead of analyzing each cancer type individually, we generated cancer TOO gene signatures using a similar method, except that we compared gene expression between LIHC, LUAD, PAAD, PRAD, whole blood, and meta-reference cfRNAs. Genes were selected for cancer TOO signatures in the following cases: 1. If a gene is found in the top 1% of genes that are overexpressed in cancer tissue compared to whole blood and meta-reference cfRNA, 2. When the gene is found in the top 5% of genes that are overexpressed in the target cancer tissue compared to other cancer types, 3. When categorized as being expressed in the target cancer tissue but not in whole blood or meta-reference cfRNA.
[0191] In all signatures, gene weights were defined as the sum of cancer tissue log2NX and the logarithmic change in cancer tissue versus meta-reference cfRNA, and the weights preserved the sign of differential expression analysis (i.e., positive for genes overexpressed in cancer tissue).
[0192] Gene signature detection using enrichment scores An enrichment score (ES) analysis framework was developed to be broadly applicable, as it allows the use of any desired gene signature. First, gene expression in the samples was standardized to Z-scores using the mean and standard deviation of meta-reference cfRNA expression. Next, the Z-scores of the desired gene signatures were aggregated using Stuffer's equation and gene weights, resulting in a single ES for each sample. Where applicable, Z-scores were aggregated separately and summed for genes with positive and negative weights. Random gene signatures of the same size (but excluding genes in the gene signature) were generated using bootstrapping (n=1,000), and empirical p-values were estimated for each signature in each sample using the bootstrapped ES distribution. If the empirical p-value was less than 0.05, the ES was set to zero.
[0193] Estimation of detection limits using H1975 mixture The 95% detection limit (LOD95) of RARE-Seq was determined using an RNA mixture in silico. Sequencing data for NCI-H1975 RNA and healthy control cfRNA were generated and analyzed for each sample type as described. Reads were randomly subsampled from each replication and merged into specific precancerous fractions ranging from 0% to 100%, and cancer was detected in each sample using an enrichment score analysis framework. Here, NCI-H1975 gene signatures (H1975 Sig) were generated by comparing NCI-H1975 samples (n=4) processed with RARE-Seq but not used for spike generation with meta-reference cfRNA (n=15). Genes were selected in the following cases: 1. In NCI-H1975, when overexpressed compared to meta-reference cfRNA, it is found in the top 1% of genes, and 2. When categorized as being expressed in NCI-H1975.
[0194] The blank limit (LOB) was calculated from the mean and standard deviation of H1975 Sig ES found in cfRNA controls (n=12) not used for gene signature generation, and set as the detection threshold for remaining spikes. Logistic regression was used to model the association between cancer fractions and H1975 Sig detection, and LOD95 was defined as the cancer fraction with sensitivity greater than 95%. This process was repeated for various RNA inputs and sequencing depths. The above in vitro NCI-H1975 mixture was treated as other cfRNA samples and detected using H1975 Sig to confirm the in silico results.
[0195] Cancer detection using elastic network classification Statistical learning models including ElasticNet (EN) logistic regression, random forest, and XGBoost were trained to classify LUAD patient cfRNAs (n=50) derived from controls (n=50). Hyperparameters were tuned and model performance estimated using nested 10-fold cross-validation (10CV) (caret R package v6.0). Features considered for model training included expression from all capture genes, AF from all recurrent somatic variants searched in ctRNA, and binarized values representing the presence of SNVs, fusions, indels, splice variants, or any type of variant detected in a given sample. Preprocessing removed features with near-zero variance in the training cohort, as well as genes not significantly differentially expressed between LUAD tumor tissue and meta-reference cfRNA as described in the enrichment score framework. Gene weights used in the enrichment score framework were also evaluated as penalty weights within the weighted ElasticNet model. To evaluate the impact of cohort size on model performance, 10CVs were repeated after subsampling 20–100% of the cohort, and the standard deviation of AUC at each subsampling level was determined across n=10 random samples. All trained models performed reasonably well in classification, but the best performance was achieved with weighted EN logistic regression. In the final weighted EN model constructed using the fully trained cohort, 279 features were selected from 1,006 considered features, including those related to variant detection in ctRNA. The final model was validated using a pending cfRNA cohort collected at an independent institution (n=28 time points from 14 LUAD patients, n=10 LDCT controls). ROC curves and metrics, e.g., sensitivity, specificity, and AUC, were generated using the pROC R package (v1.18).
[0196] Classification of cancer origin tissue Samples detected by arbitrary cancer-specific gene signatures were considered for tissue origin (TOO) analysis. TOO was classified using a modified enrichment score framework. Z-scores were aggregated by calculating the 90th percentile of all Z-scores for the target gene signature. For each sample, the scores of all cancer TOO gene signatures were ranked. Accuracy was determined based on whether the highest or second highest cancer TOO score was detected in the diagnosed cancer type.
[0197] ctRNA variant call We performed variant calls on a target list of potentially effective NSCLC alterations, compiled based on the National Comprehensive Cancer Information Network (NCCN) guidelines. For each candidate variant, we tested the null hypothesis that the variant allele frequency (AF) was consistent with the site-specific and depth-specific background error distribution. The background error rate b was calculated as the proportion of reads supporting substitutions across a training cohort (n=53) consisting of control cfRNAs. The background error distribution was obtained by sampling 10,000 positive values from a binomial distribution using candidate variant depth as the "number of trials" and b as the "success rate" (from which the p-value was estimated). The resulting p-values were corrected for multiple hypothesis testing using the Benjamini-Hochberg method, and substitutions with a q-value < 0.05 were considered significant. If present in the list of known fusion partners in lung cancer annotated on cBioPortal, candidate fusions identified by STAR-Fusion (v1.6.0; A. Dobin, et al., Bioinformatics 29, 15-21 (2013), its disclosure incorporated herein by reference) were called. If the mapping included splice-joint exons 13-15, a MET exon 14 skipping modification was called. If it exceeded 6 base pairs, indels in EGFR exons 19 and 20 were called.
[0198] ctDNA variant call Where available, a patient-specific variant list was collected from pre-clinical tumor genotyping, otherwise from pre-treatment plasma, including only variants similarly captured by the lung cancer CAPP-Seq panel. CAPP-Seq sequencing data were analyzed as follows: Briefly, reads were demultiplexed, trimmed to remove low-quality bases from the 3' end using fastp, and aligned to the hg19 human genome using BWA ALN. PCR duplication was removed using a custom barcoding method. Single nucleotide variants (SNVs) and insertion / deletion events (indels) were called using a previously reported integrated digital error suppression (iDES) pipeline (AM Newman, et al., Nat. Biotechnol. 34, 547-555 (2016), its disclosure incorporated herein by reference). Gene fusions were identified using FACTERA (AM Newman, et al., Bioinformatics 30, 3390-3393 (2014), the disclosure of which is incorporated herein by reference). The ctDNA sample allele frequencies (AFs) were calculated by averaging the AFs of the detected patient-specific variants.
[0199] MET amplification gene signature MET amplification was used to generate EGFR mutant NSCLC cell lines with acquired resistance to EGFR TKIs (PC9 / PC9-PERC17, MGH1157-1 / MGH1157-3, MGH170-1C#7 / MGH170-1D#2). The PC9-PERC17 cell line was generated by culturing PC9 cells in vitro with erlotinib. The MGH1157-1 and MGH1157-3 cell lines were developed from sequential biopsies of patients before and after treatment with a combination of gefitinib and nazartinib, respectively. Clinical profiling of MGH1157-3 tumors using MET FISH demonstrated MET amplification. The MGH170 cell line was developed from patients who acquired MET amplification after erlotinib treatment (confirmed by FISH). Single-cell clones were isolated from two separate metastatic lesions at the time of rapid autopsy after disease progression. The MGH170-1C#7 clone cell line was found to be sensitive to EGFR TKIs, while MGH170-1D #2 was found to be resistant. MET amplification was confirmed in all resistant cell lines, and MET dependence was confirmed by sensitivity to the EGFR+MET TKI combination. All tissue samples were obtained after patients signed informed consent, agreeing to participate in a protocol approved by the Dana-Farber-Harvard Cancer Center Institutional Review Board and authorizing the study to be conducted on their samples. mRNA-seq was performed on total RNA from pre-treatment and resistant cell lines (n=3 replicas or n=18 total samples at each time point for n=3 cell lines), and differential expression analysis was performed using DESeq2 with MET amplification status as a covariate. Genes were selected for MET amplification gene signatures when: 1. In MET-amplified cell lines, if it is found in the top 1% of overexpressed genes compared to meta-reference cfRNA, 2. When MET-amplified cell lines are expressed significantly differentially compared to non-amplified cell lines, and 3. When categorized as being expressed in MET-amplified cell lines.
[0200] Here, gene weights were defined as the sum of the MET-amplified cell line log2NX and the logarithmic change between MET amplification and meta-reference cfRNA, and the sign of differential expression analysis (i.e., positive for genes overexpressed in MET-amplified cell lines) was preserved by the weights.
[0201] statistical analysis All statistical analyses were performed using R (v4.1.3). Statistical tests used throughout the manuscript include the Wilcoxon rank-sum test, paired t-test, Fisher's exact test, Kruskal-Wallis ANOVA, Pearson correlation, Spearman correlation, and DeLong's AUC comparison test. Unless otherwise specified, all statistical tests comparing two groups used the two-tailed Wilcoxon rank-sum test, and all statistical tests comparing three or more groups used Kruskal-Wallis ANOVA. Unless p-values are presented, significance indicators in the figure panels are used as follows: * P<0.05; ** P<0.01; *** P<0.001, **** P<0.0001. Multiple hypothesis correction was performed using the Benjamini-Hochberg (BH) method. Unless otherwise specified, box plots represent the interquartile range (box), median (centerline), and minimum / maximum values (whiskers). Unless otherwise specified, sample size (n) presented in the text, figure panels, and figure captions refers to biologically independent individuals. Details of the statistical methods (including R packages) used for the enrichment score analysis framework, elastic net model training, and modified organism calling method are described in the respective Methods sections.
[0202] Optimization of blood collection and RNA extraction As determined using RNA-specific quantitative RT-PCR, the median concentration of cfRNA per mL of plasma was determined to be 220 pg (Figure 1B). Given such low concentrations, we first attempted to optimize blood collection and RNA extraction. Since hemolysis is known to be a significant pre-analysis factor affecting other blood-based biomarkers, we investigated its effect on cfRNA concentration. Hemoglobin levels did not correlate with cfRNA levels in the cohort, suggesting that hemolysis is not a significant pre-analysis variable in this environment (Figure 1C). In contrast, the length of time plasma was stored at -80°C negatively correlated with cfRNA levels. However, this effect was observed only during the first month of storage; thereafter, no association existed between storage time and cfRNA yield (Figure 1D). Next, we investigated the effect of various blood collection tubes on cfRNA concentration. All types of blood collection tubes (BCTs) tested, including those with and without cell-free nucleic acid preservatives, allowed for cfRNA isolation with minimal concentration differences (Figure 1E). Ten cfRNA isolation protocols were compared. Two silica column-based purification protocols yielded significantly higher concentrations (QIAamp Circulating Nucleic Acid kit and a customized version of the QIAamp Viral RNA kit; Figure 1F), and these were used for subsequent experiments.
[0203] Optimization of RNA sequencing library preparation Next, we developed and optimized a library preparation protocol that coupled cDNA synthesis based on random primers with ligation of a sequencing adapter containing a unique molecule identification (UMI) barcode, enabling accurate counting of unique cfRNA molecules recovered from plasma. After generating such a cfRNA-derived sequencing library, affinity hybridization was used to biotinylate oligonucleotides and capture the coded transcriptome. Sequencing data was processed using the custom bioinformatics pipeline described herein. High correlation was confirmed between measured cfRNA expression and the type of tube used (Figure 2A) and extraction method (Figure 2B). A protocol that preserves the orientation (i.e., strandedness) of the original RNA molecules was hypothesized to be useful for cfRNA analysis, given that this method has been reported to be useful for differential expression analysis in whole blood. However, while gene-specific expression levels were nearly identical between the stranded and non-stranded libraries, the non-stranded library contained approximately 30% more unique molecules (Figures 2C and 2D), suggesting that the stranded library preparation protocol resulted in transcript loss. We also investigated whether RNA transcript recovery could be improved using methods known to overcome the highly damaged nature of FFPE nucleic acids. Further end repair using a single-strand specific S1 nuclease more than doubled transcript recovery while faithfully preserving expression levels (Figures 2E and 2F). Therefore, although the protocol incorporated non-stranded library preparation and S1 nuclease end repair, the results could have been achieved using other protocols.
[0204] Digestion of contaminating cfDNA from extracted nucleic acids Since cfRNA analysis is known to be confounded by the presence of contaminating cfDNA, we evaluated the optimal method for DNA digestion using DNase I. Most silica column-based RNA isolation protocols recommend enzymatic DNA digestion while the RNA is bound to the column, but this was found to result in incomplete DNA removal. Performing the enzymatic digestion step on eluted nucleic acids reduced the cfDNA contamination level from 63.6% to 9.1% (Figure 2G). Expression correlations were substantially reduced between samples with and without DNA contamination, highlighting the importance of DNA removal before sequencing library preparation (Figure 2H). To further protect against DNA contamination in cfRNA analysis, cfRNA samples with high levels of putative DNA contamination were excluded from subsequent experiments.
[0205] Verification of RARE-Seq optimization Prior to application to clinical samples, the RARE-Seq method was evaluated against the widely used SMART-Seq whole transcriptome method, which includes nuclear and mitochondrial ribosomal RNA depletion but does not include other enrichment steps. Consistent with previous reports, RNA encoding protein-coding genes was the most abundant biotype in cfRNAs analyzed using SMART-Seq, accounting for 67.1% of mapped reads (range 55.6–68.2%) (Figure 2I). Protein-coding gene expression was highly matched between the SMART-Seq library and the RARE-Seq library derived from compatible cfRNAs (R=0.96, P=<2.2e-16) (Figure 2J). However, in RARE-Seq, 80.4% of coding genes had equal or higher sequencing depths, representing a 44% increase in total sequencing depth (Figure 2J). RARE-Seq reproducibility was evaluated using eight copies sequenced in five batches, collected from the same individual over four separate days. Three libraries were generated from the same cfRNA pool and used as technical copies. High mean correlations were observed: 0.988 for biological copies and 0.997 for technical copies (Figure 2K). In summary, these results suggest that RARE-Seq is a robust and reproducible method for measuring plasma cfRNA expression. Table 1. Non-cancer cohort population statistics [Table 1] Table 2. Cancer Cohort Population Statistics [Table 2] Table 3. Rare Genes (RAGs) RAG was defined as being expressed in less than 5% of healthy samples, with its mean log2NX falling in the bottom 30 percent of all genes. [Table 3-1] [Table 3-2] Table 3-3 Table 3-4 Table 3-5 Table 3-6 Table 3-7 Table 3-8 Table 3-9 Table 3-10 Table 3-11 Table 3-12 Table 3-13 Table 3-14 Table 3-15 Table 3-16 Table 3-17 Table 3-18 Table 3-19 Table 3-20 Table 3-21 Table 3-22 Table 3-23 Table 3-24 Table 3-25 Table 3-26 Table 3-27 Table 3-28 Table 3-29 Table 3-30 Table 3-31 Table 3-32 Table 3-33 Table 3-34 Table 3-35 Table 3-36 Table 3-37 Table 3-38 Table 3-39 Table 3-40 Table 3-41 Table 3-42
Claims
1. A method for preparing cell-free RNA sequencing, A step of providing a sample containing nucleic acids for sequencing, wherein the nucleic acids for sequencing are cell-free RNA, or nucleic acids derived from and representative of cell-free RNA, The sample is brought into contact with a panel of nucleic acid molecules, including those derived from or having complementary sequences to gene transcripts that are rare as cell-free RNA molecules in a control liquid biopsy, and as a result, the panel of nucleic acid molecules anneals with a subset of nucleic acids for sequencing. Methods that include...
2. The method according to claim 1, wherein the transcript, which is rarely found as a cell-free RNA molecule in a control fluid biopsy, is defined as a transcript expressed in less than 50% of the control fluid biopsy population.
3. The method according to claim 2, wherein the transcript, which is rarely found as a cell-free RNA molecule in a control fluid biopsy, is defined as a transcript expressed in less than 5% of the control fluid biopsy population.
4. The method according to any one of claims 1 to 3, wherein the transcript, which is rare as a cell-free RNA molecule in a control fluid biopsy, is defined as a transcript that falls in the bottom 60% of genes with respect to normalized expression across a population of control fluid biopsies.
5. The method according to claim 4, wherein the transcript, which is rare as a cell-free RNA molecule in a control fluid biopsy, is defined as a transcript that falls in the bottom 30% of genes in terms of normalized expression across the population of the control fluid biopsy.
6. The method according to any one of claims 1 to 5, wherein the transcript, which is rare as a cell-free RNA molecule in a control fluid biopsy, is defined as a transcript whose logarithmically transformed and normalized expression value is less than zero across a population of control fluid biopsies.
7. The method according to any one of claims 2 to 6, wherein the group of control liquid biopsies comprises at least five types of liquid biopsies.
8. The method according to claim 7, wherein the group of control liquid biopsies comprises at least 50 types of liquid biopsies.
9. The method according to any one of claims 1 to 8, wherein the control fluid biopsy is collected from an individual that, at the time the biopsy is collected, does not have one or more of the following: observed pathogenic infection, diagnosed cancer, diagnosed metabolic disorder, diagnosed neurological disorder, diagnosed immunodeficiency disorder, diagnosed autoimmune disorder, diagnosed inflammatory disorder, diagnosed cardiovascular disorder, diagnosed renal disorder, diagnosed hepatic disorder, active pregnancy, diagnosed pregnancy complications, diagnosed fetal complications, organ transplant, active rejection of organ transplant, obesity, malnutrition, cachexia, and abnormalities in clinical trials.
10. The method according to any one of claims 1 to 9, wherein the transcripts that are rare as cell-free RNA molecules in a control liquid biopsy comprise 50% of the transcripts in Table 3.
11. The method according to claim 10, wherein the transcripts that are rarely found as cell-free RNA molecules in a control liquid biopsy comprise 90% of the transcripts listed in Table 3.
12. The method according to claim 11, wherein the transcripts that are rare as cell-free RNA molecules in a control liquid biopsy comprise 100% of the transcripts in Table 3.
13. The method according to any one of claims 1 to 12, wherein the panel of nucleic acid molecules excludes at least 50% of all exome gene transcripts that are not transcripts that are rare as cell-free RNA molecules.
14. The method according to claim 13, wherein the panel of nucleic acid molecules excludes at least 90% of all exome gene transcripts that are not transcripts that are rare as cell-free RNA molecules.
15. The method according to any one of claims 1 to 14, wherein the panel of nucleic acid molecules consists of 5,000 or fewer gene transcripts in addition to transcripts that are rare as cell-free RNA molecules.
16. The method according to claim 15, wherein the panel of nucleic acid molecules consists of 500 or fewer gene transcripts in addition to transcripts that are rare as cell-free RNA molecules.
17. The method according to any one of claims 1 to 16, wherein the panel of nucleic acid molecules further comprises one or more sets of tissue-specific transcripts, cell-type-specific transcripts, clinically relevant transcripts, B-cell receptor and T-cell receptor transcripts, biomarkers, generally mutagenic transcripts, and control transcripts for inter-sample normalization.
18. The method according to claim 17, wherein the biomarker is associated with one of the following biological characteristics: medical disorder, pregnancy, fetal complications, pregnancy complications, neoplasm growth, cancer, specific cancer types, pathogenic infection, immune activation, organ transplant rejection, neurodegeneration, originating tissue, originating cell type, or activation of a biochemical pathway.
19. The method according to any one of claims 1 to 18, wherein the panel of nucleic acid molecules is a set of probes for targeted capture hybridization.
20. The method according to any one of claims 1 to 18, wherein the panel of nucleic acid molecules is a set of primers for targeted amplification.
21. The method according to any one of claims 1 to 19, wherein the cfRNA sample is derived from blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, or feces.
22. The steps include generating a sequencing library derived from the aforementioned sample, A step of performing targeted sequencing of the sequencing library to obtain sequencing results of the cell-free RNA, wherein the sequencing is targeted to the panel of nucleic acid molecules. The method according to claim 1, further comprising:
23. Steps to remove platelet expression from the sequencing results in silico. The method according to claim 22, further comprising:
24. Step 1: Perform differential transcript analysis using the sequencing results and the second sequencing results. The method according to claim 22 or 23, further comprising:
25. Step of detecting enrichment of at least one expression signature within the sequencing results. The method according to any one of claims 22 to 24, further comprising:
26. Step of detecting sequence mutagenicity within the sequencing results. The method according to any one of claims 22 to 25, further comprising:
27. Steps to infer the copy number state of one or more genes from the sequencing results. The method according to any one of claims 22 to 26, further comprising:
28. A step of using the sequencing results together with several other sequencing results to train a computational model to predict the categorical state or likelihood of a biological characteristic, wherein the cell-free RNA sample has a known categorical state of a biological characteristic. The method according to any one of claims 22 to 27, further comprising:
29. A step of using the sequencing results as input to a trained computational model to predict the categorical state or likelihood of a biological characteristic, wherein the computational model is trained using a cohort of RNA sequencing results having known categorical states of biological characteristics. The method according to any one of claims 22 to 27, further comprising:
30. A step of deriving one or more features from the sequencing results, wherein the one or more features include enrichment of one or more gene signatures, enrichment of biochemical pathways, a collection of sequence variants, and a copy number state. A step of using one or more derived features as input to a trained computational model to predict the categorical state or likelihood of a biological trait, wherein the computational model is trained using a cohort of RNA sequencing results having known categorical states of biological traits. The method according to any one of claims 22 to 27, further comprising:
31. A panel of nucleic acids for targeting transcripts that are rare as cell-free RNA molecules, Nucleic acid molecules derived from gene transcripts or having complementary sequences, which are rare as cell-free RNA molecules in control fluid biopsies. A panel containing nucleic acids.
32. The nucleic acid panel according to claim 31, wherein the transcript that is rarely present as a cell-free RNA molecule in a control fluid biopsy is defined as a transcript expressed in less than 50% of the control fluid biopsy population.
33. The nucleic acid panel according to claim 32, wherein the transcript that is rare as a cell-free RNA molecule in a control fluid biopsy is defined as a transcript expressed in less than 5% of the control fluid biopsy population.
34. A panel of nucleic acids according to any one of claims 31 to 33, wherein the transcripts that are rare as cell-free RNA molecules in a control fluid biopsy are defined as transcripts that fall in the bottom 60% of genes with respect to normalized expression across a population of control fluid biopsies.
35. The panel of nucleic acids according to claim 34, wherein the transcripts that are rare as cell-free RNA molecules in a control fluid biopsy are defined as transcripts that fall in the bottom 30% of genes with respect to normalized expression across the population of the control fluid biopsy.
36. A panel of nucleic acids according to any one of claims 31 to 35, wherein the transcripts that are rare as cell-free RNA molecules in a control fluid biopsy are defined as transcripts whose logarithmically transformed and normalized expression values are less than zero across a population of control fluid biopsies.
37. The nucleic acid panel according to any one of claims 32 to 36, wherein the group of control liquid biopsies comprises at least five types of liquid biopsies.
38. The nucleic acid panel according to claim 37, wherein the group of control liquid biopsies comprises at least 50 types of liquid biopsies.
39. The nucleic acid panel according to any one of claims 31 to 38, wherein the control fluid biopsy is collected from an individual that does not have one or more of the following conditions when the biopsy is collected: observed pathogenic infection, diagnosed cancer, diagnosed metabolic disorder, diagnosed neurological disorder, diagnosed immunodeficiency disorder, diagnosed autoimmune disorder, diagnosed inflammatory disorder, diagnosed cardiovascular disorder, diagnosed renal disorder, diagnosed hepatic disorder, active pregnancy, diagnosed pregnancy complications, diagnosed fetal complications, organ transplant, active rejection of organ transplant, obesity, malnutrition, cachexia, and abnormalities in clinical trials.
40. A panel of nucleic acids according to any one of claims 31 to 39, wherein the transcripts that are rare as cell-free RNA molecules in a control liquid biopsy comprise 50% of the transcripts in Table 3.
41. The nucleic acid panel according to claim 40, wherein the transcripts that are rare as cell-free RNA molecules in a control liquid biopsy comprise 90% of the transcripts in Table 3.
42. The nucleic acid panel according to claim 41, wherein the transcripts that are rare as cell-free RNA molecules in a control liquid biopsy comprise 100% of the transcripts in Table 3.
43. The nucleic acid panel according to any one of claims 31 to 42, wherein the panel of nucleic acid molecules excludes at least 50% of all exome gene transcripts that are not transcripts that are rare as cell-free RNA molecules.
44. The nucleic acid panel according to claim 43, wherein the panel of nucleic acid molecules excludes at least 90% of all exome gene transcripts that are not transcripts that are rare as cell-free RNA molecules.
45. The nucleic acid panel according to any one of claims 31 to 44, wherein the panel of nucleic acid molecules comprises 5,000 or fewer gene transcripts in addition to transcripts that are rare as cell-free RNA molecules.
46. The nucleic acid panel according to claim 45, wherein the panel of nucleic acid molecules comprises 500 or fewer gene transcripts in addition to transcripts that are rare as cell-free RNA molecules.
47. The panel of nucleic acids according to any one of claims 31 to 46, wherein the panel of nucleic acid molecules further comprises tissue-specific transcripts, cell-type-specific transcripts, clinically relevant transcripts, B-cell receptor and T-cell receptor transcripts, biomarkers, and generally mutagenic transcripts.
48. The nucleic acid panel according to claim 47, wherein the biomarker is associated with one of the following biological characteristics: medical disorder, pregnancy, fetal complications, pregnancy complications, neoplasm growth, cancer, specific cancer types, pathogenic infection, immune activation, organ transplant rejection, neurodegeneration, originating tissue, originating cell type, or activation of biochemical pathways.
49. The nucleic acid panel according to any one of claims 31 to 48, wherein the panel of nucleic acid molecules further comprises a set of control transcripts for inter-sample normalization.
50. The nucleic acid panel according to any one of claims 31 to 49, wherein the panel of nucleic acid molecules is a set of probes for targeted capture hybridization.
51. The nucleic acid panel according to any one of claims 31 to 49, wherein the panel of nucleic acid molecules is a set of primers for targeted amplification.
52. A method for extracting RNA from a cell-free source, (a) The step of adding glycogen to a sample containing cell-free nucleic acids and (b) The step of bringing a silica-based column into contact with a sample containing cell-free nucleic acids. Methods that include...
53. The method according to claim 52, wherein step (a) is performed before step (b).
54. The steps include: eluting cell-free nucleic acids from the silica-based column to obtain a solution of extracted cell-free nucleic acids; The step of contacting the extracted cell-free nucleic acid solution with DNase. The method according to claim 52 or 53, further comprising:
55. A method for quantifying cell-free RNA for downstream molecular applications, A step of providing a sample containing cell-free RNA, The steps include: reverse transcription of the cell-free RNA to obtain cDNA, A quantitative real-time polymerase chain reaction and a step of quantifying the concentration of cell-free RNA in the solution using the cDNA. Methods that include...
56. The method according to claim 55, wherein the step of quantifying the concentration of cell-free RNA further comprises the step of generating a standard curve based on a set of control standards having known concentrations, the control standards also being evaluated using a quantitative real-time polymerase chain reaction.
57. The sample further contains cell-free DNA, and the method is A step of quantifying the concentration of cell-free DNA in a sample using a quantitative real-time polymerase chain reaction, wherein the cell-free RNA is quantified by using a primer that spans gene introns that are relatively stable across the cell-free RNA sample, and the cell-free DNA is quantified by using a primer that anneals to a transcriptionally silent region of the genome that is relatively stable across the cell-free DNA sample. The method according to claim 55 or 56, further comprising:
58. The method according to claim 57, wherein the primer for quantifying cell-free RNA spans an intron of GAPDH, and the primer for quantifying a cell-free DNA target covers a 78 bp transcriptionally silent region of chromosome 12.
59. A method for sequencing cell-free RNA, A step of providing a library of nucleic acid molecules, wherein the library of nucleic acid molecules is derived from cell-free RNA, and the cell-free RNA is derived from a liquid biopsy. The steps include sequencing the library of nucleic acid molecules to obtain sequencing results, A step to eliminate fluctuations caused by the expression of platelets and related transcripts. Methods that include...
60. The method according to claim 59, wherein the library of nucleic acid molecules is produced by capturing or amplifying nucleic acid molecules.
61. The method according to claim 60, wherein the nucleic acid library is a whole exome library.
62. The method according to claim 60, wherein the nucleic acid library is a library that targets genes that are rare in existence.
63. The method according to any one of claims 59 to 62, wherein the cfRNA sample is derived from blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, or feces.
64. A method for generating a targeted sequencing panel for cell-free RNA sequencing, Each step involves collecting a population of control liquid biopsies containing cell-free RNA, The steps include performing sequencing of the cell-free RNA in the control liquid biopsy, The following steps include: identifying a set of genes that are rare in a population of control fluid biopsies, defined by their expression in a certain percentage of a population of control fluid biopsies or their expression level across the population of control fluid biopsies, The steps include synthesizing a set of nucleic acid molecules that are used for capturing or amplifying rare genes to obtain a targeted sequencing panel for sequencing cell-free RNA, and Methods that include...
65. The method according to claim 64, wherein the set of rare genes is defined by its expression in at least a certain percentage of a population of control fluid biopsies and its expression level across the population of control fluid biopsies.
66. The method according to claim 64 or 65, wherein the obtained targeted sequencing panel for sequencing cell-free RNA is a panel of nucleic acids according to any one of claims 31 to 51.