Systems and methods for cell free RNA sequencing

By using targeted sequencing methods and RAG-based probe or primer sets to capture probes, the problem of detecting rare abundance cell-free RNA molecules has been solved, improving the diagnostic accuracy of liquid biopsy, and showing significant advantages, especially in cancer detection.

CN121532523APending Publication Date: 2026-02-13THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480044190.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-15
Filing Date
2024-05-15
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing technologies are insufficient for effectively detecting and analyzing rare and abundant cell-free RNA molecules, resulting in inaccurate accuracy of liquid biopsy in the diagnosis of diseases such as cancer.

Method used

Using targeted sequencing, a set of capture probes or primers is used to enhance the detection of cfRNA molecules that are not commonly expressed in liquid biopsies from healthy individuals by targeting rare and abundant cell-free RNA molecules (RAG).

Benefits of technology

It improves the sensitivity and accuracy of detecting cell-free RNA, enhances the diagnostic capabilities for medical diseases, pregnancy complications, cancer, etc., and overcomes the interference of submerged signals from hematopoietic cell cfRNA.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121532523A_ABST
    Figure CN121532523A_ABST
Patent Text Reader

Abstract

The present invention provides systems and methods for cell free RNA sequencing. Targeted sequencing can be performed using a set of gene transcripts that are rare abundance in a population of cell free nucleic acid control samples.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Application Serial No. 63 / 502,368, filed May 15, 2023, entitled “Systems and Methods for Cell-Free RNA Sequencing,” the disclosure of which is incorporated by reference in its entirety.

[0003] SEQUENCE LISTING

[0004] This application contains a Sequence Listing submitted electronically in XML format and is incorporated by reference in its entirety. The XML copy, created on May 14, 2024, is named 08574 Seq Listing.xml and is 4,382 bytes in size.

[0005] Statement Regarding Federally Sponsored Research or Development

[0006] This invention was made with U.S. Government support under contract CA254179 awarded by the National Institutes of Health. The U.S. Government has certain rights in the invention. TECHNICAL FIELD

[0007] The present disclosure provides a description of improved methods of sequencing cell-free RNA. BACKGROUND

[0008] Blood-based liquid biopsies enable non-invasive characterization of health conditions, including detection of biological phenomena such as pregnancy and cancer. Liquid biopsies offer many advantages over tissue biopsies as they are minimally invasive, easily repeated over time, and more accurately reflect geographic heterogeneity between cell sources. In patients with advanced cancer disease, analysis of circulating tumor DNA (ctDNA) is clinically used for non-invasive genotyping. However, comprehensive clinical evaluation often requires expression-based analysis, such as for distinguishing tumor types or subtypes. SUMMARY

[0009] In some embodiments, a method for cell-free RNA sequencing includes performing targeted sequencing. The sequencing targets cell-free RNA molecules that are not commonly expressed in a control individual’s liquid biopsy.

[0010] In some embodiments, the nucleic acid panel is used to target transcripts of cell-free RNA molecules of rare abundance.

[0011] In some embodiments, the nucleic acid panel comprises nucleic acid molecules having sequences from or complementary to a gene transcript that is a cell-free RNA molecule of rare abundance in a control liquid biopsy.

[0012] In some embodiments, transcripts of cell-free RNA molecules of rare abundance in the control liquid biopsy are defined as transcripts expressed in less than 50% of the control liquid biopsy population.

[0013] In some embodiments, transcripts of cell-free RNA molecules of rare abundance in the control liquid biopsy are defined as transcripts expressed in less than 5% of the control liquid biopsy population.

[0014] In some embodiments, transcripts of cell-free RNA molecules of rare abundance in the control liquid biopsy are defined as transcripts in the bottom 60% genes in terms of normalized expression in the control liquid biopsy population.

[0015] In some embodiments, transcripts of cell-free RNA molecules of rare abundance in the control liquid biopsy are defined as transcripts in the bottom 30% genes in terms of normalized expression in the control liquid biopsy population.

[0016] In some embodiments, transcripts of cell-free RNA molecules of rare abundance in the control liquid biopsy are defined as transcripts with log-transformed and normalized expression values less than zero in the control liquid biopsy population.

[0017] In some embodiments, the control liquid biopsy population comprises at least 5 liquid biopsies.

[0018] In some embodiments, the control liquid biopsy population comprises at least 50 liquid biopsies.

[0019] In some embodiments, the control liquid biopsy is collected from an individual who does not have one or more of the following at the time of collection of the biopsy: observed pathogen infection, diagnosed cancer, diagnosed metabolic disease, diagnosed nervous system disease, diagnosed immunodeficiency disease, diagnosed autoimmune disease, diagnosed inflammatory disease, diagnosed cardiovascular disease, diagnosed kidney disease, diagnosed liver disease, active pregnancy, diagnosed pregnancy complication, diagnosed fetal complication, organ transplant, organ transplant active rejection, obesity, malnutrition, cachexia, and abnormality in a clinical test.

[0020] In some embodiments, transcripts of cell-free RNA molecules of rare abundance in the control liquid biopsy are defined as transcripts expressed in less than 5% of the control liquid biopsy population.

[0021] In some embodiments, transcripts of cell-free RNA molecules of rare abundance in the control liquid biopsy are defined as transcripts in the bottom 60% genes in terms of normalized expression in the control liquid biopsy population.

[0022] In some embodiments, transcripts of cell-free RNA molecules of low abundance in the control liquid biopsy are defined as transcripts in the bottom 30% of genes in terms of normalized expression in the control liquid biopsy population.

[0023] In some embodiments, transcripts of cell-free RNA molecules of low abundance in the control liquid biopsy are defined as transcripts with log-transformed and normalized expression values less than zero in the control liquid biopsy population.

[0024] In some embodiments, the control liquid biopsy population comprises at least 5 liquid biopsies.

[0025] In some embodiments, the control liquid biopsy population comprises at least 50 liquid biopsies.

[0026] In some embodiments, the control liquid biopsy is collected from an individual who does not have one or more of the following at the time of collection of the biopsy: observed pathogen infection, diagnosed cancer, diagnosed metabolic disease, diagnosed nervous system disease, diagnosed immunodeficiency disease, diagnosed autoimmune disease, diagnosed inflammatory disease, diagnosed cardiovascular disease, diagnosed kidney disease, diagnosed liver disease, active pregnancy, diagnosed pregnancy complication, diagnosed fetal complication, organ transplant, organ transplant active rejection, obesity, malnutrition, cachexia, and abnormality in clinical test.

[0027] In some embodiments, transcripts of cell-free RNA molecules of low abundance in the control liquid biopsy comprise 50% of the transcripts in Table 3.

[0028] In some embodiments, transcripts of cell-free RNA molecules of low abundance in the control liquid biopsy comprise 90% of the transcripts in Table 3.

[0029] In some embodiments, transcripts of cell-free RNA molecules of low abundance in the control liquid biopsy comprise 100% of the transcripts in Table 3.

[0030] In some embodiments, the set of nucleic acid molecules excludes at least 50% of the transcripts of the whole exome genes that are not transcripts of cell-free RNA molecules of low abundance.

[0031] In some embodiments, the set of nucleic acid molecules excludes at least 90% of the transcripts of the whole exome genes that are not transcripts of cell-free RNA molecules of low abundance.

[0032] In some embodiments, the set of nucleic acid molecules consists of 5000 or fewer transcripts of genes plus transcripts of cell-free RNA molecules of low abundance.

[0033] In some embodiments, the set of nucleic acid molecules consists of 500 or fewer transcripts of gene transcripts plus transcripts of cell-free RNA molecules of rare abundance.

[0034] In some embodiments, the set of nucleic acid molecules further comprises tissue-specific transcripts, cell type-specific transcripts, clinically relevant transcripts, B and T cell receptor transcripts, biomarkers, and common mutagenic transcripts.

[0035] In some embodiments, the biomarker is associated with one of the following biological features: a medical disease, pregnancy, a fetal complication, a pregnancy complication, tumor growth, cancer, a specific cancer type, a pathogen infection, immune activation, organ transplant rejection, neurodegeneration, tissue of origin, cell type of origin, or activation of a biochemical pathway.

[0036] In some embodiments, the set of nucleic acid molecules further comprises a set of control transcripts for inter-sample normalization.

[0037] In some embodiments, the set of nucleic acid molecules is a set of probes for targeted capture hybridization.

[0038] In some embodiments, the set of nucleic acid molecules is a set of primers for targeted amplification.

[0039] In some embodiments, a method of preparing for cell-free RNA sequencing comprises providing a sample comprising nucleic acids for sequencing.

[0040] In some embodiments, the nucleic acids for sequencing are cell-free RNA or nucleic acids derived from and representative of cell-free RNA.

[0041] In some embodiments, a method of preparing for cell-free RNA sequencing comprises contacting the sample with a set of nucleic acid molecules comprising sequences from or complementary to transcripts of genes that are cell-free RNA molecules of rare abundance in a control liquid biopsy, such that the set of nucleic acid molecules anneals to a subset of the nucleic acids for sequencing.

[0042] In some embodiments, the cfRNA sample is derived from blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, or feces.

[0043] In some embodiments, a method of preparing for cell-free RNA sequencing comprises generating a sequencing library derived from the sample.

[0044] In some embodiments, a method of preparing for cell-free RNA sequencing comprises performing targeted sequencing on the sequencing library to produce sequencing results for cell-free RNA. The sequencing targets a set of nucleic acid molecules.

[0045] In some embodiments, a preparation method for cell-free RNA sequencing comprises removing platelet expression from sequencing results in silico.

[0046] In some embodiments, a preparation method for cell-free RNA sequencing comprises performing differential transcript analysis with the sequencing results and a second sequencing results.

[0047] In some embodiments, a preparation method for cell-free RNA sequencing comprises detecting enrichment of at least one expression signature in the sequencing results.

[0048] In some embodiments, a preparation method for cell-free RNA sequencing comprises detecting sequence mutagenesis in the sequencing results.

[0049] In some embodiments, a preparation method for cell-free RNA sequencing comprises inferring copy number states of one or more genes from the sequencing results.

[0050] In some embodiments, a preparation method for cell-free RNA sequencing comprises training a computational model with the sequencing results and a plurality of other sequencing results to predict a categorical state or likelihood of a biological signature. The cell-free RNA sample has a known categorical state of the biological signature.

[0051] In some embodiments, a preparation method for cell-free RNA sequencing comprises using the sequencing results as input to a trained computational model to predict a categorical state or likelihood of a biological signature. The computational model has been trained with a cohort of RNA sequencing results having a known categorical state of the biological signature.

[0052] In some embodiments, a preparation method for cell-free RNA sequencing comprises deriving one or more features from the sequencing results. The one or more features comprise enrichment of one or more gene signatures, enrichment of biochemical pathways, a collection of sequence variants, and copy number states.

[0053] In some embodiments, a preparation method for cell-free RNA sequencing comprises using the one or more derived features as input to a trained computational model to predict a categorical state or likelihood of a biological signature. The computational model has been trained with a cohort of RNA sequencing results having a known categorical state of the biological signature.

[0054] In some embodiments, a method of extracting RNA from a cell-free source comprises (a) adding glycogen to a sample comprising cell-free nucleic acids.

[0055] In some embodiments, a method of extracting RNA from a cell-free source includes (b) contacting a silica-based column with a sample comprising cell-free nucleic acids.

[0056] In some embodiments, step (a) is performed prior to step (b).

[0057] In some embodiments, a method of extracting RNA from a cell-free source includes eluting the cell-free nucleic acids from the silica-based column to produce an extracted cell-free nucleic acid solution.

[0058] In some embodiments, a method of extracting RNA from a cell-free source includes contacting the extracted cell-free nucleic acid solution with a DNAse.

[0059] In some embodiments, a method of quantifying cell-free RNA for downstream molecular applications includes providing a sample comprising cell-free RNA.

[0060] In some embodiments, a method of quantifying cell-free RNA for downstream molecular applications includes reverse transcribing the cell-free RNA to produce cDNA.

[0061] In some embodiments, a method of quantifying cell-free RNA for downstream molecular applications includes using quantitative real-time polymerase chain reaction and the cDNA to quantify the concentration of cell-free RNA in solution.

[0062] In some embodiments, a method of quantifying cell-free RNA for downstream molecular applications includes using the amount of material determined based on the quantification of the cell-free RNA to perform one or more downstream steps of a molecular protocol.

[0063] In some embodiments, the step of quantifying the concentration of cell-free RNA further includes generating a standard curve based on a set of control standards of known concentrations. The control standards are also evaluated using quantitative real-time polymerase chain reaction.

[0064] In some embodiments, the sample further comprises cell-free DNA. A method of quantifying cell-free RNA for downstream molecular applications further includes using quantitative real-time polymerase chain reaction to quantify the concentration of cell-free DNA in the sample. Cell-free RNA is quantified by using primers that span an intron of a gene that is relatively stable in cell-free RNA samples, and cell-free DNA is quantified by using primers that anneal to a genomic transcriptionally silent region that is relatively stable in cell-free DNA samples.

[0065] In some embodiments, the primers used to quantify cell-free RNA span an intron of GAPDH, and the primers used to quantify cell-free DNA target a 78 bp transcriptionally silent region on chromosome 12.

[0066] In some embodiments, a cell-free RNA sequencing method comprises providing a library of nucleic acid molecules. The library of nucleic acid molecules is derived from cell-free RNA. The cell-free RNA is derived from a liquid biopsy.

[0067] In some embodiments, a cell-free RNA sequencing method comprises sequencing the library of nucleic acid molecules to produce sequencing results.

[0068] In some embodiments, a cell-free RNA sequencing method comprises removing variation caused by platelet-associated transcript expression.

[0069] In some embodiments, the library of nucleic acid molecules is generated by capturing or amplifying nucleic acid molecules.

[0070] In some embodiments, the library of nucleic acid molecules is a whole exome library.

[0071] In some embodiments, the library of nucleic acid molecules is a library targeting rare abundance genes.

[0072] In some embodiments, a method of generating a targeted sequencing panel for cell-free RNA sequencing comprises collecting a control population of liquid biopsies, each comprising cell-free RNA.

[0073] In some embodiments, a method of generating a targeted sequencing panel for cell-free RNA sequencing comprises sequencing the cell-free RNA of the control population of liquid biopsies.

[0074] In some embodiments, a method of generating a targeted sequencing panel for cell-free RNA sequencing comprises identifying a set of rare abundance genes in the control population of liquid biopsies, the rare abundance genes defined by at least one or more of the following: their percentage of expression in the control population of liquid biopsies or their expression level in the control population of liquid biopsies.

[0075] In some embodiments, a method of generating a targeted sequencing panel for cell-free RNA sequencing comprises synthesizing a set of nucleic acid molecules for capturing or for amplifying the rare abundance genes to produce a targeted sequencing panel for cell-free RNA sequencing.

[0076] In some embodiments, the set of rare abundance genes is defined by at least their percentage of expression in the control population of liquid biopsies and their expression level in the control population of liquid biopsies. BRIEF DESCRIPTION OF DRAWINGS

[0077] This description and claims will be more fully understood in view of the following drawings and data, which are presented as examples of the present disclosure and should not be interpreted as an exhaustive statement of all aspects of the present disclosure.

[0078] Figures 1A to 1F Schematics and data plots are provided regarding optimization of blood collection and cfRNA extraction. Figure 1A Representative bioanalyzer traces of cell-free RNA (yellow) and white blood cell RNA (red). Figure 1B cfRNA concentration per mL of plasma in healthy controls measured by quantitative PCR (see Methods; n = 117). Plasma was collected for technical experiments, centrifuged at 2,500G for 10 minutes. Figure 1C Association between hemolysis and cfRNA plasma concentration. Hemolysis amount was measured using optical density (OD) at 414 nm. Pearson and Spearman correlations are shown. Figure 1D Association between time in -80°C freezer and cfRNA plasma concentration. Pearson and Spearman correlations are shown. Figure 1E Plasma cfRNA concentration was analyzed using different blood collection tubes (BCTs). All samples were spun at 2500G during plasma isolation. Ro, Roche Cell-Free DNA BCT. D, Streck Cell-Free DNA BCT. R, Streck RNAintegra BCT. E, EDTA. Figure 1F Plasma cfRNA concentration was analyzed using different extraction methods. T-R, TRIzol-LS + Qiagen RNeasy kit. V, QIAamp Viral RNA kit. Ro, Roche High Pure Viral RNA kit. M, Qiagen miRNeasy kit. MPS, Qiagen miRNeasy Serum / Plasma kit. T-M, TRIzol-LS + Qiagen miRNeasy kit. P, mirVana PARIS kit. CCF, QIAamp ccfDNA / RNA kit. T-V, TRIzol-LS + QIAamp Viral RNA kit. CNA, QIAamp Circulating Nucleic Acids kit.

[0079] Figures 2A to 2K Data plots are provided regarding optimization of RARE-Seq library preparation and capture. Figure 2A Correlation between cfRNA expression of blood samples collected in EDTA tubes or Streck RNAintegra tubes (n = mean of 3 pairs). Figure 2B Correlation of cfRNA expression extracted using TRIzol LS + QIAamp Viral RNA method (T-V) and QIAamp Circulating Nucleic Acids kit (CNA) (n = mean of 3 pairs). Figure 2CExpression correlation between cfRNA libraries generated using strand and non-strand methods (n=average of 3 pairs). Figure 2D Sparse analysis representing the relationship between total and unique sequencing depth for cfRNA libraries. Inset depicts the percent increase in unique sequencing depth for non-strand libraries relative to paired-strand libraries. Sequencing depth was down-sampled such that pairs had equivalent depth. Figure 2E Expression correlation between cfRNA libraries generated with and without S1 nuclease end repair (n=average of 3 pairs). Figure 2F Sparse analysis representing the relationship between total and unique sequencing depth for cfRNA libraries. Log2 NX, log2 normalized expression. Figure 2G Estimated DNA contamination if DNase I digestion was performed on- or off- Qiagen columns. DNA contamination was measured by the percent of reads aligning to exon boundaries containing adjacent intronic sequence. Conditions were compared using paired t-test. Figure 2H Expression correlation between cfRNA samples digested on- and off-column (n=average of 5 pairs). Figure 2I Distribution of RNA biotypes in cfRNA using SMART-Seq whole transcriptome method (n=3). Figure 2J Expression correlation between cfRNA libraries generated using SMART-Seq and whole coding transcriptome RARE-Seq (n=average of 3 pairs). Histograms depict the log2 nx distribution for each method and Venn diagrams represent the number of coding genes that are higher (or equivalent) in log2 NX in each method. Sequencing depth was down-sampled such that pairs had equivalent depth. Blue=coding genes, gray=non-coding genes. Figure 2K Heatmap showing pairwise Pearson correlation between all technical and biological replicates sequenced for individual individuals (n=8 replicates).

[0080] Figures 3A to 3L Provided are schematic and data showing enrichment of platelet and non-hematopoietic transcripts in cfRNA. Figure 3A Schematic of RARE-Seq method. Figure 3B Differential expression analysis comparing cfRNA and matched leukocyte RNA from healthy donors using whole coding transcriptome RARE-Seq (n=10 pairs). Figure 3C Gene set enrichment analysis using cell type specific signatures from PanglaoDB (n=170 cell types) pre-ordered. Top 15 positively and top 15 negatively enriched signatures are shown and dotted lines draw signatures with significant enrichment (P < 0.05). Figure 3DRelationship between centrifugation speed and cfRNA concentration per mL of plasma during plasma separation (n=10 1200G, n=6 1800G, n=16 2500G). Comparison was made using Kruskall-Wallis test. Figure 3E Relationship between centrifugation speed and platelet-specific gene expression in cfRNA. Avg log2NX, mean of normalized expression. Comparison was made using Kruskall-Wallis test. Figure 3F Principal component analysis of cfRNA expression clustering cfRNA samples according to spin speed (PC1) and sex (PC2). Figure 3G , Figure 3E Association between mean platelet-specific gene expression and Figure 3F first principal component in 2500G samples. Pearson and Spearman correlations are shown. Figure 3H Association between mean platelet-specific gene expression and first principal component, only considering samples processed using the same centrifugation speed (2500G). Figure 3I Schematic of platelet transcript correction method. See Methods for details. Figure 3J Association between platelet-specific gene expression and first principal component after platelet transcript correction of cfRNA expression. Figure 3K Hierarchical clustering of the top 1,000 genes with the most variable expression in cfRNA. Clustering was performed using Euclidean distance. Figure 3L Hierarchical clustering of the top 1,000 genes with the most variable expression after platelet-directed correction.

[0081] Figures 4A to 4I Provided are schematics and data showing that targeting transcripts missing in healthy cfRNA improves analytical sensitivity of lung cancer detection. Figure 4A Schematic depicting selection of rare abundance genes (RAGs) in plasma. Log2NX, log2 normalized expression. Figure 4B Proportion of tissue-enriched genes overlapping with RAGs and Figure 4C Proportion of cancer-enriched genes overlapping with RAGs. Tissue-enriched and cancer-enriched genes were selected as defined in Methods. Comparison was made using Fisher’s exact test. Figure 4D Unique depth in libraries captured using either the full coding transcriptome or RAG capture panel (n=5 matched libraries). Libraries were down-sampled to have equivalent sequencing depth and filtered to include regions captured by both panels. Comparison was made using paired t-test. Figure 4E , Figure 4D Number of shared genes detected in libraries. Detection threshold was >1 unique read. Again, comparison was made using paired t-test. Figure 4F Histogram depicting Figure 4D Differences in unique depth per gene in the library. Figure 4G Schematic of the Enrichment Score (ES) analysis framework (see Methods). Figure 4H Relationship between ES detection rate and NCI-H1975 proportion for spiked-in (n=3 replicates per spike-in) captured with RAG panels. NCI-H1975 ES was used to detect the presence of cancer RNA in each spike-in. Logistic regression was used to calculate the LOD95 from computer spiked-in (shown as gray circles). Results from in vitro spiked-in are shown as red triangles. Figure 4I Empirical LODs for NCI-H1975 spiked-in (n=1 replicate per spike-in) captured with RAG panels using the same approach as Figure 4F

[0082] Figures 5A to 5I Data plots are provided that describe and validate the RARE-Seq panels targeting rare abundance genes. Figure 5A Relationship between average gene expression in healthy cfRNA and the overall percentage of healthy cfRNA samples with low or absent gene expression (n=50). Rare abundance genes (RAGs) are shown in red, other genes in gray. Log2 NX, log2 normalized expression. Figure 5B Relationship between average gene expression in healthy cfRNA and expression stability in healthy cfRNA measured by Gini coefficient (n=50). Housekeeping genes added to the capture panel focusing on RAGs are shown in yellow, other genes in gray. Figure 5C Cumulative frequency of NSCLC patients with one or more variants in genes added to the capture panel focusing on RAGs due to recurrent alterations or frequent aberrant expression in non-small cell lung cancer (NSCLC) (n=122). Top 25 genes are shown. Variant data from TCGA Lung Adenocarcinoma and Lung Squamous Cell Carcinoma projects were downloaded from cBioPortal. Figure 5D Venn diagram depicting the set of genes included in the capture panel focusing on RAGs (n=5,546 genes in total). Figure 5E and 5F Expression correlation between the full coding transcriptome and RAG capture, Figure 5E All shared genes and Figure 5F Housekeeping genes (n=5 pairs). Figure 5G Schematic of the platelet transcript correction method using meta-reference controls with variable platelet expression. This method was developed for samples captured with panels focusing on RAGs that exclude platelet genes. Figure 5H Association between platelet-specific gene expression and the first principal component after platelet transcript correction of cfRNA expression as Figure 5G shown. Platelet transcript correction was applied using meta-reference controls with variable platelet expression.​ Figure 3D to 3K The same control cohort (n=32) of healthy controls. Figure 5I Association between NCI-H1975 spike-in detection and total sequencing depth. Logistic regression was used to estimate the LOD95 of computer spike-in.

[0083] Figures 6A to 6L Data plots are provided regarding ctRNA detection in lung adenocarcinoma. Figure 6A , 6B , 6C and 6D, Relationship between sample library metrics for healthy controls (n=24), low-dose computed tomography controls (LDCT, n=26), and lung adenocarcinoma (LUAD, n=50) patients. Box plots depict Figure 6A , cfRNA concentration (ng / mL plasma), Figure 6B , cfRNA quality used for library preparation, Figure 6C , total sequencing depth and Figure 6D , unique sequencing depth. Kruskal-Wallis test was used for each comparison. Figure 6E , cfRNA detection rate for each cfRNA sample unique sequencing depth. Figure 6F , Sensitivity of cfRNA detection using LUAD Sig ES and summarized by driver oncogene subtype. Detection threshold was used that achieved >95% specificity. Figure 6G and 6H , Representative Integrative Genome Browser (IGV) images showing sequencing reads containing Figure 6G , EGFR exon 19 deletion and Figure 6H , alternative splicing of MET exon 14. Figure 6I , In a representative sample of tissue-confirmed ROS1 / CD74 gene fusion, read counts mapping to the canonical splice junction and fusion breakpoint in the ROS1 and CD74 genes. Junctions / breakpoints with >5 aligned reads are shown. Figure 6J and 6K , Coefficients of the top 30 features selected by the LUAD elastic net (EN) model Figure 6J , Negative coefficients and Figure 6K , Positive coefficients of the top 30 features selected by the LUAD elastic net (EN) model. Figure 6L , Relationship between model training cohort size and 10-fold cross-validated model performance. Dashed line represents the 10-fold cross-validation (10CV) AUC of the final LUAD EN model and error bars represent the standard deviation of the 10CV AUC at each subsampling level (n=10 samples per level).

[0084] Figures 7A to 7F Data plots are provided showing detection of ctRNA in plasma of NSCLC patients. Figure 7ADifferential expression analysis of lung adenocarcinoma (LUAD) cfRNA (n=50) and non-cancer cfRNA (n=50). Figure 7B The average expression of genes enriched in LUAD cfRNA (n=94 genes) in LUAD tumor tissue (n=385) and meta-reference cfRNA (n=15). Figure 7C Using PanglaoDB 11 Enrichment analysis of pre-sorted gene sets for cell type-specific features. The top 15 positive and negative enriched features are shown, and features with significant enrichment are highlighted by dashed lines (P < 0.05). Figure 7D Enrichment fractions (ES) of LUAD-specific gene signatures (LUAD Sig; n=72 genes) in cfRNA from LUAD patients (n=50), risk-matched controls (n=26), and healthy controls (n=24). LDCT, low-dose computed tomography. Figure 7E The sensitivity of cfRNA detection was assessed using LUAD Sig ES and summarized by stage. A detection threshold achieving ≥95% specificity was used. Figure 7F The relationship between LUAD Sig ES in cfRNA and the mean variant allele frequency (VAF) in matched cfDNA (n=15). Comparisons were performed using Pearson and Spearman correlations. ND, not detected.

[0085] Figures 8A to 8G Schematic diagrams and data charts are provided for identifying somatic alterations in NSCLC ctRNA. Figure 8A A schematic diagram depicting the ctRNA variant invocation method (see Methods). Figure 8B Oncoprints of single nucleotide variants (SNVs), insertions / deletions (Indels), gene fusions, and splicing variants found in cfRNA were analyzed. Stage IV NSCLC samples with ≥1 tissue or ctDNA-determined variant were considered (n=55), and risk-matched LDCT controls (n=36). Healthy controls were used to define variant-specific error rates, as described in the methodology. Bar graphs depict the percentage of each mutation in each gene in the samples. LDCT, low-dose computed tomography. Figure 8C The percentage of samples with ≥1 variant detected in cfRNA. Figure 8D The percentage of tissue or ctDNA variants detected in cfRNA, summarized by variant type. Figure 8E For samples with tissue or ctDNA to determine EGFR variants, the relationship between EGFR expression levels in cfRNA and variant detection was investigated. Enrichment analysis was performed using fgsea. Figure 8FWeighted elastic net (EN) models computed the predicted probability of lung cancer for each sample and were summarized by stage. Results are shown for the training cohort (LUAD n=50, control n=50) and the held-out validation cohort (LUAD n=28 time points from 14 individuals, control n=10). Figure 8G Receiver operating characteristic (ROC) curves summarizing the performance of cfRNA detection using LUAD Sig ES and EN probabilities. AUCs were compared using the DeLong test. Val, validation. AUC, area under the ROC curve. ns, not significant.

[0086] Figures 9A to 9G A schematic and data plots are provided regarding detection of EGFR tyrosine kinase inhibitor resistance mechanisms in cfRNA. Figure 9A Summary of EGFR TKI cohort (n=10 patients, n=24 time points). Resistance mechanisms were determined by tissue biopsy. METamp, MET amplification. C797S, EGFR C797S SNV, tSCLC, histological transformation to small cell lung cancer. Figure 9B Enrichment score (ES) of the small cell lung cancer (SCLC) gene signature (SCLC Sig; n=73 genes) in patients with tSCLC as evidenced by biopsy. The threshold for >95% specificity in controls is shown by the dashed line. Figure 9C Patient profile of patients with tSCLC after EGFR TKI treatment. The left axis depicts the ES or EN score and the right axis depicts the AF of tumor-determining variants detected in cfRNA. Figure 9D Enrichment score (ES) of the MET amplification-specific gene signature (METamp Sig; n=9 genes) in patients with MET amplification as evidenced by biopsy. Figure 9E Patient profile of patients with MET amplification after EGFR TKI treatment. Figure 9F EGFR C797S AF in cfRNA of patients with C797S as evidenced by biopsy. Figure 9G Patient profile of patients with EGFR C797S mutation after EGFR TKI treatment.

[0087] Figures 10A to 10H Data plots are provided regarding RARE-Seq application. Figure 10AcfRNA detection sensitivity of LUAD Sig ES in LUAD cfRNA (n=50), LIHC Sig ES in hepatocellular carcinoma (LIHC) cfRNA (n=10), PAAD Sig ES in pancreatic adenocarcinoma (PAAD) cfRNA (n=10), and PRAD Sig ES in prostate adenocarcinoma (PRAD) cfRNA (n=9). Detection thresholds were used that achieved >90% specificity. Figure 10B Accuracy of tissue of origin (TOO) determination using cancer-T00 gene signatures. Bar plot shading depicts ES ranking of true cancer type. If detected by any cancer-specific signature, the sample was considered for TOO analysis. Figure 10C Confusion matrix comparing predicted cancer type to clinically diagnosed cancer type. Figure 10D Enrichment score (ES) of normal lung tissue signature (n=5 genes) in cfRNA of patients with benign lung conditions (n=74 time points from 70 individuals). COPD, chronic obstructive pulmonary disease. COVID, COVID-19 infection. ARDS, acute respiratory distress syndrome. Figure 10E Relationship between normal lung ES and patient smoking history. All samples without known lung conditions (n=42) in Figure 10D Figure 10F Relationship between normal lung ES and ventilator status. All samples with known lung conditions (n=32 time points from 28 individuals) in Figure 10D Figure 10G Timeline plot depicting the number of unique reads mapping to COVID-19 mRNA vaccine sequences in cfRNA collected at pre-vaccination and various time points after two doses of vaccination (n=9 time points). Log2 NX, log2 normalized expression. Figure 10H Gene ontology (GO) enrichment analysis of genes significantly differentially expressed at post-vaccination time points. Top 10 gene sets from MSigDb Hallmark and C5 collections are shown.

[0088] Figure 11 A summary of cancer and non-cancer cfRNA cohorts is provided. LDCT, low-dose computed tomography. ARDS, acute respiratory distress syndrome. COVID, COVID-19 infection. VACC, post-COVID-19 mRNA vaccination. LUAD, lung adenocarcinoma. EGFR TKI, treatment with epidermal growth factor receptor tyrosine kinase inhibitor. PAAD, pancreatic adenocarcinoma. PRAD, prostate adenocarcinoma. LIHC, hepatocellular carcinoma. DETAILED DESCRIPTION

[0089] ​​Turning now to the figures and data, systems and methods for sequencing cell-free RNA (cfRNA) molecules are described. The systems and methods can include tools for performing cfRNA targeted sequencing. The systems and methods can also include various steps and / or components for enhancing cfRNA processing and input.

[0090] When sequencing cfRNA from a cell-free source, analysis of nucleic acid from the source can be difficult due to the lack of high quality nucleic acid molecules. Historically, cfRNA analysis from cell-free sources has focused on microRNAs (miRNAs). However, other RNA types of expressed nucleic acid molecules (e.g., messenger RNAs) can greatly enhance detection of biological phenomena and / or medical conditions and thus can have great benefit in the diagnostic field. Unfortunately, liquid biopsies from plasma predominantly contain cfRNA molecules derived from hematopoietic cells, drowning out signals from other potential sources. Thus, due to low signal, performing cfRNA analysis for diagnostic purposes, such as cancer detection, is difficult. Thus, new systems and methods are needed to enhance detection of cell-free source cfRNA.

[0091] In this document, systems and methods relate to sequencing of cfRNA derived from a cell-free source (also referred to as a liquid biopsy or an excretion sample). Throughout the disclosure, the term liquid biopsy is used and refers to the source of cfRNA (including excretions), and includes (but is not limited to) blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, and fecal matter. The systems and methods can perform targeted sequencing to enhance detection of some cfRNA molecules of the liquid biopsy. In some embodiments, the targeted sequencing includes targeting cfRNA molecules that are not commonly found in the liquid biopsy of a control individual (e.g., a healthy individual). In some embodiments, a set of capture probes or a set of primers are utilized to target cfRNA molecules of interest for sequencing.

[0092] Various systems and methods relate to transcript panels and their use in cfRNA molecule targeted sequencing methods. Transcript panels can refer to capture-based panels (e.g., ssDNA molecules for hybrid capture) or amplification-based panels (e.g., a set of primers for amplification). Thus, transcript panels can be utilized for targeted sequencing during sequencing library preparation.

[0093] In some embodiments, the transcriptome comprises rare abundance genes (RAGs), which are transcripts (including non-coding transcripts) of cfRNA molecules that are not commonly expressed in a liquid biopsy of a healthy individual. Targeting RAGs can have various benefits, including overcoming the overwhelming signal provided by cfRNA originating from hematopoietic cells, and can enhance detection of various biological characteristics, including medical diseases, pregnancy, fetal complications, pregnancy complications, tumor growth, cancer, specific cancer types, pathogen infection, immune activation, organ transplant rejection, neurodegeneration, tissue of origin, cell type of origin, activation of biochemical pathways, etc.

[0094] In some embodiments, RAGs are defined as genes that express below a threshold in a control liquid biopsy. A population of individuals can be utilized to obtain control liquid biopsies for analysis. In some embodiments, the control liquid biopsy is a biopsy derived from an individual that is generally healthy (a control liquid biopsy can also be referred to as a healthy liquid biopsy). Generally healthy can mean an individual that does not have one or more of the following at the time the biopsy is collected: observed pathogen infection, diagnosed cancer, diagnosed metabolic disease, diagnosed neurological disease, diagnosed immune deficiency disease, diagnosed autoimmune disease, diagnosed inflammatory disease, diagnosed cardiovascular disease, diagnosed kidney disease, diagnosed liver disease, active pregnancy, diagnosed pregnancy complication, diagnosed fetal complication, has an organ transplant, organ transplant active rejection, obesity, malnutrition, cachexia, and has a clinical test abnormality. In some embodiments, cfRNA is collected to generate a diagnosis of a particular disease. In some embodiments where a diagnosis of a particular disease (or trait) is to be generated, the control liquid biopsy is a biopsy derived from an individual that is not diagnosed with the particular disease.

[0095] To identify RAGs, any suitable number of control liquid biopsies can be used to determine which genes express below a threshold for control liquid biopsies. In various examples, the number of control liquid biopsies used to identify RAGs is 5 or more control liquid biopsies, 10 or more control liquid biopsies, 15 or more control liquid biopsies, 20 or more control liquid biopsies, 50 or more control liquid biopsies, 100 or more control liquid biopsies, 200 or more control liquid biopsies, 500 or more control liquid biopsies, 1000 or more control liquid biopsies, 2000 or more control liquid biopsies, 5000 or more control liquid biopsies, 10,000 or more control liquid biopsies, 20,000 or more control liquid biopsies, 50,000 or more control liquid biopsies, or 100,000 or more control liquid biopsies.

[0096] Various definitions of RAGs can be used based on expression of genes in control liquid biopsies. In various embodiments, RAGs are defined as transcripts expressed in less than 50% of control liquid biopsies, RAGs are defined as transcripts expressed in less than 40% of control liquid biopsies, RAGs are defined as transcripts expressed in less than 30% of control liquid biopsies, RAGs are defined as transcripts expressed in less than 20% of control liquid biopsies, RAGs are defined as transcripts expressed in less than 10% of control liquid biopsies, RAGs are defined as transcripts expressed in less than 5% of control liquid biopsies, or RAGs are defined as transcripts expressed in less than 1% of control liquid biopsies. In some embodiments, clustering techniques are utilized to classify transcripts expressed in control liquid biopsies and transcripts not expressed in control liquid biopsies.

[0097] In some embodiments, RAGs are defined by their expression level in control liquid biopsies being below a threshold. In some embodiments, expression values are normalized for comparison. In some embodiments, expression values are log-transformed (e.g., Log2NX) for comparison. In some cases, RAGs are defined as transcripts with expression values Log2NX less than a threshold (e.g., Log2NX < 0). In some embodiments, RAGs are defined as transcripts with the lowest expression values in normalized expression in control liquid biopsies. In various embodiments, RAGs are defined as transcripts in the bottom 60% of genes in normalized expression, RAGs are defined as transcripts in the bottom 50% of genes in normalized expression, RAGs are defined as transcripts in the bottom 40% of genes in normalized expression, RAGs are defined as transcripts in the bottom 30% of genes in normalized expression, RAGs are defined as transcripts in the bottom 20% of genes in normalized expression, RAGs are defined as transcripts in the bottom 10% of genes in normalized expression, or RAGs are defined as transcripts in the bottom 5% of genes in normalized expression.

[0098] In some embodiments, RAGs are defined by being detected below a threshold of control liquid biopsies and / or expressing below a threshold in control liquid biopsies. Thus, any threshold defined herein to be present in a control biopsy can be combined with any threshold defined herein for expression levels. In one example, in the Examples and Data section below, RAGs are defined as transcripts that are expressed in less than 5% of control liquid biopsies and are in the lower 30% of transcripts in terms of normalized expression. Other definitions can be combined with RAGs for various applications, such as, for example, transcripts with tissue specificity, transcripts with cell type specificity, transcripts with known clinical relevance, transcripts with known biological relevance (e.g., biomarkers), and transcripts with known and common mutational profiles, such as fusion events and variants.

[0099] Table 3 provides an example of a list of RAGs defined as being identified in less than 5% of control liquid biopsies and having an average log2NX in the lower 30% of all genes, determined by 50 replicates from whole exome sequencing (WES) using cell-free RNA (cfRNA) of 28 healthy controls and 307 whole blood gene expression data samples from the Genotype-Tissue Expression (GTEx) project. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting a plurality of RAGs in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 1% of the RAGs listed in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 5% of the RAGs listed in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 10% of the RAGs listed in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 20% of the RAGs listed in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 30% of the RAGs listed in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 40% of the RAGs listed in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 50% of the RAGs listed in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 60% of the RAGs listed in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 70% of the RAGs listed in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 80% of the RAGs listed in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 90% of the RAGs listed in Table 3. In some embodiments, the transcript panel comprises nucleic acid molecules for detecting 100% of the RAGs listed in Table 3.

[0100] Transcriptomes can target nucleic acids using capture technology or amplification technology. To target cfRNA, in some embodiments, RNA is first converted to cDNA prior to targeting. And in some embodiments, cfRNA is targeted prior to cDNA conversion. Thus, transcriptomes can comprise nucleic acid molecules complementary to either strand of RNA or cDNA, such that they can anneal and / or hybridize to RNA or cDNA. In some embodiments, transcriptomes are utilized to specifically target particular molecules of RNA and / or double-stranded cDNA. In some embodiments, capture-based panels comprise a set of single-stranded nucleic acid probes for hybridization capture of particular molecules of RNA and / or double-stranded cDNA. In some embodiments, amplification-based panels comprise a set of primers for specific reverse transcription of particular RNA and / or specific amplification of particular molecules of double-stranded cDNA.

[0101] The sequences of the capture probes and / or primers are complementary to the transcripts to be targeted. The capture probes and / or primers need not be perfectly complementary, but have sufficient complementarity to their target to be captured by hybridization or annealing to initiate amplification. The design of the probes can be based on any appropriate sequence of the target. For example, probes or primers for human transcript targets can be designed using a reference database such as hgl9 or hg38.

[0102] Particular targeting is based on sequence complementarity and genes selected for the panel. In some embodiments, the transcriptome specifically targets a panel of Rarely Abundant Genes (RAGs). In some embodiments, the transcriptome excludes non-RAG (as defined by the standard or as listed in Table 3) genes. Excluding non-RAG genes allows for facilitating a simplified sequencing protocol and improved sequencing results (e.g., greater depth of targeted sequences compared to sequencing depth provided by a whole exome). Better sequencing results allow for better sensitivity, resulting in better outcomes in various applications such as cfRNA-based diagnostics.

[0103] In some embodiments, the transcriptome excludes at least 50% of non-RAG whole exome genes. In some embodiments, the transcriptome excludes at least 60% of non-RAG whole exome genes. In some embodiments, the transcriptome excludes at least 70% of non-RAG whole exome genes. In some embodiments, the transcriptome excludes at least 80% of non-RAG whole exome genes. In some embodiments, the transcriptome excludes at least 90% of non-RAG whole exome genes. In some embodiments, the transcriptome excludes at least 95% of non-RAG whole exome genes. In some embodiments, the transcriptome excludes at least 99% of non-RAG whole exome genes.

[0104] The sequencing protocol and its applications can be enhanced by including a set of control genes, which can provide a positive assurance for the sequencing results and facilitate the standardization across samples. In some embodiments, the transcript set comprises nucleic acid molecules for detecting a set of control transcripts to provide standardization across samples. Generally, the set of control transcripts can be any set of transcripts that are commonly detected in control liquid biopsies and have relatively stable expression levels. In some embodiments, the set of control transcripts comprises one or more housekeeping transcripts (e.g., GAPDH, actin, ubiquitin).

[0105] In some embodiments, the transcript set consists of 1 control gene. In some embodiments, the transcript set consists of 5 or fewer control genes. In some embodiments, the transcript set consists of 10 or fewer control genes. In some embodiments, the transcript set consists of 20 or fewer control genes. In some embodiments, the transcript set consists of 50 or fewer control genes. In some embodiments, the transcript set consists of 100 or fewer control genes. In some embodiments, the transcript set consists of 500 or fewer control genes. In some embodiments, the transcript set consists of 1000 or fewer control genes.

[0106] The sequencing protocol and its applications can be enhanced by evaluating genes that provide further insights. For example, if a cancer diagnosis is provided, certain transcripts can provide further diagnostic insights, such as genes that are repeatedly mutagenized in cancer. Thus, the transcript set can include additional non-RAG (as defined by the standard or as listed in Table 3) genes.

[0107] In some embodiments, the transcript set consists of 10 or fewer genes in addition to the RAGs. In some embodiments, the transcript set consists of 20 or fewer genes in addition to the RAGs. In some embodiments, the transcript set consists of 50 or fewer genes in addition to the RAGs. In some embodiments, the transcript set consists of 100 or fewer genes in addition to the RAGs. In some embodiments, the transcript set consists of 200 or fewer genes in addition to the RAGs. In some embodiments, the transcript set consists of 500 or fewer genes in addition to the RAGs. In some embodiments, the transcript set consists of 1000 or fewer genes in addition to the RAGs. In some embodiments, the transcript set consists of 2000 or fewer genes in addition to the RAGs. In some embodiments, the transcript set consists of 5000 or fewer genes in addition to the RAGs. In some embodiments, the transcript set consists of 10,000 or fewer genes in addition to the RAGs.

[0108] In addition to RAGs, various types of genes can also be included, such as, for example, tissue-specific transcripts, cell type-specific transcripts, clinically relevant transcripts, biomarker transcripts, common mutational transcripts, and B and T cell clonal transcripts. In some embodiments, the transcript panel targets a panel of tissue-specific transcripts. In some embodiments, the transcript panel targets a panel of cell type-specific transcripts. In some embodiments, the transcript panel targets a panel of clinically relevant transcripts. In some embodiments, the transcript panel targets a panel of transcripts that are biomarker transcripts. In some embodiments, the transcript panel targets a panel of mutational events. In some embodiments, the transcript panel targets a panel of B and T cell clones.

[0109] In some embodiments, the transcript panel comprises nucleic acid molecules for detecting tissue-specific transcripts. A tissue-specific transcript is a transcript that is uniquely highly expressed within a particular tissue compared to its expression in all other tissues. In various embodiments, the tissue-specific transcript is at least 2-fold more expressed in the particular tissue compared to all other tissues, the tissue-specific transcript is at least 3-fold more expressed in the particular tissue compared to all other tissues, the tissue-specific transcript is at least 4-fold more expressed in the particular tissue compared to all other tissues, the tissue-specific transcript is at least 5-fold more expressed in the particular tissue compared to all other tissues, the tissue-specific transcript is at least 6-fold more expressed in the particular tissue compared to all other tissues, the tissue-specific transcript is at least 7-fold more expressed in the particular tissue compared to all other tissues, the tissue-specific transcript is at least 8-fold more expressed in the particular tissue compared to all other tissues, the tissue-specific transcript is at least 9-fold more expressed in the particular tissue compared to all other tissues, or the tissue-specific transcript is at least 10-fold more expressed in the particular tissue compared to all other tissues. The unique expression can be determined empirically or be derived from database data, such as, for example, data derived from the Genotype-Tissue Expression (GTEx) project or the Human Protein Atlas (HPA).

[0110] In some embodiments, the transcriptome comprises nucleic acid molecules for detecting cell type-specific transcripts. Cell type-specific transcripts are transcripts that are uniquely highly expressed in a particular cell type compared to their expression in other cell types. In various embodiments, the expression of a cell type-specific transcript in a particular cell type is at least 2-fold compared to other cell types, the expression of a cell type-specific transcript in a particular cell type is at least 3-fold compared to other cell types, the expression of a cell type-specific transcript in a particular cell type is at least 4-fold compared to other cell types, the expression of a cell type-specific transcript in a particular cell type is at least 5-fold compared to other cell types, the expression of a cell type-specific transcript in a particular cell type is at least 6-fold compared to other cell types, the expression of a cell type-specific transcript in a particular cell type is at least 7-fold compared to other cell types, the expression of a cell type-specific transcript in a particular cell type is at least 8-fold compared to other cell types, the expression of a cell type-specific transcript in a particular cell type is at least 9-fold compared to other cell types, or the expression of a cell type-specific transcript in a particular cell type is at least 10-fold compared to other cell types. Unique expression can be determined empirically or be derived from database data, such as, for example, data derived from PanglaoDb.

[0111] In some embodiments, the transcriptome comprises nucleic acid molecules for detecting clinically relevant transcripts (e.g., for diagnostic relevance). Clinically relevant transcripts can include transcripts that are useful for diagnosing medical conditions. For example, transcripts of EGFR, KRAS, MET, ALK, RET, and ROS1 are useful in non-small cell lung cancer (NSCLC) diagnosis.

[0112] In some embodiments, the transcriptome comprises nucleic acid molecules for detecting transcripts associated with biological features (e.g., biomarker transcripts). Biological features include, but are not limited to, medical conditions, pregnancy, fetal complications, pregnancy complications, tumor growth, cancer, specific cancer types, pathogen infection, immune activation, organ transplant rejection, neurodegeneration, tissue of origin, cell type of origin, and activation of biochemical pathways.

[0113] In some embodiments, the transcriptome comprises nucleic acid molecules for detecting transcripts associated with mutagenic events, which can be useful for identifying de novo mutagenesis and / or cancer oncogenes. Mutagenic events include, but are not limited to, gene fusion, insertion, deletion, translocation, single nucleotide variation, and splicing variation.

[0114] In some embodiments, the transcriptome includes nucleic acid molecules for detecting transcripts associated with B cell and T cell clonal detection, which is useful for tracking immune activity against a particular antigen. The transcriptome is capable of targeting V(D)J recombination of B cell receptors and T cell receptors, and thereby identifying the sequence of B cell and T cell clones. In some embodiments, the clones are detected at a single time point to identify the current activity of an immunogen (e.g., an immunogen of a pathogen infection, cancer, vaccine, or autoimmune disease). In some embodiments, the clones are detected at multiple time points to detect changes in immunogen activity (e.g., detecting a trace amount of residual disease, success of cancer treatment, waning immunogenicity from vaccination or pathogen infection, onset of autoimmunity).

[0115] Several components can be utilized and several steps can be performed to sequence the cfRNA molecules obtained in a cell-free sample. The liquid biopsy can be derived from any appropriate biological source, including, but not limited to, blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, and stool. To extract RNA, any appropriate kit can be utilized (e.g., QIAamp Circulating Nucleic Acids Kit). In some embodiments, glycogen is added to the cell-free sample containing nucleic acids prior to addition to the silica-based column. Typically, the sample also contains cell-free DNA, which can be used for other assessments.

[0116] In some embodiments, the cell-free nucleic acids are quantified using quantitative real-time polymerase chain reaction (qRT-PCR). In some embodiments, RNA and DNA are simultaneously quantified in the cell-free sample. To simultaneously quantify RNA and DNA, qRT-PCR can be performed, where primers can use a set of primers that span an intron to detect and quantify RNA, and primers can use a set of primers that target a genomic non-transcribed region to detect and quantify DNA. For quantifying RNA, an intron of a common expression gene (e.g., a housekeeping gene) can be utilized. A standard curve can be generated using control standards to quantify nucleic acids. After quantification, the cell-free sample can be split into aliquots for DNA and RNA assessment. In some embodiments, DNA is digested (e.g., DNase digestion) from the cell-free sample for RNA assessment. In some embodiments, DNA digestion is performed after elution from the silica gel column for RNA assessment. In some embodiments, one or more downstream steps of the sequencing preparation protocol use a material amount that is determined based on the cell-free RNA quantification.

[0117] In some embodiments, the cell-free RNA is converted to cDNA. In some embodiments, double-stranded cDNA is generated. In some embodiments, the double-stranded cDNA is treated with a nuclease that removes single-stranded nucleic acid molecules (e.g., S1 endonuclease).

[0118] After a set of targets is captured and / or amplified, the library can be further processed (e.g., amplified) and sequenced using high-throughput sequencing. In some embodiments, the depth of sequencing is less than 10,000 reads, the depth of sequencing is more than 10,000 reads, the depth of sequencing is more than 100,000 reads, the depth of sequencing is more than 1,000,000 reads, the depth of sequencing is more than 10,000,000 reads, or the depth of sequencing is more than 100,000,000 reads.

[0119] In some embodiments, the depth of sequencing is less than 1X genome, the depth of sequencing is more than 1X genome, the depth of sequencing is more than 5X genome, the depth of sequencing is more than 10X genome, the depth of sequencing is more than 20X genome, the depth of sequencing is more than 30X genome, the depth of sequencing is more than 40X genome, the depth of sequencing is more than 50X genome, the depth of sequencing is more than 100X genome, the depth of sequencing is more than 150X genome, or the depth of sequencing is more than 200X genome.

[0120] After sequencing, the sequencing results can be subjected to various analyses to enhance detection of targeted cell-free nucleic acid molecules. In some embodiments, only transcripts included in the transcript set are included in downstream analyses. In some embodiments, transcript counts are converted to log-transformed counts per million (CPM) to account for library size and transcriptome complexity. In some embodiments, transcript counts are normalized to account for transcript size. In some embodiments, counts are normalized to account for variation between samples. In some embodiments, unwanted variation is removed, such as, for example, variation arising from platelet-associated transcript expression.

[0121] It has been found that platelet-derived RNA can confound the analysis of cfRNA. Platelets are found in almost all types of liquid biopsies, especially when there is cell damage. Thus, platelets can confound the assessment of liquid biopsies derived from, for example, blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, and stool for assessing many biological characteristics such as medical disease, pregnancy, fetal complications, pregnancy complications, tumor growth, cancer, specific cancer types, pathogen infection, immune activation, organ transplant rejection, neurodegeneration, tissue origin, cell type origin, and activation of biochemical pathways. In some embodiments, cfRNA samples are centrifuged to remove platelets. In some embodiments, platelet expression is removed in the computer after sequencing. To remove platelet expression in the computer, a program such as the RUVseq R package is utilized to remove unwanted variation. In some embodiments, the transcript set used for targeted sequencing specifically excludes platelet-associated gene expression. Even in cases where the transcript set used for targeted sequencing specifically excludes platelet expression, control cfRNA can be used to estimate unwanted variation factors related to platelet expression to establish a correction factor related to platelet expression. In one example, platelet correction can be performed by ordinary least squares regression of log2NX on selected unwanted platelet expression factors.

[0122] In some embodiments, differential transcript expression analysis is performed using the sequencing results. In some embodiments, differential transcript expression analysis can be used for comparison of samples (e.g., medical disease vs. control; e.g., comparison from one time point to another). Any appropriate method can be utilized for performing differential transcript expression analysis (e.g., the DESeq2 R package). In one example, a generalized linear model can be constructed using sample type as a covariate. Appropriate cutoffs and significance (e.g., greater than 1 log-fold and adjusted p-value < 0.05) can be utilized to identify significantly differentially expressed genes. Further downstream analysis can be performed on differentially expressed genes to gain insight into biological phenomena related to the samples. For example, gene set enrichment analysis can be performed, which can be used to identify molecular signatures related to various phenomena.

[0123] In some embodiments, the sequencing results are utilized to detect enrichment of expression signatures associated with a biological characteristic. In general, a number of biological characteristics can be identified through expression signatures in the sequencing results. Examples of biological characteristics include, but are not limited to, medical diseases, pregnancy, fetal complications, pregnancy complications, tumor growth, cancer, specific cancer types, pathogen infection, immune activation, organ transplant rejection, neurodegeneration, tissue origin, cell type origin, and activation of biochemical pathways. In some embodiments, an enrichment score of an expression signature is computed, providing a scaled assessment of whether a biological characteristic is present in a sample. Diagnoses can be established based on the various biological characteristics present in the sequencing results. For example, liquid biopsies can be utilized to assess the health of a pregnancy or fetus by evaluating expression signatures in cfRNA associated with various health states and / or complications.

[0124] In some embodiments, the sequencing results are utilized to detect mutagenesis, including somatic or de novo mutagenic events. Examples of mutagenic events that can be detected include, but are not limited to, gene fusions, insertions, deletions, translocations, single nucleotide variants, and splicing variants. In some embodiments, the sequencing results are utilized to infer copy number states of one or more genes. In one example, MET gene amplification can be inferred from expression signatures, which is associated with targeted therapy resistance in non-small cell lung cancer. As previously described, mutagenic assessments can be diagnostic and informative for treatment selection, particularly for various cancer types.

[0125] In some embodiments, a computational model is trained to predict a categorical state or likelihood of a biological characteristic present in a cfRNA sample based on sequencing results. In some embodiments, the computational model is trained directly with sequencing results as model inputs. In some embodiments, the computational model utilizes derived features, such as, for example, standardized expression of individual genes, enrichment of one or more gene signatures, enrichment of biochemical pathways, collections of sequence variants, and copy number states. The trained computational model can then be used to evaluate a patient’s cfRNA sequencing results, such as its use as a diagnostic in a clinical setting.

[0126] For predicting biological characteristics, any appropriate computational model can be utilized. Examples of computational models include, but are not limited to, logistic regression, elastic net, LASSO, random forest, XGBoost, and neural networks. Training can be performed using expression data from samples with biological characteristics and control samples. A cohort of individuals with known categorical states of biological phenomena, whose cfRNA is sequenced and processed to train the computational model. Alternatively, because cfRNA samples are based on expression within a particular cell type, samples can be from solid tissue or any other source representative of cfRNA. After training, a feature set, weights, and hyperparameters can be selected that provide robust predictability.

[0127] Sequencing methods and diagnostics

[0128] Various methods involve sequencing cfRNA. Moreover, cfRNA sequencing methods can be used for a number of diagnostic assessments. Thus, an individual can have a liquid biopsy extracted or collected for cfRNA sequencing, which can be prepared with a RAG transcriptome. Various biomedical features can be screened for and / or diagnosed by cfRNA sequencing. Based on the diagnostic assessment, the individual can be further assessed by a clinical evaluation. The diagnostic assessment can also provide information for treatment selection, and thus, in some cases, treatment can be performed by a medical professional, such as a physician, nurse, nutritionist, or similar personnel. The cfRNA sequencing can also be used to assess treatment response, treatment outcome, and / or presence of minimal residual disease.

[0129] A number of sequencing methods can be performed to assess cfRNA. Generally, the sequencing method collects a sample containing cfRNA and prepares it for targeted sequencing with a RAG panel.

[0130] Examples of methods for cfRNA sequencing can include:

[0131] • extracting or collecting a cfRNA sample

[0132] • enriching the cfRNA sample with a targeted RAG panel

[0133] • sequencing the enriched cfRNA sample to produce sequencing results

[0134] In some embodiments, the cfRNA sample is (or is derived from) a liquid biopsy. In some embodiments, the cfRNA sample is (or is derived from) an excretion sample. In some embodiments, the cfRNA sample is (or is derived from) blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, or feces.

[0135] In some embodiments, the RAG panel comprises a set of genes that are expressed below a certain percentage in a control liquid biopsy population. In some embodiments, the RAG panel comprises a set of genes that have expression levels below a threshold in a control liquid biopsy population. In some embodiments, the RAG panel comprises a set of genes that are expressed below a certain percentage and have expression levels below a threshold in a control liquid biopsy population. In some embodiments, the RAG panel comprises a set of genes in Table 3. In some embodiments, the RAG panel comprises nucleic acid molecules for amplifying RAGs. In some embodiments, the RAG panel comprises nucleic acid molecules for capturing RAGs by hybridization.

[0136] In some embodiments, the method for cfRNA sequencing further comprises:

[0137] • collecting a population of control cfRNA samples

[0138] • sequencing each cfRNA sample to produce sequencing results

[0139] • identifying RAGs from the population of sequencing results

[0140] In some embodiments, the control cfRNA sample is extracted or collected from a healthy individual. In some embodiments, RAGs are identified by being expressed below a certain percentage in the control liquid biopsy population. In some embodiments, RAGs are identified by being expressed below a threshold level in the control liquid biopsy population. In some embodiments, RAGs are identified by being expressed below a certain percentage in the control liquid biopsy population and by being expressed below a threshold level in the control liquid biopsy population.

[0141] The sequencing of the cfRNA sample can provide information about the biological characteristics of the individual from which the cfRNA sample was extracted or collected. The diagnostic method can be used for various purposes, such as, for example, screening an individual for biological characteristics or biomedical complications, diagnosing a particular medical disease, evaluating treatment response, evaluating treatment outcome, and screening for minimal residual disease. The diagnostic method can be used to evaluate many biomedical conditions, such as, for example, medical diseases, pregnancy, fetal complications, pregnancy complications, tumor growth, cancer, specific cancer types, pathogen infection, immune activation, organ transplant rejection, neurodegeneration, tissue of origin, cell type of origin, and activation of biochemical pathways.

[0142] In some embodiments, the screening method or diagnosis comprises:

[0143] • extracting or collecting a cfRNA sample

[0144] • enriching the cfRNA sample with a set of target RAGs

[0145] • sequencing the enriched cfRNA sample to produce sequencing results

[0146] • identifying one or more biomedical conditions in the sequencing results

[0147] In some embodiments, the cfRNA sample is (or is derived from) blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, or stool. In some embodiments, the screening method is general and evaluates multiple biomedical conditions. In some embodiments, the diagnostic method is specific to a biomedical condition. In some embodiments, the biomedical condition is identified by a gene expression signature. In some embodiments, the biomedical condition is identified by a computational model trained to predict the biomedical condition using the sequencing results or features derived from the sequencing results.

[0148] In some embodiments, the biomedical condition is pregnancy. In some embodiments, the biomedical condition is a fetal complication. In some embodiments, the biomedical condition is a pregnancy complication. In assessing a fetal complication or a pregnancy complication, in some embodiments, the cfRNA sample is derived from amniotic fluid. In assessing a fetal complication or a pregnancy complication, in some embodiments, the screening method is performed at various time points throughout pregnancy. In some embodiments, when a fetal complication or a pregnancy complication is identified, a further diagnostic procedure is performed, such as, for example, a gestational diabetes assessment, a fetal genetic assessment, a fetal ultrasound, and a maternal blood test. In some embodiments, when a fetal complication or a pregnancy complication is identified, a treatment is performed, such as, for example, induction of labor, administration of anti-labor drugs, and performance of a cesarean section.

[0149] In some embodiments, the biomedical condition is a pathogen infection. In some embodiments, the biomedical condition is an immune status, such as, for example, a vaccination status or a prior pathogen infection. In some embodiments, in assessing a pathogen infection, the cfRNA sample is also enriched for pathogen sequences. In some embodiments, when a pathogen infection is identified, a treatment is performed, such as administration of an anti-pathogen drug (e.g., an antibiotic agent, an antiviral agent, an antiparasitic agent, etc.). In some embodiments, the screening method monitors a pathogen infection, an anti-pathogen treatment response, an immune response, a health status, or any combination thereof.

[0150] In some embodiments, the biomedical condition is immune activation, such as, for example, reactive activation to a pathogen, reactive activation to an immunization, or activation of an autoimmune disease. In some embodiments, the biomedical condition is inflammation. In some embodiments, when the biomedical condition is an autoimmune disease or inflammation, a treatment is performed, such as, for example, administration of an immunosuppressant and administration of an anti-inflammatory agent.

[0151] In some embodiments, the biomedical condition is organ transplant rejection. In some embodiments, the organ transplant rejection is identified by a gene expression signature associated with cytotoxicity, a gene expression signature associated with tissue of origin, a gene expression signature associated with cell type of origin, a gene sequence that distinguishes between a donor and a host, or a combination thereof. In some embodiments, the screening method is performed periodically after a host receives a transplant. In some embodiments, when organ transplant rejection is identified, a further diagnostic procedure is performed, such as, for example, a tissue biopsy of the organ and medical imaging of the organ. In some embodiments, when organ transplant rejection is identified, a treatment is performed, such as, for example, administration of an increased dosage of an immunosuppressant and administration of a stronger immunosuppressant.

[0152] In some embodiments, the biomedical condition is neurodegeneration. In assessing neurodegeneration, in some embodiments, the cfRNA sample is (or is derived from) cerebrospinal fluid. In some embodiments, neurodegeneration is identified by a gene signature associated with neurodegenerative disease, a gene signature associated with neural tissue origin, a gene signature associated with neural cell type origin, a gene signature associated with inflammation, or a combination thereof. In some embodiments, when neurodegeneration is identified, a further diagnostic procedure is performed, such as, for example, medical screening, assessment of motor activity or speech, and cognitive assessment. In some embodiments, when neurodegeneration is identified, a treatment is performed, such as, for example, a drug to reduce symptoms of neurodegeneration.

[0153] In some embodiments, the biomedical condition is cancer. In some embodiments, the screening method is performed as part of cancer monitoring efforts (e.g., prior to the presence or identification of cancer symptoms). In some embodiments, the screening method is performed during treatment to assess treatment response. In some embodiments, the screening method is performed after treatment to assess for the presence of residual cancer (e.g., MRD) after treatment, which can be performed periodically. In some embodiments, the diagnostic method provides information for cancer subtype, cancer stage, and / or treatment strategy.

[0154] Screening can be performed for many tumor types, including, but not limited to, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), anal cancer, astrocytoma, basal cell carcinoma, biliary cancer, bladder cancer, breast cancer, Burkitt lymphoma, cervical cancer, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelopro liferative neoplasms, colorectal cancer, diffuse large B-cell lymphoma, endometrial cancer, ependymoma, esophageal cancer, esthesioneuroblastoma, Ewing sarcoma, fallopian tube cancer, follicular lymphoma, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, hairy cell leukemia, hepatocellular cancer, Hodgkin lymphoma, hypopharyngeal cancer, Kaposi sarcoma, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, melanoma, Merkel cell carcinoma, mesothelioma, oral cancer, neuroblastoma, non-Hodgkin lymphoma, non-small cell lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic neuroendocrine tumor, pharyngeal cancer, pituitary tumor, prostate cancer, rectal cancer, renal cell cancer, retinoblastoma, skin cancer, small cell lung cancer, small intestine cancer, squamous cell carcinoma of the neck, T-cell lymphoma, testicular cancer, thymoma, thyroid cancer, uterine cancer, upper tract urothelial carcinoma, vaginal cancer, and vascular tumors. In some embodiments, when the cancer to be assessed is colorectal cancer or gastric cancer, the cfRNA sample is (or is derived from) a fecal sample. In some embodiments, when the cancer to be assessed is bladder cancer, kidney cancer, prostate cancer, or upper tract urothelial cancer, the cfRNA sample is (or is derived from) a urine sample.

[0155] In some embodiments, when a cancer is indicated, a number of follow-up clinical evaluations can be performed, including, but not limited to, physical examination, medical imaging, mammography, endoscopy, stool sampling, pap test, alpha-fetoprotein blood test, CA-125 test, prostate-specific antigen (PSA) test, biopsy extraction, bone marrow aspiration, and tumor marker detection test. Medical imaging includes, but is not limited to, X-ray, magnetic resonance imaging (MRI), computed tomography (CT), ultrasound, and positron emission tomography (PET). Endoscopy includes, but is not limited to, bronchoscopy, colonoscopy, colposcopy, cystoscopy, esophagoscopy, gastroscopy, laparoscopy, neuroendoscopy, proctoscopy, and sigmoidoscopy.

[0156] In some implementations, when cancer is indicated, a variety of treatments can be performed, including (but not limited to) surgery, chemotherapy, radiation therapy, immunotherapy, targeted therapy, hormone therapy, stem cell transplantation, and blood transfusion. In some implementations, anticancer and / or chemotherapeutic agents are administered, including (but not limited to) alkylating agents, platinum-based agents, taxanes, vinca alkaloids, anti-estrogens, aromatase inhibitors, ovarian suppressants, endocrine / hormone agents, bisphosphonate agents, and targeted biological agents. Medications include (but are not limited to) cyclophosphamide, fluorouracil (or 5-fluorouracil or 5-FU), methotrexate, thiotepa, carboplatin, cisplatin, taxanes, paclitaxel, protein-bound paclitaxel, docetaxel, vinorelbine, tamoxifen, raloxifene, toremifene, fulvestrant, gemcitabine, irinotecan, ixabepilone, temozolomide, topotecan, vincristine, vinblastine, eribulin, mutamycin, capecitabine, anastrozole, exemestane. emestane, letrozole, leuprolide, abarelix, buserlin, goserelin, megestrol acetate, risedronate, pamidronate, ibandronate, alendronate, zoledronate, tykerb, daunorubicin, doxorubicin, epirubicin, idarubicin, valrubicin, mitoxantrone, bevacizumab, cetuximab, ipilimumab, ado-trastuzumabemtansine), afatinib, aldesleukin, alectinib, alemtuzumab, atezolizumab, avelumab, axitinib, belimumab, belinostat, bevacizumab, blinatumomab, bortezomib, bosutinib, brentuximab vedotin, briatinib, cabozantinib, canakinumab, carfilzomib, ceritinib, cetuximab, cobimetinib, crizotinib, dabrafenib, daratumumab, dasatinib, denosumab, dinutuximab, durvalumab, elotuzumab, enasidenib, erlotinib, everolimus, gefitinib, ibritumomabtiuxetan), ibrutinib, idelalisib, imatinib, ipilimumab, ixazomib, lapatinib, lenvatinib, midostaurin, necitumumab, neratinib, nilotinib, niraparib, nivolumab, obinutuzumab, ofatumumab, olaparib, loaratumab, o simertinib, palbociclib, panitumumab, panobinostat, pembrolizumab, pertuzumab, ponatinib, ramucirumab, regorafenib, ribociclib, rituximab, romidepsin, rucaparib, ruxolitinib, siltuximab, sipuleucel-T, sonidebib, sorafenib, temsirolimus, tocilizumab, tofacitinib, tositumomab, trametinib, trastuzumab, vandetanib, vemurafenib, venetoclax, vismodegib, vorinostat, and ziv-aflibercept. The individual can be treated with a single drug or a combination of drugs as described herein. A common treatment combination is cyclophosphamide, methotrexate, and 5-fluorouracil (CMF).

[0157] Examples and supporting data

[0158] The systems and methods of the present disclosure will be better understood through a number of embodiments provided. Validation results demonstrate that the use of a targeted panel of rare-abundance gene-based sequencing greatly improves cancer detection using cfRNA samples. By using the RAG panel, the detection of expression signatures is improved 50-fold over whole transcriptome sequencing. The method is also better at detecting resistance to targeted therapies compared to cfDNA methods. The method also allows for assessment of tissue of origin, which can aid in identifying the primary site of pathology. The results can be generalized to other cell-free diagnostic technologies, including diagnosis of medical conditions, pregnancy, fetal complications, pregnancy complications, tumor growth, cancer, specific cancer types, pathogen infection, immune activation, organ transplant rejection, neurodegeneration, tissue of origin, cell type of origin, and activation of biochemical pathways.

[0159] Cell-free RNA analysis for non-invasive cancer detection and characterization

[0160] Described herein are systems and methods for cfRNA sequencing analysis, termed RARE-Seq (Random Primer & Affinity Capture of Cell-free RNA Fragments for Sequencing Enrichment Analysis). The method improves the limit of detection of circulating tumor RNA (ctRNA) relative to previous methods while maintaining high specificity for tumor-naive cancer detection. RARE-Seq has great utility in a variety of clinical applications, including non-invasive ctRNA genotyping, treatment resistance monitoring, tissue of origin analysis, and molecular characterization of non-malignant conditions.

[0161] Results

[0162] Pre-analytical factors affecting cfRNA recovery and analysis

[0163] Cell-free RNA is highly fragmented and does not contain detectable 18S and 28S rRNA peaks ( Figure 1A ), indicating high levels of degradation. In blood samples from healthy donors with no history of cancer, the median concentration of cfRNA was 220 pg per mL of plasma, representing approximately 20 cells (n = 117, Figure 1B ). To improve recovery, factors affecting cfRNA recovery and downstream expression analysis were evaluated ( Figures 1C to 1F ). Specifically, analysis of multiple blood collection and RNA extraction protocols that allow for optimal RNA recovery from plasma led to the discovery that factors such as hemolysis and time of frozen storage were minimally or not associated with cfRNA yield ( Figure 1C and 1D ). Multiple steps in the assay workflow were optimized, including removal of contaminating DNA, complementary DNA (cDNA) synthesis, and end-repair. These optimizations improved library preparation efficiency with low cfRNA input and minimized the impact of contaminating cfDNA ( Figures 2A to 2K ). Further validation of optimization led to reproducible cfRNA expression profiles in healthy controls. Optimization of the cfRNA method could be used in the RARE-Seq (Random Primer & Affinity Capture of Cell-free RNA Fragments for Sequencing Enrichment) protocol ( Figure 3A ).

[0164] Platelets and non-hematopoietic cell types are enriched in cfRNA

[0165] By applying RARE-Seq to cfRNA and matched leukocyte RNA from 10 healthy adults, expression differences between cell RNA and cell-free RNA were characterized. Differential expression analysis identified 8,336 significant genes, including 5,271 genes overexpressed in cfRNA and 3,065 genes overexpressed in leukocyte RNA ( Figure 3B ). Using gene set enrichment analysis (GSEA) and cell type-specific gene sets from single-cell RNA sequencing atlases, cfRNA was found to be significantly enriched for non-hematopoietic cell types, including endothelial cells, hepatocytes, fibroblasts, and multiple neuronal cell types (n = 71 total) ( Figure 3C ). In contrast, leukocyte RNA was significantly enriched for 12 major myeloid and lymphoid hematopoietic cell types, including T cells, NK cells, neutrophils, and monocytes. Only two hematopoietic cell types were enriched in cfRNA: erythroid precursor cells and platelets, indicating their important contribution to the cfRNA pool during erythropoiesis and megakaryopoiesis.

[0166] Given the significant enrichment of platelet transcripts in cfRNA ( Figure 3C ), it was suspected that residual platelets not removed during plasma isolation could be contributing RNA to cfRNA samples. To test this idea, cfRNA concentrations were compared between protocols using three different centrifugation speeds to isolate the cellular and cell-free plasma fractions (n = 32). Median concentrations were ~7.5-fold higher at 1200xG than at 2500xG (1.72 ng / mL vs 0.23 ng / mL; P = 0.000052; Figure 3D ). Furthermore, the average expression of platelet-specific genes (n = 130 11 ) was significantly correlated with centrifugation speed ( Figure 3E ). Principal component analysis (PCA) showed that the first principal component (PC1) captured 70% of the expression variation and was strongly correlated with platelet expression (R = -0.99, P = < 2.2e-16) ( Figure 3F and 3G). Similar results were found even in samples processed using a fixed centrifugation speed, indicating that differences in platelet counts can also lead to variation in platelet transcripts in cfRNA samples (R = 0.98, P = < 1.3e-11) ( Figure 3H ). Overall, these results indicate that platelets represent the largest source of expression variation in cfRNA and that platelet contamination is a key pre-analytical variable to control in cfRNA analysis.

[0167] While individual platelets contain lower RNA content than leukocytes, the number of platelets in blood is about two orders of magnitude greater than leukocytes. Furthermore, the physiological variation in healthy adult platelet counts ranges about three-fold, with further variability observed in the cancer context. Thus, even an effective method to minimize platelet contamination in plasma can be difficult to effectively reduce the level of platelet-associated RNA. Additionally, from a practical perspective, standardizing blood bank protocols is often not feasible, especially when analyzing previously collected samples. Therefore, it was investigated whether the expression contribution from platelet contamination could be removed by an algorithm. To this end, a computational method was developed to remove unwanted gene expression driven by different platelet contamination levels ( Figure 3I See also Figure 5G ). This method successfully removed unwanted variation from platelets, as shown by the lack of correlation between platelet expression and PC1 after correction (R = 0.07, P = 0.69) ( Figure 3J See also Figure 5H ). Hierarchical clustering further confirmed this effect, as samples clustered according to platelet gene expression before correction ( Figure 3K ), while clustering according to individual-to-individual expression differences after correction ( Figure 3L ). Thus, addressing platelet contamination in cfRNA analysis is crucial to reliably reveal gene expression contributions that are not of platelet origin.

[0168] Maximizing sensitivity of ctRNA detection

[0169] While cfRNA is highly enriched for gene expression of non-hematopoietic origin, the absolute expression levels of these genes are relatively low compared to hematopoietic-derived transcripts (e.g., globins) and highly expressed housekeeping genes (e.g., beta-actin). Therefore, it was hypothesized that selective capture of genes that are missing or lowly expressed in healthy control cfRNA can enhance the detection of non-hematopoietic tissue or disease signatures and extend the utility of RARE-Seq to a broad range of clinical applications. To test this, ‘rare abundance genes’ (RAGs) were identified using whole coding transcriptome RARE-seq data from healthy cfRNA samples (n = 50) and RNA-seq data from healthy whole blood samples from the GTEx project (n = 307) ( Figure 4A). RAGs were defined based on consistently low expression or absence of expression in both cfRNA and whole blood ( Figure 5A ; Table 3). Notably, RAGs were enriched in non-hematopoietic tissue-specific genes and tumor-specific genes from multiple cancers compared to all protein-coding transcripts ( Figure 4B and 4C ).

[0170] Given that RAGs are enriched in genes that can reflect tissue damage or presence of cancer, it was examined whether targeted sequencing of RAGs can improve the detection sensitivity of these transcripts. Therefore, a capture panel targeting 4,323 RAGs was designed. The panel was supplemented with 50 housekeeping genes prioritizing genes that express relatively uniformly in healthy cfRNA and whole blood samples while covering a broad dynamic expression range ( Figure 5B ). Genes known to be recurrently mutated in lung cancer (n = 123 genes) were also added, including genes such as TP53, EGFR, KRAS, ALK, RET, and ROS1 ( Figure 5C ). The complete gene set targeted by our RAG capture panel is summarized in Figure 5D .

[0171] Five healthy cfRNA samples were analyzed using RARE-Seq with both the whole coding transcriptome and the RAG capture panel. Gene expression was highly correlated between matched libraries (R = 0.98, P = < 2.2e-16) ( Figure 5E ), including the housekeeping gene set (R = 0.97, P = < 2.2e-16) ( Figure 5F ). In the gene set shared between the two capture panels, RAG capture resulted in significantly higher unique sequencing depth ( Figure 4D ), detection of more genes ( Figure 4E ), and higher unique sequencing depth per detected gene ( Figure 4F ). Thus, targeted capture of RAGs effectively improved the detection sensitivity of lowly expressed genes in healthy cfRNA.

[0172] The 95% limit of detection (LOD95) for the transcriptional signatures of interest was determined using either the whole coding transcriptome or the RAG targeted panel. To score the presence of a gene signature of interest in a cfRNA sample, an analytical framework was developed that computes a gene signature enrichment score (ES) by comparing the expression of the pre-defined signature genes in the sample of interest to a “meta-reference” composed of healthy control cfRNA samples and estimating its significance using bootstrapping ( Figure 4GA key advantage of this method is that, unlike machine learning-based approaches, it does not require a large training cohort of patient cfRNA samples. This is because the ES strategy relies on existing gene expression datasets, defining gene features and associated gene weights based on their expression levels in the tissues or conditions of interest.

[0173] To determine the LOD95 of RARE-Seq, RNA generated from the NCI-H1975 NSCLC cell line was serially diluted to healthy control cfRNA at various cancer proportions in a computer. To enable the use of the ES framework, NCI-H1975-specific gene signatures were generated, including the top differentially expressed genes (H1975 Sig; n=122 genes) that distinguish NCI-H1975 from the meta-reference healthy cfRNA profile. Using this signature, the LOD95 of RARE-Seq was 0.05% ( Figure 4H ), compared to the full-coding transcriptome RARE-Seq (2.8%; Figure 4I The sensitivity was >50 times greater, demonstrating the value of targeting RAG. To experimentally confirm this result, in vitro dilution experiments were performed by spiking NCI-H1975 RNA at a specified concentration into healthy donor cfRNA. These in vitro spikes showed strong agreement with the computer-generated results. Figure 4H and 4I The effect of sequencing depth (ranging from 10 million to 50 million read pairs) on the LOD95 of RAG-targeted RARE-Seq was evaluated, and it was found that LOD95 increases with sequencing depth. Figure 5I Therefore, in subsequent experiments using the RARE-Seq and RAG groups, 50 million read pairs were targeted.

[0174] Detection of non-small cell lung cancer

[0175] Following the development of RARE-Seq, the remainder of the research focused on generating proof-of-concept data for the potential utility of cfRNA analysis in several clinical applications. Given the observed higher analytical sensitivity through RAG targeting, these experiments employed RAG capture groups. First, the detection performance in non-invasive, tumor-unaware lung adenocarcinoma (LUAD, the most common type of NSCLC) was evaluated. cfRNA samples from 50 LUAD patients (stages I-IV) and 50 non-cancer controls were analyzed, including 26 risk-matched controls with a history of significant tobacco exposure collected during low-dose computed tomography (LDCT) screening for lung cancer. cfRNA concentrations, library preparation inputs, and sequencing depths were similar across groups. Figures 6A to 6C). Platelet-adjusted differential expression analysis between LUAD and controls identified 94 genes overexpressed in LUAD, including key genes known to be highly expressed in LUAD such as SFTA2, SLC34A2, NKX2-1 (also known as TTF1; Figure 7A ). LUAD DEGs were confirmed to be highly expressed in TCGA LUAD tumors, while depleted in meta-reference cfRNA controls ( Figure 7B ). Moreover, using cell type-specific gene signatures, GSEA showed that LUAD cfRNA was most enriched in epithelial cells, including alveolar type I and II cells, and most depleted in naive and memory B lymphocytes ( Figure 7C ). Thus, cfRNA from LUAD patients is enriched in transcripts highly expressed in LUAD tumors.

[0176] To further quantify the enrichment of lung adenocarcinoma-derived transcripts, the ES framework developed above was applied using a LUAD-specific gene signature (LUAD Sig; n = 72 genes) identified using RNA-seq data from LUAD tumors within TCGA, LUAD cell lines from CCLE, whole blood (GTEx), and meta-reference cfRNA controls (Methods). In 50 LUAD subjects, 39 had LUAD Sig detected (sensitivity of 78% at 95% specificity in non-cancer controls; Figure 7D ), and significantly correlated with stage (stage sensitivity: I = 40%, II = 60%, III = 80%, IV = 87%; P = 4.7e-14; Figure 7E ). Moreover, in a subset of plasma samples for which circulating tumor DNA was also detected, LUAD ES significantly correlated with ctDNA variant allele frequency (VAF) (R = 0.75, P = 0.03) ( Figure 7F ), suggesting that RARE-Seq can quantify cancer burden in cfRNA. However, only 40% of samples with both ctRNA and ctDNA data were detected by RARE-Seq, suggesting that ctRNA and ctDNA analyses can be complementary for cancer detection. Differences in sequencing depth by cohort were not associated with likelihood of detection ( Figure 6D and 6E ), suggesting that this variable does not drive detection. Moreover, successful detection did not significantly vary with changes in specific oncogenic driver mutations in LUAD ( Figure 6F ).

[0177] Genotyping of ctRNA somatic mutations

[0178] The ability of RARE-Seq to genotype somatic mutations of cancer origin was evaluated. While non-invasive tumor genotyping has been routinely performed clinically using ctDNA, ctRNA-based variant detection can be a useful complement to cfRNA analysis. Thus, a custom ctRNA genotyping method was developed to detect recurrent, clinically actionable single nucleotide variants (SNVs), insertions / deletions (indels), splicing variants, and gene fusions in NSCLC Figure 8A To ensure specificity, putative variants were discarded if their VAF did not significantly exceed the background error rate observed in healthy cfRNA controls. The performance of this method was evaluated in 55 stage IV LUAD patients and 36 LDCT controls. One or more somatic driver mutations were detected in 44% of LUAD patients and 5.6% of LDCT controls, including mutations in key NSCLC oncogenic drivers such as EGFR L858R and exon 19 deletion, KRAS G12C, ROS1 and RET fusions, and MET exon skipping events Figure 8B and 8C ; Figure 6G and 6I . Of the 74 variants identified by tumor tissue genomic DNA or ctDNA genotyping in the same patients, 22 (30%) were also detected in ctRNA, and the detection rates were similar across each variant type Figure 8D A significant relationship between variant detection rate and corresponding expression level was observed. For example, EGFR variants were significantly more frequently detected in cfRNA samples with high EGFR expression (P = 0.007; Figure 8E Thus, RARE-Seq analysis allowed for the simultaneous identification of actionable somatic mutations in cancer patients.

[0179] Given that ctRNA expression signatures and ctRNA somatic variants can individually be used to distinguish cancer patients from non-cancer controls, it was next investigated whether machine learning methods could combine these features into a single detection algorithm. A weighted elastic net (EN) classifier was trained to distinguish LUAD and control cfRNA using 10-fold nested cross-validation (CV) in a training cohort of 50 cases and 50 controls. To reduce the risk of model overfitting, gene weights based on LUAD tumor expression data were utilized during model training. The model exhibited high classification performance (AUROC = 0.9; Figure 8F and 8G ), and selected 279 features, including gene expression features (e.g., key LUAD markers such as SLC34A2, SFTPB, ROS1) and mutation-based features (e.g., detection of at least one ctRNA SNV; Figure 6J and6K ). To test whether the training cohort was large enough to estimate classification accuracy, the classifier was trained and cross-validated using subsampling, and it was found that performance was stable when using 70% or more of the training cohort ( Figure 6L ). Importantly, the EN model performed equally well on a holdout validation cohort collected at an independent institution (AUROC = 0.92) ( Figure 8F and 8G ). The performance of the EN classifier was statistically similar to the ES framework in the training cohort (P = 0.60), but improved in the validation cohort (P = 0.04). Thus, when sufficient samples are available for training, machine learning-based methods can also be used to train robust classifiers that utilize RARE-Seq data as input.

[0180] Identifying mechanisms of resistance to EGFR-targeted therapy

[0181] It was hypothesized that ctRNA analysis could allow detection of non-genetic mechanisms of treatment resistance, such as histological transformation. To explore this question, experiments focused on EGFR-mutant LUAD patients treated with EGFR tyrosine kinase inhibitors (TKIs). Acquired resistance to EGFR TKIs arises from heterogeneous mechanisms, including EGFR-dependent mechanisms (such as secondary point mutations in the EGFR gene) and EGFR-independent mechanisms (such as bypass pathway activation or histological transformation). Plasma samples from 10 stage IV EGFR-mutant LUAD patients were analyzed, who received osimertinib after at least one prior EGFR TKI progression and subsequently developed resistance to osimertinib. The mechanism of resistance present in each patient was determined by tissue biopsy at the time of progression, showing small cell histological transformation in 4 patients, MET gene amplification in 3 patients, and a de novo EGFR C797S point mutation in another 3 patients ( Figure 9A ). Each patient had blood collected prior to starting osimertinib and one or more times after progression (total n = 24 time points).

[0182] Histological transformation to small cell lung cancer (SCLC) was evaluated using the SCLC signature (SCLC Sig), which was defined using publicly available tumor gene expression data. This signature contains 73 genes, including canonical SCLC markers such as ASCL1, NEUROD1, and INSM1. At the time of radiological progression, SCLC Sig was detected in three-quarters (75%) of patients with biopsy-proven histological transformation ( Figure 9B ). In one of these patients, a decrease in SCLC Sig ES was observed in a plasma sample collected 274 days after starting SCLC-targeted chemotherapy (e.g., carboplatin / etoposide), indicating a treatment response of the small cell component. Figure 9C ).

[0183] RARE-Seq was evaluated for identifying genetic mechanisms of resistance. To detect MET pathway activation caused by MET amplification using ctRNA, a MET amplification signature (METamp Sig) was developed, which contains nine genes differentially expressed in three EGFR mutant NSCLC cell lines with acquired resistance due to MET amplification. METamp Sig was detected in two-thirds (67%) of patients with biopsy-proven MET amplification ( Figure 9D ). Notably, in one patient, METamp Sig ES decreased after the addition of the MET inhibitor savolitinib to osimertinib, coinciding with a partial response shown by imaging ( Figure 9E ). The C797S mutation read was identified in two-thirds (67%) of patients known to have the EGFR C797S mutation based on tumor DNA or cfDNA analysis ( Figure 9F and 9G ). Importantly, none of these resistance mechanisms were detected at any pre-resistance time point, confirming the specificity of the method. Altogether, these results demonstrate the proof-of-concept that RARE-Seq can identify both non-genetic and genetic mechanisms of EGFR TKI resistance, including histological transformations that are undetectable by mutation-based ctDNA methods.

[0184] Determining tissue of origin using ctRNA analysis

[0185] As the transcriptional profile of a cancer can reflect its tissue of origin (TOO), ctRNA analysis can be used to identify cancer types. The ability to non-invasively identify the tissue of origin of a malignancy can be helpful in several clinical settings, including metastatic cancer patients with unknown primary. To explore this potential application, plasma samples from patients of four advanced cancer subtypes were analyzed, including lung adenocarcinoma (LUAD, n=50), pancreatic adenocarcinoma (PAAD, n=10), prostate adenocarcinoma (PRAD, n=9), and hepatocellular carcinoma (LIHC, n=10). Tumor-specific gene signatures were generated for each cancer type using TCGA data and evaluated using the ES framework. The sensitivity of cancer-specific ES detection was 78%, 80%, 90%, and 100% for LUAD, LIHC, PAAD, and PRAD, respectively ( Figure 10A ). To distinguish between cancer types, cancer signatures were generated that included genes uniquely expressed in each cancer compared to cfRNA from all other cancers as well as non-cancer patients. Using these signatures, the top predicted TOO was correct 84% of the time, and the top two predicted TOOs identified the correct tumor type ~89% of the time ( Figure 10B and 10C ). Thus, RARE-Seq can potentially enable non-invasive TOO identification.

[0186] Application of RARE-Seq in non-malignant conditions

[0187] RARE-Seq has several applications beyond tumor assessment. It was tested whether cfRNA expression could reveal lung damage caused by benign lung conditions. cfRNA from six patients with active COVID-19 infection (10 time points), 19 patients with acute respiratory distress syndrome (ARDS), and three patients with chronic obstructive pulmonary disease (COPD) were analyzed. Using a normal lung gene expression signature derived from GTEx, lung-derived cfRNA was detected in 47% of patients across these different lung conditions ( Figure 10D ). In contrast, in adult subjects without known lung conditions (n = 42), lung cfRNA was detected in 14% and was significantly more frequent in current smokers than in former smokers or never smokers (45.5% vs 3.6%, P = 0.004; Figure 10E ). This result suggests that lung cfRNA is induced by persistent lung epithelial damage caused by recent exposure to tobacco smoke. Conversely, detection was significantly higher in patients with known lung conditions who were on a ventilator at the time of blood collection ( Figure 10F ), possibly reflecting the severity of the lung condition and the damage to the lung epithelium by mechanical ventilation. Thus, cfRNA analysis can be useful to assess benign conditions involving acute and chronic patterns of tissue damage.

[0188] In addition, given the rapidly growing interest in RNA-based vaccines and therapeutics, it was explored whether RARE-Seq can be useful to simultaneously track RNA-based treatments and their impact on the host. To this end, longitudinal plasma analysis was performed on individuals who underwent the first two mRNA-based COVID-19 vaccinations. RARE-Seq was performed on these samples using RAG capture sets supplemented with decoys targeting mRNA vaccine sequences. Vaccine-matched RNAs were detected after both vaccinations and at high levels for at least 10 days after injection ( Figure 10G ). Interestingly, differentially expressed genes after vaccination were enriched in interferon gamma response, antiviral response, chemokine activity, and leukocyte migration pathways compared to before vaccination, indicating that the host’s response to vaccination was detected ( Figure 10H ). These results highlight the potential of RARE-Seq to simultaneously measure pharmacokinetics and pharmacodynamics of RNA-based therapies.

[0189] Summary and implications of the study

[0190] A new framework for cfRNA analysis offers a variety of potential clinical applications. The level of analytical sensitivity achieved by RARE-Seq between tumor-agnostic and tumor-informed ctDNA-based methods tracking SNVs suggests its potential utility for non-invasive analysis of tumor gene expression in patients with low disease burden. A key innovation of this approach is the specific capture of transcripts that are absent or expressed at very low levels in healthy control plasma. While it is possible to analyze these rare transcripts after whole transcriptome sequencing, this approach is significantly less efficient as it primarily measures leukocyte-derived gene expression and thus would miss rare transcripts. These latter transcripts that are rare in healthy plasma cfRNA are highly enriched in cancer and organ-specific genes and thus important for detecting pathophysiology occurring in tissues outside the blood.

[0191] One important finding was the presence of a high proportion of platelet-derived transcripts in plasma cfRNA. In contrast to cfDNA shed from megakaryocytes, these transcripts appear to originate at least in part from intact platelets that are separated from plasma during blood sample processing and provide a potential confounder for cfRNA analysis. In fact, expression differences in cancer and control cfRNA reported in several previous studies appear to be largely due to differences in platelet-derived transcripts. While use of pre-analytical methods such as increased centrifugation speed can reduce platelet contamination, platelet-poor plasma preparations are rarely completely free of platelets. Moreover, such procedures are often not feasible when analysis includes historical samples from completed clinical trials. Thus, the method described herein to remove the contribution of platelets to cfRNA gene expression analysis has broad utility for future cfRNA liquid biopsy studies.

[0192] Exploration of potential applications suggests that ultra-sensitive cfRNA analysis can be useful in a variety of clinical settings. For example, because this method enables tumor-agnostic ctRNA detection, it can potentially be used for cancer screening alone or in combination with other diagnostic modalities. The fact that RARE-Seq detected lung cancer RNA in some samples that were negative by ctDNA analysis suggests that the two approaches can be complementary. A second promising application is non-invasive identification of cancer type. This approach can be used for non-invasive identification of the TOO in cancer patients with unknown primary or allow for differentiation of cancer subtypes.

[0193] In considering EGFR mutant NSCLC patients for treatment with EGFR TKIs, the results demonstrate that RARE-Seq provides a more comprehensive non-invasive analysis of mechanisms of treatment resistance, including those not driven by recurrent somatic alterations (such as histological transformation). As histological transformation is a common mechanism of resistance that occurs in multiple cancer types, and because its diagnosis currently requires an invasive tissue biopsy, non-invasive detection using cfRNA could greatly enhance the diagnostic approach. The ability of RARE-Seq to simultaneously interrogate for the presence of somatic mutations (e.g., EGFR C797S) and transcriptional effects of pathway activation by somatic alterations (e.g., MET amplification) or non-genetic mechanisms (e.g., histological transformation) enables analysis of a wider range of resistance mechanisms than is currently possible.

[0194] Finally, RARE-Seq will be used to detect cfRNA signatures of non-cancer origin in patients with malignant and benign conditions. For example, data demonstrate the presence of normal lung RNA in the plasma of patients with acute lung injury due to conditions such as COVID-19 infection or ARDS, suggesting that cfRNA analysis can also allow blood-based monitoring of tissue damage. Additionally, RARE-Seq can simultaneously track mRNA vaccines or therapeutic drugs and measure host responses, as shown by analysis of cfRNA in plasma samples following COVID-19 mRNA vaccination.

[0195] In summary, a versatile method of cfRNA analysis has been developed. The method has utility in a wide variety of clinical settings. The results also include several important biological and technical insights that enable more sensitive cfRNA analysis and are broadly applicable to cfRNA-based liquid biopsy applications. The method enables new non-invasive diagnostic applications that have the potential to advance precision medicine and improve patient care.

[0196] Methods

[0197] Human Participants and Cohorts

[0198] All samples analyzed in this study were collected under informed consent using protocols approved by the institutional review boards of the respective centers. The collection centers included Stanford University, Memorial Sloan Kettering Cancer Center (MSKCC), Massachusetts General Hospital (MGH), and University Hospital Zurich, details below. A total of 269 blood samples were collected from 201 individuals (Table 1). Figure 11 The clinical and demographic characteristics of the non-cancer cohort are in Table 1, and the cancer cohorts are in Table 2.

[0199] Cancer cohort. Blood samples were collected from non-small cell lung cancer (NSCLC) patients enrolled at Stanford University (n=33), MSKCC (n=17), and MGH (n=14). Of these, 64 (98%) were diagnosed with lung adenocarcinoma (LUAD), and 1 patient (2%) had mixed adenocarcinoma-squamous histology. Staging information was determined using the American Joint Committee on Cancer (AJCC) 8thedition. For 10 EGFR-mutant NSCLC patients enrolled at MGH, blood was collected at multiple time points, where available, prior to and following progression on the EGFR tyrosine kinase inhibitor (TKI) osimertinib (n=24 plasma samples). Tissue biopsies were collected at progression to identify resistance mechanisms. For additional exploratory analyses, blood samples were collected at Stanford University from patients diagnosed with advanced pancreatic adenocarcinoma (PAAD, n=10) and advanced prostate adenocarcinoma (PRAD, n=9) and at University Hospital Zurich from patients with hepatocellular carcinoma (LIHC, n=10).

[0200] Non-cancer cohort. Blood was collected at Stanford University from individuals without known cancer or benign lung conditions for technical experiment and method development controls (n=45). A subset of these samples comprised a ‘meta-reference’ set for defining expected expression in healthy cfRNA for various downstream analyses (n=15). Blood samples were collected at Stanford University (n=27) and at MGH (n=10) from individuals undergoing low-dose computed tomography (LDCT) screening, which these individuals were eligible for based on smoking history (>=30 pack years) and age (55-80 years). At the time of LDCT screening, it was noted that three individuals had chronic obstructive pulmonary disease (COPD). Additionally, blood samples were collected from patients treated at Stanford Hospital for acute respiratory distress syndrome (ARDS; n=20) and coronavirus disease 2019 (COVID-19; n=6). Plasma was also collected from one individual at Stanford University following two doses of Pfizer-BioNTech COVID-19 mRNA vaccination (n=8 time points).

[0201] Blood collection and plasma processing

[0202] Peripheral blood samples were collected and processed according to protocols at each center. For technical experiments performed at Stanford, whole blood was collected in EDTA tubes and plasma was isolated using centrifugation at 2,500 G for 10 minutes at 4°C. The optical density at 414 nm was measured using a Nanodrop spectrophotometer instrument to quantify hemolysis levels in plasma. Following centrifugation, all plasma was stored at -80°C until cell-free nucleic acid isolation.

[0203] cfRNA extraction

[0204] Cell-free nucleic acids were extracted from plasma using the miRNA protocol of the QIAamp Circulating Nucleic Acid Kit (Qiagen (range 0.2-8.0 mL)). Extraction was performed according to the manufacturer’s instructions with minor modifications. The resulting eluate was incubated with 14 U DNase I (RNase-free DNase Set, Qiagen) for 30 min at room temperature to digest DNA. The digested eluate was purified using the Zymo RNA Clean & Concentrator Kit and stored at -80°C.

[0205] Plasma processed before February 2021 was extracted using phenol / chloroform phase separation and purified using the QIAamp Viral RNA Kit (Qiagen). Briefly, plasma was first incubated with 3 volumes of TRI Reagent LS (Molecular Research Center) and then with 0.4 volumes of chloroform (relative to plasma input). Phase separation was performed using Maxtract 50 mL conical tubes (Qiagen) at 1,500 G for 5 min at room temperature. The aqueous phase containing RNA was carefully removed and RNA was extracted according to the QIAamp Viral RNA Kit protocol. DNA digestion and cleanup were performed as described above. For a subset of samples digested using the “on-column” protocol, 28 U DNase I was added directly to the samples bound on the silica gel column and incubated for 15 min according to the manufacturer’s recommendations. Careful analysis was performed to compare cfRNA yield and gene expression profiles between the different extraction methods (Figures 1-3). Figure 1E and 2B As no significant differences were found, cfRNA samples extracted using both methods were pooled for the analyses presented in this paper.

[0206] cfRNA quantification

[0207] Quantitative real-time polymerase chain reaction (qRT-PCR) was used for quantification of cfRNA. RNA-specific primers were designed to cover a 97-bp amplicon spanning the 2 exon boundaries in the housekeeping gene GAPDH (forward 5'-GATCATCAGCAATGCCTCCT-3' (SEQ ID NO: 1), reverse 5'-TGTGGTCATGAGTCCTTCCA-3' (SEQ ID NO: 2)). DNA-specific primers were designed to cover a 78-bp transcriptionally silent region of chromosome 12 (forward 5'-TACGGTTGGTCCTTTCTTCG-3' (SEQ ID NO: 3), reverse 5'-TTTCCTTTGGGTCTGAATGC-3' (SEQ ID NO: 4)). Reverse transcription was first performed using the High-Capacity cDNA Reverse Transcription Kit (Applied Biosystems). Quantitative PCR (qPCR) was then run using 2X Power SYBR Green PCR Master Mix (Thermo Fisher Scientific) on an Applied Biosystems 7500 Fast Real-Time PCR or QuantStudio 7 Pro instrument. Universal Human Reference RNA (Thermo Fisher Scientific) was run in parallel to generate a standard curve, and cfRNA concentration was calculated by comparing the RNA-specific Ct values of the samples to the standard curve. A similar approach was used with human genomic DNA (Promega) for quantification of DNA. If DNA was detected using the DNA-specific primers, DNA digestion, cleanup, and quantification were repeated. Size distribution of cfRNA was assessed using an Agilent Bioanalyzer RNA 6000 Pico Chip.

[0208] cfRNA library preparation and sequencing

[0209] RARE-Seq. Input mass was 424 pg cfRNA, which represents the 25th percentile of 4 mL plasma cfRNA yield in healthy controls ( Figure 1B ). For samples less than 424 pg, all extracted cfRNA was used for library preparation (range 8-424 pg). NEBNext Ultra II Directional RNA Library Prep with TMII RNA first strand synthesis module and non-directional second strand synthesis module (New England Biolabs) from cfRNA to synthesize double stranded complementary DNA (cDNA). Double stranded cDNA was treated with 100 U S1 nuclease (ThermoFisher) for 30 minutes at room temperature to hydrolyze incomplete (i.e., single stranded) regions. KAPA Hyper Prep kit (Kapa Biosystems) was used to prepare libraries for sequencing, following the manufacturer’s instructions. Whole coding transcriptome capture was performed using Roche Nimblegen SeqCap EZ MedExome Target Enrichment kit and / or Twist Biosciences Comprehensive Exon Hybridization kit, following the respective manufacturer’s instructions. Singleplex capture was performed using Twist Biosciences Hybridization kit with custom genome targeting rare abundance genes (see ‘RAG capture panel design’ below). Captured libraries were sequenced on Illumina HiSeq4000 or NovaSeq6000 instruments using 2x150-bp paired-end reads.

[0210] SMART-Seq. cfRNA from 3 controls was input into library preparation using SMART-Seq strand kit (TaKaRa) according to the manufacturer’s instructions and without additional fragmentation. Libraries were amplified and sequenced using 150-bp paired-end on Illumina NovaSeq6000.

[0211] Leukocyte RNA processing

[0212] Following plasma separation, leukocytes were isolated from plasma-depleted whole blood samples using SepMate PBMC isolation tubes (Stem Cell Technologies) according to the manufacturer’s instructions. Leukocyte pellets were stored at -80°C. RNA was extracted from leukocyte pellets by mixing the leukocyte pellet with 800 ul TRIzol (Invitrogen) and 200 ul chloroform, followed by centrifugation at 13,000 G for 15 minutes at 4°C. RNA was then purified from the aqueous supernatant using Qiagen RNeasy kit according to the manufacturer’s instructions. Total RNA was quantified with NanoDrop instrument and Agilent Bioanalyzer RNA 6000 Pico chip. As previously described, 5-10 ng RNA was used for RARE-Seq library preparation and whole coding transcriptome capture. Fragmentation was performed prior to first strand synthesis according to the manufacturer’s recommendations.

[0213] NCI-H1975 cell line RNA processing

[0214] NCI-H1975 lung adenocarcinoma cells were obtained from the American Type Culture Collection (ATCC). Cell RNA was extracted from cell pellets and quantified as described for leukocyte RNA. To create in vitro NCI-H1975 cell line spiking, NCI-H1975 RNA was serially diluted into cfRNA from a single healthy individual to create samples with cancer mass fractions of 10%, 1%, 0.1%, 0.01%, and 0.001%. Triplicate RARE-seq libraries were generated from each mixture as well as NCI-H1975 RNA (100% cancer fraction) and cfRNA only (0% cancer fraction) and captured with the full coding transcriptome panel. Additional mixtures were created using cfRNA from different healthy individuals for capture with the RAG panel (see ‘RAG capture panel design’ below).

[0215] cfDNA library preparation and sequencing

[0216] Cell-free DNA (cfDNA) was extracted from plasma using the standard cfDNA protocol of the QIAamp Circulating Nucleic Acids Kit (Qiagen). After extraction, cfDNA was quantified with the Qubit dsDNA High Sensitivity Assay Kit (Thermo Fisher Scientific) and the High Sensitivity NGS Fragment Analyzer (Agilent). Cancer Personalized Profiling by Deep Sequencing (CAPP-Seq) was used to create sequencing libraries from 32 ng cfDNA. Thereafter, hybridization-based capture (Roche NimbleGen) was used to target genes recurrently mutated in lung cancer. Libraries were sequenced on an Illumina HiSeq4000 instrument using 2x150bp paired-end reads.

[0217] Mapping, deduplication, and quality control for RARE-Seq

[0218] FASTQ files were first demultiplexed using a custom pipeline. Fastp (v0.20.0) trimmed the first 10 bases from the 5’ end of read 1 and the 3’ end of read 2, as well as removed low-quality or too short read pairs from each sample. STAR 2-pass (v2.7.0) was used to align the remaining high-quality reads to the reference transcriptome (GENCODE v27) and human genome (hg19). A custom barcode method was used to remove PCR duplicates from the transcriptome-aligned and genome-aligned files. RSEM (v1.2.28) was used to use the deduplicated reads for gene-level expression estimation.

[0219] Quality control was assessed using metrics calculated via the RNASeQC package (v2.3.5), such as read mapping quality and mapping rate, including exon, intron, and intergenic region ratios, and ribosomal RNA ratios. Furthermore, DNA contamination was estimated by calculating the percentage of reads mapped to intron sequences out of the total number of reads crossing exon boundaries. Samples were removed from downstream analysis if they had fewer than 20 million reads or if the estimated DNA contamination was greater than 20%. In summary, these QC thresholds removed three samples.

[0220] Gene expression standardization and platelet correction

[0221] RSEM expected counts of captured genes were used for expression analysis (tximport R package v1.22). Counts were first standardized using the trimmed mean M-value (TMM) method, which accounts for variations in library size and transcriptomic complexity between samples (edgeR R package v3.36). Logarithmic transformation and standardized expression values ​​are referred to as 'log2NX' in this paper. A modified version of the Removal of Unwanted Variants (RUV) method was then used to correct for platelet-induced expression variations (RUVseq R package v1.28; D. Risso et al., Nat. Biotechnol. 32, 896–902 (2014), the contents of which are incorporated herein by reference). For fully encoded transcriptomic libraries, unwanted platelet variants were estimated using RUVg with 130 platelet cell type marker genes from PanglaoDB as a negative control group. Figure 3I (O Franzen et al., Database (Oxford). 2019 Jan 1;2019:baz046, the contents of which are incorporated herein by reference). For RAG capture libraries, since most platelet genes were intentionally excluded from the groups (see 'RAG capture group design' below), healthy control meta-reference cfRNAs (n=15) were used as negative control samples within the RUVs to estimate unwanted variants. Factors most relevant to platelet expression were selected as unwanted platelet variants, defined as the mean log2NX of the 21 captured platelet cell type marker genes (e.g., PPBP). Figure 8G For RUVg and RUVs, platelet correction was performed using ordinary least squares regression with log2NX on selected unwanted platelet expression factors. RUVg and RUVs showed similar performance in the whole-coding transcriptome sample. Figure 3J and 6H). After normalization and correction, quality control was performed by calculating the Pearson correlation between the gene expression profile of each sample and the mean of all other samples from the same cohort (i.e., ‘intra-cohort correlation’). For controls with more than one biological replicate, the replicate with the highest intra-cohort correlation was selected for downstream analysis.

[0222] Differential gene expression analysis

[0223] Differential gene expression analysis was performed using DESeq2 (DESeq2 R package v1.34; M. Love et al., Genome Biol. 15, 550 (2014), the disclosure of which is incorporated herein by reference). For analysis of cfRNA and cellular RNA differential expression, a generalized linear model (GLM) was constructed using sample type as a covariate. For analysis of cancer versus control differential expression, the GLM used the condition (i.e., cancer or control) and an unwanted platelet variation factor determined by RUV as covariates. Wald tests were used to determine the significance of log-fold change estimates, followed by Benjamini-Hochberg multiple hypothesis correction. For COVID-19 vaccination time series analysis, likelihood ratio tests were used to assess the significance of differential expression, comparing the full GLM including the vaccine time points and the unwanted variation factor to a reduced model removing the vaccine terms. For all analyses, significantly differentially expressed genes were defined as having an absolute log-fold change greater than one and an adjusted p-value < 0.05. Functional analysis of differentially expressed genes used gene set enrichment analysis (GSEA) (fgsea R package v1.20) or gene ontology (GO) analysis (goseq R package v1.46) using pre-ranked log-fold change estimates. GSEA and GO analysis were performed using cell marker genes from the PanglaoDB single-cell RNA sequencing database and / or the Molecular Signature Database (MSigDb) (O Franzen et al., Database (Oxford). 2019 Jan 1;2019:baz046; A. Subramanian et al., Proc. Natl. Acad. Sci. U. S. A. 102, 15545-15550 (2005); A. Liberzon et al., Cell Syst 1, 417-425 (2015); and A. Liberzon et al., Bioinformatics 27, 1739-1740 (2011); the disclosure of each of which is incorporated herein by reference).

[0224] Focused RAG capture panel design

[0225] To identify rare abundance genes (RAGs) in cfRNA, gene expression data from controls generated with the whole coding transcriptome RARE-Seq method (n=50 biological replicates from n=28 individuals) were analyzed. In addition, gene expression data from whole blood samples generated by the Genotype-Tissue Expression (GTEx) project (n=307) were accessed through the UCSC Xena repository. Expression was TMM-normalized and k-means clustering (k=2) was used to classify genes as expressed or not expressed in each sample. RAGs were defined as expressed in less than 5% of all samples and with an average log2NX in the bottom 30% of all genes in cfRNA and whole blood. RAGs are listed in Table 3. Next, expression uniformity in cfRNA and whole blood was calculated using the Gini coefficient and endogenous control genes were selected from housekeeping genes with a Gini coefficient <0.2 in circulation. In addition, genes associated with lung cancer were manually curated for inclusion in the panel, including genes that are recurrently mutated or rearranged in lung cancer, or genes that are aberrantly expressed in lung cancer histological types. The capture panel focusing on RAGs included a total of 80,866 probes targeting 7,766,820 bp and 5,546 unique genes. Probes targeting Pfizer-BioNTech and Moderna COVID-19 mRNA vaccine sequences were also designed and used to supplement the RAG capture panel for samples collected after vaccination.

[0226] Cell type, tissue, and cancer gene signatures

[0227] For cell type-specific signatures, 7,481 marker genes for 170 human cell types were downloaded from the PanglaoDB single-cell RNA sequencing database (O Franzen et al., Database (Oxford). 2019 Jan 1;2019:baz046, the disclosure of which is incorporated herein by reference).

[0228] Genes expression data from the Genotype-Tissue Expression (GTEx) and The Cancer Genome Atlas (TCGA) projects were used to identify tissue and cancer enriched signatures. Gene-level read counts generated by the UCSC Toil RNA sequencing bioinformatics pipeline were accessed through the UCSC Xena repository. Gene expression was TMM normalized and outlier samples were removed using intra-group correlation measures. The TCGA cohort was further filtered using consensus purity estimates (CPE) to exclude tumors with tumor purity less than 60%. Normal tissue types evaluated from GTEx included bladder (n=9), brain (n=l,107), breast (n=168), colon (n=261), esophagus (n=598), kidney (n=24), liver (n=97), lung (n=241), ovary (n=79), pancreas (n=145), prostate (n=86), skin (n=497), stomach (n=167), and whole blood (n=290). A gene was defined as tissue enriched if its expression was more than 5-fold higher in each tissue than all other evaluated tissues and if the average tissue log2NX was greater than zero. Similarly, cancer tissue types evaluated from TCGA included bladder (BLCA, n=344), brain (GBM, n=144; LGG, n=183), breast (BRCA, n=914), colon (COAD, n=270), esophagus (ESCA, n=181), kidney (KICH, n=65; KIRC, n=350; KIRP, n=262), liver (LIHC, n=164), lung (LUAD, n=309; LUSC, n=365), ovary (OV, n=412), pancreas (PAAD, n=177), prostate (PRAD, n=473), melanoma (SKCM, n=94), and stomach (STAD, n=413). A gene was defined as cancer enriched if its expression was more than 5-fold higher in each cancer tissue type than all other evaluated cancer types and if the average cancer tissue log2NX was greater than zero.

[0229] Cancer-specific gene signatures were created to include genes that are most differentially expressed in each cancer tissue compared to whole blood and meta-reference cfRNA, and thus most likely to be detected in the cfRNA mixture of high hematopoietic background. For this purpose, TCGA cancer cohorts were supplemented with gene expression data from the Cancer Cell Line Encyclopedia (CCLE) accessed through Xena (LIHC, n=25; LUAD, n=76; PAAD, n=40, PRAD, n=7, SCLC, n=50) and an additional SCLC study (n=79). DESeq2 differential gene expression analysis quantified the fold change in gene-wise between cancer tissue expression and whole blood and meta-reference cfRNA expression. In cancer tissues, whole blood, and meta-reference cfRNA, each gene was also classified as expressed or not expressed using K-means clustering (k=2) of the mean gene expression in each group. Such genes were selected as cancer-specific signatures if they were:

[0230] 1. found in the top 1% of genes overexpressed in the cancer tissue of interest or in whole blood and meta-reference cfRNA, and

[0231] 2. classified as expressed in the cancer tissue of interest or in whole blood and meta-reference cfRNA.

[0232] A similar approach was used to generate cancer TOO gene signatures, except that gene expression was compared between LIHC, LUAD, PAAD, PRAD, whole blood, and meta-reference cfRNA, rather than analyzing each cancer type separately. Such genes were selected as cancer TOO signatures if they were:

[0233] 1. found in the top 1% of genes overexpressed in the cancer tissue compared to whole blood and meta-reference cfRNA,

[0234] 2. found in the top 5% of genes overexpressed in the cancer tissue of interest compared to each other cancer type, and

[0235] 3. classified as expressed in the cancer tissue of interest and as not expressed in whole blood and meta-reference cfRNA.

[0236] For all signatures, gene weights were defined as the sum of cancer tissue log2NX and the log-fold change of cancer tissue over meta-reference cfRNA, so that the sign of the differential expression analysis was preserved (i.e., genes overexpressed in the cancer tissue were positive).

[0237] Gene signature detection using enrichment scores

[0238] The Enrichment Score (ES) analysis framework was developed to be broadly applicable as it can use any gene signature of interest. Gene expression for samples was first standardized to Z-scores using the mean and standard deviation of the meta-reference cfRNA expression. Next, the Z-scores for the gene signature of interest were aggregated using the Stouffer equation and gene weights to yield a single ES for each sample. If applicable, the Z-scores for genes with positive and negative weights were aggregated separately and summed. A bootstrap method was used to create random gene signatures of the same size (but excluding genes within the gene signature of interest) (n = 1,000), and the bootstrap-processed ES distribution was used to estimate an empirical p-value for each signature in each sample. If the empirical p-value was less than 0.05, the ES was set to zero.

[0239] Estimating the limit of detection using H1975 mixtures

[0240] The 95% limit of detection (LOD95) for RARE-Seq was determined using computer RNA mixtures. Sequencing data for NCI-H1975 RNA and healthy control cfRNA were generated and analyzed as described for each sample type. Reads were randomly subsampled from each replicate and combined into pre-specified cancer fractions (ranging from 0% to 100%) and the Enrichment Score analysis framework was used to detect cancer in each sample. Here, an NCI-H1975 signature (H1975 Sig) was created by comparing NCI-H1975 samples (n = 4) processed by RARE-Seq but not used to create the spike-in and the meta-reference cfRNA (n = 15). Genes were selected if they were:

[0241] 1. found in the top 1% of genes overexpressed in NCI-H1975 compared to the meta-reference cfRNA, and

[0242] 2. classified as expressed in NCI-H1975.

[0243] The limit of blank (LOB) was calculated from the mean and standard deviation of the H1975 Sig ES found in cfRNA controls (n = 12) that were not used to create the signature and was set as the detection threshold for the remaining spike-ins. Logistic regression was employed to model the relationship between cancer fraction and H1975 Sig detection and the LOD95 was defined as the cancer fraction for which the sensitivity exceeds 95%. This process was repeated for a range of RNA inputs and sequencing depths. The in vitro NCI-H1975 mixtures described above were treated the same as the other cfRNA samples and detected using the H1975 Sig to confirm the computer results.

[0244] Cancer detection using elastic net classification

[0245] Training statistical learning models, including elastic net (EN) logistic regression, random forest, and XGBoost to classify LUAD patient cfRNA (n=50) and controls (n=50). Nested 10-fold cross-validation (10CV) was used to tune hyperparameters and estimate model performance (caret R package v6.0). Features considered for model training included expression of all captured genes, AF of all recurrent somatic variants queried in the ctRNA, and binary values representing the presence of detected SNVs, fusions, indels, splicing variants, or any type of variant in a given sample. Preprocessing removed features with near-zero variance in the training cohort, as well as genes with no significant differential expression between LUAD tumor tissues and meta-reference cfRNA, as described in the Enrichment Score Framework. Gene weights used in the Enrichment Score Framework were also evaluated as penalty weights in the weighted elastic net model. To evaluate the impact of cohort size on model performance, 10CV was repeated after subsampling the cohort from 20-100%, and the standard deviation of AUC was determined by n=10 random samples at each subsampling level. All trained models performed comparably well in classification, but the weighted EN logistic regression achieved the best performance. The final weighted EN model built using the full training cohort selected 279 genes out of the 1,006 features considered, including features related to variant detection in the ctRNA. The final model was validated using a held-out cfRNA cohort (n=28 time points from 14 LUAD patients, n=10 LDCT controls) collected at an independent institution. ROC curves and metrics such as sensitivity, specificity, and AUC were generated using the pROC R package (v1.18).

[0246] Cancer tissue-of-origin classification

[0247] Samples detected by any cancer-specific gene signature were considered for tissue-of-origin (TOO) analysis. The modified Enrichment Score Framework was used to classify the TOO. In this work, Z-scores were aggregated by computing the 90th percentile of all Z-scores in the gene signature of interest. For each sample, the scores for all cancer TOO gene signatures were ordered. Accuracy was determined based on whether the diagnosed cancer type had the highest cancer TOO score or the second highest score detected.

[0248] ctRNA variant calling

[0249] Targeted lists of actionable NSCLC alterations were compiled based on the National Comprehensive Cancer Network (NCCN) guidelines. For each candidate alteration, the null hypothesis that the variant allele frequency (AF) was consistent with the distribution of position- and depth-specific background errors was tested. The background error rate b was calculated as the proportion of reads supporting the substitution in a training cohort consisting of control cfRNA (n = 53). The background error distribution was obtained by sampling 10,000 positive values from a binomial distribution, with the candidate variant depth as the "number of trials" and b as the "success rate," from which p-values were estimated. The resulting p-values were adjusted for multiple hypothesis testing using the Benjamini-Hochberg method, and substitutions with q-values < 0.05 were considered significant. Candidate fusions identified by STAR-Fusion (vl.6.0; A. Dobin et al., Bioinformatics 29, 15-21 (2013), the disclosure of which is incorporated herein by reference) were called if they were present in the list of annotated lung cancer known fusion partners on cBioPortal. MET exon 14 skip alterations were called if the mapping contained a splice junction connecting exons 13 to 15. Indels in EGFR exons 19 and 20 were called if they were greater than 6 base pairs.

[0250] ctDNA variant calling

[0251] Patient-specific lists of alterations were compiled from prior clinical tumor genotyping, if available, or pre-treatment plasma, including only those alterations that were also captured by the lung cancer CAPP-Seq panel. CAPP-Seq sequencing data were analyzed as follows: Briefly, reads were demultiplexed, trimmed using fastp to remove low-quality bases from the 3' end, and aligned to the hg19 human genome using BWA ALN. Custom barcode methods were used to remove PCR duplicates. Single nucleotide variants (SNVs) and insertion / deletion events (indels) were called using a previously reported integrated digital error suppression (iDES) pipeline (A. M. Newman et al., Nat. Biotechnol. 34, 547-555 (2016), the disclosure of which is incorporated herein by reference). Gene fusions were identified using FACTERA (A. M. Newman et al., Bioinformatics 30, 3390-3393 (2014), the disclosure of which is incorporated herein by reference). ctDNA sample allele frequencies (AFs) were computed by the average AF of detected patient-specific alterations.

[0252] MET amplification gene signature

[0253] EGFR mutant NSCLC cell pairs with EGFR TKI acquired resistance due to MET amplification were generated (PC9 / PC9-PERC17, MGH1157-1 / MGH1157-3, MGH170-1C #7 / MGH170-1D #2). PC9-PERC17 line was generated by in vitro culturing PC9 cells with erlotinib. MGH1157-1 and MGH1157-3 cell lines were developed from serial biopsies from a patient before and after treatment with a combination of gefitinib and nazartanib. Clinical analysis of MGH1157-3 tumor using MET FISH showed MET amplification. MGH170 cell lines were developed from a patient who acquired MET amplification (confirmed by FISH) after erlotinib treatment. Single cell clones were isolated from two independent metastatic lesions at rapid autopsy after disease progression; MGH170-1C #7 clonal cell line was sensitive to EGFR TKI, while MGH170-1D #2 was found to be resistant. MET amplification was confirmed in all resistant cell lines and MET dependency was confirmed by sensitivity to EGFR + MET TKI combination. All tissue samples were obtained after patients signed informed consent for their samples to be used in a protocol approved by the Institutional Review Board of Dana-Farber-Harvard Cancer Center. Total RNA from pre-treatment and resistant cell lines was subjected to mRNA-seq (n=3 replicates per time point, n=3 cell lines, or n=18 samples total) and differential expression analysis was performed using DESeq2 with MET amplification status as a covariate. Genes were selected as MET amplification genes if they were:

[0254] 1. found in the top 1% of genes overexpressed in MET amplified cell lines compared to meta-reference cfRNA,

[0255] 2. significantly differentially expressed in MET amplified cell lines compared to non-amplified cell lines, and

[0256] 3. classified as expressed in MET amplified cell lines.

[0257] Herein, gene weight is defined as the sum of log2NX in MET amplified cell lines and log fold change of MET amplified over meta-reference cfRNA, so that the sign of weight retains the sign of differential expression analysis (i.e. genes overexpressed in MET amplified cell lines are positive).

[0258] Statistical analysis

[0259] All statistical analyses were performed in R (v4.1.3). Statistical tests used throughout the text include Wilcoxon rank-sum test, paired t-test, Fisher’s exact test, Kruskal-Wallis ANOVA, Pearson’s correlation, Spearman’s correlation, and DeLong’s test for comparing AUCs. Unless otherwise stated, all statistical tests comparing two groups use the two-sided Wilcoxon rank-sum test, and all statistical tests comparing three or more groups use the Kruskall-Wallis ANOVA. The following significance labels are used in the figures unless stated otherwise: *P < 0.05; **P < 0.01; ***P < 0.001, ****P < 0.0001. Multiple hypothesis correction was performed using the Benjamini-Hochberg (BH) method. Unless otherwise stated, the sample size (n) provided in the text, figures, and legends refers to biologically independent individuals. Details of the statistical methods (including R packages) used for the enrichment score analysis framework, elastic net model training, and variant calling methods are described in their respective sections of the Methods.

[0260] Optimizing blood collection and RNA extraction

[0261] The median concentration per mL of plasma was determined to be 220 pg cfRNA as determined using RNA-specific quantitative RT-PCR (Fig. 1A). Figure 1B Given these low concentrations, blood collection and RNA extraction were first sought to be optimized. As hemolysis has previously been shown to be an important pre-analytical factor affecting other blood-based biomarkers, its effect on cfRNA concentration was examined. Hemoglobin levels were not correlated with cfRNA amounts in the cohort, suggesting that hemolysis is not an important pre-analytical variable in this context (Fig. 1B). Figure 1C In contrast, the length of time plasma was stored at -80°C was indeed negatively correlated with cfRNA levels. However, this effect was only observed during the first month of storage, after which there was no association between storage time and cfRNA yield (Fig. 1C). Figure 1D Next, the effect of different blood collection tubes on cfRNA concentration was examined. All blood collection tube (BCT) types tested, including with and without cell-free nucleic acid protectants, were able to isolate cfRNA with minimal differences in concentration (Fig. 1D). Figure 1E Ten cfRNA isolation protocols were compared. Two silica-gel column-based purification method protocols resulted in significantly higher concentrations (QIAamp Circulating Nucleic Acids Kit and a custom version of the QIAamp Viral RNA Kit; Figure 1F ), and were used for subsequent experiments.

[0262] Optimization of RNA sequencing library preparation

[0263] Next, a library preparation protocol was developed and optimized that combines random primer-based cDNA synthesis with sequencing adapter ligation containing unique molecular identifier (UMI) barcodes, enabling accurate enumeration of unique cfRNA molecules recovered from plasma. After generating such cfRNA-derived sequencing libraries, affinity hybridization was used to capture the coding transcriptome with biotinylated oligonucleotides. Sequencing data was processed using custom bioinformatics pipelines described herein. It was demonstrated that measured cfRNA expression was highly correlated regardless of the tube type ( Figure 2A ) and extraction method ( Figure 2B ) employed. Protocols that preserve the original RNA molecule orientation (i.e., strandness) can be useful for cfRNA analysis, as this approach has been reported to be useful for analyzing differential expression in whole blood. However, while gene-specific expression levels were nearly identical between strand and non-strand libraries, non-strand libraries contained ~30% more unique molecules ( Figure 2C and 2D ), indicating that the strand library preparation protocol resulted in transcript loss. It was also examined whether the method previously shown to overcome the highly damaging nature of FFPE nucleic acids could be used to improve RNA transcript recovery. It was found that additional end-repair with single-strand specific S1 nuclease improved transcript recovery by >2-fold while faithfully preserving expression levels ( Figure 2E and 2F ). Thus, this protocol integrates non-strand library preparation and S1 nuclease end-repair, but other protocols can also be utilized to achieve the results.

[0264] Digestion of contaminating cfDNA from extracted nucleic acids

[0265] As cfRNA analysis is known to be confounded by the presence of contaminating cfDNA, the optimal method for DNA digestion using DNase I was evaluated. While most silica column-based RNA isolation protocols recommend performing enzymatic DNA digestion when RNA binds to the column, it was found that this can result in incomplete DNA removal. Performing the enzymatic digestion step on eluted nucleic acids reduced cfDNA contamination levels from 63.6% to 9.1% ( Figure 2G ). Expression correlation between samples with and without DNA contamination was significantly reduced, emphasizing the importance of DNA removal prior to sequencing library preparation ( Figure 2H ). To further prevent DNA contamination in cfRNA analysis, cfRNA samples with higher estimated DNA contamination levels were excluded in subsequent experiments.

[0266] Verification of RARE-Seq optimization

[0267] The RARE-Seq method was benchmarked against a widely used SMART-Seq whole transcriptome method that includes depletion of nuclear and mitochondrial ribosomal RNA but no other enrichment steps prior to application to clinical samples. Consistent with previous reports, protein-coding RNA was the most abundant biological type in cfRNA analyzed using SMART-Seq, accounting for 67.1% of mapped reads (range 55.6-68.2%) ( Figure 2I ). Protein-coding gene expression between SMART-Seq and RARE-Seq libraries matching cfRNA was highly concordant (R = 0.96, P = < 2.2e-16) ( Figure 2J ). However, 80.4% of coding genes had equal or higher sequencing depth in RARE-Seq, corresponding to a 44% increase in total sequencing depth ( Figure 2J ). The reproducibility of RARE-Seq was assessed using eight replicates from the same individual, collected on four separate dates and sequenced in five batches. Three libraries were prepared from the same cfRNA pool as technical replicates. High mean correlations of 0.988 and 0.997 were observed for biological and technical replicates, respectively ( Figure 2K ). Together, these results indicate that RARE-Seq is a robust and reproducible method for measuring plasma cfRNA expression.

[0268] Table 1. Non-cancer cohort demographics.

[0269] LDCT, low-dose computed tomography. ARDS, acute respiratory distress syndrome. COVID, COVID-19 infection. VACC, post-COVID-19 mRNA vaccination.

[0270] Table 2. Cancer cohort demographics.

[0271] LUAD, lung adenocarcinoma. TKI, treatment with epidermal growth factor receptor (EGFR) tyrosine kinase inhibitor. PAAD, pancreatic adenocarcinoma. PRAD, prostate adenocarcinoma. LIHC, hepatocellular carcinoma.

[0272] Table 3. Rarely abundant genes (RAGs)

[0273] RAGs were defined as expressed in less than 5% of healthy samples and with a mean log2NX in the bottom 30% of all genes.

Claims

1. A method for preparing cell-free RNA for sequencing, comprising: Provide a sample containing nucleic acids for sequencing, wherein the nucleic acids for sequencing are cell-free RNA or nucleic acids derived from and representing cell-free RNA; as well as The sample is contacted with a nucleic acid molecule containing sequences derived from or complementary to a gene transcript, which is a rare, cell-free RNA molecule in control liquid biopsy, such that the nucleic acid molecule is annealed with a subset of the nucleic acids used for sequencing.

2. The method of claim 1, wherein transcripts of cell-free RNA molecules that are rare in control liquid biopsies are defined as transcripts expressed in less than 50% of the control liquid biopsy population.

3. The method of claim 2, wherein the transcripts of cell-free RNA molecules that are rare in control liquid biopsies are defined as transcripts expressed in less than 5% of the control liquid biopsy population.

4. The method according to any one of claims 1-3, wherein the transcripts of cell-free RNA molecules that are rare in control liquid biopsies are defined as transcripts in the last 60% of genes in the control liquid biopsy population in terms of normalized expression.

5. The method of claim 4, wherein the transcripts of cell-free RNA molecules that are rare in control liquid biopsies are defined as transcripts in the last 30% of genes in the control liquid biopsy population in terms of normalized expression.

6. The method according to any one of claims 1-5, wherein the transcripts of cell-free RNA molecules that are rare in control liquid biopsies are defined as transcripts with log-transformed and normalized expression values ​​less than zero in the control liquid biopsy population.

7. The method according to any one of claims 2-6, wherein the control liquid biopsy population comprises at least 5 liquid biopsies.

8. The method of claim 7, wherein the control liquid biopsy population comprises at least 50 liquid biopsies.

9. The method according to any one of claims 1-8, wherein the control liquid biopsy is collected from an individual who does not have one or more of the following at the time of biopsy collection: observed pathogen infection, confirmed cancer, confirmed metabolic disease, confirmed neurological disease, confirmed immunodeficiency disease, confirmed autoimmune disease, confirmed inflammatory disease, confirmed cardiovascular disease, confirmed kidney disease, confirmed liver disease, active pregnancy, confirmed pregnancy complications, confirmed fetal complications, organ transplantation, active rejection of organ transplantation, obesity, malnutrition, cachexia, and abnormal clinical tests.

10. The method according to any one of claims 1-9, wherein the transcripts of the cell-free RNA molecules that are rare in control liquid biopsy comprise 50% of the transcripts in Table 3.

11. The method of claim 10, wherein the transcripts of cell-free RNA molecules that are rare in control liquid biopsies comprise 90% of the transcripts in Table 3.

12. The method of claim 11, wherein the transcripts of cell-free RNA molecules that are rare in control liquid biopsies comprise 100% of the transcripts in Table 3.

13. The method according to any one of claims 1-12, wherein the nucleic acid molecule excludes at least 50% of whole-exome gene transcripts, which are not transcripts of rare cell-free RNA molecules.

14. The method of claim 13, wherein the nucleic acid molecule excludes at least 90% of whole-exome gene transcripts, which are not transcripts of rare cell-free RNA molecules.

15. The method according to any one of claims 1-14, wherein the nucleic acid molecule consists of 5,000 or fewer gene transcripts plus rare abundance of cell-free RNA molecule transcripts.

16. The method of claim 15, wherein the nucleic acid molecule consists of 500 or fewer gene transcripts plus rare abundance of cell-free RNA molecule transcripts.

17. The method according to any one of claims 1-16, wherein the nucleic acid molecular set further comprises one or more of the following: tissue-specific transcripts, cell type-specific transcripts, clinically relevant transcripts, B cell receptor and T cell receptor transcripts, biomarkers, common mutagenic transcripts, and a set of control transcripts for intersample standardization.

18. The method of claim 17, wherein the biomarker is associated with one of the following biological characteristics: medical disease, pregnancy, fetal complications, pregnancy complications, tumor growth, cancer, specific cancer type, pathogen infection, immune activation, organ transplant rejection, neurodegeneration, tissue origin, cell type origin, or activation of biochemical pathways.

19. The method according to any one of claims 1-18, wherein the nucleic acid molecule set is a set of probes for targeted capture hybridization.

20. The method according to any one of claims 1-18, wherein the nucleic acid molecular set is a set of primers for targeted amplification.

21. The method according to any one of claims 1-19, wherein the cfRNA sample is derived from blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, or feces.

22. The method of claim 1, further comprising: Generate a sequencing library derived from the sample; as well as The sequencing library is subjected to targeted sequencing to produce sequencing results of cell-free RNA, wherein the sequencing targets the nucleic acid molecular genome.

23. The method of claim 22, further comprising: Platelet expression was removed from sequencing results in a computer (in silico).

24. The method according to claim 22 or 23, further comprising: Differential transcript analysis was performed using the sequencing results and the second sequencing results.

25. The method according to any one of claims 22-24, further comprising: Enrichment of at least one expression feature in the sequencing results.

26. The method according to any one of claims 22-25, further comprising: Detect sequence mutagenesis in the sequencing results.

27. The method according to any one of claims 22-26, further comprising: The copy number status of one or more genes can be inferred from the sequencing results.

28. The method according to any one of claims 22-27, further comprising: The sequencing results, along with several other sequencing results, are used to train a computational model to predict the classification status or probability of biological characteristics, wherein the cell-free RNA samples have a known classification status of biological characteristics.

29. The method according to any one of claims 22-27, further comprising: The sequencing results are used as input to a trained computational model to predict the classification status or probability of a biological feature, wherein the computational model has been trained using a cohort of RNA sequencing results with known classification statuses of biological features.

30. The method according to any one of claims 22-27, further comprising: One or more features are derived from the sequencing results, wherein the one or more features include enrichment of one or more gene features, enrichment of biochemical pathways, collection of sequence variations, and copy number status; as well as One or more derived features are used as input to a trained computational model to predict the classification state or probability of a biological feature, wherein the computational model has been trained using a cohort of RNA sequencing results with known classification states of biological features.

31. A nucleome for targeting transcripts of rare, cell-free RNA molecules, the nucleome comprising: Nucleic acid molecules having sequences derived from or complementary to a gene transcript, which is a cell-free RNA molecule that is rare in control liquid biopsies.

32. The nucleome of claim 31, wherein the transcript of a cell-free RNA molecule that is rare in the control liquid biopsy is defined as a transcript expressed in less than 50% of the control liquid biopsy population.

33. The nucleome according to claim 32, wherein the transcript of a cell-free RNA molecule that is rare in the control liquid biopsy is defined as a transcript expressed in less than 5% of the control liquid biopsy population.

34. The nucleic acid genome according to any one of claims 31-33, wherein the transcripts of cell-free RNA molecules that are rare in the control liquid biopsy are defined as transcripts in the last 60% of genes in the control liquid biopsy population in terms of normalized expression.

35. The nucleome of claim 34, wherein the transcripts of cell-free RNA molecules that are rare in control liquid biopsies are defined as transcripts in the last 30% of genes in the control liquid biopsy population in terms of normalized expression.

36. The nucleome according to any one of claims 31-35, wherein the transcript of a cell-free RNA molecule that is rare in control liquid biopsy is defined as a transcript having a logarithmic transformation and normalized expression value of less than zero in the control liquid biopsy population.

37. The nucleic acid group according to any one of claims 32-36, wherein the control liquid biopsy population comprises at least 5 liquid biopsies.

38. The nucleic acid genome of claim 37, wherein the control liquid biopsy population comprises at least 50 liquid biopsies.

39. The nucleic acid group according to any one of claims 31-38, wherein the control liquid biopsy is collected from an individual who does not have one or more of the following at the time of biopsy collection: observed pathogen infection, confirmed cancer, confirmed metabolic disease, confirmed neurological disease, confirmed immunodeficiency disease, confirmed autoimmune disease, confirmed inflammatory disease, confirmed cardiovascular disease, confirmed kidney disease, confirmed liver disease, active pregnancy, confirmed pregnancy complications, confirmed fetal complications, organ transplantation, active rejection of organ transplantation, obesity, malnutrition, cachexia, and abnormal clinical tests.

40. The nucleic acid genome according to any one of claims 31-39, wherein the transcripts of the cell-free RNA molecules that are rare in control liquid biopsy comprise 50% of the transcripts in Table 3.

41. The nucleome according to claim 40, wherein the transcripts of cell-free RNA molecules that are rare in control liquid biopsy comprise 90% of the transcripts in Table 3.

42. The nucleome according to claim 41, wherein the transcripts of cell-free RNA molecules that are rare in control liquid biopsy comprise 100% of the transcripts in Table 3.

43. The nucleic acid genome according to any one of claims 31-42, wherein the nucleic acid genome excludes at least 50% of whole-exome gene transcripts, which are not transcripts of rare cell-free RNA molecules.

44. The nucleic acid genome of claim 43, wherein the nucleic acid genome excludes at least 90% of whole-exome gene transcripts, which are not transcripts of rare cell-free RNA molecules.

45. The nucleic acid genome according to any one of claims 31-44, wherein the nucleic acid genome consists of 5,000 or fewer gene transcripts plus rare abundance of cell-free RNA molecular transcripts.

46. ​​The nucleic acid genome of claim 45, wherein the nucleic acid genome consists of 500 or fewer gene transcripts plus rare abundance of cell-free RNA transcripts.

47. The nucleic acid genome according to any one of claims 31-46, wherein the nucleic acid genome further comprises tissue-specific transcripts, cell type-specific transcripts, clinically relevant transcripts, B cell receptor and T cell receptor transcripts, biomarkers, and common mutagenic transcripts.

48. The nucleic acid genome of claim 47, wherein the biomarker is associated with one of the following biological characteristics: medical disease, pregnancy, fetal complications, pregnancy complications, tumor growth, cancer, specific cancer type, pathogen infection, immune activation, organ transplant rejection, neurodegeneration, tissue origin, cell type origin, or activation of biochemical pathways.

49. The nucleic acid genome according to any one of claims 31-48, wherein the nucleic acid genome further comprises a set of control transcripts for inter-sample standardization.

50. The nucleic acid set according to any one of claims 31-49, wherein the nucleic acid set is a set of probes for targeted capture hybridization.

51. The nucleic acid set according to any one of claims 31-49, wherein the nucleic acid set is a set of primers for targeted amplification.

52. A method for extracting RNA from a cell-free source, comprising: (a) Adding glycogen to a sample containing cell-free nucleic acids; as well as (b) Contact the silicon-based pillar with a sample containing cell-free nucleic acid.

53. The method of claim 52, wherein step (a) is performed prior to step (b).

54. The method according to claim 52 or 53, further comprising: Cell-free nucleic acids are eluted from a silicon-based column to produce an extracted cell-free nucleic acid solution; as well as The extracted cell-free nucleic acid solution is contacted with DNase.

55. A method for quantifying cell-free RNA for downstream molecular applications, comprising: Provide samples containing cell-free RNA; Reverse transcription of cell-free RNA to produce cDNA; as well as The concentration of cell-free RNA in solution was quantified using quantitative real-time polymerase chain reaction and cDNA.

56. The method of claim 55, wherein the step of quantifying cell-free RNA concentration further comprises generating a standard curve based on a set of control standards of known concentrations, wherein the control standards are also evaluated using quantitative real-time polymerase chain reaction.

57. The method of claim 55 or 56, wherein the sample further comprises cell-free DNA, and the method further comprises: The concentration of cell-free DNA in a sample is quantified using quantitative real-time polymerase chain reaction, wherein cell-free RNA is quantified by using primers that span gene introns that are relatively stable in cell-free RNA samples, and cell-free DNA is quantified by using primers that anneal to genomic transcriptionally silent regions that are relatively stable in cell-free DNA samples.

58. The method of claim 57, wherein the primers for quantifying cell-free RNA span the introns of GAPDH, and the primers for quantifying cell-free DNA target and cover a 78 bp transcriptionally silencing region on chromosome 12.

59. A method for sequencing cell-free RNA, comprising: Provide a nucleic acid molecular library, wherein the nucleic acid molecular library is derived from cell-free RNA, wherein the cell-free RNA is derived from liquid biopsy; The nucleic acid molecular library was sequenced to produce sequencing results; as well as Remove variations caused by platelet-related transcript expression.

60. The method of claim 59, wherein the nucleic acid molecular library is generated by capturing or amplifying nucleic acid molecules.

61. The method according to claim 60, wherein the nucleic acid library is a whole exome library.

62. The method of claim 60, wherein the nucleic acid library is a library targeting rare abundance genes.

63. The method according to any one of claims 59-62, wherein the cfRNA sample is derived from blood, plasma, lymph, cerebrospinal fluid, amniotic fluid, urine, or feces.

64. A method for generating a targeted sequence genome for sequencing cell-free RNA, comprising: Collect control liquid biopsy populations, each containing cell-free RNA; Sequencing of cell-free RNA from control liquid biopsies; A set of rare abundance genes was identified in the control liquid biopsy population, the rare abundance genes being defined by at least one or more of the following: their percentage expression or their expression level in the control liquid biopsy population; and A set of nucleic acid molecules were synthesized for capturing or amplifying the rare abundance genes to produce a targeted sequencing set for cell-free RNA sequencing.

65. The method of claim 64, wherein the set of rare abundance genes is defined at least by the percentage of them expressed in the control liquid biopsy population and their expression level in the control liquid biopsy population.

66. The method according to claim 64 or 65, wherein the resulting targeted sequencing genome for cell-free RNA sequencing is the nucleic acid genome according to any one of claims 31-51.