Methods for detecting and quantifying genomic and gene expression changes using RNA

JP2024530807A5Pending Publication Date: 2025-05-30ルーセンス ライフ サイエンシズ プライベート リミテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024514609
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-06
Filing Date
2022-05-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Current methods for detecting genomic alterations and gene expression in cancer diagnosis are limited by the instability of RNA in blood samples and the inability of DNA-based assays to effectively identify certain genomic rearrangements and gene fusions, particularly those involving long introns, leading to incomplete coverage and false negatives.

Method used

A method involving the extraction of RNA from a biological sample, conversion to cDNA, followed by multiplexed PCR using primer pairs targeting exon junctions with barcode sequences, purification, and sequencing to detect genomic alterations and gene expression levels, optimizing amplicon design for fragmented RNA.

Benefits of technology

This approach enables sensitive and accurate detection of genomic alterations and gene expression in RNA, including fusion events, with improved coverage and specificity, enhancing cancer diagnosis and treatment prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a method for detecting and quantifying genomic and gene expression changes using RNA in a biological sample. The disclosed method may include determining the presence or absence of genomic alterations and / or determining the presence or absence of gene expression and / or quantifying the level of gene expression by performing variant calling of sequence alignments obtained from the disclosed method. Variant calling may include identifying differences between consensus reads and a reference genome based on sequence alignments from the disclosed method; and determining the read count of sequence alignments that include genomic alterations. Genomic alterations may be insertions (such as duplications), deletions, single nucleotide mutations, or combinations thereof. The method also includes performing multiple multiplex PCR reactions on the converted cDNA using multiple target-specific forward and reverse primer pairs. In one embodiment, the multiple primer pairs are complementary to contiguous sequences spanning the exon-exon junctions of each target gene. The application also discloses a kit thereof. TIFF2024530807000023.tif60170
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] FIELD OF THEINVENTION The present invention relates to the detection and quantification of nucleic acids, in particular to the detection and quantification of RNA. [Background technology]

[0002] background Circulating biomarkers are promising tools used for cancer detection, prognosis prediction and prediction of cancer treatment response. These circulating biomarkers usually include DNA samples such as cell-free DNA (cfDNA) and circulating tumor cells. Various RNA molecules are also potential biomarkers for the diagnosis and prognosis prediction of various diseases such as cancer, and are known to be useful for early cancer diagnosis, tumor progression monitoring and prediction of treatment response. Cancer cells are also known to release cell-free RNA (cfRNA) into the body circulation. These cancer-related cfRNAs, also known as circulating tumor RNA (ctRNA), can be found in the serum and plasma of cancer patients. Although both cfDNA and cfRNA are promising cancer biomarkers, the measurement of cfDNA has traditionally been preferred due to the stability of cfDNA in body fluids. Despite the discovery of RNA in plasma and serum more than 20 years ago, there is still a general perception that extracellular RNA in blood is extremely unstable and highly fragmented, considering the relative instability of RNA, which becomes unstable itself when fragmented in blood, compared to DNA, due to the presence of high concentrations of ribonucleases in the blood circulation. Multiple studies have documented the presence of tumor-specific circulating RNA (ctRNA) in the serum and plasma of cancer patients. Current non-oncological clinical applications of cfRNA include the measurement of maternal and fetal cfRNA transcripts to monitor long-term phenotypic changes in both the mother and fetus, and to assess fetal gestational age. In the blood circulation, cfRNA is known to exist free, bound to proteins or lipids, or as exosomes protected in various types of membrane-derived microvesicles, making them highly stable. Plasma cfRNA is thought to be a mixture of RNA protected by RNA-binding proteins and RNA contained within extracellular vesicles. cfRNA is widely available in plasma, serum, and many other body fluids, and its paradoxical stability makes it a potential candidate for the development of biomarkers for rapid, sensitive, and inexpensive diagnosis.Moreover, detection of ctRNA provides the same mutational information as ctDNA, but in addition, it can also provide quantitative information on the expression levels of target genes of interest, potentially increasing the sensitivity of detection of variants with low allele frequency due to overexpression of tumor-specific transcripts. Finally, expression of various ctRNA species can be dysregulated due to uncontrolled cell proliferation, making ctRNA a valuable tool for cancer detection. Currently, the most common technique for detection of cfRNA is to use quantitative real-time polymerase chain reaction (qRT-PCR). However, methods involving qRT-PCR are often limited by their sensitivity when assaying low-input samples. NGS may be more suitable due to its ability to detect novel cfRNAs and distinguish RNA isoforms. In hybridization-based library preparation methods, sequence-specific bias due to enzymatic ligation during the library construction phase leads to bias in transcript representation, especially during analysis of small RNAs. Targeted NGS assays such as hybridization capture or amplicon sequencing may also allow sensitive quantification of cfRNA (as opposed to whole-transcriptome analysis, which has low conversion efficiency).

[0003] Many cancer genes exhibit genomic alterations, and these genomic alteration events have been found in a wide variety of tumors. Targeted DNA-based next-generation sequencing techniques specifically designed to detect kinase rearrangements can effectively detect oncogenic kinase fusions with high confidence. However, there are technical limitations in the ability of such DNA-based assays to detect certain genomic alterations, such as gene fusions. DNA-based assays can only identify fusions in genes where genomic rearrangements occur within typically short introns that are effectively covered in the panel. Some clinically important fusions arise from rearrangements in very long introns, the complete coverage of which would severely compromise coverage of the remaining genes on the panel. Thus, there are gaps in coverage of certain introns, resulting in blind spots in the detection of potential rearrangement breakpoints. Fusion detection using DNA does not provide direct evidence that the rearrangement produces a fusion that is expressed at the mRNA level, which is particularly problematic for rearrangements that appear non-canonical at the genomic DNA level. Indeed, one study of lung cancer tissue samples showed that RNA sequencing was used to detect alterations in 14% (36 / 254) of cases in which DNA sequencing was negative for clinically actionable mutations. For example, gene fusion events involving the neurotrophic receptor tyrosine kinase (NTRK) genes (NTRK1 / 2 / 3) and neuregulin-1 (NRG1) genes cannot be effectively covered in targeted DNA sequencing panels without compromising the cost of sequencing and coverage of the remaining genes in the sequencing panel.

[0004] Apart from the detection of genomic alteration events, being able to non-invasively and accurately quantify the genomic expression of relevant cancer biomarkers is important for predicting response to cancer therapy and for making appropriate treatment decisions. For example, the gene expression level of programmed death-ligand 1 (PD-L1) is a predictive cancer biomarker used to identify cancer patients who are likely to respond to immunotherapy. PD-L1 is also a potential predictive biomarker for measuring tumor sensitivity to immune checkpoint blockade inhibitors such as anti-PD-1 inhibitors (pembrolizumab and nivolumab), anti-cytotoxic T-lymphocyte-associated protein 4 (CTLA-4) inhibitors (ipilimumab and tremelimumab) and anti-programmed death protein 1 (PD-1) (atezolizumab, durvalumab and avelumab). Other genetic biomarkers useful for predicting the likelihood of responding to immune checkpoint inhibitor therapy include T-cell immunoglobulin and mucin domain-containing protein 3 (TIM-3), lymphocyte activation 3 (LAG-3), and cytotoxic T lymphocyte-associated protein 4 (CTLA-4). The ability to quantify the expression of these target biomarkers longitudinally and non-invasively can be very useful in monitoring treatment response and in making treatment decisions.

[0005] Conventional assays routinely detect genomic alterations at the DNA level, limiting the scope of detection to quantification of DNA genomic alterations such as mutations and genome copy number changes.

[0006] Thus, there is a need to provide a method for sensitive detection and quantification of gene expression and genomic alteration events associated with disease (such as cancer) that overcomes or at least ameliorates one or more of the above-mentioned drawbacks. There is a need to provide a method for simultaneously detecting genomic alterations, such as structural rearrangements, and gene expression using alternative sample inputs, such as RNA (such as circulating cell-free RNA (cfRNA)). Summary of the Invention

[0007] overview In one aspect, the present disclosure refers to a method of detecting genomic alterations and / or detecting gene expression and / or quantifying levels of gene expression using RNA in a biological sample, comprising the steps of: (a) extracting RNA from a biological sample and converting the RNA into complementary DNA (cDNA); (b) performing multiplex PCR reactions on the converted cDNA; (I) Target genes that can cause genomic changes a plurality of forward and reverse primer pairs specific for a plurality of Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes capable of causing genomic alterations forward The primers are aligned to approximately 50 base pairs of the exon junction of each target gene that can cause genomic alterations. Upstream is complementary to the sequence located at Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes capable of causing genomic alterations Reverse The primers are aligned to approximately 50 base pairs of the exon junction of each target gene that can cause genomic alterations. downstream is complementary to the sequence located at Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes capable of causing genomic alterations Reverse The primer comprises a barcode sequence at its 5' end, and the barcode sequence of each reverse primer corresponding to each target gene capable of causing genomic alteration is different; multiple forward and reverse primer pairs, and / or (II) Control housekeeping genes a plurality of forward and reverse primer pairs specific for a plurality of (i) each of a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes; forward the primers are complementary to sequences spanning the exon-exon junctions of each control housekeeping gene; Each of multiple forward and reverse primer pairs specific for multiple control housekeeping genes ReverseThe primers were designed to span approximately 100 base pairs of sequence across the exon-exon junctions of each control housekeeping gene. downstream is complementary to the sequence Each of multiple forward and reverse primer pairs specific for multiple control housekeeping genes Reverse the primers contain a barcode sequence at their 5' ends, and the barcode sequence of each reverse primer corresponding to each control housekeeping gene is different; (ii) each of a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes; Reverse the primers are complementary to sequences spanning the exon-exon junctions of each control housekeeping gene; Each of multiple forward and reverse primer pairs specific for multiple control housekeeping genes forward The primers were designed to span approximately 100 base pairs of sequence across the exon-exon junctions of each control housekeeping gene. downstream is complementary to the sequence Each of multiple forward and reverse primer pairs specific for multiple control housekeeping genes Reverse the primers contain a barcode sequence at their 5' ends, and the barcode sequence of each reverse primer corresponding to each control housekeeping gene is different; (iii) each forward and each reverse primer of a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes is complementary to a contiguous sequence spanning an exon-exon junction of each of the control housekeeping genes; each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of control housekeeping genes comprises a barcode sequence at its 5' end, and the barcode sequence of each reverse primer corresponding to each control housekeeping gene is different; multiple forward and reverse primer pairs, and / or (III) Target genes related to protein expression A plurality of primer sets specific to a plurality of each primer set comprises a plurality of forward and reverse primer pairs specific for each target gene associated with protein expression; (i) each of a plurality of forward and reverse primer pairs specific for each target gene associated with protein expression; forward the primers are complementary to sequences spanning exon-exon junctions of each target gene associated with protein expression; Each of multiple forward and reverse primer pairs specific for each target gene associated with protein expression Reverse Primers were designed to span approximately 100 base pairs of sequence across exon-exon junctions of each target gene relevant for protein expression. downstream is complementary to the sequence Each of multiple forward and reverse primer pairs specific for each target gene associated with protein expression Reverse the primers contain a barcode sequence at their 5' ends, and the barcode sequences of each reverse primer corresponding to each target gene associated with protein expression are different; (ii) each of a plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression; Reverse the primers are complementary to sequences spanning exon-exon junctions of each target gene associated with protein expression; Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression. forward Primers were designed to span approximately 100 base pairs of sequence across exon-exon junctions of each target gene relevant for protein expression. downstream is complementary to the sequence Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression. Primer but containing a barcode sequence at its 5' end, with the barcode sequence of each reverse primer corresponding to each target gene associated with protein expression being different; (iii) each forward and each reverse primer of the plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression is complementary to a contiguous sequence spanning an exon-exon junction of each target gene associated with protein expression; Each reverse primer of the multiple forward and reverse primer pairs specific to multiple target genes associated with protein expression comprises a barcode sequence at its 5' end, and the barcode sequence of each reverse primer corresponding to each target gene associated with protein expression is different; Multiple primer sets thereby generating a plurality of amplicons; (c) purifying the plurality of amplicons from step (b); (d) amplifying the purified products from step (c) by using universal indexed adapter primers to generate a sequencing library; (e) purifying the sequencing library obtained from step (d); (f) subjecting the purified sequencing library from step (e) to multiplex sequencing on a next-generation sequencing platform to obtain multiple sequencing reads; (g) deriving a consensus read for each sequence from the multiple sequencing reads obtained from step (f); (h) performing a sequence alignment of the consensus reads obtained from step (g) with a reference genome, (I) if the sequence alignment results in a partial alignment of an exon from a first gene with a reference genome and a partial alignment of an exon from a second gene with a reference genome, then (i) determining a sequence alignment as a split read; (ii) counting / enumerating the number of split reads from step (h)(I)(i) that support the fusion junction; and (iii) if the number of split reads from step (h)(I)(ii) is two or more, then determining the first gene and the second gene as fusion partners; (II) If the sequence alignment results in an alignment of the control housekeeping genes with the reference genome, then: (i) determining the sequence alignment to a consensus read of a control housekeeping gene; (ii) counting / enumerating the consensus read pairs of the control housekeeping genes from step (h)(II)(i); and (iii) determining the level of gene expression of a control housekeeping gene; (III) if the sequence alignment results in an alignment of the target gene with the reference genome that is relevant for protein expression; (i) determining the sequence alignment as a consensus read of the target gene associated with protein expression; (ii) counting / enumerating consensus read pairs of the target genes associated with protein expression from step (h)(III)(i); and (iii) determining the level of gene expression of the target gene related to protein expression; step; (i) determining the presence or absence of genomic alterations and / or determining the presence or absence of gene expression and / or quantifying the level of gene expression based on the sequence alignment from step (h).

[0008] In another aspect, the present disclosure provides a kit for detecting genomic alterations and / or detecting gene expression and / or quantifying gene expression levels in a biological sample using RNA by the methods disclosed herein, comprising: - as defined in the methods disclosed herein Target genes that can cause genomic changes a plurality of forward and reverse primer pairs specific for a plurality of - as defined in the methods disclosed herein Control housekeeping genesa plurality of forward and reverse primer pairs specific for a plurality of - as defined in the methods disclosed herein Target genes related to protein expression The present invention refers to a kit comprising a plurality of primer sets specific for a plurality of the above. [Brief description of the drawings]

[0009] The present invention will be better understood by reference to the detailed description, when considered in conjunction with the following non-limiting examples and the accompanying drawings.

[0010] [Figure 1] Figure 1 is a general overview of the cfRNA-based detection method for gene fusion events caused by intron DNA rearrangement between two genes described herein.Primers (represented by arrows) are designed to flank the exon junction of the gene known to undergo fusion.Primers (→) are designed to be approximately 100 base pairs long, so that if fusion product exists, the length of the resulting amplicon matches the cfRNA fragment size observed in plasma samples. [Figure 2A] Figure 2, including Figures 2A and 2B, illustrates the primer design of the disclosed method, where Figure 2A illustrates the primer design for capturing control housekeeping genes (left panel) and expressed genes (right panel) in cfRNA. At least one primer of the primer pair spans an exon-exon junction to prevent unintended amplification of cfDNA, and the length of the resulting amplicon is approximately 100 base pairs. Note that the primer pair of housekeeping genes is different from the primer pair of expressed genes, and Figure 2B illustrates the forward primer and reverse primer designed to bind to two different exons, separated by an intron that is more than 5000 base pairs in length. [Figure 2B] See legend to Figure 2A. [Figure 3A]Figure 3, including Figures 3A and 3B, shows size and concentration analysis of cfRNA from plasma total nucleic acid extracts from cancer patients and healthy individuals, where Figure 3A shows size and concentration analysis of cfRNA from plasma total nucleic acid extract from a cancer patient (sample A), Figure 3B shows size and concentration analysis of cfRNA from plasma total nucleic acid extract from another cancer patient (sample B), Figure 3C shows size and concentration analysis of cfRNA from plasma total nucleic acid extract from a healthy individual (sample C), and Figure 3D shows size and concentration analysis of cfRNA from plasma total nucleic acid extract from another healthy individual (sample D). Samples were quantified and profiled using Bioanalyzer RNA 6000 Pico kit or High Sensitivity RNA Screentape on a 4200 Tapestation. Total concentration (representative abundance) of cfRNA is generally higher in representative plasma extracted from cancer patients compared to that extracted from healthy individuals. [Figure 3B] See legend to Figure 3A. [Figure 3C] See legend to Figure 3A. [Figure 3D] See legend to Figure 3A. [Figure 4] A comparison of cfDNA and cfRNA yields in total nucleic acid extracts from plasma extracted from cancer patients and healthy individuals is shown. [Figure 5A]Figure 5, including Figures 5A, 5B and 5C, shows an example of fragmentation of extracted H2228 cell line RNA by physically shearing large sized nucleotides (>1500 nucleotides) into smaller sizes to mimic cfRNA fragment sizes. Samples were quantified and profiled using a Bioanalyzer RNA 6000 Pico kit or a High Sensitivity RNA Screentape on a 4200 Tapestation, where Figure 5A shows the fragmentation profile of extracted H2228 cell line RNA, Figure 5B shows the resulting fragmentation profile of fragmented H2228 cell line RNA, and Figure 5C shows the fragmentation profile of plasma cfRNA. The resulting fragmentation profile of H2228 cell line RNA is similar to that of plasma cfRNA, with a major RNA peak at 119 nucleotides (represented by an arrow). [Figure 5B] See legend to Figure 5A. [Figure 5C] See legend to Figure 5A. [Figure 6A] Figure 6, including Figures 6A and 6B, illustrates the detection of EML4-ALK fusion in 1 ng of fragmented H2228 RNA showing the alignment of split reads capturing the fusion breakpoints of exon 6b of EML4 and exon 20 of ALK, where Figure 6A is a visualization of the split reads on the Integrated Genome Viewer (IGV) and Figure 6B is a diagrammatic representation (from Arriba tools for detection of gene fusions) showing the exon fusion. [Figure 6B] See legend to Figure 6A. [Figure 7-1] FIG. 7 is a diagrammatic representation of the Arriba tool showing detection of various exon fusions in cancer cell lines: NCI-H660 (CRL-5813, ATCC), VCaP (CRL-2876, ATCC), MV-4-11 (CRL-9591, ATCC) and Kasumi-1 (CRL-2724, ATCC) using the multiplex amplicon sequencing method described herein on fragmented RNA samples. [Figure 7-2] See description of Figure 7-1. [Figure 8A] FIG. 8, including FIG. 8A, 8B, and 8C, shows detection of TMPRSS2-ERG gene fusion in nucleic acid extracts from metastatic prostate patients using the cfRNA-based method described herein compared to a cfDNA-based method, where FIG. 8A is an IGV graphical view showing 17 split reads that supported the presence of intronic breakpoints detected by the cfDNA-based detection method, FIG. 8B is an IGV graphical report showing 4123 split reads that supported the presence of corresponding exonic breakpoints detected by the cfRNA-based method described herein, and FIG. 8C is a diagrammatic representation of the Arriba tool showing the TMPRSS2-ERG gene fusion. [Figure 8B] See legend to Figure 8A. [Figure 8C] See legend to Figure 8A. [Figure 9A] FIG. 9, including FIG. 9A, 9B, and 9C, shows detection of CCDC6-RET gene fusion in a nucleic acid extract from a metastatic lung cancer patient using a cfRNA-based method described herein compared to a cfDNA-based method, where FIG. 9A is an IGV graph report showing 12 split reads that supported the presence of an intronic breakpoint detected by the cfDNA-based detection method, FIG. 9B is an IGV graph report showing 1474 split reads that supported the presence of a corresponding exonic breakpoint detected by the cfRNA-based method described herein, and FIG. 9C is a diagrammatic representation of the Arriba tool showing the CCDC6-RET gene fusion. [Figure 9B] See legend to Figure 9A. [Figure 9C] See legend to Figure 9A. [Figure 10A]FIG. 10, comprising FIG. 10A and FIG. 10B, shows detection of BCR-ABL1 gene fusion in RNA samples extracted from peripheral blood cell fractions of acute lymphoblastic leukemia clinical samples using the cfRNA-based methods described herein, where FIG. 10A is an IGV graph report showing the BCR-ABL1 gene fusion and FIG. 10B is a diagrammatic representation of the Arriba tool showing the BCR-ABL1 gene fusion. [Figure 10B] See legend to Figure 10A. [Figure 11] To determine the limit of detection sensitivity of the cfRNA-based methods described herein, quantification of the copy number of EML4-ALK fusion transcript per nanogram of RNA from the H2228 cell line is shown. [Figure 12] The method described herein is used to detect and quantify the expression of control genes and other target genes in cfRNA from both cancer and healthy samples.The table (upper panel) describes the cfRNA input amount of each sample tested, including two sample repeats with different input cfRNA amounts.The expression heat map (lower panel) demonstrates the distribution of expression read counts derived from the method described herein per sample.Fusion detection in the same sample is feasible, as depicted in Figure 8 and Figure 9, respectively, for C_20-347 and C_20-146, which were simultaneously positive for CCDC6-RET and TMPRSS2-ERG fusions. [Figure 13A] FIG. 13, including FIG. 13A, 13B, and 13C, shows the identification of actionable driver fusions in untreated lung cancer cases using cfRNA using the methods described herein, where FIG. 13A shows detection of LMNA-NTRK1 fusion, FIG. 13B shows detection of CD74-NRG1 fusion, and FIG. 13C shows detection of ETV6-NTRK3 fusion in cfRNA in three lung cancer cases that were negative for the presence of other driver gene mutations in cfDNA, respectively. [Figure 13B] See legend to Figure 13A. [Figure 13C]See legend to Figure 13A. [Figure 14] Figure 14, including Figures 14A and 14B, shows fusion detection in 45 lung cancer cases by cfDNA and cfRNA using the methods described herein, and shows that additional fusions were identified when using cfRNA fractions compared to cfDNA. Clinical samples processed simultaneously with cfRNA and cfDNA are compared for fusion detection, where Figure 14A shows the concordance of fusion detection based on cfDNA and cfRNA, showing that cfRNA identified additional fusions in 5 cases and cfDNA missed one detectable fusion. There were 12 cases in which concordant fusions were detected by both methods, and Figure 14B describes the range of fusions detected by both cfDNA and cfRNA methods, or one of the two methods, and the detection of multiple simultaneous fusions detected by cfRNA. (* = fusions detected by both cfDNA and cfRNA). [Figure 15] Illustrates the typical library profile of cfRNA samples converted into sequencing library seen by High Sensitivity DNA Screentape.Multiple peaks over 200 base pairs correspond to multiple products, including potential fusion products, control gene products, and other gene expression products that contain multiple forward and reverse primers.Qualified libraries will have prominent peaks with sizes over 200 base pairs. [Figure 16] IGV graphical report showing detection of an 18 bp deletion in RNA extracted from FFPE lung tumor tissue using the cfRNA-based method described herein. Expression of the EGFR c.2240_2257del p.L747_P753delinsS mutant transcript (containing the deletion) was supported by 4266 reads. [Figure 17]IGV graphical report showing detection of a single nucleotide mutation in cfRNA extracted from plasma of a metastatic lung cancer patient using the cfRNA-based method described herein. Expression of the EGFR c.2573T>G p.L858R mutant transcript (containing the single nucleotide mutation) was supported by 112 reads. [Figure 18] FIG. 18, comprising FIG. 18A and 18B, shows detection of expressed transcripts containing single nucleotide mutations, insertions (e.g., duplications) or deletion mutations using the cfRNA-based methods described herein, where FIG. 18A shows a single nucleotide mutation, insertion or deletion mutation detected in tissue RNA extracted from an FFPE tumor sample and FIG. 18B shows a single nucleotide mutation detected in cfRNA extracted from plasma. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Detailed Description The disclosed methods allow for the detection of genomic alterations and gene expression in biological samples, as well as the quantification of gene expression levels of RNA (such as cfRNA) for the purpose of detecting non-invasive cancer, predicting prognosis, and predicting treatment response. The present disclosure describes a highly multiplexed amplicon-based NGS-based method that involves tagging individual cfRNA molecules with barcode sequences and designing amplicons optimized to fit the fragmented nature of cfRNA. The methods described herein can be applied to circulating nucleic acid extracts containing both cfDNA and cfRNA, and can simultaneously detect and quantify fusion RNA transcripts and gene expression in nucleic acid extract samples. The present disclosure extends the applicability of cfRNA with a novel amplicon-based NGS assay that combines fusion detection and gene expression monitoring. In hybridization-based library preparation methods, sequence-specific bias due to enzymatic ligation during the library construction phase leads to biased representation of transcripts, especially during the analysis of low amounts of input RNA. Targeted NGS assays, such as hybridization capture or amplicon sequencing, can enable sensitive quantification of cfRNA. Although targeted NGS-based methods have a high conversion efficiency compared to whole transcriptome analysis, they have drawbacks such as cost and labor.

[0012] In a first aspect, the present disclosure refers to a method of detecting genomic alterations and / or detecting gene expression and / or quantifying levels of gene expression using RNA in a biological sample, comprising the steps of: (a) extracting RNA from a biological sample and converting the RNA into complementary DNA (cDNA); (b) performing multiplex PCR reactions on the converted cDNA; (I) Target genes that can cause genomic changes a plurality of forward and reverse primer pairs specific for a plurality of Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes capable of causing genomic alterations forwardThe primers are aligned to approximately 50 base pairs of the exon junction of each target gene that can cause genomic alterations. Upstream is complementary to the sequence located at Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes capable of causing genomic alterations Reverse The primers are aligned to approximately 50 base pairs of the exon junction of each target gene that can cause genomic alterations. downstream is complementary to the sequence located at Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes capable of causing genomic alterations Reverse The primer comprises a barcode sequence at its 5' end, and the barcode sequence of each reverse primer corresponding to each target gene capable of causing genomic alteration is different; multiple forward and reverse primer pairs, and / or (II) Control housekeeping genes a plurality of forward and reverse primer pairs specific for a plurality of (i) each of a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes; forward the primers are complementary to sequences spanning the exon-exon junctions of each control housekeeping gene; Each of multiple forward and reverse primer pairs specific for multiple control housekeeping genes Reverse The primers were designed to span approximately 100 base pairs of sequence across the exon-exon junctions of each control housekeeping gene. downstream is complementary to the sequence Each of multiple forward and reverse primer pairs specific for multiple control housekeeping genes Reverse the primers contain a barcode sequence at their 5' ends, and the barcode sequence of each reverse primer corresponding to each control housekeeping gene is different; (ii) each of a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes; Reverse the primers are complementary to sequences spanning the exon-exon junctions of each control housekeeping gene; Each of multiple forward and reverse primer pairs specific for multiple control housekeeping genes forward The primers were designed to span approximately 100 base pairs of sequence across the exon-exon junctions of each control housekeeping gene. downstream is complementary to the sequence Each of multiple forward and reverse primer pairs specific for multiple control housekeeping genes Reverse the primers contain a barcode sequence at their 5' ends, and the barcode sequence of each reverse primer corresponding to each control housekeeping gene is different; (iii) each of a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes; forward and each Reverse the primers are complementary to consecutive sequences spanning the exon-exon junctions of each control housekeeping gene; each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of control housekeeping genes comprises a barcode sequence at its 5' end, and the barcode sequence of each reverse primer corresponding to each control housekeeping gene is different; multiple forward and reverse primer pairs, and / or (III) Target genes related to protein expression A plurality of primer sets specific to a plurality of each primer set comprises a plurality of forward and reverse primer pairs specific for each target gene associated with protein expression; (i) each of a plurality of forward and reverse primer pairs specific for each target gene associated with protein expression; forward the primers are complementary to sequences spanning exon-exon junctions of each target gene associated with protein expression; Each of multiple forward and reverse primer pairs specific for each target gene associated with protein expression Reverse Primers were designed to span approximately 100 base pairs of sequence across exon-exon junctions of each target gene relevant for protein expression. downstreamis complementary to the sequence Each of multiple forward and reverse primer pairs specific for each target gene associated with protein expression Reverse the primers contain a barcode sequence at their 5' ends, and the barcode sequences of each reverse primer corresponding to each target gene associated with protein expression are different; (ii) each of a plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression; Reverse the primers are complementary to sequences spanning exon-exon junctions of each target gene associated with protein expression; Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression. forward Primers were designed to span approximately 100 base pairs of sequence across exon-exon junctions of each target gene relevant for protein expression. downstream is complementary to the sequence Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression. Reverse the primers contain a barcode sequence at their 5' ends, and the barcode sequence of each reverse primer corresponding to each target gene associated with protein expression is different; (iii) each of a plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression; forward and each Reverse the primers are complementary to contiguous sequences spanning exon-exon junctions of each target gene involved in protein expression; Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression. Reverse the primers contain a barcode sequence at their 5' ends, and the barcode sequences of each reverse primer corresponding to each target gene associated with protein expression are different; Multiple primer sets thereby generating a plurality of amplicons; (c) purifying the plurality of amplicons from step (b); (d) amplifying the purified products from step (c) by using universal indexed adapter primers to generate a sequencing library; (e) purifying the sequencing library obtained from step (d); (f) subjecting the purified sequencing library from step (e) to multiplex sequencing on a next-generation sequencing platform to obtain multiple sequencing reads; (g) deriving a consensus read for each sequence from the multiple sequencing reads obtained from step (f); (h) performing a sequence alignment of the consensus reads obtained from step (g) with a reference genome, (I) if the sequence alignment results in a partial alignment of an exon from a first gene with a reference genome and a partial alignment of an exon from a second gene with a reference genome, then (i) determining a sequence alignment as a split read; (ii) counting / enumerating the number of split reads from step (h)(I)(i) that support the fusion junction; and (iii) if the number of split reads from step (h)(I)(ii) is two or more, then determining the first gene and the second gene as fusion partners; (II) If the sequence alignment results in an alignment of the control housekeeping genes with the reference genome, then: (i) determining the sequence alignment to a consensus read of a control housekeeping gene; (ii) counting / enumerating the consensus read pairs of the control housekeeping genes from step (h)(II)(i); and (iii) determining the level of gene expression of a control housekeeping gene; (III) if the sequence alignment results in an alignment of the target gene with the reference genome that is relevant for protein expression; (i) determining the sequence alignment as a consensus read of the target gene associated with protein expression; (ii) counting / enumerating consensus read pairs of the target genes associated with protein expression from step (h)(III)(i); and (iii) determining the level of gene expression of the target gene related to protein expression; step, (i) determining the presence or absence of genomic alterations and / or determining the presence or absence of gene expression and / or quantifying the level of gene expression based on the sequence alignment from step (h).

[0013] In one example, the disclosed method is used to detect genomic changes of RNA in a biological sample. For example, the method can be used to detect known and unknown fusions and to quantify them relative to the expression levels of control housekeeping genes in a given sample. In another example, the disclosed method is used to detect gene expression of RNA in a biological sample. In yet another example, the disclosed method is used to quantify the gene expression level of RNA in a biological sample. In a further example, the disclosed method is used to simultaneously detect genomic changes of RNA and detect gene expression of RNA in a biological sample. In a further example, the disclosed method is used to simultaneously detect genomic changes of RNA and detect gene expression of RNA in a biological sample. In a further example, the disclosed method is used to simultaneously detect genomic changes of RNA, detect gene expression of RNA, and quantify gene expression of RNA in a biological sample. In a further example, the disclosed method is used to simultaneously detect genomic changes of RNA, detect gene expression of RNA, and quantify gene expression of RNA in a biological sample.

[0014] In one example, the disclosed method is used to detect genomic alterations of cfRNA in biological samples.For example, the method can be used to detect known and unknown fusions and to quantify them against the expression level of control housekeeping genes in a given sample.In another example, the disclosed method is used to detect gene expression of cfRNA in biological samples.In yet another example, the disclosed method is used to quantify gene expression level of cfRNA in biological samples.In a further example, the disclosed method is used to simultaneously detect genomic alterations of cfRNA and detect gene expression of cfRNA in biological samples.In a further example, the disclosed method is used to simultaneously detect genomic alterations of cfRNA and detect gene expression of cfRNA in biological samples.In a further example, the disclosed method is used to simultaneously detect genomic alterations of cfRNA, detect gene expression of cfRNA, and quantify gene expression of cfRNA in biological samples.In a further example, the disclosed method is used to simultaneously detect genomic alterations of cfRNA, detect gene expression of cfRNA, and quantify gene expression of cfRNA in biological samples.

[0015] In one example, the design of primers to capture fusion transcripts has two main features - 1) the presence of a random barcode sequence in the downstream (downstream to the target gene (e.g., fusion) transcript) primer to individually tag each copy of the RNA transcript, if present, and 2) the positioning of each primer approximately 50 base pairs from each exon junction in the panel, such that the expected total amplicon length is close to 90-110 base pairs. This was done to meet the observed cfRNA size distribution of samples, which peaks at 110-120 nucleotides.

[0016] In one example, the method disclosed in step (b)(I) Target genes that can cause genomic changes Multiple forward and reverse primer pairs specific for multiple of are designed as shown in FIG. each forward primer of a plurality of forward and reverse primer pairs specific to a plurality of target genes capable of causing a genomic alteration is complementary to a sequence located about 50 base pairs upstream of an exon junction of each of the target genes capable of causing a genomic alteration; each reverse primer of the plurality of forward and reverse primer pairs specific to a plurality of target genes capable of causing a genomic alteration is complementary to a sequence located about 50 base pairs downstream of an exon junction of each of the target genes capable of causing a genomic alteration; Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes capable of causing genomic alterations Reverse The primers contain a barcode sequence at their 5' ends, and the barcode sequence of each reverse primer corresponding to each target gene capable of causing genomic alteration is different. In one example, the method disclosed in step (b)(II) Control housekeeping genes A plurality of forward and reverse primer pairs specific to a plurality of (i) As shown in Figure 2A (left), multiple forward and reverse primer pairs specific for multiple control housekeeping genes were used. forward the primers are complementary to sequences spanning the exon-exon junctions of each control housekeeping gene; Each of multiple forward and reverse primer pairs specific for multiple control housekeeping genes Reverse The primers were designed to span approximately 100 base pairs of sequence across the exon-exon junctions of each control housekeeping gene. downstream is complementary to the sequence Each of multiple forward and reverse primer pairs specific for multiple control housekeeping genes Reverse the primers contain a barcode sequence at their 5' ends, and the barcode sequence of each reverse primer corresponding to each control housekeeping gene is different; (ii) each reverse primer of a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes is complementary to a sequence spanning an exon-exon junction of each of the control housekeeping genes; Each of multiple forward and reverse primer pairs specific for multiple control housekeeping genes forward The primers were designed to span approximately 100 base pairs of sequence across the exon-exon junctions of each control housekeeping gene. downstream is complementary to the sequence each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of the control housekeeping genes comprises a barcode sequence at its 5' end, and the barcode sequence of each reverse primer corresponding to each control housekeeping gene is different; (iii) each forward and each reverse primer of a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes is complementary to a contiguous sequence spanning an exon-exon junction of each of the control housekeeping genes; Each reverse primer of a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes comprises a barcode sequence at its 5' end, and the barcode sequence of each reverse primer corresponding to each control housekeeping gene is different. In one example, the method disclosed in step (b)(III) Target genes related to protein expression A plurality of primer sets specific to a plurality of each primer set comprises a plurality of forward and reverse primer pairs specific for each target gene associated with protein expression; (i) As shown in Figure 2A (right), each of multiple forward and reverse primer pairs specific for each target gene related to protein expression was forward the primers are complementary to sequences spanning exon-exon junctions of each target gene associated with protein expression; Each of multiple forward and reverse primer pairs specific for each target gene associated with protein expression Reverse Primers were designed to span approximately 100 base pairs of sequence across exon-exon junctions of each target gene relevant for protein expression. downstream is complementary to the sequence Each of multiple forward and reverse primer pairs specific for each target gene associated with protein expression Reverse the primers contain a barcode sequence at their 5' ends, and the barcode sequences of each reverse primer corresponding to each target gene associated with protein expression are different; (ii) each reverse primer of the plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression is complementary to a sequence spanning an exon-exon junction of each target gene associated with protein expression; Each of a plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression. forward Primers were designed to span approximately 100 base pairs of sequence across exon-exon junctions of each target gene relevant for protein expression. downstream is complementary to the sequence each reverse primer of a plurality of pairs of forward and reverse primers specific to a plurality of target genes associated with protein expression comprises a barcode sequence at its 5' end, and the barcode sequence of each reverse primer corresponding to each target gene associated with protein expression is different; (iii) each forward and each reverse primer of the plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression is complementary to a contiguous sequence spanning an exon-exon junction of each target gene associated with protein expression; Each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of target genes associated with protein expression comprises a barcode sequence at its 5' end, and the barcode sequence of each reverse primer corresponding to each target gene associated with protein expression is different.

[0017] In one example, as shown in FIG. 2B, the method disclosed in step (b)(II) Control housekeeping genes The forward primer of the plurality of forward and reverse primer pairs specific to a plurality of the first exon is complementary to a sequence in the first exon and is disclosed in step (b)(II). Control housekeeping genesThe reverse primer of the multiple forward and reverse primer pairs specific to multiple of is complementary to a sequence in the second exon, where the first exon and the second exon are interrupted by an intron of more than 5000 base pairs in length, thereby avoiding the unintended amplification of any genomic DNA during the multiple multiplex PCR reactions.

[0018] In one example, the method disclosed in step (b)(II) Control housekeeping genes At least one of the primers of each forward and reverse primer pair of the plurality of forward and reverse primer pairs specific to a plurality of exon-exon junctions spans an exon-exon junction. In one example, Target genes related to protein expression At least one of the primers of each forward and reverse primer pair of the plurality of forward and reverse primer pairs specific to a plurality of exon-exon junctions spans an exon-exon junction. In one example, Control housekeeping genes and / or at least one of the primers of each forward and reverse primer pair of a plurality of forward and reverse primer pairs specific to a plurality of Target genes related to protein expression At least one of the primers of each forward and reverse primer pair of the plurality of forward and reverse primer pairs specific to a plurality of exon-exon junctions spans an exon-exon junction. In one example, Control housekeeping genes and / or a forward or reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of Target genes related to protein expression In another example, the forward or reverse primer of the forward and reverse primer pairs specific to a plurality of exon-exon junctions spans an exon-exon junction. Control housekeeping genes and / or both the forward and reverse primers of the multiple forward and reverse primer pairs specific to multiple of Target genes related to protein expressionBoth forward and reverse primers of the multiple forward and reverse primer pairs specific for multiple exons span exon-exon junctions, where the exons are approximately 100 base pairs in length.

[0019] In one example, the method disclosed in step (b)(I) Target genes that can cause genetic changes Each of a plurality of forward and reverse primer pairs specific to a plurality of Reverse Primers, as disclosed in step (b)(II) Control housekeeping genes Each of a plurality of forward and reverse primer pairs specific to a plurality of Reverse The primers, as well as the nucleic acid sequence disclosed in step (b)(III) Target genes related to protein expression Each of a plurality of forward and reverse primer pairs specific to a plurality of Reverse The primer comprises a barcode sequence at its 5' end, where each barcode sequence is different. As used herein, the term "barcode sequence" refers to an encoded molecule or barcode that contains a variable amount of information within a nucleic acid sequence. For example, a barcode sequence is a tag that can be read using any of a variety of sequence identification techniques, such as nucleic acid sequencing, probe hybridization-based assays, and the like. In some examples, a barcode sequence is used in the methods described herein to tag different converted cDNA sequences of a target region of a sample, such that once the barcode sequence is tagged to the converted DNA sequence of the target region, each different converted cDNA sequence of the target region will have a unique barcode sequence attached to it, and the converted cDNA sequence of the target region from the sample is read.

[0020] The barcode sequence allows for pooled analysis of multiple unique target sequences, where sequence information obtained from the pool can later be attributed to each starting target sequence. That is, after the amplification process, the barcode sequence is used to group amplicons to form a family of amplicons with the same barcode sequence. In some instances, the barcode sequence is an overhang that is not complementary to any sequence within the target region. Each ReverseThe primers carry at their 5' ends barcode sequences that are randomly assigned as disclosed herein, allowing for unique tagging of individual cDNA molecules at the stage of forming the sequencing library.

[0021] In one example, the barcode sequence is an oligonucleotide that contains 10-16 random nucleotides, or 10-15 random nucleotides, or 10-13 random nucleotides, or 10 random nucleotides, or 11 random nucleotides, or 12 random nucleotides, or 13 random nucleotides, or 14 random nucleotides, or 15 random nucleotides, or 16 random nucleotides. In one example, the barcode sequence is an oligonucleotide that contains 10-16 random nucleotides. In one example, the barcode sequence is an oligonucleotide that contains 10 random nucleotides. In one specific example, the barcode sequence is an oligonucleotide that contains 10 random nucleotides, which can be represented as NNNNNNNNNN (SEQ ID NO: 615).

[0022] In one example, the typical length of each forward primer of the plurality of forward and reverse primer pairs disclosed in step (b), excluding the barcode sequence and the partial adapter sequence, is about 20 base pairs. In one example, the typical length of each reverse primer of the plurality of forward and reverse primer pairs disclosed in step (b), excluding the barcode sequence and the partial adapter sequence, is about 20 base pairs. In one example, the typical length of each forward primer of the plurality of forward and reverse primer pairs disclosed in step (b), including the barcode sequence and the partial adapter sequence, is about 45 base pairs, where the length of the barcode sequence is about 10 base pairs, and where the length of the partial adapter sequence is about 20 base pairs. In one example, the typical length of each reverse primer of the plurality of forward and reverse primer pairs disclosed in step (b), including the barcode sequence and the partial adapter sequence, is about 45 base pairs, where the length of the barcode sequence is about 10 base pairs, and where the length of the partial adapter sequence is about 20 base pairs.

[0023] In one example, the biological sample comprises RNA.In one example, the RNA is cfRNA.In one example, cfRNA exists freely in the biological sample, and can be directly converted into cDNA as disclosed in step (a) of the disclosed method.

[0024] In one example, cfRNA is extracted from a biological sample prior to step (a) of the disclosed method. In a further example, RNA may be naturally encapsulated within a cell and needs to be extracted prior to step (a) of the disclosed method. In one example, the cell may be any type of cell in the body. In one example, the cell is from bone, epithelium, cartilage, adipose tissue, nerve, muscle, connective tissue, esophagus, stomach, liver, gallbladder, pancreas, adrenal gland, bladder, gallbladder, large intestine, small intestine, kidney, liver, pancreas, colon, stomach, thymus, spleen, brain, spinal cord, heart, lung, eye, cornea, skin, pancreatic islet tissue or organ. In one example, the cell may be a cancer cell, stem cell, endothelial cell, or fat cell. In one example, the cell is a blood cell. The blood cell may be a white blood cell, or a platelet. In one example, the cell is selected from a cancer cell known to carry a genomic alteration. In one example, the cell is selected from a cancer cell line that is known to carry a fusion gene.In one example, the cancer cell line that carries a fusion gene can include, but is not limited to, CRL-9591, H-2228, CRL-2724, VCaP, CRL-5813, etc.Various methods for RNA extraction are known in the art and can be used for the purpose of the disclosed method.Various methods for RNA extraction are known in the art and can be used for the purpose of the disclosed method.In one example, cfRNA is extracted from the biological sample prior to step (a) using a kit, such as, but not limited to, Zymo Quick-cfRNA Serum & Plasma Kit (Zymo Research), NextPrep™ Magnazol™ cfRNA Isolation Kit (PerkinElmer), Isopure Plasma cfDNA / RNA Isolation Kit (Aline Biosciences), QIAmp Circulating Nucleic Acid Kit (Qiagen), QIAamp ccfDNA / RNA Kit (Qiagen), MagMAX™ Cell-Free Total Nucleic Acid Isolation Kit (Applied Biosystems), or the like.

[0025] In one example, the RNA extracted from cells is subjected to ultrasonic treatment, thereby making it closer to the size of cfRNA.In another example, ultrasonic treatment is performed using Covaris, Qsonica, Diagenode Bioruptor, etc.In another example, the RNA extracted from cells is subjected to heat and divalent cation-based fragmentation.In yet another example, fragmentation is performed using NEBNext® Magnesium RNA Fragmentation Module.

[0026] In one example, biological sample comprises both cfRNA and cfDNA.As used herein, cfDNA refers to the unencapsulated DNA that exists freely in the liquid sample disclosed herein and is not contained in cells.The presence of rearranged long introns prevents rearranged cfDNA from forming sequenceable products.

[0027] In the disclosed method, cfRNA present freely in or extracted from a biological sample is first converted to cDNA as disclosed in step (a) of the method of the first aspect. In one example, cfRNA is converted to cDNA by reverse transcription. The term "reverse transcription" and its grammatical variants as used herein refer to the enzyme-mediated synthesis of DNA molecules from an RNA template. The resulting DNA is known as complementary DNA (cDNA) and can be used as a template for PCR amplification. Methods of reverse transcription typically involve the use of non-target specific primers (random primers) and are well known in the art. In one example, cfRNA is converted to cDNA using a reverse transcription kit that includes a reverse transcriptase enzyme and multiple random primers. In one example, the random primer is a 6-mer primer, a 7-mer primer, an 8-mer primer, a 9-mer primer, or a combination thereof. In one example, the random primer is a 6-mer (hexamer / hexanucleotide) primer. In one example, the reverse transcription kit is selected from, but not limited to, High-Capacity cDNA Reverse Transcription Kit (Thermo Fisher Scientific), SuperScript IV One-Step RT-PCR System (Invitrogen), and the like.

[0028] In one example, the biological sample comprising RNA is a liquid sample, a tissue sample, or a cell sample. In yet another example, the tissue sample is a frozen tissue sample or a fixed tissue sample. In another example, the fixed tissue sample is a formalin-fixed paraffin-embedded (FFPE) tissue sample. In another example, the liquid sample is a body fluid. In one example, the body fluid is selected from the group consisting of blood, bone marrow, cerebrospinal fluid, peritoneal fluid, pleural fluid, lymphatic fluid, ascites, serous fluid, sputum, tears, stool, urine, saliva, lacteal fluid, gastric juice, and pancreatic juice. In one example, the body fluid is blood. In one example, the blood is plasma.

[0029] In another example, the biological sample is obtained from a subject having and / or suspected of having a disease. In another example, the disease is cancer. In yet another example, the cancer is selected from the group consisting of leukemia, lung cancer, colorectal cancer, breast cancer, pancreatic cancer, prostate cancer, nasopharyngeal cancer, liver cancer, bile duct cancer, esophageal cancer, urothelial cancer, and gastrointestinal cancer. In one example, the cancer is early stage cancer. In another example, the cancer is late stage cancer or metastatic cancer. In one example, the cancer is selected from the group consisting of metastatic prostate cancer, metastatic lung cancer, metastatic breast cancer, and leukemia.

[0030] In one example, the genomic alteration detected using the disclosed method includes structural rearrangement. In one example, the term "rearrangement" refers to a rearrangement of the order of a section of DNA. In one example, the structural rearrangement is a fusion, such as a gene fusion. In one example, the term "fusion" refers to a structural alteration caused by a structural rearrangement, such as an inter- or intra-chromosomal rearrangement. In one example, the structural rearrangement may include, but is not limited to, a deletion, an insertion (such as a duplication), an inversion, a transversion, a translocation, an alternative splicing, and the like. In one example, the structural rearrangement results in the formation of a fusion gene, such as one that is detectable using the disclosed method. In one example, a "deletion" is a sequence alteration in which at least one nucleotide is removed. In one example, a "deletion" is a sequence alteration in which more than 10 nucleotides are removed. In one example, a "deletion" is a sequence alteration in which more than 20 nucleotides are removed. In one example, a "deletion" is a sequence alteration in which more than 30 nucleotides are removed. In one example, a "deletion" is a sequence change in which more than 40 nucleotides are removed. In one example, a "deletion" is a sequence change in which more than 50 nucleotides are removed. In one example, a "deletion" can be a "small deletion" in which less than 50 nucleotides are removed. In one example, an "insertion" is a sequence change in which at least one nucleotide is inserted between two nucleotides. In one example, an "insertion" is a sequence change in which more than 10 nucleotides are inserted between two nucleotides. In one example, an "insertion" is a sequence change in which more than 20 nucleotides are inserted between two nucleotides. In one example, an "insertion" is a sequence change in which more than 30 nucleotides are inserted between two nucleotides. In one example, an "insertion" is a sequence change in which more than 40 nucleotides are inserted between two nucleotides. In one example, an "insertion" is a sequence change in which more than 50 nucleotides are inserted between two nucleotides. In one example, an "insertion" can be a "small insertion" in which less than 50 nucleotides are inserted between two nucleotides. In one example, the "insertion" is a "duplication."In one example, a "duplication" is a sequence change in which one or more copies of nucleotides are inserted directly into the 3'-flanking of the original copy. In one example, the term "inversion" refers to a sequence change in which more than one nucleotide replacing the original sequence is the reverse complement of the original sequence. In one example, the term "translocation" refers to a rearrangement of parts between non-homologous chromosomes, which can result in a "fusion". In one example, "alternative splicing" refers to the aberrant splicing of a single gene transcript, which can splice out one or more exons in a sequence from the RNA, juxtaposing normally more distant exons of the same gene. The splicing change involves the same gene, compared to a fusion, which is a definition limited to two genes. In one example, the splicing change includes MET exon 14 skipping, in which exon 14 of the MET gene is spliced ​​out, bringing exon 13 and exon 15 into close proximity, which can be detected using the methods described herein (Figure 14). In one example, the genomic alteration detected by the disclosed method comprises a single nucleotide mutation.In one example, "single nucleotide mutation" refers to a single nucleotide mutation that occurs at a specific position in a genome, which is different from the nucleotide that defines the position in a reference genome.

[0031] In one example, a "housekeeping gene" refers to a highly conserved gene essential for maintaining cell function. In one example, the control housekeeping gene is glucose-6-phosphate isomerase (GPI), FERM domain containing 8 (Multi-8), small nuclear ribonucleoprotein D3 (SNRPD3), proteasome subunit, beta type, 2 (PSMB2), TATA box binding protein (TBP), REL proto-oncogene, NF-kB subunit (REL), synaptosomal associated protein 29 (SNAP29), tubulin gamma complex associated protein 2 (TUBGCP2), receptor accessory protein 5 (REEP5), solute carrier family 4 member 1 adaptor protein (SLC4A1AP), integrin subunit beta 7 (ITGB7), protein-O-mannose kinase (POMK), ER membrane protein complex subunit 7 (EMC7), nuclear autoantigenic sperm protein (NASP), checkpoint with forkhead and ring finger domains (CHFR), ribosomal RNA processing 1 (RRP1), cytosolic iron-sulfur assembly component 1 (CIAO1), pumilio RNA binding family member 1 (PUM1), retention in endoplasmic reticulum sorting receptor 1; These include RER1, serine and arginine rich splicing factor 4 (SRSF4) (see FIG. 12B). The expression of housekeeping genes is considered to be relatively constant between samples. For example, for samples containing the same amount of RNA, the number and expression of housekeeping genes will be similar. For example, for samples containing less RNA, the number and expression of housekeeping genes will be less than that of samples containing more RNA, and vice versa. Therefore, the counting of RNA molecules of housekeeping genes can generally be used for the normalization of RNA molecules of target genes related to genetic modification targets and protein expression.

[0032] In one example, the amount of cfRNA used in the method disclosed herein is at least 6ng.In another example, the amount of cfRNA used in the method disclosed herein is about 6ng to about 100ng, or about 10ng, or about 20ng, or about 30ng, or about 40ng, or about 50ng, or about 60ng, or about 70ng, or about 80ng, or about 90ng, or about 100ng.In one example, the amount of cfRNA used in the method disclosed herein is 20ng to 50ng.

[0033] Then, on the converted cDNA as disclosed in step (b) of the first aspect, a multiplex PCR reaction is carried out as disclosed in (b)(I). Target genes that can cause genomic changes and / or a plurality of forward and reverse primer pairs specific to a plurality of the Control housekeeping genes and / or a plurality of forward and reverse primer pairs specific to a plurality of the Target genes related to protein expression wherein the nucleic acid sequence is determined using a plurality of forward and reverse primer pairs specific for a plurality of Target genes that can cause genomic changes A plurality of forward and reverse primer pairs specific for a plurality of Control housekeeping genes is different from a plurality of Target genes related to protein expression is different from the plural of .

[0034] In one example, the multiplex PCR reactions on the converted cDNA in step (b) are described in step (b)(I). Target genes that can cause genomic changes a plurality of forward and reverse primer pairs specific for a plurality of the Control housekeeping genes and a plurality of forward and reverse primer pairs specific for a plurality of the Target genes related to protein expression In one example, the multiplex PCR reactions on the converted cDNA in step (b) are carried out using a plurality of primer pairs specific to a plurality of the primers disclosed in step (b)(I). Target genes that can cause genomic changesand a plurality of forward and reverse primer pairs specific for a plurality of the Control housekeeping genes In another example, the multiplex PCR reaction on the converted cDNA in step (b) is carried out using multiple forward and reverse primer pairs specific to multiple of the sequences disclosed in step (b)(II). Control housekeeping genes and a plurality of forward and reverse primers specific for a plurality of the Target genes related to protein expression In one example, the multiplex PCR reactions on the converted cDNA in step (b) are carried out using a plurality of primer pairs specific to a plurality of the primers disclosed in step (b)(I). Target genes that can cause genomic changes and a plurality of forward and reverse primer pairs specific for a plurality of the Target genes related to protein expression The assay is carried out using multiple forward and reverse primer pairs specific for multiple of the sequences.

[0035] In one example, multiplex PCR reactions are performed on the converted cDNA using, for example, Platinum SuperFi II DNA Polymerase (Invitrogen), KAPA HiFi DNA Polymerase (Roche), Platinum Taq DNA Polymerase or Platinum SuperFi DNA Polymerase (Invitrogen) and Q5 High-Fidelity DNA Polymerase (NEB).

[0036] In one example, the multiplexed PCR reactions performed on the converted cDNA include 3 to 15 PCR cycles. In one example, the PCR amplification includes 3 PCR cycles. In one example, the PCR amplification includes 4 PCR cycles. In one example, the PCR amplification includes 5 PCR cycles. In one example, the PCR amplification includes 6 PCR cycles. In one example, the PCR amplification includes 7 PCR cycles. In one example, the PCR amplification includes 8 PCR cycles. In one example, the PCR amplification includes 9 PCR cycles. In one example, the PCR amplification includes 10 PCR cycles. In one example, the PCR amplification includes 11 PCR cycles. In one example, the PCR amplification includes 12 PCR cycles. In one example, the PCR amplification includes 13 PCR cycles.

[0037] In one example, the method disclosed in step (b)(I) Target genes that can cause genomic changes The number of forward and reverse primer pairs specific to the plurality of is at least 100. In another example, Target genes that can cause genomic changes The number of the multiple forward and reverse primer pairs specific to the multiple is 100 to 2000. In one example, Target genes that can cause genomic changes The number of forward and reverse primer pairs specific for a plurality of is between 200 and 1900, or between 300 and 1800, or between 400 and 1700, or between 500 and 1600, or between 600 and 1500, or between 700 and 1400, or between 800 and 1300, or between 900 and 1200, or between 1000 and 1100. In one example, Target genes that can cause genomic changesThe number of the plurality of forward and reverse primer pairs specific for a plurality of is about 100, about 200, about 300, or about 400, or about 500, or about 600, or about 700, or about 800, or about 900, or about 1000, or about 1100, or about 1200, or about 1300, or about 1400, or about 1500, or about 1600, or about 1700, or about 1800, or about 1900, or about 2000. In one example, Target genes that can cause genomic changes There is no upper limit to the number of multiple forward and reverse primer pairs specific for multiple of.

[0038] In one example, the method disclosed in step (b)(II) Control housekeeping genes The number of the multiple forward and reverse primer pairs specific to the multiple is at least 20. In one example, Control housekeeping genes The number of the multiple forward and reverse primer pairs specific to the multiple is 20 to 300. In one example, Control housekeeping genes The number of the plurality of forward and reverse primer pairs specific for a plurality of is between 30 and 290, or between 40 and 280, or between 50 and 260, or between 60 and 250, or between 70 and 240, or between 80 and 230, or between 90 and 220, or between 100 and 210, or between 110 and 200, or between 120 and 190, or between 130 and 180, or between 140 and 170. In one example, the method disclosed in step (b)(II) Control housekeeping genes The number of the plurality of forward and reverse primer pairs specific for a plurality of is about 20, or about 30, or about 40, or about 50, or about 60, or about 70, or about 80, or about 90, or about 100, or about 110, or about 120, or about 130, or about 140, or about 150, or about 160, or about 170, or about 180, or about 190, or about 200, or about 210, or about 220, or about 230, or about 240, or about 250, or about 260, or about 270, or about 280, or about 290, or about 300. In one example, Control housekeeping genes There is no upper limit to the number of multiple forward and reverse primer pairs specific for multiple of.

[0039] In one example, the method disclosed in step (b)(III) Target genes related to protein expression The number of the multiple forward and reverse primer pairs specific to the multiple is at least 10. In one example, Target genes related to protein expression The number of the multiple forward and reverse primer pairs specific to the multiple is 10 to 1700. In one example, the method disclosed in step (b)(III) Target genes related to protein expression The number of the multiple forward and reverse primer pairs specific to the multiple is between 10 and 1700, or between 100 and 1600, or between 200 and 1500, or between 300 and 1400, or between 400 and 1300, or between 500 and 1200, or between 600 and 1100, or between 700 and 1000. In one example, the method disclosed in step (b)(III) Target genes related to protein expression The number of the plurality of forward and reverse primer pairs specific for a plurality of is about 10, or about 100, or about 200, or about 300, or about 400, or about 500, or about 600, or about 700, or about 800, or about 900, or about 1000, or about 1100, or about 1200, or about 1300, or about 1400, or about 1500, or about 1600, or about 1700. In one example, the method disclosed in step (b)(III) Target genes related to protein expression There is no upper limit to the number of multiple forward and reverse primer pairs specific for multiple of.

[0040] In another example, the maximum total number of the plurality of forward and reverse primer pairs in the multiplex PCR reaction is about 4000, wherein the Target genes that can cause genomic changes wherein the number of forward and reverse primer pairs specific to a plurality of the above is about 2000, Control housekeeping genes wherein the number of forward and reverse primer pairs specific to a plurality of the above is about 300, and wherein Target genes related to protein expressionThe number of multiple forward and reverse primer pairs specific to multiple of is about 1700.

[0041] In one example, Target genes that can cause genomic changesThe plurality of includes an exon from a gene known to undergo fusions fused to an exon from a partner gene of the gene known to undergo fusions. In one example, the genes known to undergo fusions are ALK receptor tyrosine kinase, RET proto-oncogene, ROS proto-oncogene 1, fibroblast growth factor receptor 1 (FGFR1), fibroblast growth factor receptor 2 (FGFR2), fibroblast growth factor receptor 3 (FGFR3), neurotrophic receptor tyrosine kinase 1 (NTRK1), neurotrophic receptor tyrosine kinase 2 (NTRK2), neurotrophic receptor tyrosine kinase 3 (NTRK3), neuregulin 1 (NRG1), B-Raf proto-oncogene, serine / threonine kinase (BRAF), transmembrane serine protease 2. (TMPRSS2), MET proto-oncogene, receptor tyrosine kinase (MET), epidermal growth factor receptor (EGFR), estrogen receptor 1 (ESR1), platelet-derived growth factor receptor alpha (PDGFRA), androgen receptor (AR), BCR activator of RhoGEF and GTPases (BCR), core binding factor subunit beta (CBFB), lysine methyltransferase 2A (KMT2A), nucleophosmin 1 (NPM1), PML nucleolar scaffold (PML), and RUNX family transcription factor 1 (RUNX1).In one example, the partner genes of genes known to undergo fusions include EMAP-like 4 (EML4), kinesin family member 5B (KIF5B), coiled-coil domain containing 6 (CCDC6), CD74 molecule (CD74), transforming acidic coiled-coil containing protein 3 (TACC3), ezrin (EZR), ETS transcription factor ERG (ERG), ArfGAP with GTPase domain, ankyrin repeats and PH domain 3 (AGAP3), A-kinase anchoring protein 9 (AKAP9), KIAA1549, tropomyosin 3 (TMP3), translocation promoter region, nuclear basket protein (TPR), transport from ER to Golgi regulator (TFG), lamin A / C (LMNA), BicC family RNA-binding protein 1 (BICC1), RAD51 recombinase (RAD51 ), CD47 molecule (CD47), Yes1-associated transcription regulator (YAP1), ETS variant transcription factor 1 (ETV1), ETS variant transcription factor 4 (ETV4), ETS variant transcription factor 5 (ETV5), ETS variant transcription factor 6 (ETV6), PAPOLA and CPSF1 interacting factor (FIP1L1), centriolin (CNTRL), ABL proto-oncogene 1, non-receptor tyrosine kinase (ABL1), AF4 / FMR2 family member 1 (AFF1), MDS1 and EVI1 complex locus (MECOM), MLLT3 super elongation complex subunit (MLLT3), myosin heavy chain 11 (MYH11), PBX homeobox 1 (PBX1), retinoic acid receptor alpha (RARA), and RUNX1 partner transcriptional corepressor 1 (RUNX1T1).

[0042] The disclosed method is optimized to generate amplicons having a certain size. The selected length of 90-110 base pairs was considered optimal because shorter amplicon (less than 80 base pairs) products are less effectively retained through the multi-step library preparation method for amplicon sequencing. In one example, the length of the cDNA-derived amplicons in step (b) is 90-110 base pairs. In one example, the length of the cDNA-derived amplicons in step (b) is about 90 base pairs, or about 100 base pairs, or about 110 base pairs.

[0043] The cDNA-derived multiple amplicons in step (b) are then purified as disclosed in step (c) of the first aspect.

[0044] The disclosed method is designed to involve size-based (magnetic bead-based) separation of small primer dimer artifacts that are to be removed and desired products that are to be retained, as well as excess primers that are to be enzymatically digested (e.g., with endonucleases and exonucleases). In one example, DNA purification is performed using agents such as paramagnetic beads. In one example, the paramagnetic beads are selected from the group consisting of AMPure XP beads, SPRI beads, and Dynabeads. In one example, the paramagnetic beads are AMPure XP beads.

[0045] The purified amplicons are then amplified using universal indexed adapter primers to generate a plurality of sequencing libraries as disclosed in step (d) of the first aspect.

[0046] In one example, amplification is performed by using KAPA Hifi HotStart ReadyMix, Phusion U Hot Start DNA Polymerase (Thermo Scientific), ZymoTaq DNA Polymerase (Zymo Research), and Q5U Hot Start High-Fidelity DNA Polymerase (NEB), etc.

[0047] In one example, each universal indexed adapter primer disclosed in step (d) comprises an adapter sequence. In one example, the term "adapter sequence" refers to any nucleotide sequence that can be added to an oligonucleotide of interest to prepare the oligonucleotide of interest for various purposes. The adapter sequence is complementary to multiple oligonucleotides present on the flow cell surface of a sequencing tool, thereby allowing the DNA fragment to be attached to the sequencing tool. In some examples, the adapter sequence allows the oligonucleotide of interest to be sequenced. Sequencing platform-specific adapter sequences are known in the art, including, for example, Illumina P5 / P7 adapter sequences.

[0048] In one example, the universal indexed adapter primer disclosed in step (d) of the method of the first aspect is A forward primer containing the sequence of TIFF2024530807000002.tif11158; and containing a reverse primer containing the sequence of TIFF2024530807000003.tif11159, Here, * " represents a phosphorothioate bond, and where the underlined sequence is the barcode sequence. The formed plurality of sequencing libraries are then purified as disclosed in step (e) of the first aspect.

[0049] In one example, the purification of the multiple sequencing libraries is performed using an agent such as paramagnetic beads. In one example, the paramagnetic beads are selected from the group consisting of AMPure XP beads, SPRI beads, and Dynabeads. In one example, the paramagnetic beads are AMPure XP beads.

[0050] The purified sequencing libraries are then subjected to multiplex sequencing on a next generation sequencing platform, as disclosed in step (f) of the first aspect, to obtain multiple sequencing reads.

[0051] In one example, multiple sequencing libraries are sequenced on a NextSeq 550, NovaSeq 6000, or BGI MGISEQ-2000, DNBSEQ-G400, DNBSEQ-T7.

[0052] In one example, the sequencing libraries are qualified using an Agilent High Sensitivity DNA Screentape and quantified using a KAPA Library Quantification Kit. In one example, the sequencing libraries are qualified by determining the size profile of the sequencing library, which, if successful, will have a typical size profile of multiple prominent peaks above 200 base pairs (e.g., as shown in Figure 15).

[0053] Thereafter, a plurality of consensus reads are derived from each sequence of the plurality of sequencing reads obtained from step (f), as disclosed in step (g) of the first aspect.

[0054] In one example, step (g) of the first aspect further comprises: (g)(I) detecting the presence of a barcode sequence from each sequencing read; (g)(II) performing cluster reassignment on a plurality of sequencing reads having the same barcode sequence to generate a plurality of barcode clusters, wherein each barcode cluster comprises reads having the same barcode sequence from the same amplicon; and (g)(III) performing a consensus call on each barcode cluster to obtain a consensus read for each sequence.

[0055] The derived consensus sequence is aligned to a reference genome as disclosed in step (h) of the first aspect. In one example, the term "reference genome" refers to a DNA sequence known in the art that may be available from a public database. In one example, the term "consensus read" refers to a nucleotide sequence obtained from a consensus call. In one example, the consensus call is performed by identifying the nucleotide at each position for each sequencing result in a subgroup, comparing the identity of the nucleotide at each position across multiple sequencing results, and determining the majority nucleotide at each position. If the count of the majority nucleotide is above a threshold set to determine the majority for a particular position, the assignment of the position is the majority nucleotide. If the count of the majority nucleotide is below this threshold, no assignment of the position is made. The threshold is variable for each position and is a function of the total number of sequencing results corresponding to a particular position.

[0056] In one example, step (h) of the disclosed method further comprises, if the sequence alignment results in a partial alignment of the exons from the first gene to the reference genome and a partial alignment of the exons from the second gene to the reference genome as disclosed in step (h)(I), using the results to (i) determine the sequence alignment as split reads, (ii) count / enumerate the number of split reads from step (h)(I)(i) that support the fusion junction, and (iii) determine the first gene and the second gene as fusion partners if the number of split reads from step (h)(I)(ii) is two or more. In one example, step (h) of the disclosed method further comprises, if the sequence alignment results in an alignment of the control housekeeping gene with the reference genome as disclosed in step (h)(II), using the result to (i) determine the sequence alignment as a consensus read of the control housekeeping gene, and (ii) count / enumerate the consensus read pairs of the control housekeeping gene from step (h)(II)(i) to determine the level of gene expression of the control housekeeping gene. In one example, step (h) of the disclosed method further comprises, if the sequence alignment results in an alignment of the target gene associated with protein expression with the reference genome as disclosed in step (h)(III), using the result to (i) determine the sequence alignment as a consensus read of the target gene associated with protein expression, and (ii) count / enumerate the consensus read pairs of the target gene associated with protein expression from step (h)(III)(i) to determine the level of gene expression of the target gene associated with protein expression.

[0057] In one example, "consensus read pair" refers to the consensus sequence called after collapsing all sequencing reads that contain the same barcode sequence and primer pair. For example, each consensus read pair is presumed to belong to the original RNA molecule that was converted to cDNA. In one example, the counting / enumeration disclosed in step (h) is achieved based on consensus counting based on barcode sequence, where each RNA molecule that contains the same barcode sequence and primer pair combination represents a unique RNA molecule. In one example, all reverse primers of the multiple forward and reverse primer pairs disclosed in step (b) of the first aspect contain barcode sequences. Therefore, all RNA molecules captured by a given barcode sequence and primer pair combination can be detected and counted / enumerated.

[0058] In one example, the alignment of the derived consensus sequences with the reference genome is performed using a sequence alignment tool. In one example, the alignment tool is STAR, HISAT2, bwa, CLC, RSEM, kallisto, salmon, etc.

[0059] The sequence alignment results from step (h) are used to determine the presence or absence of genomic alterations and / or to determine the presence or absence of gene expression and / or to quantify the level of gene expression, as disclosed in step (i) of the first aspect.

[0060] In one example, the disclosed method further comprises visualizing and fusion calling the sequence alignment from step (h)(I). In one example, the visualizing is performed using Integrated Genome Viewer, or Savant Genome Browser, or the like. In one example, the fusion calling is performed using Arriba and Fusion Catcher, or the like.

[0061] In one example, determining the presence or absence of genomic alterations and / or determining the presence or absence of gene expression and / or quantifying the level of gene expression further comprises performing variant calling of the sequence alignment from step (h). In one example, determining the presence or absence of genomic alterations and / or determining the presence or absence of gene expression and / or quantifying the level of gene expression further comprises performing variant calling of the sequence alignment from step (h)(II). In one example, determining the presence or absence of genomic alterations and / or determining the presence or absence of gene expression and / or quantifying the level of gene expression further comprises performing variant calling of the sequence alignment from step (h)(III). In one example, the variant calling step comprises (i) identifying differences between the consensus reads and the reference genome based on the sequence alignment from step (h); and ii) determining the read count of the sequence alignment that includes the genomic alterations. In one example, the variant calling step includes (i) identifying differences between the consensus reads and the reference genome based on the sequence alignment from step (h)(II); and ii) determining the read count of the sequence alignment that includes the genomic variation. In one example, the variant calling step includes (i) identifying differences between the consensus reads and the reference genome based on the sequence alignment from step (h)(III); and ii) determining the read count of the sequence alignment that includes the genomic variation. In one example, the genomic variation is selected from the group consisting of an insertion (e.g., a duplication), a deletion, and a single nucleotide mutation. In one example, the variant calling is performed using Mutect2 and a custom variant caller.

[0062] In one example, the disclosed method of the first aspect is used to simultaneously detect gene expression, structural rearrangement and quantify gene expression in cfRNA from biological sample, and quantify the expression level of genes known to be overexpressed in cancer cells.In one example, the disclosed method of the first aspect is used to simultaneously detect genomic alterations in cfRNA and quantify gene expression in cfRNA from biological sample, and quantify the expression level of target genes that have undergone genomic alterations.In one example, the disclosed method of the first aspect is used to simultaneously detect gene expression and quantify gene expression in cfRNA, and quantify the expression level of target genes that are related to protein expression.

[0063] In one example, statistical modeling techniques used to visualize the expression levels of genes related to protein expression include heat map visualization, principal component analysis, hierarchical clustering, and the like.

[0064] In a second aspect, the present disclosure provides a kit for detecting genomic alterations and / or detecting gene expression and / or quantifying gene expression levels in a biological sample using RNA according to the method of the first aspect, comprising: (a) A target gene capable of causing a genomic alteration as defined in step (b)(I) of the method of the first aspect. a plurality of forward and reverse primer pairs specific for a plurality of (b) A control housekeeping gene as defined in step (b)(II) of the method of the first aspect. a plurality of forward and reverse primer pairs specific to a plurality of (c) A gene associated with protein expression as defined in step (b)(III) of the method of the first aspect. The present invention refers to a kit comprising a plurality of primer sets specific for a plurality of the above.

[0065] In one example, a person skilled in the art could design multiple primer pairs and primer sets in (a), (b) and (c) of the kit of the second aspect based on the disclosure herein, for example, as described in steps (b)(I), (b)(II) and (b)(III) of the method of the first aspect. In one example, multiple primer sets specific to multiple genes associated with protein expression defined in step (b)(III) of the method of the first aspect provided in the kit described herein can be used to determine the presence or absence of genomic alterations. In one example, multiple primer sets specific to multiple genes associated with protein expression defined in step (b)(III) of the method of the first aspect provided in the kit described herein can be used to determine the presence or absence of genomic alterations such as deletions, insertions (e.g., duplications) and single nucleotide mutations. In one example, the multiple primer sets specific to the multiple genes associated with protein expression defined in step (b)(III) of the method of the first aspect provided in the kit described herein can be used to determine the presence or absence of genomic alterations by further performing the step of variant calling described herein. In one example, the genomic alterations can be single nucleotide mutations, insertions (e.g., duplications) or deletions. In one example, the kit for detecting genomic alterations and / or detecting gene expression and / or quantifying gene expression levels of cfRNA in biological samples by the method of the first aspect further comprises a buffer, a reverse transcriptase, a DNA polymerase, and multiple deoxynucleotide triphosphates (dNTPs) for performing multiple multiplex PCR reactions. In some examples, the reagents provided in the kit described herein can be provided in separate containers with components independently distributed in one or more containers. Since the method described herein relates to sequencing (such as high-throughput sequencing), the additional components required in the sequencing process can be easily determined by those skilled in the art.

[0066] As used in this application, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "a primer" includes multiple primers, including mixtures and combinations thereof.

[0067] As used herein, the terms "increase" and "decrease" refer to the relative change in a selected trait or characteristic in a subset of a population compared to the same trait or characteristic present in the entire population. Thus, an increase indicates a change on a positive scale, and a decrease indicates a change on a negative scale. The term "change" as used herein also refers to the difference between a selected trait or characteristic of a subset of an isolated population compared to the same trait or characteristic in the population as a whole. However, this term does not involve an assessment of the difference seen.

[0068] As used herein, the term "about" in the context of a concentration of a substance, a size of a substance, a length of time, or other stated value means ±5% of the stated value, or ±4% of the stated value, or ±3% of the stated value, or ±2% of the stated value, or ±1% of the stated value, or ±0.5% of the stated value.

[0069] Throughout this disclosure, certain embodiments may be disclosed in a range format. It should be understood that the description in range format is merely for convenience and brevity, and should not be construed as an inflexible limitation on the disclosed range. Thus, the description of a range should be considered to specifically disclose all possible subranges as well as individual numerical values ​​within that range. For example, the description of a range such as 1-6 should be considered to specifically disclose subranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, etc., and individual numerical values ​​within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0070] The present disclosure illustratively described herein may be appropriately practiced in the absence of elements or limitations not specifically disclosed herein. Thus, for example, terms such as "comprising", "including", "containing" and the like are to be interpreted broadly and without limitation. Furthermore, the terms and expressions used herein are used as terms of description, not as terms of limitation, and the use of such terms and expressions is not intended to exclude the features shown and described or equivalents of any portion thereof, and it is recognized that various modifications are possible within the scope of the present disclosure as claimed. Thus, although the present disclosure has been specifically disclosed by preferred embodiments and optional features, it should be understood that modifications and variations of the present disclosure embodied herein may be utilized by those skilled in the art, and such modifications and variations are considered to be within the scope of the present disclosure.

[0071] The invention has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also forms part of this disclosure. This includes the generic description of the disclosure with a proviso or negative limitation removing any subject matter from the genus, whether or not the excluded material is specifically described herein.

[0072] Other embodiments are within the scope of the following claims and non-limiting examples. EXAMPLES

[0073] method Sample collection and processing Blood collected in the Streck Cell-free DNA BCT® was shipped at room temperature prior to plasma separation. Briefly, plasma was prepared using a two-step centrifugation process: an initial centrifugation was performed at 1500×g for 10 minutes at 4° C. to separate the plasma. The plasma layer was transferred to a separate tube and centrifuged at 15,000×g for 10 minutes at 4° C. to further remove cellular contaminants and was either immediately processed for nucleic acid extraction or stored at −80° C. until used for extraction. If frozen, plasma was completely thawed at room temperature prior to extraction.

[0074] Plasma cell-free total nucleic acids were extracted using the QIAamp Circulating Nucleic Acids kit (Qiagen). Nucleic acid extracts contained coeluted cfDNA and cfRNA fractions. cfDNA was quantified using a Qubit Fluorometer (Thermo Fisher Scientific) and sized using the Genomic DNA ScreenTape on a 4200 TapeStation (Agilent). cfRNA was quantified and profiled using the Bioanalyzer RNA 6000 Pico kit or the High Sensitivity RNA Screentape on a 4200 Tapestation.

[0075] Design of primers for fusion and expression in sequencing libraries A highly multiplexed amplicon-based NGS assay was designed to capture potential fusions in cfRNA samples. Depending on the expected orientation of the partner exons in the fusion gene, a primer upstream of the exon fusion junction ("forward" primer) or downstream of the fusion junction ("reverse" primer) was designed for the exons of the target genes. Broadly, multiple exon-flanking primers were designed for target genes known to be involved in fusion events in cancer. For all downstream primers, a random 10-base pair barcode sequence was incorporated upstream of the gene-specific sequence for consensus calling and enumeration of unique molecules. A pool of over 300 "forward" primers and over 300 "reverse" primers was prepared. To optimally capture potential fusions known to occur between genes, multiple "upstream" and "downstream" primers were included in the multiplex PCR. The primer design included exons of well-characterized genes known to undergo fusions, and the addition of the barcode sequence primers allowed for accurate enumeration of RNA transcript copies according to the enumeration method (Figure 1).

[0076] For the capture of transcripts corresponding to control genes and other genes whose expression we wished to quantify, we designed primers such that at least one primer of the pair lands on an exon-exon junction or the primer pair is within two exons interposed by an intron of more than 5000 base pairs in length. These primers were also included in the final primer pool. The specificity of cfRNA amplification was verified by performing the entire cfRNA sequencing workflow, but excluding reverse transcriptase during complementary DNA preparation. If reverse transcription was not performed, the sequencing of the intended regions, especially the control and expressed genes, could be attributed to the primers amplifying cfDNA. Any such primers were redesigned to improve specificity for RNA by reducing the 3'exon span of the exon-exon spanning primers. Primer design for expression-related target genes was similar to that of control gene targets, with at least one primer of the primer pair spanning an exon-exon junction, and two or more primer pairs per target gene were designed to cover both the 5' and 3' exons, allowing one or more amplicons to represent a given target gene, to more reliably capture expression of the target gene toward expression. A highly multiplexed primer pool was utilized with multiple upstream and downstream primers, some of which are expected to generate sequenceable targets in the majority of samples depending on expression variability, and some primers are expected to generate products only if the sample is positive for a structural rearrangement that creates a productively expressed fusion gene. Primers further possessed the appropriate extensions required to generate sequenceable libraries with sequencing adapters for Illumina sequencing (Figures 2A and 2B).

[0077] Preparation of libraries for cfRNA sequencing 20–50 ng of cfRNA was converted to complementary DNA (cDNA) using the High-Capacity cDNA Reverse Transcription Kit (Thermo Fisher Scientific) in a total volume of 20 ul with random primers. The converted cDNA was used as a template in a highly multiplexed PCR reaction for target capture using Platinum™ SuperFi II DNA Polymerase (Thermo Fisher Scientific). Briefly, cDNA was combined with primers and DNA polymerase in a single reaction and subjected to 3–15 cycles of PCR with the following conditions: 98°C, 1 min; 60°C, 1 min; 72°C, 1 min, followed by a final extension at 72°C for 5 min. The amplified products were subjected to one round of enzymatic digestion (with exonucleases, ExoI and ExoT) and two rounds of cleanup using 1.8 times the volume of AMPure XP beads and eluted with Buffer EB or nuclease-free water. The purified PCR products were then amplified with universal indexed adapter primers compatible with sequencing on the Illumina platform and primers using KAPA HiFi HotStart ReadyMix. The final amplified libraries were purified with two rounds of 0.8x AMPure XP beads to remove excess adapters and size-select the final sequencing libraries. Libraries were quantified using a High Sensitivity DNA Screentape and quantified using the KAPA Library Quantification Kit. Each library was sequenced on a Nextseq 550 to a depth of 3 million paired-end reads per sample.

[0078] Data analysis A custom pipeline was used to process the FASTQ files. First, sequenced amplicons were identified and labeled in FASTQ files based on the presence of potential primer sequences in the correct orientation, upstream or downstream (from a predefined list of primer sequences based on the panel design) of read 1 and paired read 2. Barcode sequences in read 1 were identified as primers upstream of read 1 and trimmed using cutadapt. The extracted molecular tag sequences were used to derive consensus read sequences of all overlapping reads of sequences identifiable by a predefined primer pair and unique barcode sequence. The consensus reads were then written into a new FASTQ file and aligned to the human genome reference hg19 using STAR aligner. Fusion reads in which non-contiguous regions of the genome were captured within the reads were identified as split reads, and the fusion partners were identified based on the sequence alignment. Furthermore, the presence of split read sequences mapping to two mutual partner genes was confirmed to be captured by primers specific to the identified genes. The number of split reads (read pairs) supporting the fusion junction was enumerated. Visualization and fusion calling were also performed using Arriba and FusionCatcher. At least two supporting split reads were required to call fusion and exon skipping variants (transcript variants). Molecular barcoding ensures high quality sequencing data, eliminating sequencing errors and increasing the confidence of fusion calls.

[0079] Expression level analysis was performed by enumerating consensus read pairs that supported a given amplicon predefined by the primer pair for expression. The number of read pairs was enumerated and tabulated for downstream analysis as control or target genes. Variant calling was performed on the consensus BAM file using Mutect2 and a custom variant caller to identify single nucleotide mutations, insertion mutations and deletion mutations compared to the reference sequence. Expression of mutant transcripts containing single nucleotide mutations, insertions and deletions was quantified based on the number of reads that contained a particular single nucleotide mutation, insertion mutation or deletion mutation and mapped to the intended target region. Expression of wild-type transcripts was quantified based on the number of reads that matched the reference sequence and mapped to the intended target region. The relative expression of each mutation was also determined based on the ratio of mutant reads to total reads.

[0080] result This disclosure describes methods for simultaneously detecting and quantifying clinically relevant genomic and gene expression changes using cfRNA with high sensitivity, specificity, and minimally invasive procedures.

[0081] Validation of cfRNA-based detection assays: relative abundance of cell-free nucleic acids in plasma Total cfRNA concentrations from plasma of healthy individuals and cancer patients were characterized for the presence of cfRNA and analyzed for fragment size distribution using the Bioanalyzer RNA 6000 Pico assay. cfRNA was present in all cancer samples, showing a major peak in size at 110-120 nucleotides, with a second RNA population in the 200-300 nucleotide range (Figures 3A and 3B). In terms of relative abundance, shorter fragments (110-120 nucleotides) were approximately 5-10 times more abundant than larger sized RNA fragments (200-300 nucleotides). cfRNA from healthy individuals also showed the same size distribution pattern, but cfRNA concentrations were significantly lower (Figures 3C and 3D).

[0082] We analyzed total nucleic acid extracts containing cfDNA and cfRNA from plasma of healthy and cancer individuals. Compared to the cfDNA concentration in each extract, cfRNA concentrations were generally low and differed most significantly when cfDNA concentrations exceeded 10 ng / ml plasma (Figure 4).

[0083] Technical validation of cfRNA-based multiplex amplicon sequencing detection using RNA extracted from cancer cell lines The methods described herein demonstrated the ability to detect fusions using RNA extracted from cultured cancer cell lines known to carry fusion genes, such as CRL-9591 (KMT2A-AFF1), H2228 (EML4-ALK), CRL-2724 (RUNX1-RUNX1T1), VCaP (TMPRSS2-ERG) and CRL-5813 (TMPRSS2-ERG). Because RNA from cultured cells is relatively intact compared to plasma cfDNA, cell line RNA was subjected to sonication (using Covaris) to more closely resemble the size of cfRNA. The resulting fraction was used to mimic cfRNA and demonstrate the performance of multiplex amplicon sequencing for the detection of various known fusions (Figures 5A, 5B and 5C). This was used to provide suitable material to mimic cfRNA and demonstrate the performance of multiplex amplicon sequencing for the detection of various known fusions. RNA-based fusion detection was successful in all five cancer cell lines (Figure 7). The resulting multiple sequencing libraries can be qualified using Agilent High Sensitivity DNA Screentape, as shown in Figure 15, which illustrates a typical library profile of a cfRNA sample converted into a sequencing library seen on the High Sensitivity DNA Screentape. The multiple peaks over 200 base pairs correspond to multiple products, including potential fusion products, control gene products, and other gene expression products that include multiple forward and reverse primers. A qualified library will have a prominent peak in size greater than 200 base pairs.

[0084] Sequence alignment to the reference genome showed capture of sequencing reads with partial alignment to the target exon and with partial alignment to another part of the genomic sequence corresponding to the partner gene exon, known as split reads, and confirmed detection of EML4-ALK fusion transcripts in H2228 cell line with 8364 reads supporting the split configuration with as little as 1 ng of fragmented RNA (Figure 6A and 6B). Alignment of split reads showed accurate detection of fusions in cancer cell lines: NCI-H660 cell line (CRL-5813, ATCC), VCaP cell line (CRL-2876, ATCC), human MV-4-11 cell line (CRL-9591, ATCC) and Kasumi-1 (CRL-2724, ATCC) as visualized by the Arriba tool for fusion detection in RNA sequencing data using multiplex amplicon sequencing on fragmented RNA (Figure 7).

[0085] Data comparison between cfDNA-based and cfRNA-based detection assays We tested nucleic acid extracts from plasma of two cancer patients whose cancers had previously been characterized as positive for the fusion using a DNA-based method (Liquid Hallmark).In the first case of metastatic prostate cancer, the TMPRSS2-ERG fusion was detected in cfDNA (using 70 ng of cfDNA) supported by 17 split reads mapping to intronic position chr21:42867069 within TMPRSS2 (intron 2 of TMPRSS2-NM_005656.4) and intronic position chr21:39818058 within ERG (intron 3 of ERG-NM_001291391.1) (Figure 8A). Using the same circulating nucleic acid extract, a fusion in cfRNA (corresponding to only 24 ng of cfRNA) was detected in 4123 supporting split reads, fusing exon 2 of TMPRSS2 (chr 21:42870045) with exon 4 of NM_001291391.1 (or exon 2 of ERG NM_182918.4) (chr 21:39817544) ( Figure 8B ).

[0086] In the second case of metastatic lung cancer, the CCDC6-RET fusion was detected using cfDNA (breakpoints CCDC6 intron 1 (chr10:61623181) and RET intron 11 (chr10:43611035) and cfRNA CCDC6 exon 1 (10:61665879) and RET exon 12 (10:43612031). While cfDNA was detected with 12 supporting reads, the fusion in cfRNA was supported by 13 split reads (Figures 9A and 9B).

[0087] In a third clinical sample from a hematological malignancy (acute lymphoblastic leukemia) with confirmed BCR-ABL1 rearrangement in DNA from peripheral blood cells, RNA was extracted from another fraction of the archived buffy coat and tested with the multiplex amplicon sequencing method described herein. The fusion between exon 14 of BCR and exon 2 of ABL1 was readily detectable in the RNA fraction with abundant 159,106 supporting reads. The large number of supporting reads indicates enrichment of transcripts with BCR-ABL1 fusions in the tested sample (buffy coat RNA) due to increased expression and secondary enrichment in fusion-positive cancer cells (Figure 10A and Figure 10B).

[0088] Additional fusion events are shown in Figure 13, illustrating the identification of driver fusions that may be operative in untreated lung cancer cases using cfRNA using the methods described herein. Figures 13A, 13B, and 13C show the detection of various gene fusion events, namely LMNA-NTRK1, CD74-NRG1, and ETV6-NTRK3 fusions, in cfRNA samples from three lung cancer cases, respectively. These mutations were undetectable using DNA-based assays and appear negative for the presence of other driver gene mutations in cfDNA. Furthermore, when cfDNA and cfRNA were used for fusion detection in 45 lung cancer cases using the described methods, additional fusions were identified when the cfRNA fraction was used compared to cfDNA (Figure 14). Fusion testing was performed using DNA and RNA orthogonally as sample input, and there were 12 cases in which fusion detection was concordant based on cfDNA and cfRNA as sample input. When cfRNA was used instead of cfDNA as sample input, five additional fusions were detected and one missed fusion was not detected. A list of fusion ranges detected by both cfDNA and cfRNA methods, or by one of the two methods, is shown in Figure 14B.

[0089] Detection limit Detection limit is defined as the lowest RNA concentration at which fusion events can be easily detected. The initial determination of detection limit of RNA-based fusion was performed by quantifying the number of EML4-ALK fusion transcripts present in 1 ng of H2228 cell line RNA, where EML4-ALK fusion was easily detectable using the method described herein (Figure 6 and Figure 7). The number of EML4-ALK fusion RNA transcripts was determined to be approximately 13.7 copies per 5 ng of RNA using a qRT-PCR assay specifically designed for EML4-ALK transcripts present in H2228 cells (Figure 11). Therefore, the method described herein was shown to be able to detect up to 2.72 copies of EML4-ALK fusion (in 1 ng of H2228 RNA), suggesting highly sensitive detection of RNA-based fusion.

[0090] Simultaneous detection and quantification of fusion events based on expressed cfRNA Besides the detection of fusions in cfRNA, simultaneous detection of target genes for non-invasive expression monitoring was also performed on cfRNA from cancer and healthy samples. In the same multiplex reaction, primers for 22 control genes and 13 amplicons for 6 genes related to immunotherapy response (CD274, PDCD1, CTLA4, LAG3, HAVCR2 and CD47) were included to perform combined target capture. The expression level of each target was determined based on the read counts mapped to the intended target region. The range of expression levels was visualized in an expression heatmap (Figure 12).

[0091] Typically, healthy samples had very low yields of both cfRNA and cfDNA, so as expected, expression of control and immunotherapy response genes was low across healthy samples. However, among cancer samples, a range of expression patterns was observed, with some samples showing limited expression of nearly all targets, despite comparable amounts of cfRNA material used in the method. The detection reliability and quantitative ability of the method was demonstrated by repeating the same sample with different amounts of cfRNA, which showed an increased number of expression reads, but similarity of patterns between sample repeats (Figure 12). The repeats are represented by C_20.126.1 and C_20.126.2 (repeats of sample 20.126) and C_20.1069.1 and C_20.1069.2 (repeats of sample 20.1069). In the heatmap, two repeats are closest to each other, indicating a higher similarity between two repeats of the same sample compared to the other samples.

[0092] Detection of expressed transcripts containing deletion mutations in RNA samples The method described herein demonstrated the ability to detect the 18-nucleotide deletion in RNA samples extracted from FFPE lung tumor tissues. Expression of the EGFR c.2240_2257del p.L747_P753delinsS mutant transcript (containing the deletion) was detected in 4266 supported reads ( FIG. 16 ).

[0093] Detection of expressed transcripts containing single nucleotide variations in RNA samples The methods described herein demonstrated the ability to detect single nucleotide mutations in cfRNA samples extracted from plasma of metastatic lung cancer patients. 112 reads supported expression of the EGFR c.2573T>G p.L858R mutant transcript (containing the single nucleotide mutation) (Figure 17).

[0094] Detection of expressed transcripts containing single nucleotide variations, insertions and deletions in RNA samples The method described herein demonstrated the ability to detect single nucleotide mutations, insertions and deletions in tissue RNA extracted from FFPE tumor samples (FIG. 18A) and cfRNA extracted from plasma (FIG. 18B). Simultaneous detection of target genes for the detection of expression transcripts containing single nucleotide mutations, insertions (e.g., duplications) and deletions was performed on tumor tissue RNA from four cancer samples and plasma cfRNA from three cancer samples. In the same multiplex PCR reaction, primers for desired targets were included and combined target capture was performed. Variant allele frequency (VAF) was determined based on the ratio of mutant read counts to total read counts detected from the method described herein. The effectiveness of the RNA-based method described herein is shown by the VAF percentage depicted in FIG. 18A and 18B.

[0095] Consideration This disclosure describes a method for simultaneously detecting genomic alterations, such as structural rearrangements, and gene expression using circulating cell-free RNA (cfRNA). It is envisioned that such non-invasive detection and quantification will enable cancer detection, prognosis determination, and treatment response prediction. The method is based on highly multiplexed amplicon-based NGS, involving tagging individual cfRNA molecules with barcode sequences, and designing amplicons optimized to match the fragmented nature of cfRNA. The inventors have shown that the method can be applied to circulating nucleic acid extracts containing both cfDNA and cfRNA, and can simultaneously detect and quantify fusion RNA transcripts and gene expression in such samples.

[0096] To detect structural rearrangements such as gene fusions in cfRNA analytes that result in the juxtaposition of exons from different genes resulting in fusion transcripts, we designed a targeted multiplex amplicon panel for fusion detection by next-generation sequencing (NGS). We exploited the juxtaposition of gene exons to amplify fusion transcripts with a pair of primers flanking the exon junction involved in the fusion. Primers specific to the exons of the fusion gene and partner gene known to undergo fusion were designed immediately adjacent to the exon junction site. Such juxtaposition of exons from different genes can only occur when the processed mRNA is made such that the fused exons are joined together (by splicing), making it unlikely that equivalent DNA sequences would contribute to productive amplification by the same primers due to the intervening relatively long fused introns separating the exons in the DNA.

[0097] The design of primers to capture fusion transcripts had two main features - 1) the presence of a random barcode sequence in the downstream (downstream relative to the fusion transcript) primer to individually tag each copy of the RNA fusion transcript if present, and 2) the positioning of each primer approximately 50 base pairs from each exon junction in the panel, such that the expected total amplicon length was close to 90-110 base pairs. This was done to satisfy the observed cfRNA size distribution of samples that peaked at 110-120 nucleotides. The chosen 90-110 base pair length was deemed optimal because shorter amplicon (<80 base pairs) products are less effectively retained through the multi-step library preparation method for amplicon sequencing, which involves size-based (magnetic bead-based) separation of small primer dimer artifacts that we want to remove and the desired products that we want to retain. Multiple "upstream" and "downstream" primers were included in the multiplex PCR to optimally capture potential fusions known to occur between genes. The primer design includes exons of well-characterized genes known to undergo fusions, such as ALK, RET, ROS1, FGFR2, FGFR3, among others, as well as exons of their partner genes, such as EML4, KIF5B, CCDC6, CD74, TACC3. Potential fusions between upstream and downstream exons (not limited to the gene pairs for which the design was intended) could theoretically be detected if present in the sample if the capture reaction simultaneously contains multiple primers. Broadly speaking, primers were designed to capture all exon junctions known to undergo fusions in the target gene and partner genes (as well as possible intervening exons not previously reported to be involved in fusions). Barcode sequence primers allow for accurate enumeration of copies of RNA transcripts according to the method of enumeration.

[0098] The first step in the cfRNA NGS library preparation process based on this method is to convert cfRNA (which is naturally fragmented) into complementary DNA (cDNA) using reverse transcriptase with random primers. The result of the reverse transcription reaction is the full complement of cfRNA molecules present in the sample. In addition to the exon-flanking primers for fusion detection, the multiplex reaction also included primers to several (>20) control housekeeping genes to provide a quantitative measure of the amount of cfRNA included in the reaction. The purpose of capturing transcripts of a baseline expressed gene across all sample types was to estimate the average abundance of cellular material entering the multiplex PCR reaction and to serve as a control for the entire cfRNA sequencing library preparation process, including sample extraction, reverse transcription, and PCR steps. The design of primers targeting the control target genes differed from that of the fusion targets in that at least one primer of the control gene primer pair was designed to span an exon-exon junction to prevent unintended amplification of control target gene DNA, and the resulting amplicon was approximately 100 base pairs in length (Figure 1). Primer design for expression-related target genes was similar to that of control gene targets, with at least one primer of the primer pair spanning an exon-exon junction, and two or more primer pairs per target gene were designed to cover both the 5' and 3' exons, allowing one or more amplicons to represent a given target gene, to more reliably capture expression of the target gene toward expression. A highly multiplexed primer pool was utilized with multiple upstream and downstream primers, some of which are expected to generate sequenceable targets in the majority of samples depending on the expression variability, and some primers are expected to generate products only if the sample is positive for structural rearrangements that create productively expressed fusion genes. The primers also possessed the appropriate extensions required to generate sequenceable libraries with sequencing adapters for Illumina sequencing.

[0099] In this disclosure, the use of cfRNA analysis to simultaneously detect and enhance structural rearrangements and gene expression was demonstrated. This was achieved by designing multiplex amplicon NGS assays that encompass the exons of genes involved in fusions, and designing amplicons to target gene expression by using barcode sequences and selecting the optimal amplicon size for cfRNA application. Overall abundance was quantified by the read density of accumulated read counts. In this disclosure, the problems associated with whole transcriptome sequencing, including cost and human resources, were partially overcome by applying plasma cfRNA targeted sequencing.

[0100] In this disclosure, clinically relevant altered splicing events such as MET proto-oncogene, receptor tyrosine kinase (MET) exon 14 skipping, and androgen receptor (AR) transcript variants are approached as intragenic fusion events and, if present, are designed to be captured using primer combinations that capture aberrant splicing as juxtapositions of exons of the same gene that are not normally observed but may occur in cancer. For prediction of response to various treatments, the ability to non-invasively quantify expression of relevant genes is beneficial as it allows longitudinal monitoring of response and informs clinical decisions. However, this is not routinely performed in clinical practice and is largely limited to detection of DNA-level changes such as mutations and genome copy number changes. Using sequencing techniques such as NGS, mutations are identified by comparing sequencing reads to a reference sequence (genome). Genomic copy number changes are quantified by counting the number of reads corresponding to a gene and quantifying the deviation from the normal copy number expected from a cell or sample with two copies of DNA per gene. In one example, DNA level changes include single nucleotide mutations leading to missense mutations, frameshift mutations, insertion-deletions, and splice site mutations. Non-invasive monitoring of expression changes by accessing cfRNA analytes can take advantage of overexpression of tumor-specific transcripts, resulting in amplification of tumor-derived RNA signals in blood, thereby increasing detection sensitivity. For non-invasive characterization of structural rearrangements, e.g., gene fusions in plasma, targeted cfDNA-based next-generation sequencing (NGS)-based methods are typically utilized.

[0101] To overcome the stability problem, the present disclosure applies appropriate RNA isolation procedures, removal of DNA contamination and the use of endogenous housekeeping control genes.Combined together, cfRNA can be used to provide accurate information related to cancer diagnosis, prognosis and prediction of treatment response.

[0102] The novel features of this disclosure, and why they are technically important, are as follows: 1. Special design of primers that allow the amplification of consistently short amplicons capable of amplifying targets from cfRNA, which is usually around 100 nucleotides in length when isolated from plasma. 2. Include barcode sequences in primer design for accurate enumeration of specific targets, whether they contain fusions or not. 3. Design combination for simultaneous capture of fusion (if any) and target gene expression. 4. The ability to detect novel fusions with every potential primer combination included in the multiplex panel. 5. Design of a data analysis workflow enabling parallel analysis of RNA-based fusion and expression.

[0103] The method of the present disclosure has the following advantages: 1. The methods of the disclosure use cfRNA (which lacks introns) as sample input, thereby enabling the identification of gene fusions with long introns that are typically excluded from traditional DNA-based assays. 2. The disclosed method allows for the identification of both fully characterized and novel (i.e., previously uncharacterized) genomic alteration targets. Novel genomic alteration targets can be detected with any potential primer combination included in the multiplex panel. Design of data analysis workflows that allow for parallel analysis of RNA-based fusions and expression. 3. The method disclosed herein allows for the detection of structural rearrangements of cfRNA and the determination of expression level simultaneously.For expressed cancer-related genes, ctRNA provides the same mutation information as ctDNA; Moreover, it can provide quantitative information on the expression level of target genes of interest, potentially increasing the detection sensitivity of low allele frequency variants due to overexpression of tumor-specific transcripts.The ability to non-invasively quantify the expression of these targets can be very useful for monitoring treatment response and determining treatment. 4. The disclosed method can be used in a rapid, non-invasive (requiring only one blood draw), blood-based test (e.g., to detect fusion targets in blood cfRNA). Furthermore, the method is scalable to the detection of multiple cancers in a single test, making it suitable for screening for cancer in asymptomatic populations. 5. The method disclosed herein is highly sensitive compared to conventional genome structural change detection methods. Less starting material (cfRNA) is required to obtain the same or better detection ability. For example, only 24ng cfRNA is required to detect TMPRSS2-ERG fusion in metastatic prostate cancer samples, compared to 70ng cfDNA to generate the same sequence reading. 6. The technical significance lies in the generalizable use of primers for target capture, which allows working with smaller and limited amounts of nucleic acid sample input, and unique combinations of targets are selected for sensitive and specific detection of multiple cancers. 7. The disclosed method is scalable, allowing for the capture of multiple genomic regions to identify several cancer types in a single assay. The coverage of the target genes can be expanded by the addition of forward and reverse primer pairs. 8. The methods of the present disclosure may be used in the following applications: • Detection, identification and quantification of clinically relevant and well-characterized genomic alterations (e.g. gene fusions), for example those associated with cancer. • Identification of novel cancer-specific genomic alterations. • Cancer screening in healthy individuals and individuals at high risk for the cancer examined. • Disease monitoring in cancer patients, including monitoring responses to treatments such as immunotherapy. 9. Shorter fragments are more difficult to use as starting material for sequencing-based assays due to the constraints on primer design and optimally captured sequence information.The disclosed method uses short length (about 100 nucleotides) cfRNA compared to cfDNA (about 160 base pairs).The primers described herein are optimally designed to capture fragmented cfRNA of about 100 nucleotides in length to maximize the detection sensitivity of fusion and expression change. 10. The disclosed method uses RNA, rather than DNA, as sample input for the detection of genomic alteration events. This allows the detection of genomic alteration events that would be excluded by typical DNA-based detection assays. Examples of such genomic alterations include: • Increased DNA copy number leading to overexpression of RNA; • Structural rearrangements involving very long introns in two or more genes; and • Changes in gene expression patterns that correspond to drug response or resistance.

[0104] Sequence Listing Table of forward primers specific to genes that can cause genomic alterations TIFF2024530807000004.tif147160TIFF2024530807000005.tif241160TIFF2024530807000006.tif241160TIFF2024530807000007.tif241160TIFF2024530807000008.tif241160TIFF2024530807000009.tif246160TIFF2024530807000010.tif246160TIFF2024530807000011.tif147160Table of reverse primers specific to genes that may cause genomic alterations TIFF2024530807000012.tif76161TIFF2024530807000013.tif241161TIFF2024530807000014.tif241161TIFF2024530807000015.tif241161TIFF2024530807000016.tif241161TIFF2024530807000017.tif50161Table of forward primers specific for control housekeeping genes TIFF2024530807000018.tif147161 Table of reverse primers specific for control housekeeping genes TIFF2024530807000019.tif147161 Table of forward primers specific to target genes related to protein expression TIFF2024530807000020.tif153162 Table of reverse primers specific to target genes related to protein expression TIFF2024530807000021.tif153163 Table of other sequences TIFF2024530807000022.tif45163

Claims

1. A method for detecting genomic changes and / or detecting gene expression and / or quantifying the level of gene expression using RNA in a biological sample, comprising the following steps: (a) extracting RNA from a biological sample and converting the RNA into complementary DNA (cDNA); (b) performing a plurality of multiplex PCR reactions on the converted cDNA with (I) a plurality of forward and reverse primer pairs specific for a plurality of target genes that can undergo genomic changes, wherein each forward primer of the plurality of forward and reverse primer pairs specific for the plurality of target genes that can undergo genomic changes is complementary to a sequence located approximately 50 base pairs upstream of the exon junction of each target gene that can undergo genomic changes, each reverse primer of the plurality of forward and reverse primer pairs specific for the plurality of target genes that can undergo genomic changes is complementary to a sequence located approximately 50 base pairs downstream of the exon junction of each target gene that can undergo genomic changes, each reverse primer of the plurality of forward and reverse primer pairs specific for the plurality of target genes that can undergo genomic changes contains a barcode sequence at its 5'-end, and the barcode sequences of the reverse primers corresponding to each target gene that can undergo genomic changes are different, a plurality of forward and reverse primer pairs, and / or (II) a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes, wherein (i) each forward primer of the plurality of forward and reverse primer pairs specific for the plurality of control housekeeping genes is complementary to a sequence spanning the exon-exon junction of each control housekeeping gene, each reverse primer of the plurality of forward and reverse primer pairs specific for the plurality of control housekeeping genes is complementary to a sequence approximately 100 base pairs downstream of the sequence spanning the exon-exon junction of each control housekeeping gene, each reverse primer of the plurality of forward and reverse primer pairs specific for the plurality of control housekeeping genes contains a barcode sequence at its 5'-end, and the barcode sequences of the reverse primers corresponding to each control housekeeping gene are different; (ii) Each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of the control housekeeping genes is complementary to a sequence spanning an exon-exon junction of each control housekeeping gene, each forward primer of a plurality of forward and reverse primer pairs specific to a plurality of the control housekeeping genes is complementary to a sequence about 100 base pairs downstream of the sequence spanning an exon-exon junction of each control housekeeping gene, each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of the control housekeeping genes contains a barcode sequence at its 5'-end, and the barcode sequences of the reverse primers corresponding to each control housekeeping gene are different; (iii) Each forward and each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of the control housekeeping genes are complementary to a continuous sequence spanning an exon-exon junction of each control housekeeping gene, each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of the control housekeeping genes contains a barcode sequence at its 5'-end, and the barcode sequences of the reverse primers corresponding to each control housekeeping gene are different, a plurality of forward and reverse primer pairs, and / or (III) A plurality of primer sets specific to a plurality of target genes related to protein expression, each primer set containing a plurality of forward and reverse primer pairs specific to each target gene related to protein expression, (i) Each forward primer of a plurality of forward and reverse primer pairs specific to each target gene related to protein expression is complementary to a sequence spanning an exon-exon junction of each target gene related to protein expression, each reverse primer of a plurality of forward and reverse primer pairs specific to each target gene related to protein expression is complementary to a sequence about 100 base pairs downstream of the sequence spanning an exon-exon junction of each target gene related to protein expression, Each reverse primer of a plurality of forward and reverse primer pairs specific to each target gene related to the protein expression contains a barcode sequence at its 5'-end, and the barcode sequences of the reverse primers corresponding to each target gene related to the protein expression are different. (ii) Each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of target genes related to the protein expression is complementary to a sequence spanning an exon-exon junction of each target gene related to the protein expression. Each forward primer of a plurality of forward and reverse primer pairs specific to a plurality of target genes related to the protein expression is complementary to a sequence about 100 base pairs downstream of the sequence spanning the exon-exon junction of each target gene related to the protein expression. Each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of target genes related to the protein expression contains a barcode sequence at its 5'-end, and the barcode sequences of the reverse primers corresponding to each target gene related to the protein expression are different. (iii) Each forward and each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of target genes related to the protein expression are complementary to a continuous sequence spanning an exon-exon junction of each target gene related to the protein expression. Each reverse primer of a plurality of forward and reverse primer pairs specific to a plurality of target genes related to the protein expression contains a barcode sequence at its 5'-end, and the barcode sequences of the reverse primers corresponding to each target gene related to the protein expression are different. A plurality of primer sets are used to perform, thereby producing a plurality of amplicons. (c) Purifying the plurality of amplicons from step (b). (d) Amplifying the purified product from step (c) by using an adapter primer with a universal index to prepare a sequencing library. (e) Purifying the sequencing library obtained from step (d). (f) Subjecting the purified sequencing library from step (e) to multiplex sequencing on a next-generation sequencing platform to obtain a plurality of sequencing reads. (g) Deriving a consensus read for each sequence from a plurality of sequence determination reads obtained from step (f); (h) Performing a sequence alignment between the consensus read obtained from step (g) and a reference genome, wherein (I) if the sequence alignment results in a partial alignment with the reference genome of an exon from a first gene and a partial alignment with the reference genome of an exon from a second gene, then next (i) determining that the sequence alignment is a split read, (ii) counting / enumerating the number of split reads from step (h)(I)(i) that support a fusion junction, and (iii) if the number of split reads from step (h)(I)(ii) is two or more, then in that case determining the first gene and the second gene as fusion partners, (II) if the sequence alignment results in an alignment with the reference genome of a control housekeeping gene, then next (i) determining that the sequence alignment is a consensus read of the control housekeeping gene, (ii) counting / enumerating the pairs of consensus reads of the control housekeeping gene from step (h)(II)(i), and (iii) determining the level of gene expression of the control housekeeping gene, (III) if the sequence alignment results in an alignment with the reference genome of a target gene related to protein expression, (i) determining that the sequence alignment is a consensus read of the target gene related to protein expression, (ii) counting / enumerating the pairs of consensus reads of the target gene related to protein expression from step (h)(III)(i), and (iii) determining the level of gene expression of the target gene related to protein expression, step; (i) Based on the sequence alignment from step (h), determining the presence or absence of a genomic change and / or determining the presence or absence of gene expression and / or quantifying the level of gene expression.

2. The method according to claim 1, wherein the RNA is selected from the group consisting of cell-free RNA (cfRNA) and RNA encapsulated in tissue and / or cells.

3. The method according to claim 1, wherein the biological sample is selected from the group consisting of a liquid sample, a tissue sample, and a cell sample.

4. The method according to claim 3, wherein the liquid sample is a body fluid, and optionally the body fluid is selected from the group consisting of blood, bone marrow, cerebrospinal fluid, peritoneal fluid, pleural fluid, lymph fluid, ascites, serous fluid, sputum, tear fluid, feces, urine, saliva, milk duct fluid, gastric juice and pancreatic juice, optionally the body fluid is blood, and optionally the blood is plasma.

5. The method according to claim 3, wherein the tissue sample is a frozen tissue sample or a fixed tissue sample, and optionally the fixed tissue sample is a formalin-fixed paraffin-embedded (FFPE) tissue sample.

6. The method according to claim 1, wherein the biological sample is obtained from a subject having cancer and / or suspected of having cancer.

7. The method according to claim 6, wherein the cancer is selected from the group consisting of leukemia, lung cancer, colorectal cancer, breast cancer, pancreatic cancer, prostate cancer, hypopharyngeal cancer, liver cancer, bile duct cancer, esophageal cancer, urothelial cancer, and gastrointestinal cancer.

8. The method according to claim 6, wherein the cancer is selected from the group consisting of metastatic prostate cancer, metastatic lung cancer, metastatic breast cancer, and leukemia.

9. The method according to claim 1, wherein the amount of RNA used in step (a) is 6 ng to 100 ng.

10. The method according to claim 1, wherein step (a) is performed using a reverse transcription kit, and the reverse transcription kit includes a buffer for performing reverse transcription, a reverse transcriptase, and a plurality of random primers.

11. The method according to claim 1, wherein the plurality of multiplex PCR reactions performed on the converted cDNA include 3 to 15 PCR cycles.

12. The method according to claim 1, wherein the barcode sequence is an oligonucleotide containing 10 to 16 random nucleotides.

13. The method according to claim 1, wherein the barcode sequence is an oligonucleotide containing 10 random nucleotides.

14. The target gene capable of causing genomic changes is an exon derived from a partner gene of a gene known to cause fusion, fused to an exon derived from the gene known to cause the fusion The method according to claim 1, comprising.

15. Genes known to cause fusions include anaplastic lymphoma kinase receptor tyrosine kinase, RET proto-oncogene, ROS proto-oncogene 1, fibroblast growth factor receptor 1 (FGFR1), fibroblast growth factor receptor 2 (FGFR2), fibroblast growth factor receptor 3 (FGFR3), neurotrophic receptor tyrosine kinase 1 (NTRK1), neurotrophic receptor tyrosine kinase 2 (NTRK2), neurotrophic receptor tyrosine kinase 3 (NTRK3), neuregulin 1 (NRG1), B-Raf proto-oncogene, serine / threonine kinase (BRAF), transmembrane serine protease 2 (TMPRSS2), MET proto-oncogene, receptor tyrosine kinase (MET), epidermal growth factor receptor (EGFR), estrogen receptor 1 (ESR1), platelet-derived growth factor receptor alpha (PDGFRA), androgen receptor (AR), RhoGEF and GTPase BCR activator (BCR), core-binding factor subunit beta (CBFB), lysine methyltransferase 2A (KMT2A), nucleophosmin 1 (NPM1), PML nuclear body scaffold (PML), and RUNX family transcription factor 1 (RUNX1), the method according to claim 14, selected from the group consisting of.

16. The partner genes of genes known to cause fusions are selected from the group consisting of EMAP-like 4 (EML4), kinesin family member 5B (KIF5B), coiled-coil domain-containing 6 (CCDC6), CD74 molecule (CD74), transforming acidic coiled-coil-containing protein 3 (TACC3), ezrin (EZR), ETS transcription factor ERG (ERG), ArfGAP with GTPase domain, ankyrin repeats and PH domain 3 (AGAP3), A-kinase anchoring protein 9 (AKAP9), KIAA1549, tropomyosin 3 (TPM3), translocation promoter region, nuclear basket protein (TPR), transport from ER to Golgi regulatory factor (TFG), lamin A / C (LMNA), BicC family RNA-binding protein 1 (BICC1), RAD51 recombinase (RAD51), CD47 molecule (CD47), Yes1-associated transcriptional regulator (YAP1), ETS variant transcription factor 1 (ETV1), ETS variant transcription factor 4 (ETV4), ETS variant transcription factor 5 (ETV5), ETS variant transcription factor 6 (ETV6), factor interacting with PAPOLA and CPSF1 (FIP1L1), centriolin (CNTRL), ABL oncogene 1, non-receptor tyrosine kinase (ABL1), AF4 / FMR2 family member 1 (AFF1), MDS1 and EVI1 complex locus (MECOM), MLLT3 super elongation complex subunit (MLLT3), myosin heavy chain 11 (MYH11), PBX homeobox 1 (PBX1), retinoic acid receptor alpha (RARA), RUNX1 partner transcriptional corepressor 1 (RUNX1T1), the method according to claim 14.

17. The number of a plurality of forward and reverse primer pairs specific for a plurality of target genes capable of causing genomic changes, a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes, and a plurality of primer sets specific for a plurality of target genes related to protein expression is at least 300, the method according to claim 1.

18. The length of the plurality of amplicons produced in step (b) is 90 to 110 base pairs, the method according to claim 1.

19. The method according to claim 1, wherein the purification in step (c) and / or (e) is carried out using a plurality of paramagnetic beads, and optionally the paramagnetic beads are selected from the group consisting of AMPure XP beads, SPRI beads, and dynabeads.

20. The method according to claim 1, wherein step (g) further comprises: (g)(I) detecting the presence of the barcode sequence from each sequencing read; (g)(II) performing cluster reassignment on a plurality of sequencing reads having the same barcode sequence to create a plurality of barcode clusters, wherein each barcode cluster contains reads having the same barcode sequence from the same amplicon; and (g)(III) performing a consensus call on each barcode cluster to obtain a consensus read of each sequence.

21. The method according to claim 1, wherein the step of determining the presence or absence of genomic changes and / or the presence or absence of gene expression and / or quantifying the level of gene expression further comprises performing variant calling of the sequence alignment from step (h).

22. The method according to claim 21, wherein the step of variant calling comprises: (i) identifying the differences between the consensus read and the reference genome based on the sequence alignment from step (h); and (ii) determining the read count of the sequence alignment containing genomic changes.

23. The method according to claim 21, wherein the genomic change is selected from the group consisting of insertions, deletions, and single nucleotide mutations, and optionally the insertion is a duplication.

24. A kit for detecting genomic changes and / or detecting gene expression and / or quantifying the level of gene expression using RNA in a biological sample by the method according to any one of claims 1 to 23, comprising: - a plurality of forward and reverse primer pairs specific for a plurality of target genes capable of causing genomic changes as defined in claim 1; - a plurality of forward and reverse primer pairs specific for a plurality of control housekeeping genes as defined in claim 1; and - a plurality of primer sets specific for a plurality of target genes related to protein expression as defined in claim 1.

25. The kit according to claim 24, further comprising: - Buffer for performing a plurality of multiplex PCR reactions, - Reverse transcriptase, - Buffer for performing reverse transcription, - Random primers, - DNA polymerase, and - A plurality of deoxynucleoside triphosphates (dNTPs).