Probes and methods for detecting transcripts resulting from fusion genes and / or exon skipping
The probe set for massively parallel sequencing addresses the challenge of detecting fusion genes and exon skipping by targeting junction regions, enhancing detection efficiency and applicability to low-quality RNA samples.
Patent Information
- Application Number
- JP2024106160
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-06-27
- Filing Date
- 2024-07-01
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2038-06-22
AI Technical Summary
Existing methods struggle to efficiently detect multiple fusion genes and exon skipping events in cancer genomes, particularly in low-quality RNA samples, due to the low occurrence and diversity of these events, and require specialized skills for diagnosis.
A probe and probe set designed for massively parallel sequencing that hybridizes to regions near the junctions of fusion genes or exon skipping, with specific length requirements to capture transcripts, allowing for efficient detection of these events.
Enables easy and accurate detection of fusion genes and exon skipping in various RNA samples, including low-quality FFPE samples, without the need for specialized skills, and improves detection efficiency compared to conventional methods.
Smart Images

Figure 0007777359000009 
Figure 0007777359000010 
Figure 0007777359000011
Abstract
Description
[Technical Field]
[0001] The present invention relates to a probe for determining the presence or expression level of a transcript of a fusion gene on a genome, a probe for determining the presence or expression level of a transcript produced by exon skipping, a kit containing the probe, a method for determining the presence or expression level of a transcript of a fusion gene on a genome using the probe, and a method for determining the presence or expression level of a transcript produced by exon skipping. [Background technology]
[0002] Fusion genes are known to be a cause of somatic cancer mutations, and several treatments have been developed for cancers caused by fusion genes. For example, first-line therapy using tyrosine kinase inhibitors is used for patients with cancer mutations such as the BCR-ABL1 fusion gene in chronic myeloid leukemia (Non-Patent Document 1) and the EML4-ALK fusion gene in non-small cell lung cancer (Non-Patent Document 2). This has improved the treatment outcomes for cancers caused by fusion genes.
[0003] Recent advances in sequencing technology have enabled comprehensive detection of chromosomal rearrangements in cancer genomes and transcriptomes, leading to the discovery of fusion genes such as RET, ROS1, NTRK1, NRG1, and FGRF1 / 2 / 3 genes (Non-Patent Documents 3 to 8), and these fusion genes are also being used in cancer diagnosis. Furthermore, in addition to fusion genes, it has recently been suggested that exon skipping, such as MET14 exon skipping, may also be a cause of cancer.
[0004] However, because the occurrence of these fusion genes and exon skipping is relatively low and diverse, it has been difficult to simultaneously detect multiple fusion genes that serve as target genes. Furthermore, conventional methods such as FISH, immunohistochemistry, and reverse transcription PCR require specialized skills for diagnosis, so there is a strong demand for a method that can easily detect multiple target genes for clinical application.
[0005] Targeted sequencing of cancer-related genes by target gene enrichment of gDNA using amplicon PCR or hybridization capture is one example of a method used to detect mutations such as fusion genes. However, junctions of fusion genes and other mutations are often widely distributed in the introns of each gene. Therefore, conventional hybridization capture methods require the creation of probes without bias toward introns in order to capture junctions of fusion genes and exon skipping, which requires a large number of probes.
[0006] RNA sequencing (RNA-seq) has also been proposed as an alternative method for detecting fusion transcripts from fresh frozen samples or cell lines. However, its application to low-quality RNA samples (low-quality RNA samples), such as formalin-fixed, paraffin-embedded (FFPE), is difficult because it is difficult to create reliable libraries using techniques such as polyA selection, which are commonly used for mRNA enrichment. Although cDNA capture and anchored multiplex PCR-based methods have been reported to be effective for low-quality RNA samples, these methods are clinically ineffective because the types of genes they target are very limited. Therefore, a method that can easily detect a large number of target genes even in low-quality RNA samples is needed. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] J. Erikson et al., Proc. Natl. Acad. Sci., USA 83, 1807-1811, 1986 [Non-patent document 2] M. Soda et al., Nature, 448, 561-566, 2007 [Non-patent document 3] T. Kohno et al., Nat. Med. 18, 375-377, 2012 [Non-patent document 4] K. Takeuchi et al., Nat. Med. 18, 378-381, 2012 [Non-patent document 5] D. Lipson et al., Nat. Med. 18, 382-384, 2012 [Non-patent document 6] L. Fernandez-Cuesta et al., Cancer Discov. 4, 415-422, 2014 [Non-Patent Document 7] A. Vaishnavi et al., Nat. Med., 19, 1469-1472, 2013 [Non-patent document 8] R. Wang, L et al., Clin. Cancer Res. 20,, 4107-4114, 2014 Summary of the Invention [Problem to be solved by the invention]
[0008] In one embodiment, an objective of the present invention is to provide a method that allows for easy detection of transcription products resulting from fusion genes and / or exon skipping. [Means for solving the problem]
[0009] The present inventors have created a probe that can be used to detect transcripts resulting from genomic fusion genes or exon skipping in massively parallel sequencing, and have found that this probe can be used to efficiently detect transcripts resulting from genomic fusion genes or exon skipping.
[0010] The present invention includes the following aspects. (1) A probe for determining the presence or expression level of a transcript of a fusion gene on a genome in massively parallel sequencing, the fusion gene expresses a transcription product in which a part of gene A on the 5' side and a part of gene B on the 3' side are joined at a virtual junction point; the probe hybridizes to a region derived from either gene A or B of the cDNA prepared from the transcription product, A probe in which z≧x+y is satisfied, where x is the shortest base length from the end of the probe to the virtual junction point when the probe hybridizes to the cDNA, y is the base length of the region in the probe that hybridizes to the cDNA, and z is the read length of massively parallel sequencing. (2) A probe set for determining the presence or expression level of a transcript of a fusion gene on a genome in massively parallel sequencing, the fusion gene expresses a transcription product in which a part of gene A on the 5' side and a part of gene B on the 3' side are joined at a virtual junction point; at least two different probes that hybridize to regions derived from either gene A or B of the cDNA prepared from said transcription product; A probe set in which z≧x+y is satisfied, where x is the shortest base length from the end of each probe to the virtual junction point when the probe hybridizes to the cDNA, y is the base length of the region in each probe that hybridizes to the cDNA, and z is the read length of massively parallel sequencing. (3) A probe for determining the presence or expression level of a transcript generated by exon skipping in massively parallel sequencing, In the transcript, exon A' on the 5' side and exon B' on the 3' side are joined at a virtual junction point, the probe hybridizes to a region derived from either exon A' or B' of the cDNA prepared from the transcription product; A probe in which z≧x+y is satisfied, where x is the shortest base length from the end of the probe to the virtual junction point when the probe hybridizes to the cDNA, y is the base length of the region in the probe that hybridizes to the cDNA, and z is the read length of massively parallel sequencing. (4) A probe set for determining the presence or expression level of a transcript generated by exon skipping in massively parallel sequencing, In the transcript, exon A' on the 5' side and exon B' on the 3' side are joined at a virtual junction point, at least two different probes that hybridize to regions derived from either exon A' or B' of the cDNA prepared from said transcription product; A probe set in which z≧x+y is satisfied, where x is the shortest base length from the end of each probe to the virtual junction point when the probe hybridizes to the cDNA, y is the base length of the region in each probe that hybridizes to the cDNA, and z is the read length of massively parallel sequencing. (5) The probe or probe set according to any one of (1) to (4), wherein x is 0 to 140, y is 30 to 140, and z is 100 to 300. (6) The probe set according to any one of (2), (4), and (5), which comprises at least six of the probes. (7) The probe set according to any one of (2) and (4) to (6), which consists of only probes that satisfy z≧x+y. (8) The probe set includes n probes, and the shortest base lengths of each probe are x1, x2, x3, ... x n (However, x1 <x2<x3…<x n ), x1=0, x2=x n ×1 / (n-1), x3=x n ×2 / (n-1), …x n = x n ×(n-1) / (n-1) The probe set according to any one of (2) and (4) to (7), (9) A probe for determining the presence or expression level of a transcript of a fusion gene on a genome in massively parallel sequencing, comprising: the fusion gene expresses a transcription product in which a part of gene A on the 5' side and a part of gene B on the 3' side are joined at a virtual junction point; A probe that hybridizes to a region containing the hypothetical junction point of the cDNA prepared from the transcription product. (10) A probe set for determining the presence or expression level of a transcript of a fusion gene on a genome in massively parallel sequencing, comprising: the fusion gene expresses a transcription product in which a part of gene A on the 5' side and a part of gene B on the 3' side are joined at a virtual junction point; a probe set comprising at least two different probes that hybridize to a region containing the hypothetical junction of the cDNA prepared from the transcription product; (11) A probe for determining the presence or expression level of a transcript generated by exon skipping in massively parallel sequencing, In the transcript, exon A' on the 5' side and exon B' on the 3' side are joined at a virtual junction point, A probe that hybridizes to a region containing the hypothetical junction where exon skipping can occur in cDNA prepared from the transcription product. (12) A probe set for determining the presence or expression level of a transcript generated by exon skipping in massively parallel sequencing, comprising: In the transcript, exon A' on the 5' side and exon B' on the 3' side are joined at a virtual junction point, A probe set comprising at least two different probes that hybridize to a region including the hypothetical junction at which exon skipping can occur in the cDNA prepared from the transcription product. (13) A combined probe set comprising a plurality of different probes or probe sets according to any one of (1) to (12). (14) The probe or probe set according to any one of (1) to (12) or the combined probe set according to (13), further comprising at least one probe for measuring gene expression levels. (15) The probe, probe set, or combination probe set according to any one of (1) to (14), for use with a transcription product derived from a processed biological sample. (16) A kit comprising the probe, probe set, or combination probe set according to any one of (1) to (15). (17) Preparing a transcription product from a sample derived from a subject; preparing cDNA from the transcription product; A step of concentrating target cDNA hybridized to the probe, probe set, or probe combination set according to any one of (1) to (15); a step of performing sequence analysis on the enriched target cDNA by massively parallel sequencing; and a step of determining the presence or expression level of a transcript including a transcript of the fusion gene on the genome based on the results of the sequence analysis. A method for determining the presence or expression level of a transcript containing a transcript of a fusion gene on a genome, comprising: (18) The determination is performed by the following steps: When the fusion gene expresses a transcription product in which a part of gene A on the 5' side and a part of gene B on the 3' side are joined at a virtual junction, Let the number of reads of cDNA derived from gene A where no gene fusion occurs at the virtual junction be α, the number of reads of cDNA derived from gene B be β, and the number of reads of cDNA derived from a fusion gene where gene fusion occurs at the virtual junction be γ. If 0<α or β≦γ, it is determined that the fusion gene is present; If 0<γ<α or β, it is determined that the fusion gene is present at a low expression level; The method according to (17), wherein the method is carried out by determining that the fusion gene is not present when α or β>0 and γ=0. (19) Preparing a transcription product from a sample derived from a subject; preparing cDNA from the transcription product; A step of concentrating target cDNA hybridized to the probe, probe set, or probe combination set according to any one of (1) to (15); a step of performing sequence analysis on the enriched target cDNA by massively parallel sequencing; and a step of determining the presence or expression level of a transcript, including a transcript generated by exon skipping, based on the results of the sequence analysis. A method for determining the presence or expression level of a transcript, including a transcript generated by exon skipping, comprising: (20) The determination is performed by the following steps: In the above transcript, when the 5'-side exon A' and the 3'-side exon B' are ligated at a virtual junction, Let α' be the number of reads of cDNA derived from exon A' where no gene fusion occurs at the virtual junction, β' be the number of reads of cDNA derived from exon B', and γ' be the number of reads of cDNA derived from the transcript generated by exon skipping. If 0<α' or β'≦γ', it is determined that a transcript resulting from exon skipping is present; If 0<γ'<α' or β', it is determined that a transcript resulting from exon skipping is present at a low expression level; The method according to (19), wherein the method is carried out by determining that no transcript resulting from exon skipping is present when α' or β'>0 and γ'=0. (21) The method according to any one of (17) to (20), wherein, in the determination step, when multiple probes hybridizing to the same region are present, the expression level of the transcription product is corrected based on the number of the multiple probes. (22) The method according to any one of (17) to (21), wherein the determination step comprises correcting the expression level of the transcript based on the expression level of a housekeeping gene. (23) A step of determining the presence or expression level of a transcript of a fusion gene on a genome, and / or a transcript generated by exon skipping, according to any one of the methods described in (17) to (22); A method for determining the presence or absence of a disease or the risk thereof, identifying the type of cancer, or determining the prognosis of cancer in a subject, comprising: (24) The method according to (23), wherein identifying the type of cancer comprises clustering samples derived from the subject based on the presence and / or expression levels of multiple transcripts.
[0011] This specification includes the disclosure of Japanese Patent Application No. 2017-125074, from which the present application claims priority. [Effects of the Invention]
[0012] The present invention can provide a method for easily detecting a transcription product resulting from a fusion gene and / or exon skipping. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1A is a conceptual diagram of a probe according to one embodiment of the present invention. In each of the illustrated probes, the right end is the 5' end and the left end is the 3' end. The minimum base length x from the end of the probe to the virtual junction point can be determined based on the read length z and the base length y of the region in the probe that hybridizes to the cDNA, so that a ligation-supported read including the virtual junction point can be obtained. FIG. 1B shows an example of a method for detecting transcripts resulting from fusion genes and / or exon skipping from sequencing results, according to a method according to one embodiment of the present invention. As shown in FIG. 1B, let α be the number of cDNA reads derived from gene A, which does not have a genetic mutation (gene fusion or exon skipping) at the virtual junction point; β be the number of cDNA reads derived from gene B; and γ be the number of cDNA reads derived from a fusion gene, which has a genetic mutation at the virtual junction point. If α<α or β≦γ, a mutant gene is determined to be present; if 0<γ<α or β, a mutant gene is determined to be present at a low expression level; and if α or β>0 and γ=0, a mutant gene is determined to be absent. [Figure 2]Figure 2A shows the number of ligation-supported reads per 10 million raw reads for each method shown (the Pancancer panel shows all-exon capture of synthetic cDNA derived from FFPE). Figure 2B shows the number of probes, and Figure 2C shows the target capture size, for the junction capture method of one embodiment of the present invention and the conventional coding exon capture method. V1, V2, and V3 in Figures 2B and 2C show the results for the gene panels (TOP RNA V1, TOP RNA V2, and TOP RNA V3) described in the Examples. [Figure 3] Figure 3A shows the results of mapping sequence reads to MET transcripts in cases positive for MET exon 14 skipping by RNA-seq using three different methods: poly(A) selection of RNA extracted from fresh-frozen samples (poly(A) capture), whole-exon capture of synthetic cDNA derived from FFPE samples (Pancancer panel), or junction capture of synthetic cDNA derived from FFPE samples. The region between two vertical lines indicates the region corresponding to MET exon 14; the absence of reads in this region indicates positive for exon skipping. Figure 3B shows the number of reads supporting the junction of MET exon 13 and MET exon 15 (exon skipping) per 10 million raw reads for each method. [Figure 4] Figure 4A shows a representative photograph of a bone marrow aspirate specimen stained with hematoxylin and eosin (200x magnification, scale bar 100 μm). Figure 4B shows a representative photograph of a TBLB specimen stained with hematoxylin and eosin (left, 40x magnification, scale bar 1 mm; right, 400x magnification, scale bar 100 μm). [Figure 5] Figure 5 shows the correlation between the RPKM of RNA-seq and the RPKM corrected for the number of tilings in the junction capture method. The results for the gene group used for expression level measurement are shown in A, and the results for the gene group used for fusion gene analysis are shown in B. A correlation was observed in all seven samples. [Figure 6]Figure 6 shows the results of clustering samples based on gene expression levels. The vertical axis represents each gene, and clustering was performed according to expression intensity. The horizontal axis represents each sample, and it can be seen that samples were clustered according to cancer type, such as LUAD, SARC, MUCA, and LUSC. DETAILED DESCRIPTION OF THE INVENTION
[0014] 1. Probes for determining the presence or expression level of transcripts of fusion genes on the genome In one aspect, the present invention relates to a probe for determining the presence or expression level of a transcript of a fusion gene on a genome in massively parallel sequencing.
[0015] As used herein, "massively parallel sequencing" refers to a method for performing DNA sequencing on a large scale in parallel. Massively parallel sequencing typically involves the use of 10 2 , 10 3 , 10 4 , 10 5 Or more molecules are sequenced simultaneously. Massively parallel sequencing includes, for example, next generation sequencing.
[0016] Next-generation sequencing is a method for obtaining sequence information using a next-generation sequencer, and is characterized by the ability to perform a vastly greater number of sequencing reactions simultaneously in parallel compared to the Sanger method (see, for example, Rick Kamps et al., Int. J. Mol. Sci., 2017, 18(2), p. 308 and Int. Neurourol. J., 2016, 20(Suppl. 2), S76-83). Various systems for next-generation sequencing are available, and examples that can be used include, but are not limited to, the Roche Genome Sequencer (GS) FLX System, Illumina HiSeq or Genome Analyzer (GA), Life Technologies Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator G.007 system, and Helicos BioSciences HeliScope Gene Sequencing system.
[0017] The general steps of next-generation sequencing are as follows, but are not limited to these. Next-generation sequencing begins with sample preparation. In this step, the nucleic acid to be analyzed is enzymatically or mechanically fragmented to match the read length of the next-generation sequencer. Subsequently, in many cases, an adapter sequence required for the next sequencing step is added. Furthermore, to analyze a specific gene region, the specific gene region may be enriched by PCR or the like, or a region having a specific sequence may be concentrated using a probe or the like. Enrichment of a gene region can be achieved, for example, by 4 to 12 cycles of amplification, and enrichment using a probe can be achieved by utilizing a label (e.g., biotin) attached to the probe.
[0018] Next, sequencing is performed. The details of this process vary depending on the type of next-generation sequencer, but typically the DNA is linked to the substrate via an adapter sequence, and the adapter sequence is used as a priming site for the sequencing reaction. For details of the sequencing reaction, see, for example, Rick Kamps et al. (supra).
[0019] Finally, data output is performed. This process provides a collection of sequence information (reads) obtained from the sequencing reaction. The output data can be further analyzed to derive more meaningful results, such as the number of reads, e.g., the number of ligation-supported reads per raw read.
[0020] As used herein, the term "read number" refers to the amount of amplification of an amplification product having a specific sequence. Since the read number is usually proportional to the amount of nucleic acid before sequencing, the read number can be used to estimate the expression level of a gene.
[0021] As used herein, "ligation-supporting reads" refers to junction points in a transcript resulting from gene fusion or exon skipping, or reads containing junction points on the genome resulting from gene fusion or exon skipping, and "number of ligation-supporting reads" refers to the number of ligation-supporting reads. As used herein, "raw reads" refers to the total number of reads obtained by next-generation sequencing, and the frequency of ligation-supporting reads can be evaluated by calculating the number of ligation-supporting reads per raw read.
[0022] As used herein, the term "genomic fusion gene" refers to a mutant gene formed by the joining of multiple genes as a result of chromosomal rearrangements caused by deletion, insertion, inversion, translocation, and the like. Typically, a fusion gene is transcribed to produce an RNA molecule as its expression product. Examples of RNA molecules include transcription products such as mRNA encoding a fusion protein. As used herein, the type of fusion gene is not limited, and examples include oncogenic fusion genes such as EML4-ALK, BCR-ABL1, KIF5B-RET, SLC34A2-ROS1, CD74-ROS1, SS18-SSX1, SS18-SSX2, NAB2-STAT6, EWSR1-FLI1, SYT-SSX1, FUS-CREB3L2, TPM3-ROS1, CD74-NRG1, and EWSR1-FLI1.
[0023] In the present invention, the "presence" of a transcription product of a fusion gene on the genome refers to the presence or absence of the fusion gene on the genome, and the "expression level" of a transcription product of a fusion gene refers to the expression level of transcription products such as mRNA, rRNA, and tRNA derived from the fusion gene, preferably mRNA.
[0024] In one embodiment, when a fusion gene is expressed as a transcription product in which a portion of gene A on the 5' side and a portion of gene B on the 3' side are joined at a hypothetical junction, the probe of the present invention hybridizes to a region derived from either gene A or gene B in cDNA prepared from the transcription product. The genes that can form the fusion gene and the hypothetical junction points can be determined by referring to scientific papers, patent documents, and databases such as COSMIC.
[0025] As used herein, "exon" refers to a region of the nucleotide sequence of a gene that remains in a mature transcript. Generally, in eukaryotes, a gene is transcribed as a primary transcript, and then intervening regions called introns are removed by splicing, and exons are linked to form a mature transcript. For example, in the case of a protein-coding gene, introns are removed from a precursor mRNA (pre-miRNA) generated by transcription by pre-miRNA splicing, resulting in a mature miRNA composed of linked exons.
[0026] In one embodiment, when the probes hybridize to cDNA prepared from the RNA molecules of the transcription products, the shortest base length from either the 5' or 3' end of each probe to the virtual junction is x, the base length of the region in each probe that hybridizes to the cDNA is y, and the read length of massively parallel sequencing is z, where z≧x+y. Probes that hybridize to nucleic acid regions that do not contain such virtual junctions are hereinafter also referred to as "virtual junction-free probes." Virtual junction-free probes have the advantage of being able to detect multiple fusion partners and novel fusion genes.
[0027] To facilitate understanding of the present invention, the probe design of this embodiment is shown in Figure 1A. Figure 1A shows the minimum base length x from the end of the probe to the virtual junction, the base length y of the region in the probe that hybridizes to the cDNA, and the read length z, and indicates that reads including the virtual junction can be obtained by massively parallel sequencing.
[0028] In one embodiment, the read length z is determined by the instrument and method used for massively parallel sequencing. Furthermore, if the nucleic acid derived from the sample is fragmented and / or if the nucleic acid is fragmented before sequencing, the read length may be determined based on the length of these fragments. The read length z is not limited, but may be, for example, 50 or more, 75 or more, 100 or more, 150 or more, or 160 or more, or 500 or less, 400 or less, 300 or less, 200 or less, or 180 or less, for example, 50 to 500, 100 to 300, or 150 to 200. Massively parallel sequencing includes single-read sequencing, in which sequencing is performed from only one side of the nucleic acid, and paired-end sequencing, in which sequencing is performed from both sides of the nucleic acid. The read length z is preferably the read length in paired-end sequencing.
[0029] The base length y of the region of the probe that hybridizes to the cDNA can be determined appropriately by those skilled in the art. y may be, for example, 20 or more, 30 or more, 40 or more, preferably 50 or more, 60 or more, or 80 or more, or 220 or less, 200 or less, 180 or less, preferably 160 or less, 140 or less, or 120 or less, for example, 20 to 220, 50 to 160, or 60 to 140. Preferably, the probe hybridizes to the cDNA in a region continuing from the end near the virtual junction. In one embodiment, the probe hybridizes to the cDNA over its entire length, in which case y is the same as the length of the probe.
[0030] The base length of the probe is not limited, but may be, for example, 20 or more, 40 or more, 60 or more, 80 or more, 100 or more, 110 or more, or 115 or more, or 220 or less, 200 or less, 180 or less, 160 or less, 140 or less, 130 or less, or 125 or less, for example, 20 to 220, 60 to 180, 100 to 140, 110 to 130, 115 to 125, or 120.
[0031] The minimum base length x from the end of the probe to the virtual junction point can be determined appropriately based on the read length z and the base length y of the region in the probe that hybridizes to the cDNA. For example, the minimum base length x from the end of the probe to the virtual junction point is 0, meaning that the probe is designed for the region adjacent to the virtual junction point. The upper limit of x is not limited and may be, for example, 300 or less, 250 or less, 200 or less, 150 or less, 140 or less, 130 or less, 125 or less, or 120 or less, and x may be, for example, 0 to 300, 0 to 200, 0 to 140, 0 to 125, or 0 to 120.
[0032] z≧x+y+a (a≧0) indicates that a read containing a sequence of a base or more beyond the virtual junction can be obtained. By designing multiple probes near the virtual junction in this manner, these probes can be used to efficiently enrich various types of transcripts related to the fusion gene. The value of a is not particularly limited as long as it is 0 or greater. However, since increasing the value of a increases specificity but decreases detection sensitivity, those skilled in the art can appropriately determine the value with reference to the contents of this specification. The value of a may be, for example, 5 or greater, 10 or greater, preferably 15 or greater, 20 or greater, 30 or greater, 50 or greater, or 100 or greater, or 500 or less, 400 or less, preferably 300 or less, 200 or less, or 150 or less.
[0033] A probe can be easily designed by a person skilled in the art based on the sequence of the target gene. As used herein, the term "target gene" refers to a gene that can be captured by the probe of the present invention, such as a gene that can form a fusion gene or a gene that can cause exon skipping.
[0034] Examples of such probes include (a) a base sequence of at least 20, 40, 60, 80, 100, 110, 115, or 120 consecutive bases of a complementary sequence of a target gene; (b) a base sequence in which one or more bases have been added, deleted, and / or substituted in the base sequence of (a); (c) a base sequence having, for example, 70% or more, 80% or more, preferably 90% or more, 95% or more, 97% or more, 98% or more, or 99% or more identity to the base sequence of (a); and (d) a probe containing a nucleic acid base sequence that hybridizes under stringent conditions to a base sequence of at least 20, 40, 60, 80, 100, 110, 115, or 120 consecutive bases of a target gene.
[0035] As used herein, the range of "one or more" is 1 to 10, preferably 1 to 7, more preferably 1 to 5, and particularly preferably 1 to 3, or 1 or 2. Furthermore, as used herein, the identity value for a nucleotide sequence refers to a value calculated using software that calculates identity between multiple sequences (e.g., FASTA, DANASYS, and BLAST) with default settings. For details on methods for determining identity, see, for example, Altschul et al., Nuc. Acids. Res. 25, 3389-3402, 1977 and Altschul et al., J. Mol. Biol. 215, 403-410, 1990.
[0036] As used herein, "stringent conditions" refers to conditions under which so-called specific hybrids are formed and nonspecific hybrids are not formed. Stringent conditions can be determined based on known hybridization conditions. For example, they can be determined by referring to Green and Sambrook, Molecular Cloning, 4th Ed. (2012), Cold Spring Harbor Laboratory Press. Specifically, stringent conditions can be set based on the temperature of the hybridization method, the salt concentration of the solution, and the temperature and salt concentration of the solution in the washing step of the hybridization method. More specific stringent conditions include, for example, a sodium concentration of 25 to 500 mM, preferably 25 to 300 mM, and a temperature of 42 to 68°C, preferably 42 to 65°C. More specifically, 5xSSC (83 mM NaCl, 83 mM sodium citrate) and a temperature of 42°C can be used.
[0037] The probe can be prepared based on the above sequence by a method known to those skilled in the art, including, but not limited to, chemical synthesis.
[0038] In one embodiment, the present invention relates to a probe set comprising at least two different probes. The number of probes is not particularly limited as long as it is two or more. However, if the number of probes is too small, detection sensitivity decreases, and if the number of probes is too large, costs increase. Therefore, the number of probes may be appropriately determined taking into consideration sensitivity, costs, etc., and referring to the contents of this specification. The number of probes that can be included in a probe set may be, for example, 3 or more, 4 or more, 5 or more, 6 or more, 8 or more, 10 or more, or 11 or more, or 30 or less, 25 or less, 20 or less, 15 or less, 14 or less, 13 or less, or 12 or less.
[0039] It is preferable that the shortest base length x from the end of each probe included in the probe set to the virtual junction point is not the same value and is distributed. This is because various nucleic acid fragments can be captured. For example, if a probe set includes n probes and the shortest base lengths of each probe are x1, x2, x3, ...x n (However, x1 <x2<x3…<x n ),
[0040]
number
[0041] The shortest base length of each probe can be determined so that: b is a constant, and when b is 0, it means that the shortest base length x of each probe is evenly distributed from the virtual junction point, and the larger the value of b, the more uneven the distribution from the virtual junction point. b is, for example, 50 or less, 40 or less, 30 or less, 25 or less, 20 or less, 15 or less, 10 or less, preferably 5 or less, 4 or less, 3 or less, 2 or less, 1 or less, or 0. Also, x n may be any value, for example, 20 to 500, 30 to 400, 40 to 300, 60 to 200, 80 to 180, preferably 100 to 140, 110 to 130, 115 to 125, or 120.
[0042] Furthermore, when the number of probes, n, is 3 or more, after designing the probes according to the above formula, m probes may be removed from the probe set (where m is an integer of 1 or more, for example, 1 to 5, 1 to 4, 1 to 3, 1 to 2, preferably 1, and nm≧2).
[0043] In one embodiment, the probes of the present invention can be used to enrich for specific nucleic acid sequences prior to the sequencing step of next generation sequencing.
[0044] In one embodiment, the probe of the present invention hybridizes to a nucleic acid region containing a virtual junction. Such a probe that hybridizes to a nucleic acid region containing a virtual junction is hereinafter also referred to as a "virtual junction-containing probe." Regarding a virtual junction-containing probe or a set thereof, the configuration other than the inclusion of a probe that hybridizes to a nucleic acid region containing a virtual junction, such as the base length y of the region in the probe that hybridizes to cDNA and the number of probes included in the probe set, is the same as that of the "virtual junction-free probe" described above. However, since a virtual junction-containing probe detects only one fusion gene resulting from the fusion of a part of gene A and a part of gene B, it has high specificity but is unable to detect various fusion partners.
[0045] In one embodiment, the virtual junction-containing probe hybridizes to 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, 50 or more, or 60 or more bases in regions derived from both gene A on the 5' side and gene B on the 3' side of the cDNA prepared from the transcription product of the fusion gene.
[0046] In one embodiment, the probe set of the present invention includes a "virtual junction-containing probe" in addition to the above-mentioned "virtual junction-free probe." By including both probes, detection specificity can be further improved. In one embodiment, the probe set of the present invention consists only of probes that satisfy z≧x+y and the virtual junction-containing probe. In yet another embodiment, the probe set of the present invention consists only of probes that satisfy z≧x+y.
[0047] The probe set of the present invention may be designed for the 5' and 3' ends of the exons of all target genes to be evaluated, but it is preferable to design probes only for the 5' and / or 3' ends of exons involved in gene fusions of genes known to form fusion genes.
[0048] In one embodiment, the probe or probe set of the present invention further comprises at least one probe for measuring gene expression levels. A gene expression level measurement probe is a probe used to measure gene expression levels in massively parallel sequencing. Gene expression level measurement probes can be designed to cover all genes whose expression levels are to be measured, and at a density of, for example, 2x tiling or greater. The base length of the gene expression level measurement probe is not limited, and may be, for example, 20 or more, 40 or more, 60 or more, 80 or more, 100 or more, 110 or more, or 115 or more; or 220 or less, 200 or less, 180 or less, 160 or less, 140 or less, 130 or less, or 125 or less; for example, 20 to 220, 60 to 180, 100 to 140, 110 to 130, 115 to 125, or 120. The number of probes for measuring gene expression levels for one gene is not limited, and may be, for example, 3 or more, 4 or more, 5 or more, 6 or more, 8 or more, 10 or more, or 11 or more, or 30 or less, 25 or less, 20 or less, 15 or less, 14 or less, 13 or less, or 12 or less. The probes for measuring gene expression levels may be probes for "multiple" genes, for example, 2 or more, 5 or more, 10 or more, 50 or more, 100 or more, 150 or more, 200 or more, 250 or more, preferably 300 or more, 400 or more, or 500 or more, or may be probes for 2000 or less, 1000 or less, 900 or less, preferably 800 or less, 700 or less, or 600 or less genes. Examples of target genes whose expression levels are measured include oncogenes (e.g., ALK, EGFR, ERBB2, MET) and housekeeping genes. Nucleic acids capable of binding to at least a portion of these genes can be used as probes. By including a probe for measuring the expression level, it becomes possible to measure the expression level of a gene more accurately.
[0049] In one embodiment, the present invention relates to a combination or probe set comprising a plurality of different probes or probe sets described above. Preferably, the combination probe set comprises probe sets for a plurality of different fusion genes, thereby enabling simultaneous detection of the presence or expression levels of transcripts of a plurality of fusion genes. The upper and lower limits of "multiple" are not particularly limited, and may be, for example, 2 or more, 5 or more, 10 or more, 50 or more, 100 or more, 150 or more, 200 or more, 250 or more, preferably 300 or more, 400 or more, or 500 or more, or 2000 or less, 1000 or less, 900 or less, preferably 800 or less, 700 or less, or 600 or less.
[0050] In one embodiment, the probes, probe sets, or combination probe sets described herein are suitable for use with samples containing degraded or degraded RNA, such as transcription products derived from biological samples that have been processed, including heat treatment, freezing, acid treatment, base treatment, and preferably fixation such as FFPE (formalin-fixed paraffin-embedded). 2. Effects of the probe of the present invention As described above, the probes of the present invention can capture and enrich nucleic acid fragments that yield reads containing hypothetical junctions through massively parallel sequencing. Therefore, by performing massively parallel sequencing on the enriched sample, fusion genes can be efficiently detected. Furthermore, in one embodiment, the probe set of the present invention is used for cDNA prepared from transcription products such as mRNA. Since the probes can be concentrated near hypothetical junctions, it can have the advantage of requiring fewer probes than intron capture methods, which capture intronic portions of genomic DNA, and coding exon capture methods, which capture all exon portions. Furthermore, in one embodiment, the probe set of the present invention contains probes concentrated near hypothetical junctions, allowing for the acquisition of various nucleic acid fragments containing hypothetical junctions. Ryan Tewhey et al. (Genome Biology, 2009, 10, R116) showed that probe tiling at a density of 2× or higher does not improve coverage. Therefore, it was surprising that the detection efficiency of fusion genes or exon skipping can be improved by concentrating probes near hypothetical junctions. As used herein, "tiling" refers to the density at which probes are designed for a target gene, and the tiling multiple value n means that the probes are designed at intervals of w / n, where w is the length of the probe.
[0051] In one embodiment, the probe of the present invention does not require the polyA sequence contained in mRNA for transcription or enrichment, and therefore can efficiently detect fusion genes, particularly in samples where RNA has been degraded or degraded. 3. Probes for determining the presence or expression level of transcripts resulting from exon skipping In one aspect, the present invention relates to a probe for determining the presence or expression level of a transcript resulting from exon skipping in massively parallel sequencing, or a probe set comprising at least two different probes. Assuming that exon A' on the 5' side and exon B' on the 3' side are ligated at a virtual junction in the transcript, the probe of this aspect hybridizes to a region derived from either exon A' or B' in cDNA prepared from the transcript. In one embodiment, when the probes hybridize to the cDNA prepared from the transcript, the shortest base length from the end of each probe to the virtual junction is x, the base length of the region in each probe that hybridizes to the cDNA is y, and the read length of the massively parallel sequencing is z, where z is equal to or greater than x+y.
[0052] In one aspect, the present invention relates to a probe for determining the presence or expression level of a transcript resulting from exon skipping in massively parallel sequencing, wherein the transcript has exon A' on the 5' side and exon B' on the 3' side joined at a hypothetical junction, and the probe hybridizes to a region including the hypothetical junction where exon skipping may occur in cDNA prepared from the transcript, or a probe set including at least two different probes of the present invention.
[0053] As used herein, "exon skipping" refers to a phenomenon in which a splicing error results in the removal of some exons along with introns, resulting in abnormal exon ligation. For example, if a wild-type gene contains exons A', B', and C', a splicing error may result in exon B' being skipped, resulting in exon A' and exon C' being ligated, rather than exons A', B', and C' being ligated together. Because the products resulting from exon skipping are abnormal products, they often cause disease. For example, skipping of exon 14 of the mesenchymal-epithelial transition (MET) is known to be associated with the incidence of non-small cell lung cancer.
[0054] The configuration of the probe of this embodiment, other than that for determining the presence or expression level of a transcript resulting from exon skipping, such as the number of probes, the shortest base length x from the end of each probe to the virtual junction point, the base length y of the region in each probe that hybridizes with the cDNA, the read length z for massively parallel sequencing, and the sequence and design of each probe, are similar to those described in "1. Probes for determining the presence or expression level of a transcript of a fusion gene on a genome." The fact that the probe may further include a probe for measuring gene expression levels is also similar to that described in "1. Probes for determining the presence or expression level of a transcript of a fusion gene on a genome." Furthermore, the effects of the probe of this embodiment are similar to those described in "2. Effects of the probe of the present invention" above.
[0055] In one aspect, the present invention relates to a probe set including both the above-mentioned "1. Probe for determining the presence or expression level of a transcript of a fusion gene on a genome" and the "probe for determining the presence or expression level of a transcript resulting from exon skipping" of this aspect. By using this probe set, both the fusion gene and exon skipping can be detected simultaneously. 4. Probe-containing kit In one aspect, the present invention relates to a kit comprising the probe, probe set, or combination probe set described in "1. Probe for determining the presence or expression level of a transcript of a fusion gene on a genome" above and / or "3. Probe for determining the presence or expression level of a transcript resulting from exon skipping" above.
[0056] In addition to the probe, the kit may contain, for example, buffers, enzymes, instructions for use, and the like.
[0057] This kit can be used to determine the presence or expression level of a transcript of a fusion gene, and / or the presence or expression level of a transcript resulting from exon skipping. 5. Method for determining the presence or expression level of a transcript including a transcript of a fusion gene In one aspect, the present invention relates to a method for determining the presence or expression level of a transcript containing a transcript of a fusion gene on a genome. The method of this aspect includes, in this order, a step of preparing a transcript from a sample derived from a subject (transcript preparation step), a step of preparing cDNA from the transcript (cDNA preparation step), a step of enriching target cDNA hybridized to the probes of the probe set or combination probe set described above in "1. Probes for determining the presence or expression level of a transcript of a fusion gene on a genome" (enrichment step), a step of sequencing the enriched target cDNA by massively parallel sequencing (sequencing step), and a step of determining the presence or expression level of a transcript containing a transcript of a fusion gene on a genome based on the sequencing results (determination step).
[0058] Each step of the method will now be described in detail. (1) Transcription product preparation process In the transcription product preparation step, transcription products are prepared from a sample derived from a subject. As used herein, the subject is not limited to a particular species, but is preferably a mammal, for example, a primate such as a human or chimpanzee, a laboratory animal such as a rat or mouse, a livestock animal such as a pig, a cow, a horse, a sheep, or a goat, or a pet animal such as a dog or a cat, preferably a human.
[0059] As used herein, the term "sample" refers to a biological sample subjected to the method of the present invention. Examples of samples that can be used in the present invention include, but are not limited to, bodily fluids, cells, or tissues isolated from a living organism. Examples of bodily fluids include blood, sweat, saliva, milk, and urine. Examples of cells include peripheral blood cells, lymph and tissue fluids containing cells, hair matrix cells, oral cells, nasal cells, intestinal cells, vaginal cells, mucosal cells, and sputum (which may contain alveolar cells or liver cells). Examples of tissues include cancer lesions such as the brain, pharynx, thyroid gland, lung, breast, esophagus, stomach, liver, pancreas, kidney, small intestine, large intestine, bladder, prostate, uterus, and ovaries, preferably the lung. Biopsy samples of these tissues can be used. When a biopsy sample is used, histological pathological diagnosis and detection of the fusion gene by the method of the present invention can be performed simultaneously, allowing for more accurate identification of the subject's pathological condition.
[0060] In one embodiment, the sample used is a sample containing degraded or degraded RNA, such as a biological sample that has been subjected to processing, such as heat treatment, freezing, acid treatment, base treatment, or preferably fixation such as FFPE (formalin-fixed paraffin embedding).
[0061] The transcription product (total RNA) may include rRNA, tRNA, and mRNA, but is preferably mRNA.
[0062] Transcription products can be prepared from samples using any known method. For example, transcripts can be extracted by mixing a sample with a solubilizing solution containing guanidine thiocyanate and a surfactant, and then subjecting the resulting mixture to physical treatment (e.g., stirring, homogenization, sonication, etc.). Preferably, a method (AGPC method) can also be used in which phenol and chloroform are further added, followed by stirring and centrifugation to recover an aqueous layer containing the transcripts. Subsequently, transcripts can be obtained from the aqueous layer by alcohol precipitation or other methods. Alternatively, commercially available kits such as RNA-Bee (Tel-Test Inc.) and TRIZOL (Thermo Fisher Scientific) can be used to extract RNA. For specific procedures, please refer to protocols in the field, such as Green and Sambrook, Molecular Cloning, 4th Ed. (2012), Cold Spring Harbor Laboratory Press. For other biological techniques described herein, such as the following cDNA preparation and concentration steps, please also refer to Green and Sambrook (cited above). (2) cDNA preparation process cDNA can be produced from the transcript obtained in the transcript preparation step by reverse transcription using a reverse transcriptase. Those skilled in the art can appropriately select known primers, reverse transcriptases, and reaction conditions for use in the reverse transcription reaction. In the method of the present invention, the target nucleic acid fragment is concentrated by the concentration step described below, so there is no need to reverse transcribe only mRNA using a poly(A) sequence; for example, total RNA can be reverse transcribed using a random primer or the like. (3) Concentration process In the enrichment step, target cDNA hybridized to the probe, probe set, or combination probe set described herein is enriched. Enrichment can be performed using any method known to those skilled in the art. For example, a label can be attached to the probe, and the target cDNA hybridized to the probe can be enriched by interaction between the label and another substance. For example, biotin can be attached to the probe, and the cDNA hybridized to the probe can be enriched by interaction with avidin. Alternatively, enrichment can be performed by affinity chromatography using a substrate or antigen-antibody reaction. Alternatively, magnetic beads can be attached to the probe, and the cDNA hybridized to the probe can be enriched by magnetic attraction.
[0063] Before or after the enrichment step using the probe set, the cDNA may be enzymatically or mechanically fragmented to match the read length in massively parallel sequencing. Furthermore, adapter sequences necessary for the subsequent sequencing step may be added. To analyze a specific gene region before or after the enrichment step, the specific gene region may be enriched by PCR or other methods. Enrichment of the gene region can be achieved, for example, by an amplification step of 4 to 12 cycles. (4) Sequencing step In the sequencing step, the enriched target cDNA is sequenced by massively parallel sequencing. The details of the sequencing step vary depending on the type of instrument used for massively parallel sequencing, but typically the cDNA is linked to a substrate via an adapter sequence, and the adapter sequence is used as a priming site for the sequencing reaction. For details of the sequencing reaction, see, for example, Rick Kamps et al. (supra).
[0064] This step yields a collection of sequence information (reads) obtained by the sequencing reaction. The output data can be further analyzed to derive more meaningful results, such as the number of reads, e.g., the number of ligated support reads per raw read. Massively parallel sequencing devices are commercially available from various manufacturers and can be used. For example, but not limited to, Roche's Genome Sequencer (GS) FLX System, Illumina's HiSeq or Genome Analyzer (GA), Life Technologies' Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator's G.007 system, and Helicos BioSciences' HeliScope Gene Sequencing system can be used. (5) Judgment process In the determination step, the presence or expression level of a transcript containing a transcript of the fusion gene on the genome is determined based on the results of the sequencing step. An example of the determination step is shown in Figure 1B. The specific method for the determination step is not limited, but can be performed, for example, according to the following criteria.
[0065] Assuming that a fusion gene expresses a transcription product in which a part of gene A on the 5' side and a part of gene B on the 3' side are joined at a virtual junction, the number of reads of cDNA derived from gene A where no gene fusion occurs at the virtual junction is defined as α, the number of reads of cDNA derived from gene B is defined as β, and the number of reads of cDNA derived from a fusion gene where gene fusion occurs at the virtual junction is defined as γ. If 0<α or β≦γ, it is determined that the fusion gene is present; If 0<γ<α or β, it is determined that the fusion gene is present at a low expression level; If α or β>0 and γ=0, it can be determined that the fusion gene is not present.
[0066] When α and / or β = 0 and γ = 0, it is considered that either the transcript of the fusion gene is not present or the transcript is degraded due to poor sample quality. In this case, it is possible to accurately determine which is correct by more precisely counting the reads near the virtual junction of the wild-type transcripts of both genes of the putative fusion gene.
[0067] Since the number of reads is usually proportional to the amount of nucleic acid before sequencing, the expression level can be determined based on the number of reads of the gene. The expression level can be determined as a relative value, for example, by comparing the number of reads with that of a wild-type gene or with that of a healthy subject, or it can be determined as an absolute value, such as the number of reads measured under specific conditions.
[0068] In one embodiment, when multiple probes hybridize to the same region, the determination step includes correcting the expression level of the transcript based on the number of probes. Because the probe set of the present invention contains probes concentrated near the virtual junction, probes may be designed to overlap in the same region. Accordingly, the number of reads for the transcript corresponding to the region may be calculated as higher depending on the number of probes. Therefore, to more accurately determine the expression level based on the number of reads, it is preferable to correct the number of reads by the number of probes hybridizing to the same region. The method for correcting the number of reads based on the number of probes is not limited. For example, the number of reads can be corrected by dividing the number of reads by the number of probe tilings (e.g., for 5x tiling, the number of reads can be divided by 5, and for 10x tiling, the number of reads can be divided by 10).
[0069] In one embodiment, the determining step includes correcting the expression level of the transcript based on the expression level of at least one housekeeping gene. Correction based on a housekeeping gene is particularly preferred for more accurate comparison of expression levels when using different probe sets and / or different samples. Housekeeping genes known in the art can be used, for example, at least one, at least two, at least three, at least five, or all of ACTB, B2M, GAPDH, GUSB, H3F3A, HPRT1, HSP90AB1, NPM1, PPIA, RPLP0, TFRC, and UBC. The method for correcting the number of reads based on housekeeping genes is not limited. For example, the number of reads can be corrected by dividing the number of reads of the transcript whose expression level is to be measured by the number of reads of the housekeeping gene.
[0070] The method of this embodiment can diagnose diseases by determining the presence or expression level of a fusion gene in the genome, and can also select appropriate therapies, such as drugs, based on the genetic background of a subject, such as information on the presence or expression level of a fusion gene in the genome. 6. Method for determining the presence or expression level of transcripts, including transcripts resulting from exon skipping In one aspect, the present invention relates to a method for determining the presence or expression level of a transcript, including a transcript resulting from exon skipping. The method of this aspect comprises, in this order, a step of preparing a transcript from a sample derived from a subject (transcript preparation step), a step of preparing cDNA from the transcript (cDNA preparation step), a step of enriching target cDNA hybridized to the probe, probe set, or probe combination set described in "3. Probes for determining the presence or expression level of a transcript resulting from exon skipping" above (enrichment step), a step of sequencing the enriched target cDNA by massively parallel sequencing (sequencing step), and a step of determining the presence or expression level of a transcript, including a transcript resulting from exon skipping, based on the sequencing results (determination step).
[0071] Except for the fact that this method is for determining the presence or expression level of a transcript, including a transcript generated by exon skipping, and that the probes used are different, the configuration of this method, such as the transcript preparation step, cDNA preparation step, concentration step, sequencing step, and determination step, is similar to the above-mentioned "5. Method for determining the presence or expression level of a fusion gene transcript." Therefore, the following description will focus on the differences from the above-mentioned "5. Method for determining the presence or expression level of a fusion gene transcript."
[0072] In one aspect, the present invention relates to a method for performing a cDNA enrichment step using both the above-mentioned "1. Probe for determining the presence or expression level of a transcript of a fusion gene on a genome" and the above-mentioned "3. Probe for determining the presence or expression level of a transcript resulting from exon skipping," which allows for simultaneous detection of both the fusion gene and exon skipping.
[0073] The determination step can be carried out as described above in "5. Method for determining the presence or expression level of a transcript of a fusion gene." That is, assuming that exon A' on the 5' side and exon B' on the 3' side are ligated at a hypothetical junction in a transcript, the number of reads of cDNA derived from exon A' where no gene fusion occurs at the hypothetical junction is defined as α', the number of reads of cDNA derived from exon B' is defined as β', and the number of reads of cDNA derived from a transcript generated by exon skipping is defined as γ'. If 0<α' or β'≦γ', it is determined that a transcript resulting from exon skipping is present; If 0<γ'<α' or β', it is determined that a transcript resulting from exon skipping is present at a low expression level; When α' or β'>0 and γ'=0, this can be achieved by determining that no transcript resulting from exon skipping is present. 7. Methods for determining the presence or absence of a disease or the risk of such a disease, identifying the type of cancer, or determining the prognosis of cancer In one aspect, the present invention relates to a method for determining the presence or risk of a disease in a subject, identifying the type of cancer (e.g., primary cancer), or determining the prognosis of cancer (or a cancer patient), comprising a step of determining the presence or expression level of a transcript of a fusion gene on the genome, and / or a transcript resulting from exon skipping, according to the methods described herein. The determination step can be performed as described above in "5. Method for determining the presence or expression level of a fusion gene transcript" and / or "6. Method for determining the presence or expression level of a transcript resulting from exon skipping." The method of this aspect differs from the methods described above in "5. Method for determining the presence or risk of a fusion gene transcript" or "6. Method for determining the presence or expression level of a transcript resulting from exon skipping" in that the method determines the presence or risk of a disease, identifies the type of cancer, or determines the prognosis of cancer.
[0074] In the method of this embodiment, the type of disease is not limited as long as the presence or risk of the disease can be determined by fusion genes or exon skipping, and examples include cancers such as brain tumors, pharyngeal cancer, thyroid cancer, lung cancer, breast cancer, esophageal cancer, gastric cancer, liver cancer, pancreatic cancer, kidney cancer, small intestine cancer, colon cancer, bladder cancer, prostate cancer, cervical cancer, ovarian cancer, sarcoma, lymphoma, or melanoma, preferably lung cancer or sarcoma.
[0075] In addition to the determination step, the method of this embodiment may include a step of assessing the presence or absence of a disease or the risk of a disease in a subject (evaluation step), a step of identifying the type of cancer (identification step), or a step of determining the prognosis of cancer (determination step) based on the presence or expression level of a transcript of the fusion gene on the genome and / or the presence or expression level of a transcript generated by exon skipping. Evaluation process The evaluation step can be performed by utilizing known associations between fusion genes or exon skipping and diseases. For example, EML4 (echinoderm microtubule associated protein like 4)-ALK (anaplastic lymphoma kinase) can be used to determine the presence or risk of non-small cell lung cancer, BCR (B cell receptor)-ABL1 (Abelson murine leukemia viral oncogene homolog 1) can be used to determine the presence or risk of non-small cell lung cancer, TAF15 (TATA-box binding protein associated factor 15)-NR4A3 (nuclear receptor subfamily 4 group A member 3) can be used to determine the presence or risk of non-small cell lung cancer, AHRR (aryl-hydrocarbon receptor repressor)-NCOA2 (nuclear receptor coactivator 2) can be used to determine the presence or risk of non-small cell lung cancer.
[0076] In the evaluation step, if the presence of a transcription product of the fusion gene or the presence of a transcription product resulting from exon skipping is detected, or if the expression level of the fusion gene or the expression level of the transcription product resulting from exon skipping is higher than, for example, that of a healthy subject, it can be evaluated that the subject is suffering from the disease or has a high risk of having the disease. Identification process and judgment process Identification of cancer type and prognosis of cancer can be performed by utilizing the association between a disease and a transcript, including a transcript of a fusion gene on the genome and / or a transcript generated by exon skipping. The association between the transcript and the disease may be known or unknown.
[0077] As used herein, "prognosis" refers to reduction in tumor burden, suppression of tumor growth, the course or outcome of a disease (e.g., presence or absence of recurrence, life or death, etc.), preferably the length of survival time, and the level of risk of recurrence, after therapeutic treatment such as chemotherapy. Determination of prognosis may be, for example, a prediction of survival time or survival rate after a certain period of time after therapeutic treatment.
[0078] In one embodiment, the identification and determination step includes clustering samples derived from subjects based on the presence and / or expression levels of multiple transcripts. This embodiment is particularly advantageous when the association between the transcripts and disease is unknown. The number of transcripts in this embodiment is not limited, but may be, for example, 2 or more, 5 or more, 10 or more, 20 or more, 30 or more, 50 or more, 100 or more, 200 or more, 300 or more, 400 or more, or 500 or more, or 20,000 or less, 10,000 or less, 5,000 or less, preferably 3,000 or less, 2,000 or less, or 1,000 or less. When clustering samples based on the presence and / or expression levels of multiple transcripts, a standard sample derived from a subject whose cancer type has been identified or whose prognosis has been predicted can be added. This enables more accurate clustering based on cancer type or prognosis. The clustering method is not limited, but for example, samples can be clustered based on gene expression levels using the statistical analysis software R heatmap.3.
[0079] The type of cancer in the identifying step is not limited, but may be, for example, brain tumor, pharyngeal cancer, thyroid cancer, lung cancer (e.g., lung adenocarcinoma), breast cancer, esophageal cancer, stomach cancer, liver cancer, pancreatic cancer, kidney cancer, small intestine cancer, colon cancer, bladder cancer, prostate cancer, cervical cancer, ovarian cancer, sarcoma, lymphoma, or melanoma, preferably lung cancer (e.g., lung adenocarcinoma) or sarcoma.
[0080] The method of this embodiment for determining the presence or absence of a disease or the risk thereof, for identifying the type of cancer, or for determining the prognosis of cancer may be performed in combination with other methods, such as histological pathological diagnosis, biomarker detection by FISH, RT-PCR, immunohistochemistry, etc., and imaging diagnosis such as CT, MRI, and nuclear medicine examination. Combining with other methods can improve the accuracy of disease detection. [Example]
[0081] Materials and Methods gDNA targeted sequencing Genomic DNA (500 ng) was isolated from FFPE samples using the GeneRead DNA FFPE Kit (Qiagen), and target fragments were enriched using the SureSelectXT Custom Kit (Agilent). Custom-made probes were designed to hybridize and capture the gDNA of the target genes. Massively parallel sequencing of the isolated fragments was performed using the HiSeq2500 platform (Illumina) with the paired-end option. From the large dataset, only sequence reads with a Q value ≥ 20 for each base were selected and mapped to the reference human genome sequence (hg19) using the bowtie 2 algorithm (http: / / bowtie-bio.sourceforge.net / bowtie2 / index.shtml). Somatic mutations were identified using MuTect (http: / / www.broadinstitute.org / cancer / cga / mutect). In addition, mutation candidates were selected based on the following criteria: judgment = KEEP (KEEP indicates positive somatic mutations due to mutect), tumor read depth ≥ 20x, mutation rate ≥ 10%, and normal read depth ≥ 10x. RNA-seq with polyA selection Total RNA was extracted from fresh-frozen samples using RNA-Bee (Tel-Test Inc., # CS-104B), treated with DNase I (Life Technology), and then subjected to polyA-RNA selection. RNA-seq libraries were prepared using the NEBNext Ultra Directional RNA Library Prep Kit (New England Bio Labs) according to the manufacturer's protocol. Next-generation sequencing (NGS) was performed from both ends of each cluster using the HiSeq2500 platform (Illumina). RNA-seq by cDNA capture Total RNA was extracted from FFPE samples using the RNeasy FFPE Kit (Qiagen) and treated with DNase I (Life Technologies). cDNA synthesis, probe capture, and library preparation for coding exon capture were performed using the TruSight RNA Pan-Cancer Panel (Illumina) according to the manufacturer's protocol.
[0082] cDNA synthesis and library preparation for junction capture were performed using the SureSelect RNA Capture kit (Agilent Technologies) according to the manufacturer's protocol. Custom probes for junction capture were designed to hybridize and capture sequences near the hypothetical junctions of target genes. Specifically, considering the 170-bp read length of the massively parallel sequencing used, and assuming that a hybridization region of the probe with cDNA of 50 or more bases would yield reads containing the hypothetical junctions, the probes were designed so that the minimum length from the end of each probe to the hypothetical junction when hybridized to the cDNA was 120 bases or less. All probes were 120 bp long. To obtain as many different reads as possible, probes were designed with 5x or 10x tiling for junction capture. NGS sequencing was performed from both ends of each cluster using the HiSeq2500 platform (Illumina). As an example, the sequence numbers of the probe sets used to identify exon 13 of EML4, exon 20 of ALK, and the EML4-ALK fusion gene are shown in Table 1 below.
[0083] [Table 1]
[0084] Example 1: Detection of fusion genes by junction capture method result The sequencing data were analyzed by counting the number of sequence reads supporting the presence of the fusion transcript junction and examining whether the fusion transcript was significantly expressed compared to the wild-type gene transcript.
[0085] Furthermore, if each gene transcript was present but the fusion gene transcript was absent, it indicated that the fusion transcript was absent. However, if the number of reads for each gene was 0, we carefully evaluated whether the mRNA was not expressed or whether it was due to mRNA degradation based on the quality of the sample.
[0086] As a pilot study, we developed a small target panel (TOP RNA V1) targeting 67 fusion genes based on the junction capture method and compared it with a panel obtained by the conventional intron capture method (TOP DNA), which detects the genomic junctions of fusion genes, and the TruSight RNA Pan-Cancer Panel (Illumina), which is based on the coding exon capture method.
[0087] The TOP RNA V1 panel obtained by the junction capture method was able to detect fusion genes more accurately than the TOP DNA panel obtained by the intron capture method, and also had a higher value of junction-supported reads / 10 million raw reads (Table 2, Figure 2A). These results suggest that the junction capture method is a superior method for detecting fusion genes.
[0088] [Table 2]
[0089] Next, we designed a larger target panel (TOP RNA V2) using the junction capture method, covering sarcoma fusion genes, and a panel (TOP RNA V3) covering all fusion genes reported in the COSMIC database. Although the RNA integrity score (RIN) of the FFPE samples from which RNA was extracted ranged from 1.1 to 2.3, indicating severe degradation, all fusion transcripts were detectable (Table 3). Furthermore, the junction capture method significantly reduced both the number of expected probes and the target capture size (the length of the nucleic acid sequence captured by the probe) compared to panels designed using the coding exon capture method (Figure 2B and Figure 2C). This suggests that the junction capture method is highly cost-effective.
[0090] The quality of RNA-seq can be assessed by calculating housekeeping gene coverage and coverage percentage. The following criteria were used to define excellent RNA-seq quality: average housekeeping gene coverage >500X and 100X, and housekeeping gene coverage >70%. The absence of ligation-supporting reads may be due to the advanced degradation of FFPE-derived RNA. Therefore, to ensure true fusion gene negatives, we developed a pipeline to count ligation-supporting reads of wild-type transcripts of both genes in the putative fusion gene reported in the COSMIC database. The results of this analysis for case #31 (EML4-ALK-positive lung adenocarcinoma) confirmed that this tumor was truly negative for the analyzed fusion transcripts (data not shown). Example 2: Detection of exon skipping by junction capture method Next, we investigated whether the junction capture method can detect transcripts such as MET exon 14 skipping, which has been reported to be oncogenic in lung adenocarcinoma. RNA was extracted from five FFPE samples from lung adenocarcinoma cases identified as having MET exon 14 skipping by RNA-seq using fresh frozen samples. The number of junction-supporting reads supporting the junction of exon 13 to exon 15, i.e., exon 14 skipping, was counted. The junction capture method identified MET exon 14 skipping in all five FFPE samples with exon skipping, but no junction-supporting reads were observed in any of the other 34 cases without MET exon skipping (Figure 3, Table 3). This indicates that the junction capture method can also detect exon skipping.
[0091] [Table 3]
[0092] Example 3: Application of the junction capture method to biopsy samples We also evaluated the applicability of the junction capture method to small biopsy samples. We prepared RNA from fusion-positive FFPE specimens, including core needle biopsies, fine needle aspiration biopsies, and transbronchial lung biopsies (TBLB). Surprisingly, in all RNA-seq experiments, we detected numerous junction-supporting reads supporting the correct fusion transcripts specific to each specimen (Figure 4, Table 4).
[0093] [Table 4]
[0094] Example 4: Clinical utility of the junction capture method The clinical utility of this method was evaluated by testing junction capture on FFPE specimens obtained from surgical resections of 40 cases of stage II or III NSCLC negative for KRAS and EGFR mutations. MET exon 14 skipping, EML4-ALK fusion gene, and RET fusion gene were detected in 3, 2, and 1 cases, respectively (data not shown). To evaluate the clinical utility of junction capture for the diagnosis of sarcoma, junction capture was also performed on sarcoma patients in a prospective study. The results are shown in Table 5 below.
[0095] [Table 5]
[0096] One case (#44) was diagnosed as myxofibrosarcoma due to proliferation of spindle cells with atypical nuclei adjacent to the myxoid stroma. However, this case was found to be a soft tissue angiofibroma because the AHRR-NCO2A gene fusion gene, which is specific to angiofibroma, was detected by junction capture analysis. Another case (#48) was positive for TAF15-NR4A3, which is consistent with a diagnosis of extraskeletal chondrosarcoma.
[0097] These results indicate that the junction capture method can be used for the diagnosis of diseases. Example 5: Measurement of gene expression levels In this example, gene expression levels were measured using the junction capture method. (Materials and Methods) Gene expression measurement Total RNA was extracted from FFPE samples and RNA-seq using cDNA capture (junction capture) was performed for 11 housekeeping genes (ACTB, B2M, GAPDH, GUSB, H3F3A, HPRT1, HSP90AB1, PPIA, RPLP0, TFRC, and UBC) according to Example 1. For comparison, total RNA was also extracted from fresh-frozen samples according to Example 1 and RNA-seq using poly(A) selection was performed.
[0098] However, in this example, in addition to the custom probe (TOP RNA V3) for the junction capture method shown in Example 1, a probe for measuring gene expression levels was added for enrichment. The probes used for measuring gene expression levels were probes designed with 2x tiling for 125 genes, including oncogenes such as ERBB2. All probes were 120 bases long. Correction of the number of reads based on the number of tilings As described in Example 1, in order to obtain as many types of reads as possible in the junction capture method, probes were designed with 5x or 10x tiling, concentrating them near the virtual junction points. Therefore, when estimating gene expression levels based on the number of reads, there is a risk that the calculated expression level will be higher depending on the number of probes. Therefore, in the junction capture method, the number of reads was corrected by dividing the number of reads by the number of probe tilings (for example, for 5x tiling, the number of reads was divided by 5, and for 10x tiling, the number of reads was divided by 10). Read number correction based on housekeeping genes Because FFPE samples (Group A) were used for the junction capture method and fresh-frozen samples (Group B) were used for RNA-Seq with poly(A) selection, differences in sample quality were corrected by equalizing the expression levels of housekeeping genes in both groups. Specifically, coefficients were calculated to correct the expression levels of Group B so that the log_2 average ratio of the expression levels of 11 housekeeping genes in Group A and Group B was equal, and these coefficients were used to correct the expression levels of all genes. (result) Expression levels of 11 housekeeping genes (ACTB, B2M, GAPDH, GUSB, H3F3A, HPRT1, HSP90AB1, NPM1, PPIA, RPLP0, TFRC, and UBC) were measured in seven samples from lung cancer patients using poly(A)-selected RNA-Seq and junction capture.
[0099] As a result, for housekeeping genes, a correlation was observed between the RPKM (Reads Per Kilobase of exon model per Million mapped reads) values for polyA-selected RNA-Seq and the junction capture method (data not shown).
[0100] Next, we calculated the correlation coefficients between the RPKM of RNA-seq and the RPKM corrected based on the number of tilings in the junction capture method for the gene group used for expression level measurement and the gene group used for fusion gene analysis. Here, the gene group used for expression level measurement is the gene group whose expression was measured using a probe for gene expression level measurement, and the gene group used for fusion gene analysis is the gene group whose expression was measured using a custom probe for the junction capture method.
[0101] The results for the gene group used for expression level measurement are shown in Figure 5A and Table 6, and the results for the gene group used for fusion gene analysis are shown in Figure 5B and Table 7. A correlation was observed between the RPKM of RNA-seq and the RPKM of the junction capture method for both the gene group used for expression level measurement and the gene group used for fusion gene analysis, with a stronger correlation observed for the gene group used for expression level measurement. These results indicate that while probes used for gene expression level measurement are more suitable for measuring expression levels, custom probes for junction capture can also be used for measuring expression levels. These results also indicate that accurate gene expression measurements are possible even when custom probes for junction capture are used in addition to probes used for gene expression level measurement.
[0102] [Table 6]
[0103] [Table 7]
[0104] Example 6: Cancer clustering based on gene expression Gene expression measurements were performed on samples from patients with LUAD (lung adenocarcinoma), SARC (sarcoma), MUCA (multiple carcinoma), and LUSC (lung squamous cell carcinoma) using the junction capture method, with the addition of probes for gene expression measurement, as described in Example 5. Specifically, expression values were determined for a total of 467 genes, both for expression measurement and fusion gene analysis, by correcting the number of reads based on the number of tilings and the number of reads based on housekeeping genes, as described in Example 5. The determined expression values (xn, n = 1,...,N, where N is the number of genes) were logarithmically transformed (log_2(xn+1)), and clustering was performed based on the values using heatmap.3 in the statistical analysis software R.
[0105] As a result, LUAD, SARC, MUCA, and LUSC were clustered based on the gene expression levels, as shown in Figure 6. This indicates that the type of primary cancer can be identified by measuring gene expression levels using the method of the present invention. [Industrial Applicability]
[0106] The present invention provides a method for easily detecting a transcription product resulting from a fusion gene and / or exon skipping, which enables the diagnosis of a disease and the selection of an appropriate drug based on the genetic background of a subject, and thus has great industrial applicability.
[0107] All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety.
Claims
1. In massively parallel sequencing, The presence or expression level of the transcript of the fusion gene on the genome, and / or Presence or expression level of transcripts resulting from exon skipping A probe set for determining comprising three or more capture probes designed at a density of 2x tiling or greater; cDNA prepared from a transcription product expressed from the fusion gene, in which a portion of gene A on the 5' side and a portion of gene B on the 3' side are joined at the junction, and / or cDNA prepared from a transcript in which exon A' on the 5' side and exon B' on the 3' side are joined at the junction by exon skipping wherein when each capture probe hybridizes to the cDNA, x is the shortest base length from the end of the capture probe to the junction point, y is the base length of the region in each capture probe that hybridizes to the cDNA, and z is the read length of massively parallel sequencing, z≧x+y and x is 300 or less.
2. The probe set according to claim 1 , which consists of three or more capture probes.
3. The probe set according to claim 1 or 2, wherein x is 0 to 250, y is 20 to 220, and z is 50 to 500.
4. The probe set according to any one of claims 1 to 3, wherein the three or more capture probes hybridize to a region of the cDNA that includes the junction point.
5. A combined probe set consisting of the probe sets according to any one of claims 1 to 4 for a plurality of different fusion genes.
6. The probe set according to any one of claims 1 to 4 or the combined probe set according to claim 5, further comprising at least one probe for measuring gene expression levels.
7. A probe set according to any one of claims 1 to 4 or a combination probe set according to claim 5, for use with transcription products derived from processed biological samples.
8. A kit for determining the presence or expression level of a transcription product of a fusion gene, and / or the presence or expression level of a transcription product resulting from exon skipping, comprising the probe set according to any one of claims 1 to 4 or the combination probe set according to claim 5.
Citation Information
Patent Citations
Enrichment and sequencing of targeted DNA
JP2015516814A
Enrichment and next-generation sequencing of total nucleic acids, including both genomic DNA and cdan
JP2016510992A
Target sequence enrichment
JP2016515384A
Capture probes and primers for genotyping of chronic myeloid leukemia, and uses thereof
KR1020130083185A
Locked nucleic acids for capturing fusion genes
WO2017015513A1