A method, system, device, and medium for analyzing EML4-ALK fusion gene types

By combining RNA sequencing with multiplex PCR and high-throughput sequencing, and using known type fusion gene reference sequences and primer combinations, the accuracy problem of fusion gene type detection at the RNA level was solved, and the identification of known and unknown types of fusion genes was achieved, supporting personalized treatment.

CN119360963BActive Publication Date: 2025-10-10HANGZHOU LC BIOTECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411493985.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2025-10-10
Estimated Expiration
2044-06-26

AI Technical Summary

Technical Problem

Existing technologies lack effective methods to accurately analyze fusion gene types at the RNA level, resulting in inaccurate detection of fusion gene subtypes, which affects the effectiveness of personalized molecular targeted therapy.

Method used

Through an RNA sequencing-based method, multiplex PCR amplification and high-throughput sequencing are performed using reference sequence files and primer combinations of known types of fusion genes. Combined with data quality control and comparison analysis, sequencing data that can be aligned to specific exon sequences are identified and counted to determine the type of gene fusion.

Benefits of technology

It achieves accurate detection of known types of fusion genes and helps discover unknown types of fusion genes, supporting more accurate individualized treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360963B_ABST
    Figure CN119360963B_ABST
Patent Text Reader

Abstract

The application discloses a kind of methods, systems, equipment and media for analyzing EML4-ALK fusion gene type, belong to the technical field of fusion gene detection.The method comprises obtaining RNA sequencing data of a sample to be tested;According to the primer combination coverage area, obtain the primer coverage area sequence file;The RNA sequencing data is compared with the primer coverage area sequence file for the second time;The number of sequencing data that can be aligned to the exon sequence of two genes at the same time is counted.Using the method and system of the application, unknown type of fusion gene can be accurately detected, which is more conducive to the development of downstream research on the related mechanism and function of fusion gene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related patents

[0002] This application is a divisional application of the Chinese invention patent application with application number 2024108311341, application date June 26, 2024, and invention name “A method, system, device and medium for analyzing fusion gene types based on RNA sequencing”. Technical Field

[0003] The present invention belongs to the technical field of fusion gene detection, and in particular, relates to a method, system, device and medium for analyzing the type of EML4-ALK fusion gene. Background Art

[0004] Lung cancer ranks first in both morbidity and mortality among malignant tumors in my country. Non-small cell lung cancer (NSCLC) accounts for approximately 85% of all cases, with >70% of patients diagnosed at an advanced stage. Gene fusion (rearrangement) is a driver of NSCLC. Fusion of the echinoderm microtubule-associated protein link 4 (EML4) gene with the anaplastic lymphoma kinase (ALK) gene is the most common type of ALK fusion. All EML4-ALK fusion genes exhibit certain biological activity, leading to sustained high expression of the ALK tyrosine kinase, thereby activating downstream signaling pathways such as PI3K / AKT and MAPK, thereby driving tumor development and metastasis. The prevalence of ALK rearrangement (fusion) in NSCLC patients is approximately 3%-7%, and in Chinese patients, the prevalence is approximately 3%-11%. In particular, the prevalence of ALK rearrangement (fusion) is as high as 30%-42% in lung adenocarcinoma patients without EGFR or KRAS mutations.

[0005] Numerous studies have demonstrated that personalized molecular targeted therapy targeting driver genes has significant efficacy and a high safety profile for NSCLC patients. Currently recommended targeted therapies for ALK fusions include the ALK inhibitors ceritinib, alectinib, crizotinib, and brigatinib. Crizotinib, a targeted therapy for ALK fusions, is already available in China.

[0006] Currently, the techniques for detecting fusion genes mainly include fluorescence in situ hybridization (FISH), reverse transcription polymerase chain reaction (RT-PCR), immunohistochemistry (IHC), and next-generation sequencing (NGS). Traditional methods such as fluorescence in situ hybridization, immunohistochemistry, and RT-PCR have the disadvantages of low throughput and high operation requirements. In recent years, next-generation sequencing (NGS) has gradually been valued and promoted due to its advantages of high throughput, wide detection range, and ability to analyze unknown types of fusion genes.

[0007] The detection of fusion genes can be performed at the DNA level or at the RNA level. Structural variations occur at the DNA level, which may not affect the transcription of cancer-driving RNA and thus affect protein synthesis. RNA-level detection can accurately determine the type of fusion gene that occurs during transcription and can also confirm the results of DNA-level fusion detection. Combining multiple PCR amplification and high-throughput sequencing methods can detect multiple types of fusion types at one time.

[0008] The raw data obtained by high-throughput sequencing need to be analyzed to obtain the RNA-derived fusion gene type. The analysis process of high-throughput sequencing data generally includes data quality control, sequence alignment, and alignment result statistics. The key point of fusion gene analysis is to obtain the genes that have undergone fusion and the sites where fusion has occurred. Gene fusion has multiple subtypes, which may have certain similarities and differences in sequence. Different subtypes may have different biological and clinical significance, and should be distinguished as much as possible during analysis. However, there is currently a lack of effective analysis methods to make the detection of fusion gene subtypes more accurate. SUMMARY

[0009] To solve at least one of the above technical problems, the technical solutions adopted by the present application are as follows:

[0010] The first aspect of the present application provides a method for analyzing fusion gene types based on RNA sequencing, comprising the following steps:

[0011] Obtaining RNA sequencing data of a sample to be tested;

[0012] Obtaining a fusion gene reference sequence file of known types of fusion genes, at least including the names of known types of fusion genes and fusion reference sequences, wherein the fusion reference sequence is obtained by extending 60-75 bases forward and backward from the breakpoint, and thus the length of the fusion reference sequence is 120bp-150bp;

[0013] Performing a first comparison between the RNA sequencing data and the fusion reference sequence file, and performing the following processing:

[0014] (1) if the base mismatching ratio of a sequencing data to the fusion reference sequence is less than 5%, the sequencing data is reserved;

[0015] (2) if a sequencing data crosses the breakpoint on the fusion reference sequence and the number of bases covered on both sides of the breakpoint is greater than 20, the sequencing data can be aligned to the fusion reference sequence,

[0016] (3) if a sequencing data can be aligned to multiple fusion reference sequences in the fusion reference sequence file, only the alignment result of the priority is reserved,

[0017] The number of sequencing data that can be aligned to each fusion reference sequence is counted, and if the number of sequencing data that can be aligned to a certain fusion reference sequence is greater than a preset threshold, the sample to be tested has a corresponding type of gene fusion.

[0018] In the present application, gene fusion refers to the fusion of a part of a first gene and a part of a second gene, wherein the first gene is abnormally broken at a certain point and fused with the second gene to form a hybrid gene composed of two originally distant genes, i.e. a fusion gene. Gene fusion is also known as gene rearrangement, i.e. the order of genes on the chromosome is changed.

[0019] In some embodiments of the present application, the name of the known type of fusion gene in the fusion gene reference sequence file at least includes the name of the first gene, the exon number of the first gene that is fused; and the name of the second gene, the exon number of the second gene that is fused.

[0020] In some specific embodiments of the present application, the first gene is EML4 gene and the second gene is ALK gene. The fusion gene composed of EML4-ALK is composed of two parts, EML4 provides a sequence from the start to a certain middle exon, and the middle position is not unique. ALK provides a sequence from a certain middle exon to the end, and the middle position is not unique, and the connection site of the two genes is called fusion site or breakpoint.

[0021] EML4 belongs to the echinoderm microtubule-associated protein-like family, including WD-repeat region, N-terminal Basic region and HELP domain. They are all related to the potential of EML4-ALK tumorigenesis, and the Basic region plays an important role in the process of EML4-ALK dimerization. Removing the Basic region can reduce the catalytic activity of EML4-ALK fusion protein by 84%, which can be inferred that the Basic region may be the key region of EML4-ALK fusion protein to play a role.

[0022] ALK is a member of the insulin receptor superfamily, which can be fused with ATIC-, TFG-, CLTC-, and other genes. Its structure includes extracellular ligand binding region, transmembrane region, and intracellular tyrosine kinase region.

[0023] EML4 and ALK are located on the short arm of chromosome 2, but in opposite directions, separated by 12 Mb. When fused, one of EML4 or ALK must be reversed and connected to the other. The EML4-ALK gene subtype depends on the position of the two breakpoints. The EML4 breakpoint shows variability, and the breakpoints of exons 2, 6, 13, 14, 15, 17, 18, 20, and 23 have been detected. These form at least eight types of EML4-ALK fusion genes: EML4-ALK.E13:A20.COSF408, EML4-ALK.E6a:A20.COSF411, EML4-ALK.E6:A20.COSF1544, EML4-ALK.E6:A19.COSF1296, EML4-ALK.E20:A20.COSF409, EML4-ALK.E18:A20.COSF487, EML4-ALK.E14:A20.COSF491, and EML4-ALK.E2:A20.COSF478.

[0024] In some embodiments of the present application, the reference sequences of the above eight types of EML4-ALK fusion genes are shown in SEQ ID No. 1~SEQ ID No. 8, respectively.

[0025] In some embodiments of the present application, the RNA sequencing data of the test sample is obtained by the following steps:

[0026] obtaining the RNA sample of the test sample;

[0027] performing multiplex PCR amplification on the RNA sample using a primer combination;

[0028] performing high-throughput sequencing on the amplification product to obtain the RNA sequencing data.

[0029] In some embodiments of the present application, the test sample is a biological sample of a subject, which includes but is not limited to paraffin-embedded tissue samples, pathological tissue samples, blood samples, pleural effusion samples, and cerebrospinal fluid samples.

[0030] In some embodiments of the present application, further comprising the step of quality control of the RNA sequencing data:

[0031] (1) removing adapter sequences;

[0032] (2) removing low-quality sequences,

[0033] If the remaining data is greater than 200,000, the quality control is qualified.

[0034] In some embodiments of the present application, the method for analyzing fusion gene types based on RNA sequencing further comprises:

[0035] According to the primer combination coverage region of the first gene and the second gene of the fusion gene, a primer coverage region sequence file is obtained, including at least the gene name, exon number and exon sequence of the first gene and the second gene; saved in a fa file, the saved format is gene name_exon number, for example, G1_E4 represents the 4th exon of G1 gene.

[0036] The RNA sequencing data is secondly aligned with the primer coverage region sequence file, and the following processing is performed:

[0037] (1) If a sequencing data can be aligned to multiple exon sequences in the primer coverage region sequence file, all alignment results are retained;

[0038] (2) If a sequencing data is aligned to an exon sequence of the first gene and an exon sequence of the second gene at the same time, the sequencing data is a fusion gene sequence:

[0039] Further, if the sequencing data is aligned to multiple exon sequences of the first gene, the exon with the largest number is recorded, for example, aligned to exon 4 (exon4, E4) and exon 5 (exon5, E5) of the first gene at the same time, recorded as G1_E5; similarly, if the sequencing data is aligned to multiple exon sequences of the second gene, the exon with the smallest number is recorded, for example, aligned to exon 18 (exon18, E18) and exon 19 (exon19, E19) of the second gene at the same time, recorded as G2_E18.

[0040] The number of sequencing data that can be aligned to the exon sequences of two genes at the same time is counted, and if the number of sequencing data that can be aligned to the exon sequences of two genes at the same time is greater than a preset threshold, the sample to be tested has occurred a corresponding type of gene fusion. For example, if the sequencing data that can be aligned to G1_E5 and G2_E18 at the same time is greater than a preset threshold, the sample to be tested has occurred G1_E5-G2_E18 gene fusion, i.e. the 5th exon of G1 gene and the 18th exon of G2 gene have occurred gene fusion.

[0041] In some embodiments of the present application, if the type of the obtained fusion gene is known, the name of the corresponding known type of fusion gene is outputted. As previously described, the name of the known type of fusion gene at least includes the first gene name and the exon number of the first gene where fusion occurs; and the second gene name and the exon number of the second gene where fusion occurs.

[0042] In some embodiments of the present application, the primer combination includes the nucleotide sequences shown in SEQ ID No. 9, 10, 12, 13, 15, 16, 18, 19, 21, 22, and preferably, further includes the nucleotide sequences shown in SEQ ID No. 11, 14, 17, 20, 23, for blocking the amplification of wild type genes.

[0043] According to the design of the primer combination amplification region, the primer coverage sequence file is designed, specifically, first, all types of fusion genes that can be amplified by the primer are obtained, and 50-70 bases are extended upstream and downstream of the breakpoint of a certain type of fusion gene as the primer coverage region sequence of the type of fusion gene. The second aspect of the present application provides a system for analyzing fusion gene types based on RNA sequencing, which includes the following modules:

[0044] A data input module for inputting RNA sequencing data of a sample to be tested;

[0045] A first storage module for storing fusion gene reference sequence files obtained according to known types of fusion genes, the fusion gene reference sequence file at least including the name of the known type of fusion gene and the fusion reference sequence, wherein the fusion reference sequence is obtained by extending 60-75 bases forward and backward from the breakpoint;

[0046] A first alignment module connected with the data input module and the first storage module, respectively, for first comparing the RNA sequencing data with the fusion reference sequence file and performing the following processing:

[0047] (1) If the base mismatching ratio of a sequencing data to the fusion reference sequence is less than 5%, the sequencing data is retained;

[0048] (2) If a sequencing data crosses the breakpoint on the fusion reference sequence and the number of bases covered on both sides of the breakpoint is >20, the sequencing data can be aligned to the fusion reference sequence;

[0049] (3) If a sequencing data can be aligned to multiple fusion reference sequences in the fusion reference sequence file, only the preferentially aligned result is retained,

[0050] The fusion gene known type analysis module is connected with the first alignment module, and is used for counting the number of sequencing data that can be aligned to each fusion reference sequence, and if the number of sequencing data that can be aligned to a certain fusion reference sequence is greater than a preset threshold, the sample to be tested has a corresponding type of gene fusion.

[0051] In some embodiments of the present application, the system further comprises:

[0052] The second storage module is used for storing primer coverage region sequence files of the first gene and the second gene of the fusion gene according to primer combination coverage regions, and the primer coverage region sequence files at least include gene names, exon sequence numbers and exon sequences.

[0053] The second alignment module is connected with the data input module and the second storage module respectively, and is used for performing second alignment of the RNA sequencing data and the primer coverage region sequence files and performing the following processing:

[0054] (1) If one piece of sequencing data can be aligned to multiple exon sequences in the primer coverage region sequence files, all alignment results are retained;

[0055] (2) If one piece of sequencing data is aligned to an exon sequence of the first gene and an exon sequence of the second gene at the same time, the sequencing data is a fusion gene sequence:

[0056] If the sequencing data is aligned to multiple exon sequences of the first gene, the exon with the largest sequence number is recorded; similarly, if the sequencing data is aligned to multiple exon sequences of the second gene, the exon with the smallest sequence number is recorded,

[0057] The fusion gene all type analysis module is connected with the second alignment module, and is used for counting the number of sequencing data that can be aligned to two gene exon sequences at the same time, and if the number of sequencing data that can be aligned to two gene exon sequences at the same time is greater than a preset threshold, the sample to be tested has a corresponding type of gene fusion.

[0058] In some embodiments of the present application, the system further comprises a fusion gene type integration analysis module connected with the fusion gene known type analysis module and the fusion gene all type analysis model respectively, and if the type of the obtained fusion gene is known, the name of the corresponding known type fusion gene is output.

[0059] The third aspect of the present application provides a computer device, comprising:

[0060] The memory is used for storing a computer program.

[0061] a processor for implementing the steps of the method for analyzing fusion gene types based on RNA sequencing according to any one of the first aspect of the present application when the computer program is executed.

[0062] The fourth aspect of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the method for analyzing fusion gene types based on RNA sequencing according to any one of the first aspect of the present application.

[0063] The fifth aspect of the present application provides a fusion gene, which is EML4_E6-ALK_E18, i.e., a gene in which the 6th exon of EML4 and the 18th exon of ALK gene are fused.

[0064] In some embodiments of the present application, the nucleotide sequence of the fusion gene EML4_E6-ALK_E18 is shown in SEQ ID No. 24.

[0065] In the present application, primers can be designed based on the sequences of 60-75 bases upstream and downstream of the breakpoint of the fusion gene for amplification of the fusion gene, thereby completing the detection of the fusion gene. In some embodiments of the present application, the upstream and downstream primers are designed based on the nucleotide sequence shown in SEQ ID No. 25. Of course, those skilled in the art understand that the upstream and downstream primers are respectively located before and after the breakpoint.

[0066] Advantages of the present application

[0067] Compared with the prior art, the present application has the following advantages:

[0068] By using the method and system of the present application, known types of fusion genes can be accurately detected, and unknown types of fusion genes can also be detected, which is more conducive to the development of downstream research on the related mechanisms and functions of fusion genes. BRIEF DESCRIPTION OF DRAWINGS

[0069] Figure 1 A flow chart of the method for analyzing fusion gene types based on RNA sequencing according to Example 1 of the present application is shown.

[0070] Figure 2 A judging method for judging whether a read sequence is from a fusion gene according to Example 1 of the present application is shown.

[0071] Figure 3 An alignment result of RNA sequencing data and a primer coverage region file according to Example 1 of the present application is shown.

[0072] Figure 4A system diagram for analyzing known types of fusion genes based on RNA sequencing analysis in the embodiment 2 of the present application is shown.

[0073] Figure 5 A system diagram for analyzing all types of fusion genes based on RNA sequencing analysis in the embodiment 3 of the present application is shown. DETAILED DESCRIPTION

[0074] In order to make the technical problems solved by the present application, technical solutions and beneficial effects clearer, the present application will be further described in detail below in combination with embodiments.

[0075] The following examples are presented to demonstrate preferred embodiments of the present application. Those skilled in the art will appreciate that the techniques disclosed in the following examples represent techniques that can be used to practice the present application, and thus can be considered to be preferred methods of practicing the present application. However, it is understood that those skilled in the art will be able to devise many alternative methods that will be appreciated in light of the present disclosure, and yet still fall within the spirit and scope of the present application.

[0076] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs, and the materials described herein will be referred to by the citation of the reference that first introduced such term or material into the art. In case of conflict, the present specification, including that of any immediately preceding claims, will control.

[0077] Those skilled in the art will appreciate that numerous equivalents to the specific procedures described herein will be apparent in light of the present disclosure, and thus will recognize the applicability of the certain embodiments of the application and equivalent constructions to other equivalent procedures. Such equivalents are intended to be encompassed by the claims.

[0078] The experimental methods in the following examples are all conventional methods unless otherwise specified. The instruments and equipment used in the following examples are all conventional laboratory instruments and equipment unless otherwise specified; the test materials used in the following examples are all purchased from conventional biochemical reagent stores unless otherwise specified.

[0079] Example 1 A method for analyzing EML4-ALK fusion gene types based on RNA high-throughput sequencing data

[0080] The analysis process of the present embodiment is shown in Figure 1

[0081] 1. EML4-ALK fusion gene reference sequence file

[0082] ​The COSMIC database (https: / / cancer.sanger.ac.uk / cosmic) was accessed and screened for EML4-ALK fusion gene information for preparing a fusion gene reference sequence file (fa file) including a fusion gene type name (GeneName), a fusion gene chromosome position (chr), fusion gene strand information (strand), a fusion gene ID from the COSMIS database (transID), a fusion start position (varStart), a fusion end position (varEnd), a breakpoint (breakPoint), a length (len), and base information of a fusion reference sequence (fusionSeq) obtained by extending 60-75 bases forward and backward from the breakpoint, with a length of 120-150 bp. Among them, the fusion gene name includes which exons of EML4 and ALK constitute the reference sequence, such as EML4-ALK.E13:A20.COSF408, which represents a fusion of the 13th exon of EML4 and the 20th exon of ALK. The known EML4-ALK fusion gene reference sequence file is shown in Table 1.

[0083] Table 1 Known EML4-ALK fusion gene reference sequence file

[0084]

[0085] The nucleotide sequence of each fusion reference sequence (fusionSeq) is as follows:

[0086] EML4-ALK.E13:A20.COSF408 (SEQ ID No. 1)

[0087] GGAGATGTTCTTACTGGAGACTCAGGTGGAGTCATGCTTATATGGAGCAAAACTACTGTAGAGCCCACACCTGGGAAAGGACCTAAAGTGTACCGCCGGAAGCACCAGGAGCTGCAAGCCATGCAGATGGAGCTGCAGAGCCCT

[0088] EML4-ALK.E6a:A20.COSF411 (SEQ ID No. 2)

[0089] CCCAAATTAATACCAAAAGTTACCAAAACTGCAGACAAGCATAAAGATGTCATCATCAACCAAGTGTACCGCCGGAAGCACCAGGAGCTGCAAGCCATGCAGATGGAGCTGCAGAGCCCT

[0090] EML4-ALK.E6: A20.COSF1544 (SEQ ID No. 3)

[0091] CCCAAATTAATACCAAAAGTTACCAAAACTGCAGACAAGCATAAAGATGTCATCATCAACCAAGCTGACCACCCACCTGCAGTGTACCGCCGGAAGCACCAGGAGCTGCAAGCCATGCAGATGGAGCTGCAGAGCCCTGAGTACAAGC

[0092] EML4-ALK.E6: A19.COSF1296 (SEQ ID No. 4)

[0093] CCCAAATTAATACCAAAAGTTACCAAAACTGCAGACAAGCATAAAGATGTCATCATCAACCAAGTGTCACCCACCCCGGAGCCACACCTGCCACTCTCGCTGATCCTCTCTGTGGTGACCTCTGCCCTCGTGGCCGCCCTGGTCCTGG

[0094] EML4-ALK.E20: A20.COSF409 (SEQ ID No. 5)

[0095] ACATTCCAGCTACATCACACACCTTGACTGGTCCCCAGACAACAAGTATATAATGTCTAACTCGGGAGACTATGAAATATTGTACTTGTACCGCCGGAAGCACCAGGAGCTGCAAGCCATGCAGATGGAGCTGCAGAGCCCT

[0096] EML4-ALK.E18: A20.COSF487 (SEQ ID No. 6)

[0097] GGCAGGTGGTTTGTTCTGGATGCAGAAACCAGAGATCTAGTTTCTATCCACACAGACGGGAATGAACAGCTCTCTGTGATGCGCTACTCAATAGTGTACCGCCGGAAGCACCAGGAGCTGCAAGCCATGCAGATGGAGCTGCAGAGCCCT

[0098] EML4-ALK.E14: A20.COSF491 (SEQ ID No. 7)

[0099] AATAATTCTGTGGGATCATGATCTGAATCCTGAAAGAGAAATAGAGATATGCTGGATGAGCCCTGAGTACAAGCTGAGCAAGCTCCGCACCTCGACCATCATGACCGACTACAACCCCAACTACTGCTTTGCTGGCAAGACCT

[0100] EML4-ALK.E2: A20.COSF478 (SEQ ID No. 8)

[0101] CGTCTTGCAATCTCTGAAGATCATGTGGCCTCAGTGAAAAAATCAGTCTCAAGTAAAGTGTACCGCCGGAAGCACCAGGAGCTGCAAGCCATGCAGATGGAGCTGCAGAGCCCT

[0102] 2. Primer design

[0103] The primer design of the present application not only meets the targeting of known types of fusion genes, but also seeks to find unknown types of fusion genes. The EML4_E13_ALK_E20 (i.e. EML4 gene 13 exon and ALK gene 20 exon fusion), EML4_E6_ALK_E20, EML4_E20_ALK_E20, EML4_E18_ALK_E20, EML4_E2_ALK_E20 fusion gene type design primer (can cover the above 8 types of fusion genes), specifically, extend 50-70 bases upstream and downstream of the breakpoint as primer coverage region, design upstream and downstream primers for the two genes in the primer coverage region. With the primer combination designed in this embodiment, the least number of primers can be used to obtain as many fusion gene types as possible, for example, the above-mentioned EML4-ALK.E6:A19.COSF1296 can be amplified by using the primer designed in the primer coverage region of EML4_E6_ALK_E20; for example, EML4-ALK.E14:A20.COSF491 can be amplified by using the primer designed in the primer coverage region of EML4_E13_ALK_E20, and conversely, EML4_E6_ALK_E20 can obtain EML4_E6_ALK_E20, EML4_E7_ALK_E20, EML4_E8_ALK_E20, EML4_E9_ALK_E20, EML4_E6_ALK_E19, EML4_E6_ALK_E18, EML4_E6_ALK_E17, etc. The final primer combination is shown in Table 2:

[0104] Table 2 Targeted amplification primer combination

[0105]

[0106] In the table, the primer with the name ending with P1 is the Blocker primer for wild type amplification of the blocking gene.

[0107] According to the above-mentioned designed primer, the primer coverage region sequence file is prepared, including the EML4 gene and ALK gene exon sequence covered by each primer.

[0108] 3. RNA sequencing

[0109] The sample to be tested is obtained from the paraffin-embedded sample of the surgical resection tissue of a non-small cell lung cancer patient, and the RNA sample is extracted. Further, the above-mentioned designed primer is used for multiplex PCR amplification to obtain a sequencing library. The NexSeq CN500 sequencer is used for high-throughput sequencing data.

[0110] Data quality control is performed on the high-throughput sequencing data: remove the adapter sequence, remove the low-quality sequence, and the remaining data is greater than 200,000, then the quality control is qualified.

[0111] 4. Analysis of known types of EML4-ALK fusion gene based on RNA sequencing data

[0112] This step is based on RNA high-throughput sequencing data to analyze the known types of EML4-ALK fusion gene, specifically:

[0113] The quality-controlled high-throughput sequencing data is first aligned with the reference sequence file using the alignment software BWA, and the following processing is performed:

[0114] (1) If the base mismatch rate of a sequencing data (reads) to the fusion reference sequence is less than 5%, the sequencing data is retained;

[0115] (2) If a sequencing data (reads) spans the breakpoint on the fusion reference sequence, and the number of bases covered on both sides of the breakpoint is > 20, it is considered that the sequencing data can be aligned to the fusion reference sequence, and the alignment principle of the known type of fusion gene is as shown in Figure 2

[0116] (3) If the sequencing data (reads) can be aligned to multiple fusion reference sequences in the fusion reference sequence file, only the preferred alignment (highest alignment score) result is retained, and the sub-optimal alignment result is not counted,

[0117] The number of sequencing data (reads) aligned to a certain fusion reference sequence is counted, and if the normalized value is greater than the reference threshold (16), it is considered that the sample has occurred this type of gene fusion, and the statistical results are shown in Table 3.

[0118] Table 3 Number of reads aligned to known types of EML4-ALK fusion gene

[0119]

[0120] As shown in Table 3, the number of reads aligned to EML4-ALK.E6a:A20.COSF734 is 3768, which exceeds the reference threshold, and the sample has occurred EML4-ALK.E6a:A20.COSF7340 type of gene fusion.

[0121] 5. Analysis of all types of EML4-ALK fusion gene based on RNA sequencing data

[0122] Using the above steps, only the known types of EML4-ALK fusion gene can be analyzed, and the unknown types cannot be analyzed. Therefore, the inventors further implemented the following steps to complete the analysis of all types of fusion gene. ​

[0123] The qualified sequencing data is secondly aligned with the primer coverage region sequence file using the alignment software BWA, and the alignment result is obtained, as shown in Figure 3 The specific steps are as follows:

[0124] (1) If the sequencing data (reads) can be aligned to the multiple exon sequences of the primer coverage region sequence file, all the alignment results are retained;

[0125] (2) The alignment results are simplified: if a sequencing data (reads) is aligned to the exon sequences of EML4 and ALK, the sequencing data is a fusion gene sequence. If the sequencing data is aligned to multiple exon sequences of multiple EML4s, the exon with the largest serial number is recorded; if the sequencing data is aligned to multiple exon sequences of multiple ALKs, the exon with the smallest serial number is recorded.

[0126] (3) The number of fusion gene sequences of the primer coverage region sequence file alignment result is counted, and is marked with the exon combination. If the normalized value is greater than the reference threshold value, it is considered that the sample has this type of gene fusion. The statistical results are shown in Table 4.

[0127] Table 4 Number of reads aligned to all types of EML4-ALK fusion genes

[0128]

[0129] As shown in Table 4, the number of reads that can be aligned to EML4_E6_ALK_E20 reaches 3740, and the number of reads that can be aligned to EML4_E6_ALK_E18 reaches 546, both of which exceed the reference threshold value, and the sample has EML4_E6_ALK_E20 and ML4_E6_ALK_E18 type gene fusion.

[0130] 6. Analysis result merging

[0131] The results of the known type analysis of the EML4-ALK fusion gene and the results of the analysis of all types of EML4-ALK fusion genes are integrated, and the analysis results are output.

[0132] If the predicted results contain known types of gene fusion, the known type of gene fusion ID is output; if the known type of gene fusion is not contained, the newly predicted type of gene fusion is output. The final analysis results are shown in Table 5.

[0133] Table 5 Number of reads aligned to all types of EML4-ALK fusion genes

[0134]

[0135] In this example, a new fusion gene, EML4_E6-ALK_E18, is obtained, and its full-length sequence is as follows:

[0136]

[0137] The sequence near the breakpoint is as follows:

[0138] CCCAAATTAATACCAAAAGTTACCAAAACTGCAGACAAGCATAAAGATGTCATCATCAACCAAGTGATGGAAGGCCACGGGGAAGTGAATATTAAGCATTATCTAAACTGCAGTCACTGTGAGGTAGACGAATGTCACATGGACCCTG (SEQ ID No. 25)

[0139] The sequence near the breakpoint can be further used to design primers to detect the above-mentioned fusion gene.

[0140] Example 2: System for analyzing known types of fusion genes based on RNA sequencing

[0141] Based on the method of Example 1, this embodiment provides a system for analyzing fusion gene types based on RNA sequencing, such as Figure 4 As shown, it includes the following modules:

[0142] Data input module, used to input RNA sequencing data of samples to be tested;

[0143] A first storage module is used to store a fusion gene reference sequence file obtained based on a known type of fusion gene, the fusion gene reference sequence file including the name of the known type of fusion gene and a fusion reference sequence, wherein the fusion gene is formed by fusing the first part of the first gene and the second part of the second gene, and the fusion reference sequence is obtained by extending 60 to 75 bases forward and backward from the breakpoint;

[0144] The first comparison module is connected to the data input module and the first storage module respectively, and is used to perform a first comparison between the RNA sequencing data and the fusion reference sequence file and perform the following processing:

[0145] (1) If the base mismatch ratio between a sequencing data and the fusion reference sequence is less than 5%, the sequencing data will be retained;

[0146] (2) If a sequencing data piece spans the breakpoint on the fusion reference sequence and the number of bases covered on both sides of the breakpoint is greater than 20, then the sequencing data piece can be aligned to the fusion reference sequence;

[0147] (3) If a sequencing data can be aligned to multiple fusion reference sequences in the fusion reference sequence file, only the result of the priority alignment is retained.

[0148] The fusion gene known type analysis module is connected with the first alignment module, and is used for counting the number of sequencing data that can be aligned to each fusion reference sequence. If the number of sequencing data that can be aligned to a certain fusion reference sequence is greater than a preset threshold, the corresponding type of gene fusion occurs in the sample to be tested.

[0149] Embodiment 3: System for analyzing all types of fusion genes based on RNA sequencing

[0150] Also based on the method of embodiment 1, this embodiment is improved on the basis of embodiment 2, as shown in Figure 5 Specifically, the system of embodiment 2 is further included:

[0151] The second storage module is used for storing primer coverage region sequence files of the first gene and the second gene of the fusion gene corresponding to the primer combination coverage region, and the primer coverage region sequence files include gene names, exon sequence numbers and exon sequences.

[0152] The second alignment module is connected with the data input module and the second storage module respectively, and is used for performing second alignment between the RNA sequencing data and the primer coverage region sequence files and performing the following processing:

[0153] (1) If one piece of sequencing data can be aligned to multiple exon sequences in the primer coverage region sequence files, all alignment results are retained;

[0154] (2) If one piece of sequencing data is aligned to the exon sequence of the first gene and the exon sequence of the second gene at the same time, the piece of sequencing data is a fusion gene sequence:

[0155] If the piece of sequencing data is aligned to multiple exon sequences of the first gene, the exon with the largest sequence number is recorded. Similarly, if the piece of sequencing data is aligned to multiple exon sequences of the second gene, the exon with the smallest sequence number is recorded,

[0156] The fusion gene all type analysis module is connected with the second alignment module, and is used for counting the number of sequencing data that can be aligned to the exon sequences of two genes at the same time. If the number of sequencing data that can be aligned to the exon sequences of two genes at the same time is greater than a preset threshold, the corresponding type of gene fusion occurs in the sample to be tested.

[0157] The fusion gene type integration analysis module is connected with the fusion gene known type analysis module and the fusion gene all type analysis model respectively. If the type of the obtained fusion gene is known, the name of the corresponding known type fusion gene is output.

[0158] All documents referred to in the present application are incorporated herein by reference as if each were individually incorporated. In addition, it is to be understood that the application can be carried out by specifically different embodiments and that each disclosed embodiment can be implemented with or without the corresponding use of the other embodiments. Other embodiments will occur to those skilled in the art upon consideration of this disclosure or can be learned from practice of the application. The present embodiments are therefore to be considered in all respects as illustrative and not restrictive, and all changes coming within the meaning and equivalency range of the appended claims are intended to be embraced therein.

Claims

1. A method for analyzing EML4-ALK fusion gene types, characterized in that: The following steps are involved: The RNA sequencing data of the sample to be tested is obtained by the following steps: Obtain RNA samples of the samples to be tested, The RNA samples were subjected to multiplex PCR amplification using a primer combination designed using the following method: EML4_E13_ALK_E20, EML4_E6_ALK_E20, EML4_E20_ALK_E20, EML4_E18_ALK_E20, and EML4_E2_ALK_E20 fusion gene types were selected, 50 to 70 bases were extended upstream and downstream of the break site of each fusion gene as a primer coverage region, and upstream and downstream primers were respectively designed on the two genes in the primer coverage region to obtain the primer combination. The nucleotide sequences of the primer combination are shown in SEQ ID No. 9 to SEQ ID No.

23. Performing high-throughput sequencing on the amplified product to obtain the RNA sequencing data; According to the exon sequences of the first gene and the second gene of the fusion gene corresponding to the primer combination coverage region, a primer coverage region sequence file is obtained, including the gene name, exon number and exon sequence; Perform a second alignment of the RNA sequencing data with the primer coverage region sequence file and perform the following processing: (1) If a sequencing data can be aligned to multiple exon sequences in the primer coverage region sequence file, all alignment results are retained; (2) If a piece of sequencing data is aligned to the exon sequence of the first gene and the exon sequence of the second gene at the same time, then the sequencing data is a fusion gene sequence: If the sequencing data is aligned to multiple exon sequences of the first gene, then the exon with the largest sequence number is recorded; similarly, if the sequencing data is aligned to multiple exon sequences of the second gene, then the exon with the smallest sequence number is recorded. The number of sequencing data that can be simultaneously aligned to two gene exon sequences is counted. If the number of sequencing data that can be simultaneously aligned to two gene exon sequences is greater than a preset threshold, the corresponding type of gene fusion occurs in the sample to be tested.

2. The method for analyzing the EML4-ALK fusion gene type according to claim 1, characterized in that: The steps for quality control of RNA sequencing data are further included: (1) Removal of linker sequences; (2) Remove low-quality sequences, If the remaining data is greater than 200,000, the quality control is qualified.

3. A system for analyzing EML4-ALK fusion gene types, characterized in that: Includes the following modules: The data input module is used to input the RNA sequencing data of the sample to be tested, and the RNA sequencing data of the sample to be tested is obtained by the following steps: Obtain RNA samples of the samples to be tested, The RNA samples were subjected to multiplex PCR amplification using a primer combination designed using the following method: EML4_E13_ALK_E20, EML4_E6_ALK_E20, EML4_E20_ALK_E20, EML4_E18_ALK_E20, and EML4_E2_ALK_E20 fusion gene types were selected, 50 to 70 bases were extended upstream and downstream of the break site of each fusion gene as a primer coverage region, and upstream and downstream primers were respectively designed on the two genes in the primer coverage region to obtain the primer combination. The nucleotide sequences of the primer combination are shown in SEQ ID No. 9 to SEQ ID No.

23. Performing high-throughput sequencing on the amplified product to obtain the RNA sequencing data; A second storage module is used to store a primer coverage region sequence file obtained according to the exon sequences of the first gene and the second gene of the fusion gene corresponding to the primer combination coverage region, wherein the primer coverage region sequence file includes the gene name, exon number and exon sequence; A second comparison module is connected to the data input module and the second storage module, and is used to perform a second comparison between the RNA sequencing data and the primer coverage region sequence file and perform the following processing: (1) If a sequencing data can be aligned to multiple exon sequences in the primer coverage region sequence file, all alignment results are retained; (2) If a piece of sequencing data is aligned to the exon sequence of the first gene and the exon sequence of the second gene at the same time, then the sequencing data is a fusion gene sequence: If the sequencing data is aligned to multiple exon sequences of the first gene, then the exon with the largest sequence number is recorded; similarly, if the sequencing data is aligned to multiple exon sequences of the second gene, then the exon with the smallest sequence number is recorded. The fusion gene all types analysis module is connected to the second comparison module and is used to count the number of sequencing data that can be simultaneously compared to two gene exon sequences. If the number of sequencing data that can be simultaneously compared to two gene exon sequences is greater than a preset threshold, the corresponding type of gene fusion has occurred in the sample to be tested.

4. A computer device, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of a method for analyzing EML4-ALK fusion gene types according to any one of claims 1 to 2 when executing the computer program.

5. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for analyzing the EML4-ALK fusion gene type according to any one of claims 1 to 2 are implemented.

Citation Information

Patent Citations

  • Method and kit for screening fusion gene of lung cancer by virtue of double-color FISH (Fluorescence In Situ Hybridization)

    CN104212903A

  • Detection method of human ROS1 fusion gene

    CN107541550A