Variant filtering method and device for tumor whole transcriptome sequencing

By constructing whitelists and blacklists, combining sequencing coverage, variation frequency and population frequency for mutation filtering of tumor whole transcriptome sequencing, the problems of long variation screening time and high error rate in the prior art are solved, and efficient and accurate variant recognition and filtering are achieved.

CN118824362BActive Publication Date: 2025-06-10SHANGHAI CINOPATH MEDICAL TESTING CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410824902.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2025-06-10
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

The prior art has problems of high time consumption and high error rate in the variation filtration of tumor whole transcriptome sequencing, resulting in inefficiency in clinical diagnosis and treatment decisions.

Method used

A comprehensive variation filtering method, including Snv/Indel variant filtering and fusion gene filtering, was used to filter by constructing whitelists and blacklists, combining sequencing coverage, variation frequency and population frequency, and classification and labeling of the results to improve accuracy.

Benefits of technology

It significantly shortens the time for variant screening, reduces the error rate, improves the efficiency of clinical report interpretation, and ensures accurate identification and filtering of pathogenic variant sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118824362B_ABST
    Figure CN118824362B_ABST
Patent Text Reader

Abstract

The present invention discloses a variant filtering method and device for whole transcriptome sequencing of tumors. By constructing a local database of Snv / Indel variants and fusion genes from RNA-seq sequencing of historical tumor samples, constructing local detection frequencies of Snv / Indel variants and fusion genes, a whitelist of tumor-related genes, and blacklists of Snv / Indel variants and fusion genes; performing bioinformatics analysis and filtering on the detected Snv / Indel variants and fusion genes in the original data, and one-key classification and tagging on the reporting system; the method makes full use of the data information constructed locally and the quality of the sample data itself, significantly reducing the number of invalid sites that need to be analyzed during the interpretation of clinical reports, shortening the time for site screening in tumor report interpretation, while filtering out benign and high-frequency variants and avoiding false-negative pathogenic variant sites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sequencing bioinformatics processing, and specifically to a variant filtering method and device for tumor whole transcriptome sequencing. Background Art

[0002] RNA-seq sequencing converts all RNA species into complementary DNA fragments (complementary DNA libraries), that is, constructs a library supporting sequencing with cDNA as a template, and performs whole transcriptome level detection on it through a sequencing platform. The RNA-seq sequencing technology does not require pre-designed specific probes, can directly determine the sequence of each transcript fragment, can detect single-base differences, similar genes in a gene family, and different transcripts caused by alternative splicing, and can detect rare transcripts and new transcripts with as few as several copies in a cell. Compared with the whole genome and whole exome, the transcriptome has rich gene expression and sequence information, and has unique advantages in analyzing gene fusions, splicing variations, and gene expression profiles in tumors.

[0003] The raw data of RNA-seq undergoes bioinformatics analysis of sequence preprocessing, sequence alignment, variant identification, information annotation, and variant filtering, and genetic analysts comprehensively consider information such as population frequency, variant quality, and the degree of influence of variants on proteins, and find the number of fusion genes and Snv / Indel variants with clinical significance from millions of variants for auxiliary clinical diagnosis, treatment guidance, and prognosis evaluation of tumor patients. The above is the currently common RNA-seq data analysis scheme.

[0004] The existing filtering schemes may have the following problems: (1) Not setting a whitelist will cause Snv / Indel variants and fusion genes with clinical significance to be missed; (2) Second-generation sequencing technology has false positive sites in specific genomic regions caused by the sequencing process, which are difficult to filter; (3) Variant screening is a complex and time-consuming process with a low error tolerance. Excessive variant filtering may result in the removal of positive variants related to disease diagnosis, treatment, or prognosis, and insufficient variant filtering will leave a large number of variants for manual operation, which is time-consuming and error-prone. Summary of the Invention

[0005] The present invention provides a variant filtering method and device for tumor whole transcriptome sequencing, which is used to solve at least one of the defects of long time consumption and high error rate in existing technology variant screening and filtering.

[0006] In view of this, the solution of the present invention is as follows:

[0007] A variant filtering method for tumor whole transcriptome sequencing, including Snv / Indel variant filtering and fusion gene filtering; wherein:

[0008] The Snv / Indel variant filtering step includes: filtering low-quality and high-frequency variants, non-coding region variants, and filtering exon and splicing region variants by combining population frequency and the Snv / InDel variant whitelist in sequence for the Snv / InDel variants generated from the original data; then classifying and tagging the filtered results by combining the tumor-related gene whitelist and the Snv / InDel variant black and white lists, and the tag types include: non-hotspot, low-frequency repetitive region, non-classical transcript variant, and blacklist Snv / InDel;

[0009] The fusion gene filtering step includes: filtering the fusion genes detected from the original data by combining the fusion gene blacklist and the fusion gene whitelist in sequence, classifying and tagging the filtered results, and the tag types include: non-coding gene, gene interval, repetitive fusion gene pair, and fusion gene blacklist.

[0010] Further, the construction process of the Snv / InDel variant whitelist includes:

[0011] Obtaining a temporary Snv / InDel variant whitelist one based on the variants rated as Tier I and Tier II after manual review of historical tumor whole-transcriptome data;

[0012] Extracting the pathogenic and suspected pathogenic variants of tumor-related genes to generate a temporary Snv / InDel variant whitelist two;

[0013] Merging the temporary Snv / InDel variant whitelist one and the temporary Snv / InDel variant whitelist two to obtain the Snv / InDel variant whitelist;

[0014] The construction process of the Snv / InDel variant blacklist includes:

[0015] Obtaining a temporary Snv / InDel variant blacklist based on the variants not rated as Tier I, Tier II, or Tier III after manual review of historical tumor whole-transcriptome data;

[0016] Selecting the variants in the temporary Snv / InDel variant blacklist that are not in the Snv / InDel variant whitelist to generate the Snv / InDel variant blacklist.

[0017] Furthermore, the filtering of low-quality and high-frequency variants includes filtering based on the sequencing coverage of Snv / Indel variants, the product of sequencing coverage and variant frequency, and the threshold setting of tumor local detection frequency. Preferably, the tumor local detection frequency threshold is based on historical tumor full transcriptome positive samples, multiple sample subsets are selected according to the gradient, different historical local detection frequency thresholds are set for each gradient, the number of variants retained in each subset and each gradient corresponding to a positive detection rate of 100% is sorted from small to large, and the threshold near the first quartile is selected as the historical local detection frequency threshold.

[0018] Furthermore, the filtering of non-coding region variations includes: filtering and removing variations belonging to intergenic regions, ncRNA, and upstream and downstream regions of genes, and filtering variations in introns and UTR regions with a population frequency of more than 0.1 or that do not affect splicing; and / or,

[0019] The filtering of exon and splicing region variants includes: filtering variants belonging to exons and splicing regions with a population frequency of more than 0.1, filtering variants that are not in the Snv / Indel variant whitelist or database with a population frequency of more than 20 and more than 0.01, filtering variants that are synonymous mutations and do not affect splicing, and filtering variants with a local detection frequency of more than 5% and less than 3 included in COSMIC for splicing variants.

[0020] Furthermore, in the Snv / Indel variation filtering step, the non-hotspot types include: not in the candidate gene list, high local detection rate, and population frequency higher than 0.001.

[0021] Preferably, the non-hotspot types are labeled based on the following criteria:

[0022] For genes with Snv / Indel mutations that are not in the tumor-related gene whitelist, if they meet the requirements of not being pathogenic or suspected pathogenic in the Clinvar database, or meet the requirements of ≤10 or null mutations in the COSMIC database, add the label "not in the candidate gene list, non-hotspot";

[0023] For Snv / Indel variants with a local detection frequency of ≥5%, which meet the requirements of not being pathogenic or suspected pathogenic in the Clinvar database, or meet the requirements of ≤10 or null values ​​in the COSMIC database, add the label "Local detection frequency is high, non-hotspot";

[0024] For Snv / Indel variants with a maximum population frequency ≥ 0.001 and a COSMIC database inclusion ≤ 10 or a null value, a “population frequency higher than 0.001, non-hotspot” label was added.

[0025] Further, in the Snv / Indel variant filtering step, the low-frequency repetitive region type is marked based on the condition that the short tandem repeat count of the Snv / Indel variant is ≥5 and the variant frequency of the Snv / Indel variant in the tumor sample is ≤10%.

[0026] Further, the construction process of the fusion gene whitelist includes:

[0027] Based on historical tumor whole-transcriptome data, generating a temporary fusion gene whitelist one for fusion genes that have been manually reviewed and rated as level 1 or 2;

[0028] Based on fusion genes recorded in the literature and databases, generating a temporary fusion gene whitelist two;

[0029] Taking the union of the temporary fusion gene whitelist one and the temporary fusion gene whitelist two as the fusion gene whitelist;

[0030] The construction process of the fusion gene blacklist includes:

[0031] Based on the monitoring results of several patients in historical tumor whole-transcriptome data and fusion genes verified as false positives in the control cell line, generating a temporary fusion gene blacklist one;

[0032] Using healthy human controls to detect fusion genes and generating a temporary fusion gene blacklist two;

[0033] Taking the fusion genes in the temporary fusion gene blacklist two that are not in the fusion gene whitelist to generate a temporary fusion gene blacklist three;

[0034] Taking the union of the temporary fusion gene blacklist one and the temporary fusion gene blacklist three to generate a fusion gene blacklist one;

[0035] Based on historical tumor whole-transcriptome data, generating a temporary fusion gene blacklist four for fusion genes that have not been manually reviewed and rated as level 1, 2, or 3;

[0036] Taking the fusion genes in the fusion gene blacklist four that are not in the fusion gene whitelist to generate a fusion gene blacklist two;

[0037] In the fusion gene filtering step, for the fusion genes whose breakpoint pairs detected are in the fusion gene blacklist two during the classification and tagging process, add the tag "fusion gene blacklist two".

[0038] Further, in the fusion gene filtering step, after filtering the fusion genes detected in the original data, the reference sequence is aligned. According to whether one of split_reads1 + split_reads2 and discordant_mates is 0, the fusion genes with split_reads1 + split_reads2 + discordant_mates greater than different thresholds are retained.

[0039] Further, it also includes the step of rechecking the Snv / Indel variations and fusion gene tags. The rechecking process includes confirming the tags of the data directly filtered out and the pending tags.

[0040] Another object of the present invention is to provide a variant filtering device for tumor whole transcriptome sequencing, including:

[0041] Snv / Indel variant filtering module: used to sequentially filter the low-quality and high-frequency variations, non-coding region variations of the Snv / InDel variations generated from the original data, and filter the exonic and splicing region variations by combining the population frequency and the Snv / Indel variant whitelist; then classify and label the filtered results by combining the tumor-related gene whitelist, Snv / Indel variant black and white lists. The label types include: non-hotspot, low-frequency repetitive region, non-classical transcript variant and blacklist Snv / Indel;

[0042] The fusion gene filtering module: used to sequentially filter the fusion genes detected in the original data by combining the fusion gene blacklist and the fusion gene whitelist, classify and label the filtered results. The label types include: non-coding gene, gene interval, repetitive fusion gene pair.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] The variant filtering method provided by the present invention constructs the local detection frequencies of Snv / Indel variations, fusion genes, the tumor-related gene whitelist, Snv / Indel variations, and fusion gene black / white lists; performs bioinformatics analysis filtering and classification labeling on the Snv / Indel variations and fusion genes detected in the original data; the method makes full use of the locally constructed data information and the quality of the sample data itself, significantly reducing the number of invalid sites that need to be analyzed during the clinical report interpretation process, shortening the site screening time in tumor report interpretation, and avoiding false-negative pathogenic variant sites while filtering out benign and high-frequency variations. Description of the Drawings

[0045] Figure 1 It is a schematic diagram of the overall process of the variant filtering method described in the present invention.

[0046] Figure 2 Schematic diagram of the construction process of the Snv / Indel variant black and white lists in the embodiments of the present invention.

[0047] Figure 3 Schematic diagram of the construction process of the fusion gene black and white lists in the embodiments of the present invention.

[0048] Figure 4 Schematic diagram of the bioinformatics filtering and tagging process of Snv / Indel variants in the embodiments of the present invention.

[0049] Figure 5 Schematic diagram of the bioinformatics filtering and tagging process of the fusion gene shown in the embodiments of the present invention. Specific embodiments

[0050] The technical solutions of the present invention will be clearly and completely described below in conjunction with the preferred embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] In one embodiment, referring to the appendix Figure 1 , a variant filtering method for tumor whole transcriptome sequencing is provided, and the steps include:

[0052] 1. Construct the historical detection frequencies of Snv / Indel variants and fusion genes in the local samples;

[0053] 2. Construct a white list of tumor-related genes;

[0054] 3. Construct black / white lists of Snv / Indel variants and fusion genes in the whole transcriptome;

[0055] 4. Select quality thresholds and frequency thresholds to filter low-quality and high-frequency variants;

[0056] 5. Perform filtering of other conditions according to the annotation information;

[0057] 6. Generate an excel result file for the Snv / Indel variants and fusion genes after the bioinformatics filtering in step 5;

[0058] 7. For the results obtained in step 6, perform one-key classification and tagging in the reporting system, and finally obtain a result file for further rating processing.

[0059] The mutation filtering method described in the above embodiments constructs Snv / Indel mutations, local detection frequencies of fusion genes, a whitelist of tumor-related genes, Snv / Indel mutations, and black / white lists of fusion genes; performs bioinformatics analysis filtering and classification tagging on the detected Snv / Indel mutations and fusion genes in the original data; the method makes full use of the locally constructed data information and the quality of the sample data itself, significantly reducing the number of invalid sites that need to be analyzed during the interpretation of clinical reports, shortening the time for site screening in tumor report interpretation, and avoiding false-negative pathogenic variant sites while filtering out benign and high-frequency mutations.

[0060] In one embodiment, the specific steps for constructing the local Snv / Indel mutation and historical detection frequency of fusion genes of a sample are as follows:

[0061] 1. Store the Snv / Indel mutation data of the whole transcriptome of 10,168 historical tumor cases in a CSV file named mutation_data.csv, which includes sample ID, mutation position, mutated base, and read depth. Then use the Pandas library for data processing, count the number of occurrences of each Snv / Indel mutation in the sample set, and calculate the local frequency of each Snv / Indel mutation.

[0062] 2. Store the detection results of fusion genes in the whole transcriptome of 10,168 historical tumor cases in a CSV file named fusion_data.csv, which includes sample ID, gene name, breakpoint 1, breakpoint 2, and reads supporting the fusion. Use the Pandas library to read this file, and then perform a deduplication operation on the gene names in the same sample. The same fusion gene is only counted once for detection in a sample. Finally, count the number of occurrences of each fusion gene in the sample set, and calculate the local frequency of each fusion gene.

[0063] In one embodiment, by sorting out the NCCN guidelines, expert consensus, literature reviews, and COSMIC tumor-related genes of each tumor, a "whitelist of tumor-related genes" is generated, totaling 2,016 genes.

[0064] Refer to Appendix Figure 2 , in a specific embodiment, the construction process of the black and white lists of Snv / Indel mutations is as follows:

[0065] 1. Generate a "temporary white list 1 of Snv / Indel mutations" with a total of 546 Snv / Indel mutations from the Snv / Indel mutations of Tier I and Tier II with clinical significance in the artificial evaluation reports of the whole transcriptome data of 10,168 historical tumors;

[0066] 2. Download the ClinGen and Clinvar databases, extract the pathogenic / suspected pathogenic variants of tumor-related genes, and generate the "Temporary Snv / Indel Variant Whitelist 2", with a total of 1,976 Snv / Indel variants;

[0067] 3. Use the open() function to read the whitelist files "Temporary Snv / Indel Variant Whitelist 1" and "Temporary Snv / Indel Variant Whitelist 1", and then use the union() function to generate the "Snv / Indel Variant Whitelist 1", with a total of 2,522 Snv / Indel variants;

[0068] 4. For the historical 10,168 cases of tumor whole transcriptome data, generate the "Temporary Snv / Indel Variant Blacklist 1" for the variants that are not rated as Tier I, Tier II, or Tier III after manual review;

[0069] 5. Use the open() function to read the files "Temporary Snv / Indel Variant Blacklist 1" and "Snv / Indel Variant Whitelist 1", and then use the difference() function to obtain the variants that are not in the "Snv / Indel Variant Whitelist 1" but are in the "Temporary Snv / Indel Variant Blacklist 1", generating the "Snv / Indel Variant Blacklist 1", with a total of 1,006 Snv / Indel variants.

[0070] Refer to Appendix Figure 3 , and in a specific embodiment, the construction process of the fusion gene black and white lists is as follows:

[0071] 1. Generate the "Temporary Fusion Gene Whitelist 1" for the fusion genes of levels 1 and 2 with clinical significance evaluated manually from the historical 10,168 cases of tumor whole transcriptome data, with a total of 353 fusion genes;

[0072] 2. Organize the fusion genes in each tumor NCCN guideline, expert consensus, literature review, and COSMIC data to generate the "Temporary Fusion Gene Whitelist 2", with a total of 1,000 fusion genes;

[0073] 3. Use the open() function to read the whitelist files "Temporary Fusion Gene Whitelist 1" and "Temporary Fusion Gene Whitelist 2", and use the union() function to take the union to generate the "Fusion Gene Whitelist 1", with a total of 1,246 fusion genes;

[0074] 4. Generate the "Temporary Fusion Gene Blacklist 1" based on the monitoring results of 100 patients among the newly diagnosed patients in the historical 10,168 cases of tumor whole transcriptome data and the fusion genes verified as false positives in the control Hela cell line, with a total of 42 fusion genes;

[0075] 5. Using 100 healthy individuals as controls to detect fusion genes and generate "Temporary Fusion Gene Blacklist 2", with a total of 839 fusion genes;

[0076] 6. Taking the fusion genes in the difference set of "Temporary Fusion Gene Blacklist 2" that are not in "Fusion Gene Whitelist 1" as "Temporary Fusion Gene Blacklist 3", with a total of 839 fusion genes;

[0077] 7. Taking the union of "Temporary Fusion Gene Blacklist 1" and "Temporary Fusion Gene Blacklist 3" to generate "Fusion Gene Blacklist 1", with a total of 871 fusion genes;

[0078] 8. Based on the historical 10,168 cases of tumor whole transcriptome data, generating "Temporary Fusion Gene Blacklist 4" for the fusion genes that are not rated as level 1, level 2, or level 3 after manual review;

[0079] 9. Taking the variants in the difference set of "Temporary Fusion Gene Blacklist 4" that are not in "Fusion Whitelist 1" to generate "Fusion Gene Blacklist 2", with a total of 3,717 fusion genes;

[0080] In one embodiment, referring to Appendix Figure 4 , the specific process of bioinformatics Snv / Indel variant filtering is as follows:

[0081] 1. Filtering low-quality and historical local high-frequency variants

[0082] For the Snv / Indel variants originally detected in the sample, judge whether the variant sites in the Snv / Indel variant data "vcf file 1" of the sample meet any of the following conditions. If so, delete them and output the remaining Snv / Indel variants, "Snv / Indel Variant Bioinformatics Filtering Result 1":

[0083] A. The DP (sequencing coverage) of the Snv / Indel variant < 2;

[0084] B. The DP (sequencing coverage) of the Snv / Indel variant * AF (variant frequency) ≤ 2;

[0085] C. The local detection frequency in 10,168 historical tumor cases > 20%;

[0086] For the local detection frequency threshold, multiple sample subsets were selected from the historical 10,168 tumor full transcriptome positive samples according to different gradients such as 5, 10, and 50 samples. Different historical local detection frequency thresholds were set for each gradient. The number of variants retained in each gradient of each subset corresponding to a positive detection rate of 100% was sorted from small to large, and the threshold near the first quartile was selected as the historical local detection frequency threshold. By filtering low-quality and historical local high-frequency variants, the reporting of false positive sites can be reduced.

[0087] 2. The process of filtering based on the Snv / Indel mutation annotation information, the region where the mutation is located, the population frequency, the mutation type and other conditions is as follows:

[0088] 1) For Snv / Indel variant bioinformatics filtering results 1, the variants belonging to intergenic regions, ncRNA, and upstream and downstream regions of genes are filtered out to reduce the interference of non-coding region variants on variant screening;

[0089] 2) For the Snv / Indel variant bioinformatics filtering results, "1" variants that belong to gene introns and non-coding regions (UTRs) and have a population frequency greater than or equal to 0.1 or variants that do not affect splicing are filtered out, and variants in non-coding regions that may affect gene function are retained;

[0090] 3) For Snv / Indel variant bioinformatics filtering results 1, the variant filtering process for variants belonging to exons and splicing regions is as follows:

[0091] 3.1 Filter the Snv / Indel variant bioinformatics filtering results 1" variants belonging to exons and splicing regions with a population frequency greater than or equal to 0.1, reducing the number of benign variants entering the subsequent manual interpretation process and increasing the workload of screening;

[0092] 3.2 For the variants retained after filtering in step 3.1, those that are not in the "Snv / Indel variant whitelist 1" or the COSMIC included variants with a population frequency greater than or equal to 20 and a population frequency greater than or equal to 0.01 are filtered out;

[0093] 3.3 Filter and remove synonymous mutations that do not affect splicing among the variants retained after filtering in step 3.2;

[0094] 3.4 Filter and remove the splicing variants whose local detection frequency is greater than or equal to 5% and whose COSMIC inclusion number is less than 3 among the variants retained after filtering in step 3.2;

[0095] After the above filtering process, the "Bioinformatics Filtering Result 2" Excel is output, displayed in the reporting system and used for subsequent analysis.

[0096] In a specific embodiment, referring to the appendix Figure 4 , the one-key classification and labeling process of Snv / Indel mutations is as follows:

[0097] 1. For the "bioinformatics filtering result 2" excel result, add the label of "not in the candidate gene list, non-hotspot" for mutations where the gene of the Snv / Indel mutation is not in the "whitelist of tumor-related genes" and the Clinvar database entry is not "Pathogenic" or "Likely_Pathogenic" or the COSMIC database entry is ≤ 10 or is a null value.

[0098] 2. For the "bioinformatics filtering result 2" excel result, add the label of "high local detection frequency, non-hotspot" for mutations where the local detection frequency of the Snv / Indel mutation ≥ 5% and the Clinvar database entry is not "Pathogenic" or "Likely_Pathogenic" or the COSMIC database entry is ≤ 10 or is a null value.

[0099] 3. For the "bioinformatics filtering result 2" excel result, add the label of "population frequency higher than 0.001, non-hotspot" for mutations where the maximum population frequency of the Snv / Indel mutation ≥ 0.001 and the COSMIC database entry is ≤ 10 or is a null value.

[0100] 4. For the "bioinformatics filtering result 2" excel result, add the label of "low-frequency repeat region mutation" for mutations where the number of short tandem repeats of the Snv / Indel mutation ≥ 5 and the mutation frequency of the Snv / Indel mutation in the tumor sample ≤ 10%.

[0101] 5. Add the label of "non-classical transcript variant" for mutations where the transcript of the Snv / Indel mutation is a non-classical transcriptome variant.

[0102] 6. For the "bioinformatics filtering result 2" excel result, add the label of "blacklist Snv / Indel" for mutations in the "Snv / Indel mutation blacklist 1".

[0103] 7. According to the review of the Snv / Indel mutations after one-key classification and labeling, after confirmation, the data with added labels can be directly filtered and removed. At the same time, for Snv / Indel mutations that need to be further determined, click to remove the label, and they can continue to be retained for subsequent rating analysis.

[0104] In a specific embodiment, referring to the appendix Figure 5 , the specific process of bioinformatics filtering of fusion genes is as follows:

[0105] 1. For the fusion gene results detected from the original data, first filter the fusion genes according to the fusion genes in the "Fusion Gene Blacklist 1" to obtain the "Fusion Bioinformatics Filtering Result 1".

[0106] 2. Perform quality filtering on the fusion genes that are not in the "Fusion Gene Whitelist 1" for fusion gene detection to obtain the "Fusion Bioinformatics Filtering Result 2".

[0107] 3. The quality filtering threshold is as follows: if either (split_reads1 + split_reads2) or (discordant_mates) of the fusion gene is 0, then keep it when split_reads1 + split_reads2 + discordant_mates ≥ 5; if both (split_reads1 + split_reads2) and (discordant_mates) are not 0, then keep it when split_reads1 + split_reads2 + discordant_mates ≥ 3.

[0108] In a specific embodiment, referring to the appendix Figure 5 , the one - key classification and tagging process of fusion genes is as follows:

[0109] 1. Add tags according to the gene annotation content of gene 1 or gene 2 of the fusion gene:

[0110] A. Fusion genes in which gene 1 or gene 2 contains lncRNA such as RP4 - 761J14.8;

[0111] B. Fusion genes in which gene 1 or gene 2 contains ncRNA such as MIR181A1HG;

[0112] If it meets any one of the above two conditions, the "non - coding gene" tag can be added.

[0113] 2. Add tags according to the region of the location annotation of gene 1 or gene 2 of the fusion gene: For fusion genes in which the location annotation of gene 1 or gene 2 is located in the intergenic region, add the "intergenic region" tag.

[0114] 3. Add tags according to the fusion gene pair: For fusion genes with the same gene 1::gene 2 location, keep the one with the retained fusion gene read_frame result of "in - frame" and the most Total_reads, and add the "duplicate fusion gene pair" tag to the remaining identical fusion gene pairs.

[0115] 4. Match the breakpoints detected for the fusion genes with the fusion gene blacklist 2, and add the label "fusion gene blacklist 2" to the fusion genes within the fusion gene blacklist 2.

[0116] 5. Based on the review of the one-click classification and labeling of the fusion genes, after confirmation, the data with added labels can be directly filtered and removed. At the same time, for the fusion genes that need to be further determined, click to remove the label, and they can continue to be retained for subsequent rating analysis.

[0117] To evaluate the effect of the present invention, the present invention will hereinafter explain the running time and filtering advantages of the Snv / Indel variation and fusion gene filtering in the RNA-seq data detection of the present invention in combination with detection cases.

[0118] Case 1 RNA-seq data filtering results of T-lymphoblastic lymphoma

[0119] 1. After the RNA-seq sequencing data of this sample was downloaded and aligned, the number of original detected results of Snv / Indel variations and fusion genes was 314,070 and 20 respectively;

[0120] 2. After the Snv / Indel variations were filtered by bioinformatics annotation, 305 Snv / Indel variations remained. After the remaining 305 variations went through the one-click classification and labeling process, 19 Snv / Indel variations remained;

[0121] 3. After the fusion genes were filtered by the fusion gene blacklist, 15 fusion genes remained. After the remaining 15 fusion genes went through the one-click classification and labeling process, 10 fusion genes remained;

[0122] 4. The results of this example were interpreted manually according to the cancer variant interpretation and reporting standards jointly issued by AMP, ASCO, and CAP in 2017. The reported SET-NUP214 fusion gene and the NOTCH1 NM_017617:exon27:c.5033T>A:p.L1678Q and JAK3 NM_000215:exon13:c.1718C>T:p.A573V gene variations were all retained in the filtered results.

[0123] Case 2 RNA-seq data filtering results of acute promyelocytic leukemia

[0124] 1. After the RNA-seq sequencing data of this sample was downloaded and aligned, the number of original detected results of Snv / Indel variations and fusion genes was 332,110 and 26 respectively;

[0125] 2. After the Snv / Indel variations were filtered by bioinformatics annotation, 360 Snv / Indel variations remained. After the one-key classification and tagging process for the remaining 360 variations, 17 Snv / Indel variations remained;

[0126] 3. After the fusion genes were filtered by the fusion gene blacklist, 8 fusion genes remained. After the one-key classification and tagging process for the remaining 8 fusion genes, 5 fusion genes remained;

[0127] 4. The results of this case were interpreted manually according to the cancer variant interpretation and reporting standards jointly issued by AMP, ASCO, and CAP in 2017. The reported PML-RARA, RARA-PML fusion genes and the FLT3 NM_004119:exon20:c.2503G>T:p.D835Y gene variant were all retained in the filtered results.

[0128] RNA-seq Data Filtering Results of Case 3 Acute Lymphoblastic Leukemia

[0129] 1. After the RNA-seq sequencing data of this sample was downloaded and aligned, the original detected numbers of Snv / Indel variations and fusion genes were 327,033 and 69 respectively;

[0130] 2. After the Snv / Indel variations were filtered by bioinformatics annotation, 288 Snv / Indel variations remained. After the one-key classification and tagging process for the remaining 288 variations, 22 Snv / Indel variations remained;

[0131] 3. After the fusion genes were filtered by the fusion gene blacklist, 25 fusion genes remained. After the one-key classification and tagging process for the remaining 25 fusion genes, 5 fusion genes remained;

[0132] 4. The results of this case were interpreted manually according to the cancer variant interpretation and reporting standards jointly issued by AMP, ASCO, and CAP in 2017. The reported TAF15-ZNF384 fusion gene and the KRAS NM_033360:exon4:c.437C>T:p.A146V, KRAS NM_033360:exon3:c.182A>G:p.Q61R, ZEB2

[0133] NM_014795:exon10:c.3113A>G:p.H1038R, NF1

[0134] NM_001042492:exon12:c.1392+1delinsCC gene variants were all retained in the filtered results.

[0135] RNA-seq Data Filtering Results for Case 4: Acute Myeloid Leukemia

[0136] 1. After alignment of the RNA-seq sequencing data of this sample, the number of original detected Snv / Indel variations and fusion genes were 340,419 and 44 respectively;

[0137] 2. After bioinformatics annotation filtering of Snv / Indel variations, 377 Snv / Indel variations remained. After the one-key classification and labeling process for the remaining 377 variations, 16 Snv / Indel variations remained;

[0138] 3. After filtering the fusion genes by the fusion gene blacklist, 19 fusion genes remained. After the one-key classification and labeling process for the remaining 19 fusion genes, 8 fusion genes remained;

[0139] 4. The results of this case were interpreted manually according to the cancer variant interpretation and reporting standards jointly issued by AMP, ASCO, and CAP in 2017. The reported CBFB-MYH11 fusion gene and the KIT NM_000222:exon8:c.1250_1256delinsG:p.T417_D419delinsS gene variant were both retained in the filtered results.

[0140] RNA-seq Data Filtering Results for Case 5: Chronic Myeloid Leukemia

[0141] 1. After alignment of the RNA-seq sequencing data of this sample, the number of original detected Snv / Indel variations and fusion genes were 352,095 and 58 respectively;

[0142] 2. After bioinformatics annotation filtering of Snv / Indel variations, 452 Snv / Indel variations remained. After the one-key classification and labeling process for the remaining 452 variations, 19 Snv / Indel variations remained;

[0143] 3. After filtering the fusion genes by the fusion gene blacklist, 15 fusion genes remained. After the one-key classification and labeling process for the remaining 15 fusion genes, 4 fusion genes remained;

[0144] 4. The results of this case were interpreted manually according to the cancer variant interpretation and reporting standards jointly issued by AMP, ASCO, and CAP in 2017. The reported BCR-ABL1 fusion gene was retained in the filtered results.

[0145] RNA-seq Data Filtering Results for Case 6: Acute Myeloid Leukemia

[0146] 1. After the RNA-seq sequencing data of this sample was downloaded and aligned, the numbers of original detected Snv / Indel variations and fusion genes were 237,116 and 87 respectively;

[0147] 2. After the Snv / Indel variations were filtered by bioinformatics annotation, 366 Snv / Indel variations remained. After the remaining 366 variations went through the one-key classification and tagging process, 16 Snv / Indel variations remained;

[0148] 3. After the fusion genes were filtered by the fusion gene blacklist, 31 fusion genes remained. After the remaining 31 fusion genes went through the one-key classification and tagging process, 7 fusion genes remained;

[0149] 4. The results of this case were interpreted manually according to the cancer variant interpretation and reporting standards jointly issued by AMP, ASCO, and CAP in 2017. The reported KMT2A-ELL fusion gene and the WT1 NM_024426:exon7:c.1124delinsCC:p.R375Pfs*15, CCND3 NM_001760:exon5:c.811dupC:p.R271Pfs*53, MRE11

[0150] NM_005590:exon15:c.1773_1774del:p.G593Kfs*4 gene variations were all retained in the filtered results.

[0151] RNA-seq Data Filtering Results of Case 7 Acute Myeloid Leukemia

[0152] 1. After the RNA-seq sequencing data of this sample was downloaded and aligned, the numbers of original detected Snv / Indel variations and fusion genes were 237,116 and 22 respectively;

[0153] 2. After the Snv / Indel variations were filtered by bioinformatics annotation, 267 Snv / Indel variations remained. After the remaining 267 variations went through the one-key classification and tagging process, 19 Snv / Indel variations remained;

[0154] 3. After the fusion genes were filtered by the fusion gene blacklist, 8 fusion genes remained. After the remaining 8 fusion genes went through the one-key classification and tagging process, 4 fusion genes remained;

[0155] 4. The results of this case were interpreted manually according to the cancer variant interpretation and reporting standards jointly issued by AMP, ASCO, and CAP in 2017. The reported KMT2A-MLLT3 fusion gene and the KRAS NM_033360:exon2:c.35G>A:p.G12D gene variant were both retained in the filtered results.

[0156] RNA-seq Data Filtering Results of Case 8 Acute Myeloid Leukemia

[0157] 1. After the RNA-seq sequencing data of this sample was downloaded and aligned, the numbers of original detected Snv / Indel variants and fusion genes were 306,905 and 20 respectively;

[0158] 2. After the Snv / Indel variants were filtered by bioinformatics annotation, 353 Snv / Indel variants remained. After the remaining 353 variants underwent a one-key classification and tagging process, 27 Snv / Indel variants remained;

[0159] 3. After the fusion genes were filtered by the fusion gene blacklist, 18 fusion genes remained. After the remaining 18 fusion genes underwent a one-key classification and tagging process, 5 fusion genes remained;

[0160] 4. The results of this case were interpreted manually according to the cancer variant interpretation and reporting standards jointly issued by AMP, ASCO, and CAP in 2017. The reported NUP98-NSD1 and NSD1-NUP98 fusion genes, as well as the FLT3 FLT3-ITD, IDH2 NM_002168:exon4:c.419G>A:p.R140, and POLQ NM_199420:exon16:c.3214_3215del:p.L1072Vfs*2 gene variants, were all retained in the filtered results.

[0161] RNA-seq Data Filtering Results of Case 9 Acute Lymphoblastic Leukemia

[0162] 1. After the RNA-seq sequencing data of this sample was downloaded and aligned, the numbers of original detected Snv / Indel variants and fusion genes were 351,373 and 60 respectively;

[0163] 2. After the Snv / Indel variants were filtered by bioinformatics annotation, 242 Snv / Indel variants remained. After the remaining 242 variants underwent a one-key classification and tagging process, 12 Snv / Indel variants remained;

[0164] 3. After filtering the fusion genes through the fusion gene blacklist, 30 fusion genes remained. After the one-key classification and tagging process for the remaining 30 fusion genes, 10 fusion genes remained;

[0165] 4. The results of this example were interpreted manually according to the cancer variant interpretation and reporting standards jointly issued by AMP, ASCO, and CAP in 2017. The ETV6-RUNX1 and HRF1-KDM4B1 fusion genes were reported and retained in the filtered results.

[0166] RNA-seq data filtering results for Case 10 of acute lymphoblastic leukemia

[0167] 1. After the RNA-seq sequencing data of this sample was downloaded and aligned, the numbers of original detected Snv / Indel variants and fusion genes were 346,240 and 28 respectively;

[0168] 2. After filtering the Snv / Indel variants through bioinformatics annotation, 296 Snv / Indel variants remained. After the one-key classification and tagging process for the remaining 296 variants, 18 Snv / Indel variants remained;

[0169] 3. After filtering the fusion genes through the fusion gene blacklist, 21 fusion genes remained. After the one-key classification and tagging process for the remaining 21 fusion genes, 6 fusion genes remained;

[0170] 4. The results of this example were interpreted manually according to the cancer variant interpretation and reporting standards jointly issued by AMP, ASCO, and CAP in 2017. The TCF3-PBX1 fusion gene was reported and retained in the filtered results.

[0171] Comparative analysis of each link in the filtering of Snv / Indel variants and fusion genes for the above 10 samples shortened the time for variant screening. While filtering out benign and high-frequency variants, it avoided missing Snv / Indel variants and fusion genes with clinical significance, as shown in Table 1 below.

[0172] Table 1: Time comparison of variant screening

[0173]

[0174] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for filtering variants in whole transcriptome sequencing of tumors, characterized in that: Including Snv / Indel variation filtering and fusion gene filtering; among them: The Snv / Indel variation filtering step includes: filtering the Snv / InDel variation generated by the original data in sequence for low-quality and high-frequency variation, filtering the non-coding region variation, and filtering the exon and splicing region variation in combination with the population frequency and the Snv / InDel variation whitelist; then classifying and labeling the filtered results in combination with the tumor-related gene whitelist, the Snv / InDel variation blacklist and whitelist, and the label types include: non-hotspot, low-frequency repeat region, non-classical transcript variation and blacklist Snv / InDel; The fusion gene filtering step includes: filtering the fusion genes detected from the original data in combination with the fusion gene blacklist and the fusion gene whitelist in sequence, and classifying and labeling the filtering results, wherein the label types include: non-coding genes, gene intervals, repeated fusion gene pairs, and fusion gene blacklist; The construction process of the fusion gene blacklist includes: Based on the historical tumor transcriptome data of several patients' monitoring results and the fusion genes verified as false positives in the control cell lines, a temporary fusion gene blacklist was generated; Using healthy human controls to detect fusion genes, a temporary fusion gene blacklist II was generated; Take the fusion genes in the temporary fusion gene blacklist 2 that are not in the fusion gene whitelist to generate the temporary fusion gene blacklist 3; Taking the union of temporary fusion gene blacklist 1 and temporary fusion gene blacklist 3 to generate fusion gene blacklist 1; A temporary fusion gene blacklist 4 was produced based on the historical tumor transcriptome data and manually reviewed and rated as non-level 1, 2, or 3 fusion genes; Take the fusion genes in the fusion gene blacklist 4 that are not in the fusion gene whitelist to generate the fusion gene blacklist 2; The fusion gene filtering process is based on the fusion gene blacklist 1 and the fusion gene whitelist to obtain the fusion gene filtering result, and the classification and labeling process is based on the breakpoints detected by the fusion gene in the filtering result and the fusion gene blacklist 2 to match, and add the label of "fusion gene blacklist 2".

2. The variation filtering method according to claim 1, characterized in that: The Snv / InDel mutation whitelist construction process includes: Based on historical tumor transcriptome data, the mutations rated as Tier I and Tier II after manual review were included in the temporary Snv / InDel mutation whitelist 1; Extract pathogenic and suspected pathogenic mutations in tumor-related genes and generate a temporary Snv / InDel mutation whitelist II; Merge the temporary Snv / InDel mutation whitelist 1 and the temporary Snv / InDel mutation whitelist 2 to obtain the Snv / InDel mutation whitelist; The Snv / InDel mutation blacklist construction process includes: Based on historical tumor transcriptome data, a temporary Snv / InDel mutation blacklist was obtained after manual review and rating of non-Tier I, Tier II, and Tier III mutations; The Snv / InDel mutation blacklist is generated by taking the mutations in the temporary Snv / InDel mutation blacklist that are not in the Snv / InDel mutation whitelist.

3. The variation filtering method according to claim 1, characterized in that: The filtering of non-coding region variations includes: filtering and removing variations belonging to intergenic regions, ncRNA, and upstream and downstream regions of genes, and filtering variations in introns and UTR regions with a population frequency of more than 0.1 or that do not affect splicing; and / or, The filtering of exon and splicing region variants includes: filtering variants belonging to exons and splicing regions with a population frequency of more than 0.1, filtering variants that are not in the Snv / Indel variant whitelist or database with a population frequency of more than 20 and more than 0.01, filtering variants that are synonymous mutations and do not affect splicing, and filtering variants with a local detection frequency of more than 5% and less than 3 included in COSMIC for splicing variants.

4. The variation filtering method according to claim 1, characterized in that: In the Snv / Indel variation filtering step, the non-hotspot types include: not in the candidate gene list, high local detection rate, and population frequency higher than 0.

001.

5. The variation filtering method according to claim 4, characterized in that: The non-hotspot types are labeled based on the following criteria: For genes with Snv / Indel mutations that are not in the tumor-related gene whitelist, if they meet the requirements of not being pathogenic or suspected pathogenic in the Clinvar database, or meet the requirements of ≤10 or null mutations in the COSMIC database, add the label "not in the candidate gene list, non-hotspot"; For Snv / Indel variants with a local detection frequency of ≥5%, which meet the requirements of not being pathogenic or suspected pathogenic in the Clinvar database, or meet the requirements of ≤10 or null values ​​in the COSMIC database, add the label "Local detection frequency is high, non-hotspot"; For Snv / Indel variants with a maximum population frequency ≥ 0.001 and a COSMIC database inclusion ≤ 10 or a null value, a "population frequency higher than 0.001, non-hotspot" label was added.

6. The variation filtering method according to claim 1, characterized in that: In the Snv / Indel variation filtering step, the low-frequency repeat region type marker is based on the number of short tandem repeats of the Snv / Indel variation being ≥5 and the Snv / Indel variation frequency of the tumor sample being ≤10%.

7. The variation filtering method according to claim 1, characterized in that: The construction process of the fusion gene whitelist includes: A temporary fusion gene whitelist 1 is generated based on the historical tumor transcriptome data and manually reviewed and rated as level 1 or 2 fusion genes; Generate temporary fusion gene whitelist 2 based on fusion genes recorded in literature and databases; The union of temporary fusion gene whitelist 1 and temporary fusion gene whitelist 2 is taken as the fusion gene whitelist.

8. The variation filtering method according to claim 1, characterized in that: In the fusion gene filtering step, after the fusion genes detected in the original data are filtered, the reference sequence is compared, and according to whether one of split_reads1+split_reads2 and discordant_mates is 0, the fusion genes whose split_reads1+split_reads2+discordant_mates is greater than different thresholds are retained.

9. The variation filtering method according to claim 1, characterized in that: The process also includes a step of reviewing Snv / Indel variants and fusion gene tags, wherein the review process includes confirming tags for direct filtering and removal of data, and confirming pending tags.

10. A variation filtering device for tumor whole transcriptome sequencing, characterized in that: include: Snv / Indel variation filtering module: used to filter Snv / InDel variations generated by raw data in turn for low-quality and high-frequency variations, filtering non-coding region variations, and filtering exon and splicing region variations based on population frequency and Snv / InDel variation whitelist; Then, the filtered results are classified and labeled by combining the whitelist of tumor-related genes and the blacklist and whitelist of Snv / InDel mutations. The label types include: non-hotspots, low-frequency repeat regions, non-classical transcript mutations, and blacklist Snv / InDel; The fusion gene filtering module is used to filter the fusion genes detected from the original data in combination with the fusion gene blacklist and the fusion gene whitelist, and classify and label the filtering results. The label types include: non-coding genes, gene intervals, and repeated fusion gene pairs; The construction process of the fusion gene blacklist includes: Based on the historical tumor transcriptome data of several patients' monitoring results and the fusion genes verified as false positives in the control cell lines, a temporary fusion gene blacklist was generated; Using healthy human controls to detect fusion genes, a temporary fusion gene blacklist II was generated; Take the fusion genes in the temporary fusion gene blacklist 2 that are not in the fusion gene whitelist to generate the temporary fusion gene blacklist 3; Taking the union of temporary fusion gene blacklist 1 and temporary fusion gene blacklist 3 to generate fusion gene blacklist 1; A temporary fusion gene blacklist 4 was produced based on the historical tumor transcriptome data and manually reviewed and rated as non-level 1, 2, or 3 fusion genes; Take the fusion genes in the fusion gene blacklist 4 that are not in the fusion gene whitelist to generate the fusion gene blacklist 2; The fusion gene filtering process is based on the fusion gene blacklist 1 and the fusion gene whitelist to obtain the fusion gene filtering result, and the classification and labeling process is based on the breakpoints detected by the fusion gene in the filtering result and the fusion gene blacklist 2 to match, and add the label of "fusion gene blacklist 2".

Citation Information

Patent Citations

  • Gene fusion and mutation detection method and system based on RNAseq data

    CN110021346A

  • Variation filtering method based on multi-sample whole exon sequencing

    CN114664375A

  • RNA-seq-based data analysis, variation rating and report generation system and method

    CN116453591A