A method for detecting large germline gene rearrangements
By processing NGS sequencing data and CNV algorithms, combined with Zscore value and breakpoint analysis, the detection problem of large-fragment rearrangements of the germline RB1 gene was solved, and efficient and low-cost diagnosis of hereditary retinoblastoma and genetic susceptibility gene detection of other cancers were achieved.
Patent Information
- Application Number
- CN202310071294.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-31
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-01-31
AI Technical Summary
Existing multiplex amplicon quantification (MAQ) and multiplex ligation-dependent probe amplification (MLPA) technologies are difficult to effectively detect large rearrangements of the germline RB1 gene, and have problems such as high cost, strict experimental conditions, and insufficient detection sensitivity and specificity.
Using NGS sequencing data, the sequencing data of the exon region is processed through the CNV algorithm to screen candidate exons. Combined with Zscore value and breakpoint analysis, possible large fragment rearrangement regions are screened to improve the accuracy and sensitivity of detection.
It achieves high-sensitivity and high-specificity detection of large-fragment rearrangements of the germline RB1 gene, reduces detection costs, and improves detection efficiency. It is suitable for early diagnosis of familial hereditary retinoblastoma and genetic susceptibility gene detection of other cancers.
Smart Images

Figure CN116013410B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a germline RB1 The invention relates to a method for detecting and analyzing large fragment rearrangements, and in particular to a technology for detecting large fragment rearrangements based on high-throughput sequencing data, and belongs to the field of bioinformatics technology. Background Art
[0002] Retinoblastoma gene (Retinoblastomal transcriptional corepressor 1, RB1 ) is located on human chromosome 13 and is the first tumor suppressor gene discovered in humans. This gene is associated with the occurrence of retinoblastoma (RB). Retinoblastoma is the most common intraocular cancer in children, with more than 90% of bilateral cases and 10-25% of unilateral cases being hereditary. RB1 Mutation [1]. Hereditary RB is related to germline mutations, which are inherited from parents or acquired during fetal development. The incidence of Rb in newborns is 1 / 20,000 [2]. About 40% of RB subjects are mainly caused by genetic susceptibility genes. RB1 It is caused by germline mutation of the gene [3] and is generally considered to be familial (hereditary) retinoblastoma, and the inheritance mode is autosomal dominant inheritance. Beijing Children's Hospital affiliated to Capital Medical University enrolled 263 unrelated unilateral RB subjects between February 2014 and August 2020, and used NGS and ddPCR technology to detect RB1 Point mutations and large deletions of the gene were detected in 39 (14.8%) subjects. RB1 Germline mutations, of which 11 subjects (28.2%) had missense mutations, 10 subjects (25.6%) had nonsense mutations, 2 subjects (5.1%) had frameshift mutations, 1 subject (2.6%) had synonymous mutations, 7 subjects (17.9%) had large deletions, 2 subjects (5.1%) had splice site mutations, and 6 subjects (15.4%) had variants of uncertain significance [1]. RB1 Among the subjects with mutations, 7 (17.9%) subjects with large deletions had pathogenic mutations, indicating that the germline RB1 Large rearrangement mutations are closely associated with tumorigenesis. Therefore, developing an NGS-based technology to detect large rearrangements is crucial for early diagnosis and treatment of familial genetic diseases to prevent further tumor progression and metastasis.
[0003] Large genomic rearrangements (LGRs) in germline mutations refer to duplications or deletions of hundreds to millions of bases, often involving one or more exons. Compared with conventional point mutations or small insertions and deletions, they are more difficult to detect. Currently, the main technologies used to detect large rearrangements include multiplex ligation probe amplification (MLPA), multiplex amplicon quantification (MAQ) (multiplex PCR), and next-generation sequencing (NGS)-CNV.
[0004] The multiplex ligation-linked probe amplification (MLPA) method is the gold standard for detecting LGR. It is a method based on multiplex polymerase chain reaction (PCR technology) that can detect changes in gene copy number, point mutations, and DNA methylation. It can semi-quantitatively amplify up to 50 MLPA probes simultaneously in a single reaction using a very small amount of DNA, and the relative amount of its amplified products reflects the relative copy number of the target sequence. Because commercial probe kits are relatively expensive, more and more laboratories use homemade MLPA probe mixtures, which require the design of a new mixture of up to 50 MLPA probes. The preparation process of MLPA probes includes probe design, probe synthesis, probe library optimization, performance verification, etc., which is time-consuming [4] and has high requirements for laboratory conditions, making it difficult to popularize in clinical applications. In addition, MLPA probes can only detect known point mutations and will miss unknown and low-frequency point mutations.
[0005] Multiplex amplicon quantification (MAQ), also known as multiplex PCR technology, is a new low-cost and high-throughput PCR-based technology. A single amplification reaction can determine the copy number changes of multiple genes. Compared with the MLPA method, MAQ can reliably detect important genomic aberrations and has broad application prospects [5]. This technology also has some limitations. For example, as the number of multiplex PCR detection increases, dimers or even polymers are easily formed, nonspecific amplification occurs, amplification efficiency decreases, and detection sensitivity and specificity decrease. In severe cases, nonspecific amplification dominates, target amplification fails, resulting in false positives or false negatives, affecting the accuracy of the test results. MAQ technology can also only be used to detect known point mutations.
[0006] References:
[0007] 1. Fang X, Chen J, Wang Y, et al. Gallie BL. RB1单侧视网膜母细胞瘤患者的种系突变谱和临床特征。眼科遗传学。2021年10月;42(5):593 - 599。
[0008] 2. 费尔南德斯AG、波洛克BD、拉比托FA。美国的视网膜母细胞瘤:40年发病率和生存分析。小儿眼科与斜视杂志。2018年5月1日;55(3):182 - 188。
[0009] 3. 上原J、布尔德奥特F、福克斯WD等。视网膜母细胞瘤和神经母细胞瘤的易感性及监测。临床癌症研究。2017年7月1日;23(13):e98 - e106。
[0010] 4. 赫米格 - 赫尔泽尔C、萨沃拉S。多重连接依赖探针扩增(MLPA)在肿瘤诊断和预后评估中的应用。诊断分子病理学。2012年12月;21(4):189 - 206。
[0011] 5. 昆普斯C、范罗伊N、海尔曼L等。多重扩增子定量(MAQ),一种同时检测神经母细胞瘤拷贝数改变的快速有效方法。BMC基因组学。2010年5月12日;11:298。 Summary of the Invention
[0012] The problem to be solved by the present invention is that multiplex amplicon quantification (MAQ) and multiplex ligation probe amplification (MLPA) are currently difficult to infer gene mutations of large fragment rearrangements. A method for detecting germline RB1 The present invention is a method for rearranging large fragments. RB1 NGS sequencing data of genes, RB1 The exon is used as the unit, and the sequencing data is processed using the CNV algorithm to obtain the RB1 The CNV map of each exon is quality-controlled and filtered to select candidate exons. Some candidate exons with low evidence are filtered out according to the amplification or deletion threshold, and finally the sample to be tested is obtained. RB1 Large exon rearrangements.
[0013] A method for detecting large germline gene rearrangements comprises the following steps:
[0014] Step 1: Amplify or capture the exon region of the target gene using primers or probes to obtain a library for high-throughput sequencing;
[0015] Step 2: Perform high-throughput sequencing on the library to obtain RB1 Sequencing depth data of all exon reads, calculate the Zscore value of each exon;
[0016] Step 3: filter and obtain sample data with noise intensity less than a set value;
[0017] Step 4: Make the following determinations:
[0018] Step 4-1, when it is necessary to determine RB1 When an exon is deleted, all SNPs in the target interval must be homozygous;
[0019] Step 4-2, when the Zscore of the exon region is less than the set threshold, the exon is classified into the first exon set;
[0020] Step 4-3: Find the two breakpoints in the exon region based on the read data and obtain the positions of the breakpoints on the reference genome. If the sequences outside the two breakpoints on the reference genome are identical or complementary to the sequencing data after splicing, all exons between the two breakpoints are included in the second exon set.
[0021] Step 4-4: performing a union process on the first exon set and the second exon set to obtain exons with deletions.
[0022] In step 1, the sequencing depth data needs to be normalized, GC corrected or noise reduced.
[0023] The Zscore value is calculated by the following steps:
[0024] Z = (x - μ) / σ;
[0025] x is the Log2Ratio value of a single exon, μ is the mean Log2Ratio of the baseline sample, and σ is the standard deviation of the baseline sample.
[0026] In the third step, the screening process includes the following steps:
[0027] Calculate the VP value, ; ;
[0028] It refers to the log2Ratio value of a certain exon, and n is the total number of exons;
[0029] And it must meet the following requirements: for blood samples, VP≤0.5; for FFPE, fresh tissue samples or other body fluid samples other than blood samples, VP≤0.8.
[0030] In step 3, the criteria for determining heterozygous SNPs are: 0.4 <Vaf<0.6。
[0031] In step 3, the criteria for determining a homozygous SNP are: Vaf<0.1 or Vaf>0.9.
[0032] In step 4-2, the Zscore reaching the set threshold means that: when one exon is deleted, the conditions that need to be met are that the exon Zscore is <-6 and the CNV ratio is <0.70; when 2-8 exons are deleted, the conditions that need to be met are that the Zscore of the target interval is ≤-6; when ≥9 exons are deleted, the conditions that need to be met are that the Zscore of the target interval is ≤-4.
[0033] The above methods are applied to non-therapeutic and diagnostic procedures.
[0034] A device for detecting large germline gene rearrangements, comprising:
[0035] Sequencing module, used for high-throughput sequencing of the library to obtain RB1 The sequencing depth data of all exon reads is used to calculate the Zscore value of each exon; the library is obtained by amplifying or capturing the exon region of the target gene with primers or probes and then processing;
[0036] The denoising module is used to filter the read data obtained by the sequencing module to obtain sample data with noise intensity less than a set value;
[0037] SNP screening module, used to screen out target intervals where all SNPs are homozygous;
[0038] A first exon deletion determination module is configured to determine that when the Zscore of an exon region is less than a set threshold, the exon is included in the first exon set;
[0039] The second exon deletion judgment module is used to find the two breakpoints in the exon region based on the reads data and obtain the position of the breakpoints on the reference genome. If the sequences outside the two breakpoints on the reference genome are identical or complementary to the sequencing data after splicing, all exons between the two breakpoints are classified as the second exon set;
[0040] The union processing module is used to perform union processing on the first exon set and the second exon set to obtain the exons with deletions.
[0041] The screening process of the denoising module includes the following steps:
[0042] Calculate the VP value, ;
[0043] It refers to the log2Ratio value of a certain exon, and n is the total number of exons;
[0044] And it must meet the following requirements: for blood samples, VP≤0.5; for FFPE, fresh tissue samples or other body fluid samples other than blood samples, VP≤0.8.
[0045] A computer-readable medium records computer instructions for executing the above-mentioned method for detecting large germline gene rearrangements. Beneficial effects
[0046] The present invention provides a method for detecting germline based on NGS data. RB1 Bioinformatics analysis methods for large fragment rearrangements can be used for genetic testing of subjects with familial hereditary retinoblastoma.
[0047] Compared with existing technologies, NGS testing RB1 Gene mutation types include point mutations, small insertions and deletions, and large rearrangements. NGS testing is used for the detection of subjects with familial hereditary tumors, showing a more comprehensive range of mutation types and higher sensitivity and specificity. In the routine diagnosis of hereditary retinoblastoma, a sample can simultaneously obtain a more comprehensive RB1 Mutation information can save detection time and reduce detection costs.
[0048] NGS-CNV technology not only detects various germline RB1Providing early diagnosis and treatment services, NGS testing can also be expanded to other cancer susceptibility genes RB1 Pathogenic germline mutations of the LGR gene can even be used to detect LGR of other germline susceptibility genes. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Figure 3 is the AUC curve under different Z-score threshold conditions: A is the AUC curve of single exon deletion Z-score under different thresholds, B is the AUC curve of 2-8 exon deletion Z-score under different thresholds, and C is the AUC curve of ≥9 exon deletion Z-score under different thresholds;
[0050] Figure 2 NGS testing RB1 An example of IGV analysis of large gene deletions using breakpoint sequence pairing.
[0051] Figure 3 This is the CNV result graph before and after sample noise reduction;
[0052] Figure 4 It is an example of IGV with a higher VP sample;
[0053] Figure 5 NGS testing RB1 Examples of IGVs with large gene deletions;
[0054] Figure 6 Retinoblastoma germline RB1 Sector diagram of gene mutation types;
[0055] Figure 7 Other cancers originate in the germline RB1 Distribution of subjects with large fragment rearrangements;
[0056] Figure 8 This is a flow chart of the method. DETAILED DESCRIPTION
[0057] NGS testing of germline cells in the following implementation scheme RB1 The detection of large gene fragment rearrangement is obtained by genetic testing using the Shihe No. 1 panel independently developed by Shihe Gene. This sequencing panel has comprehensive coverage RB1 The whole exon and important intron regions of the gene, the specific gene list of the Shihe No. 1 panel is shown in Table 1. In general, the method of the present invention can detect the tumor susceptibility gene germline through the NGS sequencing data obtained by Shihe No. 1 panel. RB1 Large gene rearrangement, used for early diagnosis of retinoblastoma and other solid tumors.
[0058] Table 1: 425 panel gene list
[0059]
[0060]
[0061]
[0062]
[0063] The above detection process uses the gene panel as an example, which is only used to obtain RB1 Gene sequencing data does not necessarily mean that a specific gene panel is required for testing.
[0064] In the present invention, RB1 LGR often manifests as heterozygous copy number loss of one or more exons.
[0065] The NGS detection process involves sample extraction, library preparation, and sequencing, which can be processed using existing high-throughput sequencing methods. The present invention is not particularly limited to this. When processing the data off the machine, the following steps are based on:
[0066] Obtain CNV analysis data:
[0067] The sequencing data is processed based on the CNV algorithm to obtain the CNV map. The algorithm is described as follows: RB1 The exon read counts were normalized, the average sequencing depth of the samples to be tested was calculated, GC correction was performed and the samples were relocated to the diploid state, the log2Ratio value was calculated, noise reduction was performed based on the healthy sample pool and PCA, the noise and true CNV signals were removed from the samples to be tested, and the Zscore value of the sequencing depth of each exon was calculated based on the Log2Ratio after noise reduction.
[0068] The Zscore value of each exon is calculated as follows:
[0069] Z = (x - μ) / σ;
[0070] x is the Log2Ratio value of a single exon, μ is the mean Log2Ratio of the baseline sample, and σ is the standard deviation of the baseline sample.
[0071] Next, VP quality control is required. The VP quality control is used to display the noise intensity of the sample CNV map. It is specifically described as the average square of the log2Ratio difference between each exon and the previous exon. The calculation formula is as follows:
[0072] ; If the VP value is too high, it is considered that the sample has excessive noise and the quality control is unqualified, which will lead to false positive or false negative results in the sample copy number variation.
[0073] exon k 、exon k-1 are the previous exon of each exon and each exon respectively.
[0074] Based on RB1 the CNV map of exons for quality control filtering, the threshold of the sample needs to meet the following requirements for the quality control to be qualified: for blood samples, VP ≤ 0.5; for FFPE, fresh tissue samples or other body fluid samples, VP ≤ 0.8.
[0075] When the VP of the blood sample ≤ 0.5, it indicates that the quality control of the exons in this sample is qualified;
[0076] When the VP of FFPE, fresh tissue samples or other body fluid samples ≤ 0.8, it indicates that the quality control of the exons in this sample is qualified;
[0077] Next, it is necessary to combine the deletion threshold to filter and screen the candidate exons with qualified quality control to remove some exons with low evidence:
[0078] After screening out the exons with qualified quality control, SNPs are introduced as the basis for auxiliary judgment to exclude false positives. The selected SNP sites are SNP sites with a mutation frequency ≥ 1% in the total population or East Asian population in databases such as 1000g / ExAC / gnomAD, that is, high-frequency germline mutation sites, as shown in Table 2.
[0079] Table 2: RB1 Gene candidate SNP sites
[0080]
[0081] If all SNPs in the target interval are homozygous, it is possible to further judge whether CNV deletion has occurred according to the Zscore threshold; if the SNPs in the target interval are heterozygous or offset SNPs, it is judged that the possibility of deletion is very small; since when a large fragment deletion occurs, one strand of the heterozygous SNP will be lost and it is very likely that it cannot be detected, so the accuracy can be improved through this verification step. The following judgment rules are used during verification: the abundance ranges of the homozygous SNP and offset SNP are: homozygous SNP (Vaf < 0.1 or Vaf > 0.9), heterozygous SNP (0.4 < Vaf < 0.6), offset SNP (0.1 ≤ Vaf ≤ 0.4 or 0.6 ≤ Vaf ≤ 0.9).
[0082] Specifically:
[0083] The necessary condition for determining the deletion threshold is that the SNP type is homozygous. When one exon is deleted, the conditions that need to be met are that the exon Zscore is less than -6 and the exon CNV ratio is less than 0.70. When 2-8 exons are deleted, the conditions that need to be met are that the Zscore of each target interval is less than or equal to -6. When ≥9 exons are deleted, the conditions that need to be met are that the Zscore of each target interval is less than or equal to -4.
[0084] Among them, CNV ratio=2 Log 2 Ratio ;
[0085] The rationality of the Zscore threshold in the above steps was tested by calculating the AUC curves of single exon deletions under the conditions of -2, -3, -4, -5, -6, -7, and -8 for 135 patients with LGR mutation status of RB1 gene verified by ddPCR. Figure 1 A, when the Zscore threshold is -6, the detection accuracy is higher. For 2-8 exon deletions, the AUC curves of Zscore values under the conditions of -2, -3, -4, -5, -6, -7, and -8 are shown in Figure 2. Figure 1 B, when the Zscore threshold is -6, the detection accuracy is higher. ≥9 exon deletions, the AUC curves of Zscore values under the conditions of -2, -3, -4, -5, and -6 are shown in Figure 2. Figure 1 When the Zscore threshold of C is -4, the detection accuracy is higher.
[0086] Furthermore, in the above steps, deletions are determined on an exon-by-exon basis by determining sequencing read depth. However, individual exons may have Z scores or CNV ratios slightly above the threshold due to factors such as GC content, DNA quality, probe coverage, and probe specificity, potentially leading to false negatives or false positives for those exons. To further improve detection sensitivity, this patent uses breakpoint analysis to supplement the test results.
[0087] For target interval breakpoint sequence analysis, we first perform breakpoint analysis across all exon regions based on sequencing data. This can be obtained using existing methods, such as Delly software. This software uses two algorithms to detect gene structural variation: discordantly mapped read-pairs and split-read (https: / / github.com / dellytools / delly.git).
[0088] First, calculate the direction of the read-pair and the insert size distribution. Based on these data, determine whether all discordantly mapped read-pairs have abnormal mapping directions or the insert size is larger than expected. Next, split-read analysis can help find the fusion breakpoint information with single-base accuracy. The analysis process can be divided into multiple steps: 1) Retrieve candidate split-reads; 2) Extract reference genes in the structural variation region; 3) Index and k-mer calculation; 4) Find the best breakpoint support; 5) Split-read consistency calculation; 6) The consensus sequence of the split-read is mapped to the fusion reference gene region to generate a table of candidate RB1 gene structural variations. According to the chromosomal position of the internal breakpoint of RB1, the sample bam format raw data file is loaded through IGV to further check whether the sequences of the two breakpoints match each other, such as Figure 2 Schematic diagram. If the sequences of the two breakpoints within the RB1 gene match each other (the sequences are identical or complementary), we believe that there is a large deletion in the target interval. This method can help determine the specific chromosomal location of the germline large deletion.
[0089] The following testing process is based on the above statistical method, more specifically:
[0090] The sequencing data was processed for copy number variation (CNV) detection. First, the read count of the panel coverage area was calculated and normalized. The average sequencing depth of the sample to be tested was calculated. After GC correction, the sample was relocated to the diploid state. The log2Ratio value was calculated. Noise reduction was performed based on a large sample pool of healthy samples and PCA. The noise and true CNV signal were removed from the sample to be tested. The Zscore value of each exon was calculated based on the Log2Ratio after noise reduction. The CNV signal of the sample after noise reduction was stronger. Figure 3 .right RB1 The CNV map of the exon is subjected to noise reduction and quality control filtering to make the analysis results more reliable. The sequencing data quality control parameter VP can show the noise intensity of the sample CNV. When the sample VP is higher, such as Figure 4 As shown, these characteristics are often accompanied by poor uniformity, low Q30, or sequencing depth warnings. The quality of sequencing data directly affects the accuracy of downstream data analysis. Sufficient data quality assurance is a prerequisite for data analysis, so thorough quality control of sequencing data is necessary to improve data accuracy.
[0091] Six blood samples from subjects with retinoblastoma were included. RB1The gene status has been verified by ddPCR, and the ddPCR verification results are used as the standard. These 6 samples were sequenced using the Shihe No. 1 large panel shown in Table 1. We first tested whether the inclusion of the target region breakpoint sequence pairing relationship affects the performance of the sequencing results. According to the bioinformatics analysis results, the IGV was used to check for large fragment deletions. RB1 Gene, RB1 The large deletions are all located on the short arm of chromosome 13. Figure 5 V1 sample is shown in RB1 The large deletion occurred on both sides of the chromosome at positions 13:49,034,823 and 13:49,053,867, and the breakpoint sequences before and after matched each other, further indicating that a large deletion existed at e21-e26, which was consistent with the Zscore analysis results. Figure 2 Display V5 sample RB1 The large deletion occurs on both sides of the chromosome at positions 13:48,948,963 and 13:48,954,246, respectively. The breakpoint sequences before and after also match each other, further indicating that a large deletion exists at e13-16. However, the Zscore analysis results of this sample only show the deletion of e13-14. Adding matching analysis of the breakpoint sequences can improve the sensitivity of detection.
[0092] Finally, we used the analysis process of the present invention and Atlas-CNV (PMID: 30890783) software to analyze the sequencing data. RB1 The gene status is analyzed and the detection of this method is calculated based on the detected exons. RB1 The sensitivity and specificity of the gene LGR are analyzed. RB1 The specific information of gene LGR is shown in Table 3:
[0093] Table 3: Samples under different analysis process methods RB1 Results of large fragment rearrangement analysis ( n = 6)
[0094]
[0095] Note: ID represents the name of each sample; ddPCR- RB1 Representative ddPCR detection RB1 Gene LGR; Process of the present invention- RB1 Representatives tested and analyzed by the bioinformatics analysis process of the present invention RB1 Gene LGR results; Atlas-CNV- RB1 Representatives were detected and analyzed by Atlas-CNV (PMID: 30890783) software RB1Gene LGR results; del represents gene copy number loss.
[0096] Verification of confirmed germline using ddPCR RB1 The results of gene LGR mutation were used as the standard, with deletion RB1 The performance of two high-throughput LGR analysis methods was analyzed on an exon-by-exon basis, as shown in Table 4. The results demonstrate that both methods achieved 100% specificity, with the proposed method demonstrating a 59% improvement in sensitivity and a 25% improvement in accuracy compared to the Atlas-CNV method. In terms of analytical performance, the proposed analytical process offers superior performance.
[0097] Table 4: Samples under different analysis processes RB1 Large fragment rearrangement analysis performance (in exon units, n=162)
[0098]
[0099] Testing and validation of other sample sets
[0100] A total of 188 cases of retinoblastoma germline cells were collected based on whole blood samples. RB1 The mutation types of each subject were sorted out. The results showed that 130 subjects (69%) had missense mutations, 13 subjects (7%) had large fragment deletions, 31 subjects (17%) had frameshift mutations caused by small fragment deletions, and 14 subjects (7%) had frameshift mutations caused by small fragment duplications. Figure 6 shown.
[0101] The samples of subjects who were sent to Shihe No. 1 for testing from November 2020 to February 2022 were screened, and 15 cases were found to have germline RB1 Large gene rearrangements were all pathogenic or likely pathogenic mutations, including 6 cases (40%) of ovarian cancer, 5 cases (33%) of lung cancer, 3 cases (20%) of breast cancer, and 1 case (6.7%) of pancreatic cancer. Figure 7 shown.
Claims
1. A method for detecting large germline gene rearrangements, characterized in that: The steps include: Step 1: Amplify or capture the exon region of the target gene using primers or probes to obtain a library for high-throughput sequencing; Step 2: Perform high-throughput sequencing on the library to obtain RB1 Sequencing depth data of all exon reads, calculate the Zscore value of each exon; Step 3: filter and obtain sample data with noise intensity less than a set value; Step 4: Make the following determination: Step 4-1, when it is necessary to determine RB1 When an exon is deleted, all SNPs in the target interval must be homozygous; Step 4-2, when the Zscore of the exon region is less than the set threshold, the exon is classified into the first exon set; Step 4-3: Find the two breakpoints in the exon region based on the read data and obtain the positions of the breakpoints on the reference genome. If the sequences outside the two breakpoints on the reference genome are identical or complementary to the sequencing data after splicing, all exons between the two breakpoints are included in the second exon set. Step 4-4, performing a union process on the first exon set and the second exon set as the exons with large deletions; In the third step, the screening process includes the following steps: Calculate the VP value, ; It refers to the log2Ratio value of a certain exon, and n is the total number of exons; And it must meet the following requirements: for blood samples, VP≤0.5; for FFPE or fresh tissue samples, VP≤0.
8.
2. The method for detecting large germline gene rearrangements according to claim 1, wherein: In step 1, the sequencing depth data needs to be normalized, GC corrected or noise reduced.
3. The method for detecting large germline gene rearrangements according to claim 1, wherein: The Zscore value is calculated by the following steps: Z = (x - μ) / σ; x is the Log2 Ratio value of a single exon, μ is the mean Log2 Ratio of the baseline sample, and σ is the standard deviation of the baseline sample.
4. The method for detecting large germline gene rearrangements according to claim 1, wherein: In step 3, the criteria for determining whether the target interval SNP is homozygous are: Vaf<0.1 or Vaf>0.
9.
5. The method for detecting large germline gene rearrangements according to claim 1, wherein: In step 4-2, the Zscore reaching the set threshold means that: when one exon is deleted, the conditions that need to be met are that the exon Zscore is <-6 and the CNV ratio is <0.70; when 2-8 exons are deleted, the conditions that need to be met are that the Zscore of the target interval is ≤-6; when ≥9 exons are deleted, the conditions that need to be met are that the Zscore of the target interval is ≤-4.
6. A device for detecting large germline gene rearrangements, characterized in that: include: Sequencing module, used for high-throughput sequencing of the library to obtain RB1 Sequencing depth data of all exon reads, calculate the Zscore value of each exon; The library is obtained by amplifying or capturing the exon region of the target gene using primers or probes and then processing; The denoising module is used to filter the read data obtained by the sequencing module to obtain sample data with noise intensity less than a set value; SNP screening module, used to screen out target intervals where all SNPs are homozygous; A first exon deletion determination module is configured to determine that when the Zscore of an exon region is less than a set threshold, the exon is included in the first exon set; The second exon deletion judgment module is used to find the two breakpoints in the exon region based on the reads data and obtain the position of the breakpoints on the reference genome. If the sequences outside the two breakpoints on the reference genome are identical or complementary to the sequencing data after splicing, all exons between the two breakpoints are classified as the second exon set; a union processing module, configured to perform union processing on the first exon set and the second exon set to obtain exons with deletions; The screening process of the denoising module includes the following steps: Calculate the VP value, ; It refers to the log2Ratio value of a certain exon, and n is the total number of exons; And it must meet the following requirements: for blood samples, VP≤0.5; for FFPE or fresh tissue samples, VP≤0.
8.
7. A computer-readable medium, characterized in that It records computer instructions that can run the method for detecting large germline gene rearrangements according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method, device and system for screening and diagnosing liver cancer based on high-throughput sequencing method
CN110736834A
Method and device for screening, diagnosing or risk grading of ovarian cancer
CN110880356A