A method for eliminating background and improving detection sensitivity of target sites in NGS data processing
Patent Information
- Application Number
- CN202211658759.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-22
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2042-12-22
AI Technical Summary
[0004]本发明的目的是克服目前测序过程中由于存在较多背景噪音从而导致分析得到的突变位点出现较多假阳性,降低了检测灵敏度的缺陷
[0034] (1) When using primers to amplify the target region, mutations may be introduced. That is, the mutation frequency of sites that are not mutated in the standard can sometimes reach 0.1%, thus generating high background noise. This invention can eliminate the defect of high background signal, improve the detection sensitivity, and reduce false positive detection results.
Smart Images

Figure CN115798583B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene sequencing data analysis, and specifically to a method for improving the sensitivity of target site detection by eliminating background in NGS data processing. Background Technology
[0002] For multiplex gene detection, the most commonly used technology is Thermo Fisher's IonPGM / Proton platform. This platform features a simple targeted sequencing workflow, flexible throughput options, extremely fast sequencing speed, and micro-sample input. Therefore, it is particularly suitable for parallel detection of multiple mutation types in multiple genes. However, for some complex sample types, such as cell-free DNA (cfDNA) samples and formalin-fixed and paraffin-embedded (FFPE) samples, it is currently not possible to effectively and accurately identify target sites. Background noise introduced during sequencing, or the presence of background noise unrelated to tumor development in these samples, leads to a large number of false positive results among the analyzed mutation sites, resulting in the omission of some low-frequency (<1%) mutations and some false positive mutations. In addition, the background of each chip in each batch is likely to be different. For detection with a precision level of 0.1%, the background fluctuation between batches has a significant impact, making background control particularly important. All these factors affect the detection results and sensitivity of NGS-based target site detection methods.
[0003] Current methods for eliminating background noise and improving detection sensitivity in data processing include directly using call mutation software to obtain mutation sites. However, enhancing the reliability of mutation sites increases sequencing throughput, leading to increased detection costs and subsequent computational and time burdens. Therefore, there is an urgent need for a method to improve the sensitivity of NGS gene mutation detection. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of current sequencing processes, which are characterized by a large amount of background noise, leading to a high number of false positives at the analyzed mutation sites and reducing detection sensitivity.
[0005] To achieve the above objectives, this invention provides a method for improving target site detection sensitivity by eliminating background in NGS data processing, comprising the following steps:
[0006] Step 1: Amplicon technology is used to amplify tumor-related genes and regions in the DNA of the sample to be tested;
[0007] Step 2: Provide a negative sample as a control, and use NGS sequencing to sequence the tumor-related genes to obtain the raw sequencing data of the target gene of the sample to be tested;
[0008] Step 3: Align the raw sequencing data with the human reference genome to obtain a bam file containing sequence alignment information. Use samtools mpileup to process the bam file to obtain an mpileup file. When a site appears in the mpileup file whose base is different from that in the human reference genome, the site is determined to be a mutation site, and the mutation frequency AF of the mutation site is calculated.
[0009] Step 4: Determine whether the mutation site meets the predetermined conditions to determine whether the mutation site is a positive site;
[0010] The method for determining whether the mutation site meets the predetermined conditions specifically includes:
[0011] 4.1 Perform preliminary filtering based on the following conditions:
[0012] a.AFTargetCutNeg>AFNonTargetMean+2SD;
[0013] b. AFTargetCutNeg > 0.1%;
[0014] Positive candidate sites in the sample to be tested must be obtained by simultaneously meeting both of the above conditions;
[0015] The AFTargetCutNeg method selects a single point as the target site of the sample to be tested, calculates the mutation frequency AF of all sites in the sample to be tested, and subtracts the mutation frequency of the corresponding site in the negative sample from the mutation frequency AF of all sites.
[0016] The AFNonTargetMean refers to the average frequency of the mutation frequency AF of all non-target sites, using the mutation frequency AF of all non-target sites as the background signal.
[0017] The SD is the standard deviation of the mutation frequency AF of all non-target sites;
[0018] 4.2 Retaining the positive candidate sites in the samples after the initial filtering, further filtering is performed using two parameters, AVscore and RatioScore:
[0019] The AVscore = 2**((AFTargetCutNeg-0.001)*100);
[0020] The RatioScore=2**((AFTargetCutNeg / (AFNonTargetMax+0.00001)-1.5)*10); TotalScore=AVscore*RatioScore;
[0021] Wherein, AFNonTargetMax refers to the maximum value among all mutation frequencies AF of non-target sites, * indicates multiplication, and ** indicates exponential operation;
[0022] 4.3 Calculate the TotalScore of the standard sample:
[0023] Provide positive and negative standards, calculate the TotalScore of the positive standard and the TotalScore of the negative standard respectively according to the calculation method in 4.2, and determine the threshold based on the range of the TotalScore of the positive standard and the TotalScore of the negative standard;
[0024] 4.4. Based on different sample types, calculate the TotalScore of each sample to determine the positive or negative status of the target site:
[0025] For cfDNA samples: TotalScore>2.5, then the target site of the sample is positive;
[0026] For FFPE samples: TotalScore>10, then the target site of the sample is positive.
[0027] Preferably, the raw sequencing data is preprocessed before the alignment in step 3.
[0028] Preferably, the preprocessing of the raw sequencing data involves using FastP software to initially filter the obtained raw sequencing data, removing reads of low quality and shorter than 50 bp, and then comparing the remaining reads with the human reference genome.
[0029] Preferably, the processing described in step 3 includes removing reads caused by PCR duplications in the bam file using the default parameters of the picard software.
[0030] Preferably, in section 4.3, the positive standard is a positive standard whose positive mutation abundance is confirmed to be not less than 0.1% by droplet digital PCR.
[0031] Preferably, the sample type to be tested includes any one of cfDNA sample, FFPE sample, fresh tissue sample, and blood cell sample.
[0032] Preferably, the sampling method for the sample to be tested includes biopsy or puncture.
[0033] The beneficial effects of this invention are:
[0034] (1) When using primers to amplify the target region, mutations may be introduced. That is, the mutation frequency of sites that are not mutated in the standard can sometimes reach 0.1%, thus generating high background noise. This invention can eliminate the defect of high background signal, improve the detection sensitivity, and reduce false positive detection results.
[0035] (2) It can be used to detect various ultra-low frequency mutations in nucleotides. In practical applications, the samples used can be formalin-fixed paraffin-embedded tissue (FFPE) samples, cell lines, and tissue fluid. Especially for mutation detection of single cfDNA samples, it improves the resolution of false positives, thereby ensuring the ability to detect ultra-low frequency mutations in cfDNA and avoiding the defects of existing cfDNA detection methods.
[0036] (3) When the target site is detected using the method of the present invention, the target site detection results are 100% consistent with the standard, that is, the positive site detection rate is 100% and the negative site detection rate is 0%, which greatly reduces the sequencing and analysis costs and has high application value. Attached Figure Description
[0037] Figure 1 This is a flowchart of the method for improving target site detection sensitivity by eliminating background in NGS data processing according to the present invention.
[0038] Figure 2 ROC curve for cfDNA sample.
[0039] Figure 3 The ROC curve for the FFPE sample. Detailed Implementation
[0040] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments. Experimental operations, reagents, and instruments not mentioned in this invention are all commonly used in the art.
[0041] Terminology Explanation:
[0042] NGS: Next Generation Sequencing.
[0043] Bwa: A comparison method software used to find the position of reads in Refseq and finally obtain a BAM format file.
[0044] Samtools: A tool for processing BAM files.
[0045] Igvtools: A tool for visualizing sequencing depth distribution, which can display the distribution of various data such as sequence alignment, copy number variation, and mutation sites.
[0046] Picard is a set of command-line tools for processing high-throughput sequencing (HTS) data and formats such as SAM / BAM / CRAM and VCF, which are defined in the Hts-specs repository.
[0047] Fastp is a data quality control and filtering software that can filter low-quality sequences and sequences with a large number of N (empty base).
[0048] Sequencing read: Each DNA sequence read during high-throughput sequencing is called a read.
[0049] Allele frequency (AF): The ratio of mutated reads to all reads.
[0050] AF Target: Mutation frequency of the target site in the sample.
[0051] Droplet digital PCR (ddPCR) is a third-generation PCR (Polymerase Chain Reaction) technique, a method for absolute quantification of nucleic acid molecules. The principle of ddPCR is to microdropletize the sample before PCR amplification, dividing the reaction system containing nucleic acid molecules into thousands of nanoliter-sized droplets. Each droplet may contain no target nucleic acid molecule or contain one to several target nucleic acid molecules. After PCR amplification, each droplet is detected individually. Droplets with fluorescent signals are interpreted as 1, and droplets without fluorescent signals are interpreted as 0. Based on the Poisson distribution principle and the number and proportion of positive droplets, the initial copy number or concentration of the target molecule can be determined.
[0052] The ROC curve, or Receiver Operating Characteristic curve, is a curve plotted using a series of different binary classification methods (cutoff values or decision thresholds), with the true positive rate (sensitivity) on the ordinate and the false positive rate on the x-axis. The AUC (Area Under Curve) is defined as the area under the ROC curve. The AUC value is often used as an evaluation criterion for the model. The AUC value ranges between 0.5 and 1. The closer the AUC is to 1, the higher the realism of the detection method; when it equals 0.5, the realism is the lowest, and it has no application value.
[0053] Circulating cell-free DNA (cfDNA), also known as cell-free DNA, refers to DNA circulating outside of cells in the bloodstream. Under normal physiological conditions, cfDNA mainly originates from the degradation of genomic DNA in senescent and apoptotic cells. However, when the body experiences disease, such as malignant tumors, trauma, organ transplant rejection, organ failure, or serious infectious diseases, abnormally necrotic cells release large amounts of DNA into the bloodstream. The amount of cfDNA in the blood is trace. With continuous accumulation and support from clinical research, it has been widely used in liquid biopsy, non-invasive prenatal screening, medication guidance, and the diagnosis of infectious diseases. However, when using next-generation sequencing technology to detect cfDNA in single samples, the inherent sequencing error of 0.1% to 1% in high-throughput sequencing instruments makes it particularly important to employ a method that can improve the sensitivity of target site detection to address the low-frequency mutation detection in single cfDNA samples.
[0054] Because amplification errors occur with a certain probability during amplicon sequencing library construction, negative control samples must be added for each library construction and sequencing run. The frequency of sites corresponding to the negative control samples needs to be subtracted during data analysis to reduce false positives. In this invention, the negative control samples are fragmented DNA extracted from 293T cells purchased by the Chinese Academy of Sciences, and were verified by droplet digital PCR to be negative at the target sites.
[0055] The sample type to be tested in this invention can be any one of cfDNA samples, FFPE samples, fresh tissue samples, and blood cell samples, and the sampling method can be biopsy or puncture. Taking cfDNA samples as an example, this invention provides a detailed explanation of a method for improving the sensitivity of target site detection by eliminating background in NGS data processing.
[0056] Example
[0057] according to Figure 1 The flowchart of a method for improving target site detection sensitivity by eliminating background in NGS data processing is described in detail.
[0058] 1. Sample Acquisition:
[0059] 1.1: Select a single cfDNA sample and extract DNA from it. The sample used is plasma cfDNA.
[0060] 1.2: 10 ml of whole blood was collected from a colorectal cancer patient and centrifuged to obtain approximately 4 ml of plasma. cfDNA was obtained using a nucleic acid extraction kit (Changzhou Tongshu Biotechnology Co., Ltd., catalog number: 150170). A library was constructed from the cfDNA using a human KRAS / NRAS / BRAF gene mutation detection kit (Changzhou Tongshu Biotechnology Co., Ltd., catalog number: 330140), and the library was sequenced using an IonProton sequencer.
[0061] 2. Gene sequencing data preprocessing:
[0062] like Figure 1 In step 102, the tumor-related genes are sequenced using NGS sequencing to obtain the raw sequencing data of the target gene of the sample to be tested. The raw sequencing data is preprocessed by using the default parameters of the FastP software to filter out low-quality reads and reads shorter than 50 bp. The remaining reads are then compared with the human reference genome using the default parameters of the BWA software to obtain a BAM file containing sequence alignment information.
[0063] 3. Data comparison:
[0064] The raw sequencing data was compared with the human reference genome. Based on the selected mutation site information, and according to the chromosome number, chromosome start point location, and chromosome end point location, the BAM file data for each mutation site within its corresponding region was extracted. Figure 1 In step 104, the bam file is processed using samtools mpileup to obtain an mpileup file, and the reads caused by PCR duplication in the bam file are removed using the default parameters of the picard software.
[0065] 4. Mutation analysis:
[0066] The igvtools algorithm is used to count the four bases at the target site. Bases that differ from those in the human reference genome are considered mutations at that target site. The mutation frequency at the target site is calculated, representing the number of ACTG reads, denoted as Na, Nc, Nt, and Ng, respectively. If the base in the human reference genome is A, then mutations at this position are denoted as A->C, A->T, and A->G, with corresponding mutation frequencies AF being Nc / (Na+Nc+Nt+Ng), Nt / (Na+Nc+Nt+Ng), and Ng / (Na+Nc+Nt+Ng). For example, in this embodiment, the KRAS gene mutation G12D is detected at chr12:25398284. The reference base is C. If the detected read is T at the corresponding position, the mutation changes C to T, and the corresponding amino acid change is G to D. Let Nt represent the number of reads for mutation T, and let N be the total number of reads. The mutation frequency is Nt / N, denoted as AF Target.
[0067] 5. Determine whether the mutation site meets predetermined conditions to determine whether the mutation site is a positive site, wherein the method for determining whether the mutation site meets predetermined conditions specifically includes:
[0068] 5.1, such as Figure 1 Step 106: Perform preliminary filtering based on the following conditions:
[0069] a.AFTargetCutNeg>AFNonTargetMean+2SD;
[0070] b. AFTargetCutNeg > 0.1%;
[0071] Positive candidate sites in the sample to be tested must be obtained by simultaneously meeting both of the above conditions;
[0072] The AFTargetCutNeg refers to selecting a single point as the target site of the sample to be tested, calculating the mutation frequency AF of all sites in the sample to be tested, and subtracting the mutation frequency of the corresponding site in the negative sample from the mutation frequency AF of all sites.
[0073] The AFNonTargetMean refers to the average frequency of the mutation frequency AF of all non-target sites, using the mutation frequency AF of all non-target sites as the background signal.
[0074] The SD represents the standard deviation of the mutation frequency AF of all non-target sites.
[0075] 5.2, such as Figure 1In step 108, retaining the positive candidate sites in the filtered test samples, two more parameters, AVscore and RatioScore, are introduced for further filtering:
[0076] AVscore=2**((AFTargetCutNeg-0.001)*100);
[0077] RatioScore=2**((AFTargetCutNeg / (AFNonTargetMax+0.00001)-1.5)*10);
[0078] TotalScore=AVscore*RatioScore;
[0079] Wherein, AFNonTargetMax refers to the maximum value among all mutation frequencies AF of non-target sites. * indicates multiplication, and ** indicates exponentiation.
[0080] 5.3 Calculate the TotalScore of the standard sample:
[0081] Positive and negative standards were prepared by mixing fragmented DNA positive at the corresponding mutation sites with negative standards that were negative at all mutation sites in a specific ratio. The positive mutation abundance was confirmed to be no less than 0.1% by ddPCR. After sequencing, raw data were obtained, and the TotalScore was calculated according to the above procedure. A specific threshold was determined based on ROC curve analysis.
[0082] Based on the positive and negative standards determined by ddPCR, a TotalScore > 4 was initially established as the threshold for determining the positivity or positivity of the locus. Further, 150 clinical cfDNA samples with locus positivity confirmed by other methods and 100 FFPE samples were selected. Negative samples were labeled 0, and positive samples were labeled 1. ROC curves were plotted for each. Figure 2 , Figure 3 As shown, the AUC of the ROC curve for the cfDNA clinical sample was 0.955, and the AUC of the ROC curve for the FFPE sample was 0.967.
[0083] 5.4. Based on different sample types, calculate the cuto-ff value of the TotalScore for each sample to determine the positive or negative status of the target site.
[0084] like Figure 1 In step 110, based on the ROC results in section 5.3, the thresholds for cfDNA and FFPE samples were determined:
[0085] For cfDNA samples: TotalScore>2.5, then the target site of the sample is positive;
[0086] For FFPE samples: TotalScore>10, then the target site of the sample is positive.
[0087] The specific calculation process for TotalScore is as follows:
[0088] As shown in Table 1 below, taking chr7:140453136 as an example, the chr7:140453136 site is the target site for detection, with a corresponding frequency of 0.0022. The mutation frequency of the corresponding site in the negative sample is 0.0001, so the corresponding AFTargetCutNeg is 0.0021 (0.0022-0.0001). The average value of the non-target region in the sample is 0.00045, SD = 8.85e-8. The AFTargetCutNeg of this site satisfies the filtering conditions a and b in 5.1. Further calculation shows AVscore = 1.07 and RatioScore = 5.66, so TotalScore = AVscore * RatioScore = 6.05. The sample type is cfDNA, satisfying the condition TotalScore > 2.5. Therefore, the target site of this sample is judged to be positive. Other test sample sites are judged using the same method. The detected sites are compared with the standard, and the positive site detection rate is 100%.
[0089] Table 1 shows the mutation frequency of the target sites detected and the corresponding mutation frequency in negative samples.
[0090]
[0091]
[0092]
[0093] In summary, this invention calculates the mutation frequency of all mutation sites, subtracts the mutation frequency of the negative control sample from the mutation frequency of the target site, and uses the mutation frequency of non-target sites as a background signal. A predetermined condition is set, and the background signal is used for filtering to retain the filtered target sites. Furthermore, the mutation frequency of the target site itself and its difference from the non-target site with the highest mutation frequency are considered. These two factors are quantified into a TotalScore. Different thresholds are defined according to different sample types to determine the positive or negative status of the site. This method is applicable to mutation detection in cfDNA and FFPE samples, eliminating the high background signal deficiency of amplicon sequencing in NGS sequencing, improving detection sensitivity, and significantly reducing false positive results.
[0094] Although the present invention has been described in detail through the preferred embodiments above, it should be understood that the above description should not be considered as a limitation of the present invention. Various modifications and substitutions to the present invention will be apparent to those skilled in the art after reading the above description. Therefore, the scope of protection of the present invention should be defined by the appended claims.
Claims
1. A method for improving the sensitivity of target site detection by eliminating background in NGS data processing for non-diagnostic purposes, characterized in that, Includes the following steps: Step 1: Amplicon technology is used to amplify tumor-related genes and regions in the DNA of the sample to be tested; Step 2: Provide a negative sample as a control, and use NGS sequencing to sequence the tumor-related genes to obtain the raw sequencing data of the target gene of the sample to be tested; Step 3: Align the raw sequencing data with the human reference genome to obtain a BAM file containing sequence alignment information. Use samtools mpileup to process the BAM file to obtain an mpileup file. When a site appears in the mpileup file whose bases are different from those in the human reference genome, the site is identified as a mutation site, and the mutation frequency AF of the mutation site is calculated. The processing includes removing reads caused by PCR duplications in the BAM file using the default parameters of picard software. Step 4: Determine whether the mutation site meets the predetermined conditions to determine whether the mutation site is a positive site; The method for determining whether the mutation site meets the predetermined conditions specifically includes: 4.1 Perform preliminary filtering based on the following conditions: a. AFTargetCutNeg > AFNonTargetMean+2SD; b. AFTargetCutNeg > 0.1%; Positive candidate sites in the sample to be tested must be obtained by simultaneously meeting both of the above conditions; The AFTargetCutNeg method selects a single point as the target site of the sample to be tested, calculates the mutation frequency AF of the target site of the sample to be tested, and obtains the result by subtracting the mutation frequency of the corresponding site in the negative sample from the mutation frequency AF of the target site of the sample to be tested. The AFNonTargetMean refers to the average frequency of the mutation frequency AF of all non-target sites, using the mutation frequency AF of all non-target sites as the background signal. The SD is the standard deviation of the mutation frequency AF of all non-target sites; 4.2 Retaining the positive candidate sites in the samples after the initial filtering, two parameters, AVscore and RatioScore, are introduced for further filtering: The AVscore is calculated as 2**((AFTargetCutNeg-0.001)*100). The RatioScore = 2**((AFTargetCutNeg / (AFNonTargetMax+0.00001)-1.5)*10); TotalScore=AVscore * RatioScore; Wherein, AFNonTargetMax refers to the maximum value among all mutation frequencies AF of non-target sites, * indicates multiplication, and ** indicates exponential operation; 4.3 Calculate the TotalScore of the standard sample: Provide positive and negative standards, calculate the TotalScore of the positive standard and the TotalScore of the negative standard respectively according to the calculation method in 4.2, and determine the threshold based on the range of the TotalScore of the positive standard and the TotalScore of the negative standard; 4.
4. Based on different sample types, calculate the TotalScore of each sample to determine the positive or negative status of the target site: For cfDNA samples: TotalScore > 2.5, the target site of the sample is positive; For FFPE samples: TotalScore > 10, then the target site of the sample is positive.
2. The method as described in claim 1, characterized in that, The process before the alignment in step 3 includes preprocessing of the raw sequencing data.
3. The method as described in claim 2, characterized in that, The preprocessing of the raw sequencing data involves using FastP software to initially filter the obtained raw sequencing data, removing low-quality reads and reads shorter than 50 bp, and then comparing the remaining reads with the human reference genome.
4. The method as described in claim 1, characterized in that, In section 4.3, the positive standard is a positive standard whose positive mutation abundance is confirmed to be not less than 0.1% by droplet digital PCR.
5. The method as described in claim 1, characterized in that, The sample type to be tested includes any one of cfDNA sample, FFPE sample, fresh tissue sample, and blood cell sample.
6. The method as described in claim 1, characterized in that, The sampling methods for the samples to be tested include biopsy or puncture.
Citation Information
Patent Citations
Point mutation detection method and device based on amplicon next-generation sequencing
CN106282356A