High-throughput genome sequencing variation detection system and method based on low input starting amount
By using software such as Mutect2, Vcftools, Manta and Strelka2 in a high-throughput genome sequencing variant detection system, combined with weighted scoring to filter mutations, the problem of low-frequency mutation detection at low investment starting volume is solved, and the accuracy and reliability of the detection are improved.
Patent Information
- Application Number
- CN202411204402.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-08-29
AI Technical Summary
Under the conditions of low input starting volume, high-throughput genome sequencing variant detection is difficult to effectively identify low-frequency mutations, and false negative and false positive detection are prone to occur.
A high-throughput genome sequencing variant detection system based on low input starting amount is adopted, which includes a data preprocessing module, a variation detection module, a variation integration filtering module and a detection output module. Mutect2 and Vcftools were identified, and structural mutations were detected by Manta and Strelka2, and the mutations were filtered by weighted scoring to improve the low-frequency detection capability.
It effectively improves the detection ability of low-frequency mutations, reduces the rate of false negative and false positive detection, and ensures that mutation detection with low detection limits is performed under the conditions of low input DNA starting volume.
Smart Images

Figure CN119152934B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of genome variation detection, and specifically relates to a high-throughput genome sequencing variation detection system and method based on low input starting amount. Background Art
[0002] In the early stages of production and development, application scenarios, technical details and costs are all key factors that need to be considered. Second-generation technology is usually used to perform low-input starting libraries and bioinformatics software to detect low-frequency mutations.
[0003] High-throughput genome sequencing technology (also known as second-generation sequencing technology) has made significant progress in the past decade, and has also made major breakthroughs in single nucleotide variation detection:
[0004] 1. More comprehensive variation: High-throughput genome sequencing technology can sequence the entire genome or DNA in a specific region, providing comprehensive single nucleotide variation information (including single nucleotide mutations, insertion / deletion variations, copy number variations, etc.)
[0005] 2. Lower cost: Current sequencers can sequence at a higher throughput, with data at the terabyte level.
[0006] 3. Data analysis support: The amount of data generated by high-throughput genome sequencing is huge, which requires strong data analysis and interpretation capabilities. Currently, a large number of bioinformatics tools and databases have been developed for variant annotation, function prediction and function interpretation of high-throughput sequencing data, providing more support and guidance for scientists.
[0007] The second-generation sequencing technology was first applied in the field of scientific research. Its sequencing depth is often not high, and the sequencing raw materials are relatively easy to obtain. The starting amount of DNA input is generally above 100ng. However, there are additional challenges for tumor samples in the field of clinical special examinations: First, tumor raw materials are not easy to obtain; second, the extremely high sequencing depth in clinical special examinations often leads to high DUP. This is because the frequency of somatic mutations in tumors has requirements for the detection limit (generally below 0.5%), and the higher the sequencing depth, the easier it is to obtain mutated reads. In order to achieve the goal of extremely high sequencing depth, we will use PCR to repeatedly amplify the library in order to amplify the signal, even if redundant reads are introduced in the process. Generally, the more rounds of PCR, the more redundant reads are introduced. The more starting amount of sequencing library is invested, the fewer PCR rounds are required.
[0008] Generally speaking, for higher mutation frequencies (≥20%), sequencing depth ≥200X is sufficient to identify 95% of mutations; for lower mutation frequencies (≤10%), the system should be improved rather than simply increasing the sequencing depth. This also shows that for high-throughput detection of low-input starting sequencing libraries, it is still difficult and necessary to ensure the accuracy of detection. Summary of the invention
[0009] The purpose of the present invention is to propose a high-throughput genome sequencing variation detection system and method based on low input starting amount, so as to effectively improve the low-frequency detection capability while avoiding false negative / positive caused by experimental or sequencing influences.
[0010] In view of this, the scheme of the present invention is as follows:
[0011] The first aspect of the present invention provides a high-throughput genome sequencing variation detection system based on low input starting amount, comprising:
[0012] Data preprocessing module, used to align sequencing data to reference genes to obtain alignment results;
[0013] A variation detection module, comprising a first detection module and a second detection module, for obtaining variation data based on the comparison results respectively; the first detection module is used to obtain first variation data including mutation sites, and the second detection module is used to detect structural variation and identify second variation data including single nucleotide mutations and insertion and deletion results;
[0014] The variation integration and filtering module is used to integrate and filter the variation data; in the integration process, the union of the first variation data and the second variation data is taken to obtain the third variation data, and each variation in the third variation data is weighted and scored according to the contribution of the variation information parameter to the variation credibility; the filtering process includes filtering the variation with low score;
[0015] The detection output module is used to output the final filtered variant data.
[0016] Furthermore, the first detection module uses Mutect2 and Vcftools to identify and process mutation sites; and / or, the second detection module uses Manta and Strelka2 to detect and identify single nucleotide mutations and insertions and deletions.
[0017] Furthermore, the variation information parameters are divided into supporting parameters and opposing parameters according to their contribution to the variation, and are respectively divided into positive and negative; the supporting parameters include mutation depth, mutation frequency, and number of reads; the opposing parameters include sequencing depth, variation source, alignment quality, unique alignment, and regional complexity; the variation source refers to the situation where the variation data is derived from the first and second variation data.
[0018] Preferably, the weighted scoring also includes additional overall scoring for the situation where the variation comes from three sets of variation data, alignment quality, unique alignment, and regional complexity penalty.
[0019] Furthermore, the sequencing depth of the sequencing data is above 200X, and the mutation frequency is below 10%; it is particularly suitable for a sequencing depth greater than 10000X, and a mutation frequency below 1%.
[0020] Furthermore, the detection system is suitable for sequencing when the starting amount of DNA input is less than 100 ng, and more preferably the starting amount is less than 50 ng.
[0021] Furthermore, the sequencing data is free of low-quality data before alignment, including but not limited to removing uncleaned adapter sequences, removing continuous low-quality base sequences, discarding low-quality sequences, and discarding sequences that are too short.
[0022] Furthermore, the filtering process also includes: using a local frequency library for filtering; and / or filtering mutations that have no actual impact; and / or checking bam for filtering based on the authenticity of the mutation.
[0023] The second aspect of the present invention provides a method for detecting variation in high-throughput genome sequencing based on low input starting amount for non-diagnostic purposes, comprising:
[0024] Align the sequencing data to the reference gene to obtain the alignment results;
[0025] Obtaining variation data based on the comparison results, including obtaining first variation data including the mutation site, and identifying second variation data including single nucleotide mutations and insertion / deletion results by detecting structural variation;
[0026] Integrate and filter the variant data; in the integration process, take the union of the first variant data and the second variant data to obtain the third variant data, and weight and score each variant in the third variant data according to the contribution of the variant information parameter to the variant credibility; the filtering process includes filtering variants with low scores;
[0027] Output the final filtered variant data.
[0028] According to a third aspect of the present invention, an electronic device is provided, comprising a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, the variation detection method as described in the second aspect is implemented.
[0029] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, and a processor implements the variation detection method as described in the second aspect when executing the computer program.
[0030] Compared with the prior art, the beneficial effects of the present invention include but are not limited to:
[0031] The variation detection system of the present invention is based on identifying mutation sites and detecting structural variations to ensure that low-frequency variations are detected as much as possible, thereby reducing the false negative detection rate; by merging two sets of variation data, for each variation, a weighted score is given to the credible contribution of the variation based on the variation information parameters to screen credible variations, thereby avoiding false positive detections. Overall, it can effectively improve the low-frequency detection capability while avoiding false negatives / positives caused by experiments or sequencing.
[0032] The variation detection system of the present invention can improve the detection precision and recall rate as much as possible when performing molecular biological diagnosis, and can ensure mutation detection with a low detection limit under the condition of low input DNA starting amount. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0034] Figure 1 This is an overall flow chart of the high-throughput genome sequencing variation detection system based on low input starting amount described in the present invention. DETAILED DESCRIPTION
[0035] The following provides definitions of some terms used in this specification. Unless otherwise specified, all terms used herein are generally understood by those skilled in the art to which this solution belongs.
[0036] Terminology explanation:
[0037] Low input starting amount library: The starting DNA input amount of a general library is above 100 ng. A DNA input amount of around 50 ng is called a low input starting amount library.
[0038] Sequencing depth: also called DP; in VCF files, it is also expressed as DP, representing the "total sequence coverage depth" of the position, that is, how many times the site has been sequenced, unit: multiplication (X).
[0039] Mutation frequency: also known as AF; at a certain site in the genome, the coverage depth supporting a certain "mutation sequence" accounts for the ratio of the "total sequence coverage depth" of this site. Taking the fields in the VCF file as an example, DP represents the "total sequence coverage depth" of the position; AD represents the "mutation sequence coverage depth"; the calculation of AF is equal to = AD (mutation sequence coverage depth) / DP (total sequence coverage depth).
[0040] VCF: VCF (Variant Call Format) is a common DNA sequence variation recording format, commonly used in genomics research and genetic variation analysis. It is a text format, mainly used to store and describe SNPs (single nucleotide polymorphisms), Indels (insertions or deletions), and other types of DNA sequence variation information in single or multiple samples. The VCF format usually consists of the following parts: metadata: including file format version, sample information, reference genome and other information; header: composed of a series of comment lines starting with "#", describing the meaning of each column of the VCF file; variant information: arranged in columns, including chromosome position, reference sequence, variant sequence, quality score, filter status, annotation information, etc. The advantage of the VCF format is that it can record the DNA sequence variation information of multiple samples at the same time, and supports the comparison and statistical analysis of variant sites between different samples. In addition, the VCF format also provides a standard file format for genomics research, which is convenient for data sharing and processing. The VCF format is an important DNA sequence variation recording format, which is widely used in genomics research, genetic variation analysis and bioinformatics.
[0041] False positive call: A false positive call is a situation where a variant or mutation is incorrectly flagged as present when it is not. This may be the result of experimental error, data processing errors, sample degradation, or other factors.
[0042] False-negative call: A false-negative call is a failure to detect or missed occurrence of a variant or mutation that actually exists.
[0043] Mutation rating: According to the interpretation criteria for tumor gene mutations jointly issued by the Association for Molecular Pathology (AMP), the American Society of Clinical Oncology (ASCO) and the College of American Pathologists (CAP), the mutation levels are divided into the following three categories: Level 1: refers to mutations with clear clinical significance; Level 2: refers to mutations with potential clinical significance; Level 3: refers to mutations with unknown clinical significance.
[0044] BAM: It is the most commonly used comparison data storage format in genetic data analysis. It is a binary file format used to store large-scale sequencing data, especially for storing the comparison results between sequences and reference genomes.
[0045] In order to ensure the accuracy of high-throughput gene variation detection in low-input starting library (about 50 ng), a method and system for variation detection in low-input starting library based on high-throughput genome sequencing data were proposed, which can effectively improve the low-frequency detection capability while avoiding false negative / positive results caused by experimental or sequencing effects.
[0046] The inventors of the present invention have found that for variants with a sequencing depth of more than 200X and a mutation frequency of less than 10%, simply using mutect2 cannot guarantee the detection effect, and there are a large number of false positive detections. False positive detections may mislead diagnosis and treatment plans, so when analyzing and interpreting NGS test results, it is necessary to pay attention to excluding the impact of false positives; when the sequencing depth is greater than 10000X and the mutation frequency is less than 1%, not only are there a large number of false positive detections, but false negative detections also become more common. False negative detections may cause patients to miss key diagnostic information or treatment opportunities, so in NGS testing, it is necessary to pay attention to reducing the false negative detection rate to ensure the accuracy and reliability of the results.
[0047] In order to systematically solve this problem, we need to first ensure that low-frequency variants are detected as much as possible, thereby reducing the false negative detection rate. To achieve this goal, various variant detection software (Mutect2 software) need to be adjusted to the most sensitive possible. However, after preliminary testing, we tried to adjust more than a dozen combinations of parameters, but all the variants in the system could not be detected. After investigation, we speculated that this was because the default statistical test of the Mutect2 software could not accurately distinguish whether a low-quality variant was a true low-quality variant or a false low-quality variant. Under more conservative considerations, the software would lose some ultra-low-frequency variants (mutation frequency below one thousandth). Therefore, we introduced a second software, strelka2, into the bioinformatics detection system. This software is an open source software developed by Illumina. The software is characterized by an average running speed of about 17-22 times that of the Mutect2 software. This can make low-frequency variants easier to detect by combining the results of the software Manta. It is also possible to train features by training true (false) negative data sets to ensure that mutations in the reaction system are detected as much as possible. In addition, it should be noted that the current bioinformatics software used to detect mutations is not only Mutect2 and Strelka2, but also includes Varscan, Vardict, DeepVariant and other well-known somatic mutation detection software. However, based on the published literature and industry experience, other software is far less efficient (including accuracy and time) than Mutect2 and Strelka2. In order to reduce the redundancy of the system as much as possible, we only use Mutect2 and Strelka2 for mutation detection.
[0048] By taking the intersection and union of the detected mutations of the two mutation software according to the site (chromosome position), the mutations can be divided into three sets: Mutect2, Strelka2 and common mutations. The mutations are scored according to the corresponding scoring equations according to the different sets, and this score will be used for mutation filtering later.
[0049] After roughly solving the false negative detection phenomenon (it is difficult to eliminate in theory, and we can only iterate and tune as much as possible), the subsequent analysis will focus on solving the false positive detection process. Use the mutation scoring, open source database, and self-built database mentioned above to perform multiple filtering and annotations on the variants and add labels. Finally, output the results in a unified header style and upload them to the report system for long-term maintenance and iterative upgrades of the self-built database and Strelka2's positive / negative detection training set.
[0050] The process of gene variation detection based on low input starting amount and high throughput can be summarized into three modules: preprocessing module, variation result integration module, variation (mutation) filtering module and detection module, which are used for genome sequencing data preprocessing, variation result integration, variation (mutation) filtering and outputting filtered reliable variation results respectively. The flowchart is as follows Figure 1 The details are as follows:
[0051] 1. Data preprocessing
[0052] (1) Genome sequencing data preprocessing
[0053] Use bcl2fastq software to provide the index sequence information of each sample before sequencing and split the original data into fastq format data.
[0054] (2) Raw data processing
[0055] The original genome Fastq data contains some low-quality data, which will affect subsequent analysis, so low-quality data needs to be removed here. Use the software fastp to process the data quality, remove the uncleaned adapter sequence, remove the continuous low-quality base sequence, discard the low-quality sequence, and discard the sequence that is too short.
[0056] (3) Alignment of fastq sequences with reference genome
[0057] According to the consistency of the base sequence with the human reference genome hg19, the sequence is mapped to the reference genome and a bam file of the alignment result is generated.
[0058] 2. Integrate the mutation results
[0059] (4) Detect somatic mutations using Mutect2+Vcftools+VEP
[0060] The process of using Mutect2 combined with Vcftools and VEP for somatic mutation detection is a classic process. First, Mutect2 is used to compare sample and reference genome data to identify potential mutation sites. Then, Vcftools is used to process and filter VCF files to improve the reliability of variants. Finally, Variant Effect Predictor (VEP) is used to functionally annotate and interpret mutations to help determine which mutations may have biological significance.
[0061] (5) Detecting somatic mutations using Manta+Strelka2+VEP
[0062] Manta is used to detect structural variations, such as insertions, deletions, and inversions. Then, Strelka2 is used to identify single nucleotide mutations and small insertions and deletions. Subsequently, VEP is used to functionally annotate and interpret these mutations to help determine their potential impact. The integration of these three tools can comprehensively capture the information of somatic mutations, thereby effectively analyzing and understanding the variations in the genome.
[0063] 3. Filter mutations
[0064] (6) Integrate and score variants
[0065] ① Read the mutation files from two sources (both in VCF format) separately to extract the "basic information of mutation" (including: chromosome position, bases before mutation, and bases after mutation). In this way, two sets with "basic information of mutation" as elements are obtained; then calculate the intersection and union of the two sets, and finally mark the sources of all mutations in the two software in turn.
[0066] Labeling method: All variants mentioned in the two software are labeled as three sources: if the variant comes only from Mutect2, it is labeled as A; if the variant comes only from Strelka2, it is labeled as B; if the variant is detected by both software, it is labeled as C.
[0067]
[0068]
[0069] ② Collect evidence items according to BAM
[0070] First, based on the "basic information of the variation" obtained in the previous step, use Python's pandas module to obtain the reads record (i.e., alignment) corresponding to the variation in BAM (Note: all required evidence items should be obtained from the BAM file); then, parse each read record in turn, collect the evidence items required to assess the credibility of the variation, and give each evidence item a logical plus or minus score mechanism, with supportive evidence having a positive score and opposing evidence having a negative score; the higher the final score, the higher the credibility of the variation, and the variation is divided into reasonable levels.
[0071] Evidence items include supporting evidence and opposing evidence, as follows:
[0072] I. Supporting evidence:
[0073] Mutation frequency (AF): The fraction of alleles in a tumor sample that is mutated, indicating the relative abundance of the mutation. The higher the proportion of reads that detect the mutation, the higher the authenticity of the mutation. The score should be a positive number and must conform to a monotonically increasing function. The smaller the AF, the smaller the score, that is, the slope is proportional to the AF.
[0074] Logarithm of reads for detection: The more logarithms of reads for detection of mutation, the higher the authenticity of the mutation. The value should be a positive number and must conform to a monotonically increasing function. The higher the number, the smaller the bonus ratio, that is, the slope is inversely proportional to the logarithm of reads for detection.
[0075] Allele Depth (AD): The more reads that detect a variant, the more realistic the variant is. Therefore, the score for this item should be a positive number and should conform to a monotonically increasing function. When AD is small (AD<=8), the score should be smaller, that is, the slope is proportional to AD.
[0076] II. Opposing evidence:
[0077] Sequencing Depth (DP): The depth of sequencing coverage at this position. Higher coverage generally means more reliable variant detection.
[0078] Alignment quality (high_MQ_reads): The average MQ of the reads where the variant is located is not less than 50% of the average MQ of other reads, and the average MQ of the variant is > 10, then no points will be deducted for this item. Otherwise, points will be deducted;
[0079] Uniq: The variant is in a unique alignment region. If 80% of the reads do not have XA, this label is True;
[0080] Complexity of the region: The mutation is in a complex sequence region (whether the sequences around the mutation are complex sequences).
[0081] Source of variation: If the same variation is detected in two softwares at the same time, the higher the authenticity of the variation, the more likely this label is True; otherwise, a penalty is imposed.
[0082] ③ Score the evidence items mentioned in the above steps. The scoring rules are as follows:
[0083]
[0084] ④ After scoring the evidence items mentioned in the above steps in turn, determine whether the four items of high_MQ_reads, Uniq, complex, and variant sources are all unsatisfactory, and then perform additional scoring;
[0085] Evidence Item Extra points All four items are not satisfied Additional penalty points 30 Any three items are not satisfied Additional penalty points 20 Any two items do not satisfy Additional penalty 10 points No deduction for any of the four items Extra points 10
[0086] ⑤ Divide mutations into five levels based on the scores:
[0087] Below 0 points: Grade E, false variation;
[0088] 0-40 points: D level, possibly a false mutation;
[0089] 40-80 points: Grade C, variants of uncertain authenticity;
[0090] 80-120 points: Grade B, possibly a true mutation;
[0091] 120 points and above: Grade A, credible variation;
[0092] (7) Annotate local frequency according to local frequency library
[0093] Construction of local frequency library:
[0094] In the above steps, we describe the location of each mutation (including chromosome number and base location) for a certain sample, as well as the base before mutation (REF) and the base after mutation (ALT) at that location. For a large cohort of samples (such as 500 samples), we can count all mutations in all samples and calculate the number of samples where each mutation occurs (such as 250) in turn. Then the local frequency of the mutation in the local frequency library of this cohort is 250 / 500=0.5; that is, for a fixed local frequency library, one mutation corresponds to one local frequency.
[0095] (8) Filtering variations
[0096] ① Use local frequency library for filtering
[0097] The local frequency database records the frequency of local detection of a mutation in the total number of times the analysis is performed. If the local frequency detection is greater than 20% and is not in the common whitelist and blacklist, the mutation is considered to be a false mutation introduced by the reaction system and is marked as LocalDB.
[0098] ②Filter by score
[0099] Variants with mutation ratings of A and B are retained, variants with mutation ratings of C, D, and E are filtered out, and low-confidence variants are marked as LowQ.
[0100] ③Filter based on mutation function
[0101] Remove mutations that have no actual effect on the function; mark them as DontAffectFunc;
[0102] ④ Check bam to verify the authenticity of the mutation
[0103] 4. Output results
[0104] Filter out the unreliable variants labeled in step (8), report and grade the remaining reliable variants, and interpret the variants based on the literature reports of the variants. The interpretation content includes but is not limited to providing reference for disease diagnosis, prognosis, recurrence, and treatment.
[0105]
[0106] Example
[0107] After receiving DNA sequencing data from a pediatric tumor patient, perform the following operations:
[0108] (1) Genome sequencing data preprocessing
[0109] The original data statistics are as follows:
[0110] Sample Raw_reads Raw_bases Test1 146,875,740 22,178,236,740
[0111] (2) Raw data processing
[0112] After quality control, high-quality sequences were obtained, and the data statistics are as follows:
[0113] Samples Clean_reads Clean_bases Q20(%) Q30(%) clean Bases% Test1 144,485,354 20,945,984,384 98.17% 94.91% 94.44
[0114] (3) Fastq and reference genome alignment
[0115] The alignment of the sequence data with the human reference genome hg19 is as follows:
[0116]
[0117] (4) Detect somatic mutations using Mutect2+Vcftools+VEP
[0118] CHR POS REF ALT AD AF Number of reads detected DP 12 112888189 G A 103 0.373 176 276 16 50827573 C A 87 0.026 103 193 X 41201995 T C 32 0.1 221 321
[0119] (5) Detecting somatic mutations using Manta+Strelka2+VEP
[0120] CHR POS REF ALT AD AF Number of reads detected DP 10 871189 T C 56 0.172 226 326 12 112888189 G A 103 0.373 176 276 X 41201995 T C 32 0.1 221 321
[0121] (6) Integrate and score variants
[0122] CHR 10 12 16 X POS 871189 112888189 50827573 41201995 REF T G C T ALT C A A C AD 40 40 40 40 AF 50 50 50 50 Number of reads detected 45 45 45 45 DP 0 0 0 0 Sources of variation -30(B) 0(C) -30(A) 0(C) High MQ reads 0 0 0 -30 Uniq -30 0 0 -30 Complex 0 0 0 0 Extra points -10 0 0 -10 Total score 65 135 105 65 Rating C A B C
[0123] (7) Filter mutations
[0124] CHROM POS REF ALT Total score Rating Mutation Label 10 871189 T C 65 C LowQ 12 112888189 G A 135 A PASS 16 50827573 C A 105 B LocalDB X 41201995 T C 65 C LowQ
[0125] (8) Output results
[0126] CHROM POS REF ALT Total score Rating Mutation Label 12 112888189 G A 135 A PASS
[0127] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will fall within the protection scope of the present invention.
Claims
1. A high-throughput genome sequencing variation detection system based on low input starting amount, characterized in that: include: Data preprocessing module, used to align sequencing data to reference genes to obtain alignment results; A variation detection module, comprising a first detection module and a second detection module, for obtaining variation data based on the comparison results respectively; the first detection module is used to obtain first variation data including mutation sites, and the second detection module is used to detect structural variation and identify second variation data including single nucleotide mutations and insertion and deletion results; The variation integration and filtering module is used to integrate and filter the variation data; in the integration process, the union of the first variation data and the second variation data is taken to obtain the third variation data, and each variation in the third variation data is weighted and scored according to the contribution of the variation information parameter to the variation credibility; the filtering process includes filtering the variation with low score; The detection output module is used to output the final filtered variant data; The variation information parameters are divided into supporting parameters and opposing parameters according to their contribution to the variation, and are respectively divided into positive and negative; the supporting parameters include mutation depth, mutation frequency, and number of reads; the opposing parameters include sequencing depth, variation source, alignment quality, unique alignment, and regional complexity; the variation source refers to the situation where the variation data is derived from the first and second variation data.
2. The detection system according to claim 1, characterized in that: The first detection module uses Mutect2 and Vcftools to identify and process mutation sites; and / or, the second detection module uses Manta and Strelka2 to detect and identify single nucleotide mutations and insertions and deletions.
3. The detection system according to claim 1, characterized in that: The weighted scoring also includes additional overall scoring for each penalty situation of variation source, alignment quality, unique alignment, and regional complexity.
4. The detection system according to claim 1, characterized in that: The sequencing depth of the sequencing data is greater than 200X, and the mutation frequency is less than 10%; and / or, the starting amount of DNA input during sequencing is less than 100 ng.
5. The detection system according to claim 1, characterized in that: The sequencing data were cleaned of low-quality data before alignment.
6. The detection system according to claim 1, characterized in that: The filtering process also includes: filtering using a local frequency library; and / or filtering variants that have no actual impact; and / or filtering based on the authenticity of the variants by checking bam.
7. A method for detecting variation based on high-throughput genome sequencing with low input starting amount for non-diagnostic purposes, characterized in that: include: Align the sequencing data to the reference gene to obtain the alignment results; Obtaining variation data based on the comparison results, including obtaining first variation data including the mutation site, and identifying second variation data including single nucleotide mutations and insertion / deletion results by detecting structural variation; Integrate and filter the variant data; in the integration process, take the union of the first variant data and the second variant data to obtain the third variant data, and weight and score each variant in the third variant data according to the contribution of the variant information parameter to the variant credibility; the filtering process includes filtering variants with low scores; Output the final filtered variant data; The variation information parameters are divided into supporting parameters and opposing parameters according to their contribution to the variation, and are respectively divided into positive and negative; the supporting parameters include mutation depth, mutation frequency, and number of reads; the opposing parameters include sequencing depth, variation source, alignment quality, unique alignment, and regional complexity; the variation source refers to the situation where the variation data is derived from the first and second variation data.
8. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to claim 7 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the processor executes the computer program to implement the method according to claim 7.
Citation Information
Patent Citations
Identification method and system of embryonic line SNV and InDel variation and readable storage medium
CN117711487A