Sequencing data analysis method, device and equipment and computer readable storage medium
By generating customized reference sequences and using better mutation detection software and filtering methods, the problem of omission of false positive and low chimeric proportion mutations in the analysis of amplicon sequencing mutation data is solved, and more accurate mutation detection and genetic disease diagnosis are achieved.
Patent Information
- Application Number
- CN202510354216.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
Existing amplicon sequencing mutation data analysis methods are prone to introducing false positive mutations or missing mutations with low chimeric ratios, resulting in inaccurate analysis results.
Generate customized reference sequences, block regions outside the target genomic region, and use better mutation detection and filtering methods to perform mutation detection and filtering.
It improves the accuracy of mutation detection, reduces false positive mutations, can detect mutations with low chimeric ratios, and improves the diagnosis rate of genetic diseases.
Smart Images

Figure CN120299518A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of sequencing data analysis, and particularly to an analysis method, device, equipment and computer-readable storage medium for sequencing data. Background Art
[0002] Amplicon sequencing technology is a technology based on high-throughput sequencing. It uses specific primers to perform PCR amplification on target DNA, and then sequences the PCR products. Compared with whole-genome sequencing or whole-exome sequencing, amplicon sequencing technology can sequence only a part of the genome, thus saving costs. Due to the large-scale amplification of sequences, amplicon sequencing technology can provide a deeper understanding of the information of this part of the genome. In addition, the PCR reaction uses target-DNA-specific primers, which can avoid the amplification of non-specific DNA fragments and improve data quality. Amplicon sequencing can achieve ultra-high-depth coverage in the target region, detect low-proportion chimeric mutations, and obtain the mutation proportion of chimeric mutations. The detected mutations are more reliable than those detected by whole-exome and whole-genome sequencing. However, in the current methods for analyzing amplicon sequencing mutation data, the sequencing data is directly aligned with a reference sequence (such as hg19 (human genome version 19)), which is likely to introduce a large number of false-positive mutations or easily miss mutations with low chimeric proportions, thus imposing a burden on the analysis. Summary of the Invention
[0003] Based on the above problems, the present application provides an analysis method, device, equipment and computer-readable storage medium for sequencing data.
[0004] The present application adopts the following technical solutions:
[0005] The first aspect of the present application discloses a method for analyzing sequencing data, including: generating a customized reference sequence: obtaining a target genomic region, and masking the regions outside the target genomic region in the initial reference sequence to obtain the customized reference sequence; sequence alignment: aligning the sequencing data of the sample to be analyzed with the customized reference sequence to obtain an alignment result; mutation analysis: performing mutation detection and mutation filtering based on the alignment result to obtain the mutation result of the sample to be analyzed. It should be noted that in the current sequencing data analysis method, hg19 is directly selected as the reference sequence, while in the present application, by obtaining the target genomic region of interest and masking the genomic regions outside the target region, a new reference sequence (i.e., the customized reference sequence) is generated for subsequent alignment. This can help avoid the situation where genomic mutations cause the sequence to be multi-aligned to other regions, enabling the sequencing data to be aligned to the target region to the greatest extent, thereby capturing various mutation information, including single nucleotide mutations, DNA sequence insertions, and DNA sequence deletions, and obtaining the gene mutation result more precisely and accurately. The present application can also improve the data analysis efficiency.
[0006] In one implementation manner of the present application, the target genomic region includes any one or more of the following: the genomic region related to the disease of the sample to be analyzed, the regions of all genes in the OMIM database.
[0007] In one implementation manner of the present application, the sequencing data is the sequencing data of amplicon sequencing.
[0008] In one implementation manner of the step of generating the customized reference sequence in the present application, it further includes removing the multi-alignment regions of the target genomic region: masking the regions with non-unique alignment in the target genomic region.
[0009] In one implementation manner of the present application, before the alignment step, it further includes removing adapters and performing quality control on the sequencing data of the sample to be analyzed.
[0010] In one implementation manner of the present application, the quality control conditions meet the following two situations to pass the quality control: the base ratio of Q30 is greater than or equal to 80%, and the GC content is greater than or equal to 30% and less than or equal to 60%.
[0011] In one implementation manner of the present application, VarDict is used for mutation detection.
[0012] In one implementation of the present application, among the filtering conditions for mutation filtering, mutations that meet at least one of the following conditions are filtered: the mutation frequency is lower than 0.1%, the minimum base quality value of the mutation is lower than 20, the minimum alignment quality of the mutation is lower than 10, the unit depth quality value of the mutation is lower than 2, the positions of the mutation in the read segments where it appears are all the same, and the mutation belongs to the common SNPs in dbSNP.
[0013] The second aspect of the present application discloses an apparatus for analyzing sequencing data, including: a customized reference sequence generation module: used to obtain a target genomic region, and mask the regions outside the target genomic region in the initial reference sequence to obtain the customized reference sequence; a sequence alignment module: used to align the sequencing data of the sample to be analyzed with the customized reference sequence to obtain an alignment result; a mutation analysis module: used to perform mutation detection and mutation filtering according to the alignment result to obtain the mutation result of the sample to be analyzed.
[0014] The third aspect of the present application discloses a device, including a memory and a processor, where the memory is used to store a program, and the processor realizes the analysis method of sequencing data as described in the first aspect of the present application by executing the program stored in the memory.
[0015] The fourth aspect of the present application discloses a computer-readable storage medium, in which a program is stored, and the program can be executed by a processor to realize the analysis method of sequencing data as described in the first aspect of the present application.
[0016] The beneficial effects of the present application are as follows:
[0017] The present application can analyze and obtain gene mutation results more precisely and accurately. Description of the Drawings
[0018] Figure 1 is a schematic flowchart of the analysis method of sequencing data involved in the present application.
[0019] Figure 2 is a schematic diagram of the apparatus for analyzing sequencing data involved in the present application.
[0020] Figure 3 is a schematic diagram of the device for analyzing sequencing data involved in the present application.
[0021] Figure 4 is a Sanger sequencing result diagram of the sample involved in the embodiment of the present application. Detailed Embodiments
[0022] The present invention will be further described in detail below in conjunction with the accompanying drawings through specific embodiments. In the following embodiments, many detailed descriptions are provided to enable a better understanding of the present application. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other materials or methods. In some cases, some operations related to the present application are not shown or described in the specification, which is to avoid the core part of the present application being overwhelmed by excessive descriptions. For those skilled in the art, it is not necessary to describe these related operations in detail, and the related operations can be fully understood based on the description in the specification and the general technical knowledge in the art.
[0023] In addition, the features, operations or characteristics described in the specification can be combined in any appropriate manner to form various embodiments. At the same time, the steps or actions in the method description can also be reordered or adjusted in an obvious manner by those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for clearly describing a certain embodiment, and do not mean that they are the necessary sequences, unless it is stated that a certain sequence must be followed.
[0024] Amplicon sequencing has ultra-high depth coverage in the target region, can detect low-proportion chimeric mutations, and can also obtain the mutation proportion of chimeric mutations. The detected mutations are more reliable than whole exome and whole genome sequencing. Currently, the methods for analyzing amplicon sequencing mutation data directly align to the untreated reference sequence, use non-optimal mutation detection software and filtering methods, which are prone to missing low-chimeric proportion mutations and introducing a large number of false positive mutations, causing a burden on the analysis.
[0025] In view of this, the present application generates a customized reference sequence by removing the multi-alignment regions in the target genome, and uses a specifically better mutation detection software and filtering method to more accurately analyze the genomic mutation information of patients. In the present application, shielding the genomic regions outside the target region in the initial reference sequence helps to avoid the situation where genomic mutations cause the sequence to be multi-aligned to other regions, enabling the sequencing data to be aligned to the target region to the greatest extent, thereby capturing various mutation information (including single nucleotide mutations, insertions, deletions); and then using software more suitable for amplicon sequencing mutation detection to perform mutation detection and filter the mutations, so as to accurately obtain the mutation data. The advantages of the present application include: 1. It can detect low-chimeric proportion mutations and calculate the chimerism of point mutations. 2. Using a customized reference sequence to minimize the omission of positive mutations. 3. Using better mutation detection software and filtering parameters to reduce noise data. 4. Effectively reducing false positives while accurately detecting true mutations.
[0026] In summary, in the present application, a customized reference sequence is generated according to the target region, and the sequencing data of the sample to be analyzed is quality-controlled and filtered, so as to capture the mutation signal as much as possible and filter out false positives, thereby achieving accurate judgment of mutations. This method can be applied to the analysis of sequencing data in the case where the patient phenotype is known and the phenotype is related to a specific genomic region. For amplicon sequencing data, the analysis method of the present application is used for mutation analysis, and mutations with extremely low chimeric ratios can be detected without a control sample, thereby improving the diagnosis rate of genetic diseases and providing assistance for the analysis and treatment of genetic diseases.
[0027] The present application provides an analysis method, device, equipment and computer-readable storage medium for sequencing data.
[0028] As described above, the present application relates to an analysis method for sequencing data.
[0029] In some examples, the sequencing data may include the sequencing data of amplicon sequencing.
[0030] Figure 1 It is a schematic flowchart of the analysis method for sequencing data involved in the present application.
[0031] As Figure 1 shown, in some examples, the analysis method for sequencing data may include: step S100, generating a customized reference sequence.
[0032] In some examples, generating a customized reference sequence may include: obtaining a target genomic region, and masking the region outside the target genomic region in the initial reference sequence to obtain the customized reference sequence. It should be noted that the initial reference sequence is related to the species. For example, the initial reference sequence of humans can be hg19 (human genome version 19). Currently, when analyzing the sequencing data of human samples, hg19 is usually directly selected as the reference sequence for alignment. In the present application, by masking the region outside the target genomic region in hg19 as the new reference sequence, the alignment efficiency can be improved and the introduction of false positive mutations can be reduced.
[0033] In some examples, the target genomic region may refer to the genomic region of interest. For example, it may be the genomic region related to the disease of the sample to be analyzed, the region of all genes in the OMIM database. The present application is not limited to this. For example, the target genomic region may be the genomic region of a certain chromosome or chromosomal region. To improve the generality of the customized reference sequence, the target genomic region may be the region of all genes in the OMIM database. Thus, this customized reference sequence can be applied to the analysis of sequencing data related to most genetic diseases.
[0034] In some examples, generating a customized reference sequence may include: obtaining a target genomic region, removing multi-alignment regions, and masking regions outside the target genomic region to obtain a customized reference sequence.
[0035] In some examples, removing multi-alignment regions may include: masking regions with non-unique alignments in the target genomic region. Specifically, removing multi-alignment regions may include: selecting the sequence of the target region, aligning it to the reference genome (hg19), and masking regions with non-unique alignments in the target region. Thereby, it is beneficial to ensure that all regions are uniquely aligned in the genome, thus improving the alignment efficiency. This application is not limited thereto, and other methods may also be used to remove multi-alignment regions. For example, the sequence of the target region may be selected, aligned to the reference genome (hg19), and the uniquely aligned regions may be selected. It should be noted that "alignment" refers to the process of mapping short read sequences (reads) to the reference genome. Alignment tools (such as BWA, Bowtie2, etc.) will attempt to find the best matching position of each read on the reference genome. The alignment result may be unique or multiple. A "uniquely aligned interval" refers to an interval where a read can be uniquely aligned to a specific interval on the reference genome, rather than multiple positions. In other words, the best alignment position of this read is only one, rather than multiple possible positions.
[0036] In some examples, removing multi-alignment regions may specifically include: obtaining a file of uniquely aligned intervals of the human genome from the UCSC website, calculating the uniquely aligned score values of all regions in the target genomic region, and removing regions with uniquely aligned score values other than 1, that is, only retaining regions with uniquely aligned score values of 1. It should be noted that the "uniquely aligned score value" can be used to evaluate whether a read is uniquely aligned to a certain interval on the reference genome. The uniquely aligned score value = 1 / the number of times a read completely matches across the entire genome. For example, if a read matches only once across the entire genome, the uniquely aligned score value is 1; if a read can match twice in the genome, the uniquely aligned score value is 0.5. The value range of the uniquely aligned score value is from 0 to 1. A uniquely aligned score value of 1 indicates that the sequence of this interval is uniquely matched across the entire genome and will not be aligned to other regions. In this application, the uniquely aligned score values of each genomic region can be analyzed using the bigWigAverageOverBed tool. A uniquely aligned score value of 1 indicates that this region is a uniquely aligned region.
[0037] In some examples, masking regions outside the target genomic region may include: obtaining the human genome sequence of the hg19 version as the initial reference sequence, masking regions more than 150 bp outside the extended regions upstream and downstream of each target genomic region in the initial reference sequence, and converting the reference bases of these regions into "N". Thus, by processing the initial reference sequence, a new reference sequence (i.e., a customized reference sequence) is generated for sequence alignment.
[0038] In some examples, the method for analyzing sequencing data may further include steps of adapter trimming and quality control.
[0039] In some examples, the steps of adapter trimming and quality control include: performing adapter trimming and quality control on the sequencing data of the sample to be analyzed.
[0040] In some examples, performing adapter trimming on the sequencing data of the sample to be analyzed may include: using the fastp tool to trim the adapter sequences in the sequencing data.
[0041] In some examples, performing quality control on the sequencing data of the sample to be analyzed may include: if any of the following conditions is met, the sample fails the quality control and needs to be re-sequenced: 1) the Q30 base ratio of the sample sequence after trimming the adapter sequences is less than 80%; 2) the GC content of the sample sequence after trimming the adapter sequences is less than 30% or higher than 60%.
[0042] In some examples, the method for analyzing sequencing data may include: step S200, sequence alignment.
[0043] In some examples, sequence alignment may include: aligning the sequencing data of the sample to be analyzed with the customized reference sequence to obtain an alignment result. Among them, the sequencing data of the sample to be analyzed may be the sequencing data that has undergone adapter trimming and quality control.
[0044] In some examples, sequence alignment may include: using BWAmem or other tools, with default parameters, aligning the sequencing data of the sample to be analyzed to the customized reference sequence, and sorting and generating an index for the alignment result.
[0045] In some examples, the method for analyzing sequencing data may include: step S300, mutation analysis.
[0046] In some examples, mutation analysis may include: performing mutation detection and mutation filtering based on the alignment result to obtain the mutation result of the sample to be analyzed.
[0047] In some examples, VarDict is used for mutation detection. It should be noted that this software is more suitable for mutation detection in amplicon sequencing.
[0048] In some examples, mutation filtering may include: if a mutation meets any of the following conditions, the mutation will be filtered: 1) the allele frequency AF of the mutation is less than 0.1%; 2) the minimum base quality value BQ of the mutation is less than 20; 3) the minimum alignment quality MQ of the mutation is less than 10; 4) the variant depth quality value VD of the mutation is less than 2; 5) the positions of the mutation in the reads where it appears are the same. 6) If the mutation appears in dbSNP common SNPs. It should be noted that the minimum base quality value BQ refers to the lowest quality score that supports the mutated base. The higher the BQ, the higher the credibility of the mutated base. BQ represents the probability that the base is misidentified, usually in the form of a Phred quality score. For example, if the quality value of a base is 20, it means that the base has a 99% probability of being correct; the minimum alignment quality MQ refers to the lowest alignment quality score that supports the mutated read, which is a value used to measure the reliability of the alignment of short read sequences (reads) to the reference genome. It reflects the probability that the read is uniquely aligned to a specific position. The higher the MQ, the higher the accuracy of the read alignment to the genome; the variant depth quality value VD refers to the quality score of the mutation at the unit depth, which is the number of reads that support the mutation at a certain mutation site, usually expressed as the ratio of the number of reads supporting the mutation to the total coverage. The higher the VD, the higher the credibility of the mutation at a specific depth; the same position of the mutation in the reads means that the mutation appears at the same position in all reads that support it, which is used to increase the authenticity of the mutation and reduce the impact of sequencing or alignment errors; dbSNP (Database of Single Nucleotide Polymorphisms) is a public database that contains a large number of known single nucleotide polymorphisms (SNPs) and other small-scale genetic variations. If a mutation has been recorded as a common SNP in the dbSNP database, it means that the mutation may be a known polymorphic site, rather than a rare or de novo mutation.
[0049] In some examples, dbSNP common SNPs refer to the human common SNP sites downloaded from the dbSNP database (address: https: / / ftp.ncbi.nih.gov / snp / organisms / human_9606_b151_GRCh37p13 / VCF / common_all_20180423.vcf.gz).
[0050] In a specific embodiment, the overall steps of the method for analyzing sequencing data can be as follows: 1) Obtain the target region and remove the multi-alignment region; 2) From the initial reference sequence, mask the sequences outside the target region obtained in step 1 as the new reference sequence. 3) Perform adapter removal and quality control on the sequencing data to obtain the processed sequencing data; 4) Align the processed sequencing data to the new reference sequence; 5) Perform mutation analysis, including mutation detection and mutation filtering.
[0051] This application also relates to an apparatus for analyzing sequencing data.
[0052] Figure 2 It is a schematic diagram of the apparatus 800 for analyzing sequencing data according to this application. As Figure 2 shown, the apparatus 800 for analyzing sequencing data may include a customized reference sequence generation module 801 and a sequence alignment module 802. The apparatus 800 for analyzing sequencing data may also include a mutation analysis module 803. The functions of each module can be as follows:
[0053] In a specific embodiment, the customized reference sequence generation module 801 may be used to obtain the target genomic region and mask the regions outside the target genomic region to obtain the reference sequence. For the specific functions, reference can be made to the corresponding content in the aforementioned method for analyzing sequencing data, which will not be elaborated here.
[0054] In a specific embodiment, the sequence alignment module 802 may be used to align the sequencing data of the sample to be analyzed with the reference sequence to obtain the alignment result. For the specific functions, reference can be made to the corresponding content in the aforementioned method for analyzing sequencing data, which will not be elaborated here.
[0055] In a specific embodiment, the mutation analysis module 803 may be used to perform mutation detection and mutation filtering based on the alignment result to obtain the mutation result of the sample to be analyzed. For the specific functions, reference can be made to the corresponding content in the aforementioned method for analyzing sequencing data, which will not be elaborated here.
[0056] This application also relates to a device for analyzing sequencing data.
[0057] Figure 3 It is a schematic diagram of the device 100 for analyzing sequencing data according to this application. As Figure 3 shown, the device 100 for analyzing sequencing data may include a processor 10, a memory 20, and a computer program 21 (which may also be referred to as a computer-readable storage medium 21) stored in the memory 20. The computer program 21 can run on the processor 10, and when the processor 10 executes the computer program 21, the above method for analyzing sequencing data is implemented, for example, implemented Figure 1The steps shown. Alternatively, when the processor 60 executes the computer program 62, it implements the functions of the various modules in the apparatus 800 for analyzing sequencing data, such as Figure 2 the functions of the modules 801, 802, and 803 shown.
[0058] The device 100 for analyzing sequencing data can be a computing device such as a desktop computer, notebook, handheld computer, and cloud server. The device 100 for analyzing sequencing data can include but is not limited to the processor 10 and the memory 20. Those skilled in the art can understand that Figure 3 merely examples of the device 100 for analyzing sequencing data, and do not constitute a limitation on the device 100 for analyzing sequencing data. For example, it can include more or fewer components than shown, or combine certain components, or different components. For example, the device 100 for analyzing sequencing data can also include input / output devices, network access devices, buses, etc.
[0059] The processor 10 can be a central processing unit (CPU), or can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0060] The memory 20 can be an internal storage unit of the device 100 for analyzing sequencing data, for example, a hard disk or memory. The memory 20 can also be an external storage device of the device 100 for analyzing sequencing data, such as a plug-in hard disk equipped on the device 100 for analyzing sequencing data, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 20 can include both the internal storage unit of the device 100 for analyzing sequencing data and the external storage device.
[0061] The memory 20 can be used to store the computer program 21 and other programs and data required by the device 100 for analyzing sequencing data.
[0062] The memory 20 can also be used to temporarily store the data that has been output or will be output.
[0063] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0064] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0065] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the module or unit is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0066] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0067] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0068] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, all or part of the processes in the above-mentioned embodiment methods of the present application can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0069] The present application will be further described in detail below through specific examples and comparative examples. The following examples only further illustrate the present application and should not be construed as a limitation of the present application. In these examples and comparative examples, unless otherwise specified, the instruments and software used are all commercially available, and the operation steps are carried out in accordance with the product instructions and conventional operation specifications.
[0070] Example 1:
[0071] (1) The sample to be analyzed comes from our company. The sample number is: AS24175775 sample. The subject has hereditary epilepsy. Amplicon sequencing is used to detect whether there are disease-related mutations. The NM_001040142.2:c.2480A>T(p.Asp827Val) site of the target gene and the approximately 100bp region upstream and downstream thereof are captured by target probes, and sequencing is performed through the illumina second-generation sequencing platform.
[0072] (2) Analyze the amplicon sequencing data:
[0073] 1) Obtain the genomic coordinates of the hg19 version of all genes (a total of 6,236 genes) in the OMIM database and save them as a bed file named target.bed.
[0074] 2) Obtain the human genome uniquely mappable region file from the UCSC website (address: https: / / hgdownload.cse.ucsc.edu / goldenPath / hg19 / encodeDCC / wgEncodeMapability / wgEncodeCrgMapabilityAlign50mer.bigWig). Subsequently, use the bigWigAverageOverBed tool (address: https: / / hgdownload.cse.ucsc.edu / admin / exe / linux.x86_64 / bigWigAverageOverBed) to calculate the uniquely mappable score values for all regions in step 1), and only retain the regions with a uniquely mappable score value of 1. The generated file is named uniq_map.bed. The purpose of this step is to ensure that all regions are uniquely mappable in the genome, thereby improving the mapping efficiency.
[0075] 3) Use the bedtools maskfasta tool with the human reference genome hg19.fa downloaded from ucsc as the input file to mask the genomic regions other than uniq_map.bed (convert the bases to be masked into 'N'). Thus, a customized reference sequence is generated and named hg19.masked.fa. It should be noted that the customized reference sequence generated in this embodiment contains the uniquely mappable regions of all genes (a total of 6,236 genes) in the OMIM database and has good generality, and can be used for the analysis of sequencing data of amplicon sequencing of most single-gene genetic diseases / multigene genetic disease-related samples.
[0076] 4) Install the fastp tool (address: https: / / github.com / OpenGene / fastp), and use the parameter --detect_adapter_for_pe to remove the adapter sequences from the sequencing data of the sample to be analyzed and generate quality control results. If any of the following conditions is met, the sample fails the quality control and needs to be re-sequenced: ① The Q30 base ratio of the sample sequence after removing the adapter sequences is less than 80%; ② The GC content of the sample sequence after removing the adapter sequences is less than 30% or higher than 60%. In this example, fastp is used to remove the adapter sequences from the sample sequence, and the files after removing the adapter sequences are named AS24175775.clean.1.fq.gz and AS24175775.clean.2.fq.gz; the Q30 base ratio of the sample after removing the adapter sequences is 94%, and the GC content is 37%, so the quality control is qualified.
[0077] 5) Use BWAmem or other tools with default parameters to align the sample sequence to the reference sequence generated in step 2), and sort the alignment results and generate an index. Use bwamem as the sequence alignment tool, the reference sequence is hg19.masked.fa generated in step 3), and the input files are AS24175775.clean.1.fq.gz and AS24175775.clean.2.fq.gz generated in step 4) to generate the alignment result, named AS24175775.bam. Use samtools sort to sort AS24175775.bam to generate AS24175775.sorted.bam, and use samtools index to generate the index
[0078] AS24175775.sorted.bam.bai.
[0079] 6) Use VarDict (address: https: / / github.com / AstraZeneca-NGS / VarDictJava) for mutation detection. This software is more suitable for amplicon sequencing mutation detection. The input files include: the bam file (AS24175775.sorted.bam) generated in step 5), the reference sequence (chr2.masked.fa) generated in step 3), and the bed file (target.bed) containing the coordinates of the target amplification region (hg19 version). Subsequently, set the sample name to AS24175775, and use other default parameters for mutation detection.
[0080] 7) Download human common SNP sites from dbSNP (address: https: / / ftp.ncbi.nih.gov / snp / organisms / human_9606_b151_GRCh37p13 / VCF / common_all_20180423.vcf.gz) as a reference set for mutation filtering.
[0081] 8) Vardict will output quality control indicators for mutations. If the quality control indicators meet any of the following conditions, the mutation will be filtered: ① The allele frequency AF of the mutation is less than 0.1%; ② The minimum base quality value BQ of the mutation is less than 20; ③ The minimum mapping quality MQ of the mutation is less than 10; ④ The per-depth quality value VD of the mutation is less than 2; ⑤ The positions of the mutation in the reads where it appears are all the same. ⑥ If the mutation appears in dbSNP common SNPs.
[0082] (3) Analysis results:
[0083] Finally, a mutation file AS24175775.vcf is generated. The mutation records in AS24175775.vcf are as follows:
[0084] Chromosome number Position Reference base Mutated base Chimeric ratio 2 166198897 A T 39.97%
[0085] According to the mutation file, there is only one mutation in the mutation analysis result, which is chr2:166198897:A>T. The chimeric ratio of this point mutation is 39.97%. This mutation can also be expressed as NM_001040142.2:c.2480A>T(p.Asp827Val), and the gene involved is SCN2A. This gene is recorded in the OMIM database as being related to progressive epilepsy encephalopathy (OMIM ID: *182390). This mutation has not been included in the ClinVar database yet, but the site NM_001040142.2(SCN2A):c.2480A>G(p.Asp827Gly) with the same mutation position but different bases is included in the ClinVar database (number: RCV003786904.3) and is rated as pathogenic. Therefore, the gene mutation analyzed in this example is very likely the cause of the epilepsy in the subject.
[0086] Subsequently, Sanger sequencing was performed on this sample to verify the mutation. Figure 4 It is the Sanger sequencing result diagram of the sample involved in the embodiment of the present application. As Figure 4 shown, Sanger sequencing also detected this mutation, which can prove that the amplicon sequencing data analysis method of this example can accurately detect mutations related to diseases.
[0087] Comparative Example 1: The amplicon sequencing data of the AS24175775 sample was also used as the analysis object. However, when analyzing the amplicon sequencing data, steps 1), 2), and 3) were omitted. That is, the human reference genome hg19.fa was directly used as the reference sequence. The remaining steps were the same as those in Example 1. The analysis results showed that there were 68 more mutations than in Example 1 (see the table below). These 68 mutations were verified by viewing with IGV and confirmed to be false positives:
[0088]
[0089]
[0090] Comparative Example 2: The amplicon sequencing data of the AS24175775 sample was also used as the analysis object. However, when analyzing the amplicon sequencing data, step 2) was omitted. That is, the operation of removing multiple alignment regions in step 2 was not performed, and the genomic regions other than target.bed were directly masked in step 3 to be used as the customized reference sequence. The remaining steps were the same as those in Example 1. The analysis results showed that there were 3 more mutations than in Example 1 (see the table below). These 3 mutations were verified by viewing with IGV and confirmed to be false positives:
[0091] Chromosome number Position Reference base Mutated base 2 166198862 A T 2 166198914 C A 2 166198925 G A
[0092] Comparative Example 3: The amplicon sequencing data of the AS24175775 sample was also used as the analysis object. However, when analyzing the amplicon sequencing data, in step 8), mutations with AF lower than 0.1% were not filtered, and the remaining filtering conditions remained the same. The remaining steps were the same as those in Example 1. The analysis results showed that there were 2 more mutations than in Example 1 (see the table below). These 2 mutations were verified by viewing with IGV and confirmed to be false positives:
[0093] Chromosome number Position Reference base Mutated base 2 166198859 T A 2 166198911 A C
[0094] Comparative Example 4: The amplicon sequencing data of the AS24175775 sample was also used as the analysis object. However, when analyzing the amplicon sequencing data, in step 8), mutations with BQ lower than 20 were not filtered, and the remaining filtering conditions remained the same. The remaining steps were the same as those in Example 1. The analysis results showed that there were 6 more mutations than in Example 1 (see the table below). These 6 mutations were verified by viewing with IGV and confirmed to be false positives:
[0095] Chromosome number Position Reference base Mutated base 2 166198856 G T 2 166198868 T C 2 166198915 T G 2 166198932 G T 2 166198940 A T 2 166198943 T G
[0096] Comparative Example 5: The amplicon sequencing data of the AS24175775 sample was also used as the analysis object. However, when analyzing the amplicon sequencing data, in step 8), mutations with MQ lower than 10 were not filtered, and the remaining filtering conditions remained the same. The remaining steps were the same as those in Example 1. The analysis results showed that there were 20 more mutations than in Example 1 (see the table below). These 20 mutations were verified by viewing with IGV and confirmed to be false positives:
[0097] Chromosome number Position Reference base Mutated base 2 166198853 C G 2 166198854 A G 2 166198856 G T 2 166198859 T A 2 166198861 C A 2 166198868 T C 2 166198873 T A 2 166198879 A G 2 166198884 T A 2 166198886 G T 2 166198902 T A 2 166198909 T G 2 166198911 A C 2 166198920 T A 2 166198928 A T 2 166198933 G T 2 166198936 T C 2 166198941 A G 2 166198943 T G 2 166198946 G A
[0098] Comparative Example 6: The amplicon sequencing data of the AS24175775 sample was also used as the analysis object. However, when analyzing the amplicon sequencing data, in step 8), mutations with VD lower than 2 were not filtered, and the remaining filtering conditions remained the same. The remaining steps were the same as those in Example 1. The analysis results showed that there were 7 more mutations than in Example 1 (see the table below). These 7 mutations were verified by viewing with IGV and confirmed to be false positives:
[0099] Chromosome number Position Reference base Mutated base 2 166198860 C T 2 166198862 A T 2 166198884 T A 2 166198888 A T 2 166198914 C A 2 166198928 A T 2 166198939 C A
[0100] Comparative Example 7: The amplicon sequencing data of the AS24175775 sample was also used as the analysis object. However, when analyzing the amplicon sequencing data, in step 8), mutations with the same position in the read segments where they appeared were not filtered, and the remaining filtering conditions remained the same. The remaining steps were the same as those in Example 1. The analysis results showed that there were 2 more mutations than in Example 1 (see the table below). These 2 mutations were verified by viewing with IGV and confirmed to be false positives:
[0101] Chromosome number Position Reference base Mutated base 2 166198852 C T 2 166198868 T C
[0102] Comparative Example 8: The amplicon sequencing data of the AS24175775 sample was also used as the analysis object. However, when analyzing the amplicon sequencing data, in step 8), SNPs common in dbSNP were not filtered, and the remaining filtering conditions remained the same. The remaining steps were the same as those in Example 1. The analysis results showed that there were 3 more mutations than in Example 1 (see the table below). These 3 mutations were verified by viewing with IGV and confirmed to be false positives:
[0103] Chromosome number Position Reference base Mutated base 2 166198854 A G 2 166198858 A C 2 166198870 A T
[0104] In addition, in this example, amplicon sequencing was also performed on another 10 samples, and the same data analysis as in Example 1 was performed on their sequencing data, and Sanger verification was also carried out. The results are shown in the table below. As shown in the table below, the mutations detected by the amplicon sequencing data analysis method of the present application all have good accuracy:
[0105]
[0106]
[0107] In summary, the amplicon sequencing data analysis method of this embodiment can reduce noise, filter out most false positive mutations, and significantly improve the analysis accuracy.
[0108] The above content is a further detailed description of the present application in combination with specific implementation manners, and it cannot be determined that the specific implementation of the present application is only limited to these descriptions. For those of ordinary skill in the technical field to which the present application belongs, without departing from the concept of the present application, several simple deductions or substitutions can also be made.
Claims
1. A method for analyzing sequencing data, characterized in that Comprising: Generating a customized reference sequence: Obtaining a target genomic region, and masking regions outside the target genomic region in an initial reference sequence to obtain the customized reference sequence; Sequence alignment: Aligning the sequencing data of a sample to be analyzed with the customized reference sequence to obtain an alignment result; Mutation analysis: Performing mutation detection and mutation filtering based on the alignment result to obtain the mutation result of the sample to be analyzed.
2. The analysis method according to claim 1, characterized in that, The target genomic region includes any one or more of the following: genomic regions related to the disease of the sample to be analyzed, regions of all genes in the OMIM database.
3. The analysis method according to claim 1, characterized in that, The sequencing data is the sequencing data of amplicon sequencing.
4. The analysis method according to claim 1, wherein In the step of generating the customized reference sequence, it further includes removing the multi-alignment regions of the target genomic region: masking the regions with non-unique alignment in the target genomic region.
5. The analysis method according to claim 1, wherein Before the alignment step, it further includes adapter removal and quality control on the sequencing data of the sample to be analyzed; Preferably, when the quality control conditions meet the following two cases, it is considered that the quality control passes: the proportion of bases with Q30 is greater than or equal to 80%, and the GC content is greater than or equal to 30% and less than or equal to 60%.
6. The analysis method according to claim 1, characterized in that, Using VarDict for mutation detection.
7. The analysis method according to any one of claims 1 to 6, characterized in that Among the filtering conditions for the mutation filtering, mutations that meet at least one of the following conditions are filtered: the mutation frequency is lower than 0.1%, the minimum base quality value of the mutation is lower than 20, the minimum alignment quality of the mutation is lower than 10, the unit depth quality value of the mutation is lower than 2, the positions of the mutations in the reads where they appear are the same, and the mutation belongs to common SNPs in dbSNP.
8. An apparatus for analyzing sequencing data, characterized in that, Comprising: Customized reference sequence generation module: Obtaining a target genomic region, and masking regions outside the target genomic region in an initial reference sequence to obtain the customized reference sequence; Sequence alignment module: Used to align the sequencing data of a sample to be analyzed with the customized reference sequence to obtain an alignment result; Mutation analysis module: Used to perform mutation detection and mutation filtering based on the alignment result to obtain the mutation result of the sample to be analyzed.
9. An apparatus for analyzing sequencing data, characterized in that, Comprising a memory and a processor, the memory is used to store a program, and the processor realizes the analysis method of the sequencing data as described in any one of claims 1 to 7 by executing the program stored in the memory.
10. A computer-readable storage medium, characterized in that, A program is stored in the storage medium, and the program can be executed by a processor to realize the analysis method of the sequencing data as described in any one of claims 1 to 7.