Method for determining SNP detection region and method for correcting sequencing sample

By screening SNP sites and consistency alignment in specific regions of the reference genome, the SNP detection region is optimized, and the problem of sample confusion in high-throughput sequencing is solved, and fast and accurate sample correction and detection are achieved.

CN116705153BActive Publication Date: 2025-07-22BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310341881.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-09-16
Filing Date
2023-03-31
Publication Date
2025-07-22
Estimated Expiration
2043-03-31

AI Technical Summary

Technical Problem

In multiomics studies, high-throughput sequencing faces a huge sample size, the risk of sample confusion increases, resulting in inaccurate subsequent analysis results, and it is difficult for the prior art to effectively verify sample consistency.

Method used

By screening SNP sites in specific regions in the reference genome, combining common sequencing and methylation sequencing results for consistency comparison, SNP detection regions were determined, and sequencing samples were corrected using the consistency comparison results. The dbSNP library and low CG region screening method were used to optimize the SNP detection region, and SNP detection and alignment were used using GATK and BisSNP tools.

Benefits of technology

The SNP detection speed is improved, from 2.5 days to 5 hours, improving the accuracy and speed of sample verification and ensuring the accuracy of subsequent analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116705153B_ABST
    Figure CN116705153B_ABST
Patent Text Reader

Abstract

The present application relates to a method for correcting sequencing samples, including: respectively performing ordinary sequencing and methylation sequencing on a plurality of samples to obtain ordinary sequencing results and methylation sequencing results, obtaining a first set of SNP sites based on the ordinary sequencing results, and a second set of SNP sites obtained based on the methylation sequencing results, taking the overlapping sites of the first set of SNP sites and the second set of SNP sites as the third set of SNP sites, obtaining a consistency alignment result based on the detection results of the ordinary sequencing at the third set of SNP sites and the detection results of the methylation sequencing at the third set of SNP sites, and correcting the sequencing samples based on the comparison between the consistency alignment result and a specified threshold. The SNP detection speed of the method of the present application is increased from 2.5 days to 5 hours, improving the SNP detection efficiency of methylation data. The present application also improves the consistency alignment method, increasing the speed from 30 minutes to 10 seconds, comprehensively improving the accuracy and speed of sample verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of biotechnology. Specifically, it relates to a method for determining SNP detection regions and a method for correcting sequencing samples. Background Art

[0002] DNA methylation analysis using methylation sequencing (e.g., whole genome bisulfite sequencing (WGBS)) is increasingly regarded as a valuable diagnostic tool for detecting, diagnosing, and / or monitoring diseases such as cancer. DNA methylation has been shown to be tissue-specific, can be used for early cancer detection, and can trace the primary tumor site based on the methylation characteristics of circulating tumor DNA (ctDNA).

[0003] Whole genome sequencing (WGS) is a high-throughput sequencing technology used to rapidly and cost-effectively determine the complete genomic sequence of an organism. Deep sequencing of the genome is of great significance for clinical research. WGS sequencing is the cornerstone of precision medicine in understanding the importance of genomic mutations in health and disease.

[0004] With the development of high-throughput sequencing technology, multi-omics research has been continuously deepened. By performing high-throughput sequencing on each omics and integrating and studying the data, the interrelationships between substances in the fields of basic research, disease diagnosis, and drug development can be comprehensively and systematically understood.

[0005] In multi-omics research, high-throughput sequencing often has to face a huge sample size, which usually increases the risk of sample confusion. Sample information symmetry is the basis for subsequent information analysis. Therefore, before analysis, the samples need to be verified first to ensure that the samples for subsequent analysis match the known information and obtain more accurate results in the research and analysis. Summary of the Invention

[0006] This application provides a method for determining SNP (single nucleotide polymorphism) detection regions and a method for correcting sequencing samples. The method of this application can accurately and efficiently verify test samples.

[0007] Specifically, this application relates to the following:

[0008] 1. A method for determining SNP detection regions, the method comprising the following steps:

[0009] Screen the reference gene loci and mutant gene loci in the database that are A or T to obtain a first locus set,

[0010] Screen the sites in the first locus set where the minimum allele frequency of A mutated to T or T mutated to A is greater than 0.3 and less than 0.55 to obtain the second locus set.

[0011] Screen the sites in the second locus set that are located in the low CG regions of the reference genome to obtain the third locus set. Based on the third locus set, obtain the SNP detection region.

[0012] 2. The method according to item 1, wherein the number of sites in the third locus set is more than 1M.

[0013] 3. The method according to item 1, wherein based on the third locus set, the SNP detection region is obtained as follows:

[0014] Extend the sites in the third locus set 100 - 200bp forward and backward, and merge the overlapping regions to obtain the SNP detection region.

[0015] 4. The method according to item 1, wherein the length of the SNP detection region is more than 300M bp.

[0016] 5. The method according to item 1, wherein the database is the dbSNP database.

[0017] 6. The method according to item 1, wherein the low CG region refers to the region without CG bases in the 200bp bin interval of the reference genome.

[0018] 7. A method for correcting a sequencing sample, the method comprising the following steps:

[0019] Sequentially perform ordinary sequencing and methylation sequencing on multiple samples to obtain ordinary sequencing results and methylation sequencing results.

[0020] Obtain the first SNP locus set based on the ordinary sequencing results, and the second SNP locus set obtained based on the methylation sequencing results.

[0021] Take the overlapping sites of the first SNP locus set and the second SNP locus set as the third SNP locus set.

[0022] Based on the detection results of ordinary sequencing at the third SNP locus set and the detection results of methylation sequencing at the third SNP locus set, obtain a consistency alignment result.

[0023] When the consistency alignment result is greater than the specified threshold, determine that the results of ordinary sequencing and methylation sequencing are the ordinary sequencing results and methylation sequencing results of the same sample.

[0024] When the consistency comparison result is less than or equal to the specified threshold, it is determined that the results of the ordinary sequencing and the methylation sequencing are the results of ordinary sequencing and methylation sequencing of different samples.

[0025] 8. The method according to item 7, wherein the ordinary sequencing is whole-genome sequencing, targeted genome sequencing or microarray sequencing, and the methylation sequencing is bisulfite whole-genome sequencing or targeted genome methylation sequencing.

[0026] 9. The method according to item 7, wherein obtaining the first set of SNP sites based on the ordinary sequencing results includes:

[0027] Performing SNP detection on the ordinary sequencing results to obtain preliminary SNP sites,

[0028] Screening the SNP sites with a variant site reads depth greater than 8 and the number of reads with mutations greater than 3 among the preliminary SNP sites to obtain the first set of SNP sites.

[0029] 10. The method according to item 9, wherein performing SNP detection on the ordinary sequencing results is to perform SNP detection on the results of the ordinary sequencing in the SNP detection region.

[0030] 11. The method according to item 7, wherein obtaining the second set of SNP sites based on the methylation sequencing results includes:

[0031] Performing SNP detection on the results of the methylation sequencing in the SNP detection region to obtain preliminary SNP sites,

[0032] Screening the SNP sites with a variant site reads depth greater than 8 and the number of reads with mutations greater than 3 among the preliminary SNP sites to obtain the second set of SNP sites.

[0033] 12. The method according to item 10 or 11, wherein the SNP detection region is determined by the method described in any one of items 1-5.

[0034] 13. The method according to item 7, wherein the consistency comparison result is the ratio of the number of sites with the same genotype in the third set of SNP sites between the ordinary sequencing and the methylation sequencing to the total number of sites in the third set of SNP sites.

[0035] 14. The method according to item 7, wherein the ordinary sequencing result is the sequencing result after removing low-quality sequencing.

[0036] 15. The method according to item 7, wherein the methylation sequencing result is the sequencing result after removing low-quality sequencing.

[0037] 16. The method according to item 7, wherein the specified threshold is between the maximum consistency alignment result of different samples and the minimum consistency alignment result of the same sample.

[0038] In multi-omics research, high-throughput sequencing often has to face a huge sample size, which usually increases the risk of sample confusion. The symmetry of sample information is the basis for subsequent information analysis. Therefore, before analysis, the samples need to be corrected first to ensure that the samples for subsequent analysis match the known information, and to ensure that more accurate results can be obtained in subsequent research and analysis such as modeling prediction.

[0039] This application provides a method for determining SNP detection regions and a method for correcting sequencing samples using the determined SNP detection regions. The method of this application uses the SNP sites jointly detected by ordinary sequencing and methylation sequencing omics data to replace the use of all SNP sites for sample consistency comparison. On the basis of ensuring accuracy, the detection efficiency is greatly improved, and the SNP detection speed is increased from 2.5 days to 5 hours.

[0040] The method of this application also improves the consistency alignment method, which not only improves the accuracy of the consistency alignment, but also increases the speed from 30 minutes to 10 seconds, comprehensively improving the accuracy and speed of sample verification. Description of the Drawings

[0041] Figure 1 It is a schematic flow chart for determining SNP detection regions;

[0042] Figure 2 It is a schematic diagram for setting the specified threshold;

[0043] Figure 3 It is a schematic diagram for SNP detection and alignment of sequencing data. Detailed Embodiments

[0044] The following further illustrates this application with reference to embodiments. It should be understood that the embodiments are only used to further illustrate and explain this application and are not used to limit this application.

[0045] Unless otherwise defined, the technical and scientific terms in this specification have the same meaning as commonly understood by those skilled in the art. Although methods and materials similar or equivalent to those described herein can be used in experiments or practical applications, the materials and methods are described below. In case of conflict, the present specification, including its definitions, shall prevail. In addition, the materials, methods, and examples are for illustrative purposes only and are not restrictive. The following further illustrates this application with specific embodiments, but does not limit the scope of this application.

[0046] Single nucleotide polymorphisms (SNPs) mainly refer to DNA sequence polymorphisms caused by single nucleotide variations at the genomic level, including base transversions, transitions, insertions, and deletions. It is the most common type of human heritable variation, accounting for more than 90% of all known polymorphisms. As the third-generation molecular marker, SNPs are widely used in many fields such as molecular genetics, forensic physical evidence testing, and disease diagnosis and treatment.

[0047] This application provides a method for determining SNP detection regions, as Figure 1 shown, the method includes the following steps:

[0048] Step 1: Screen the sites in the database where the reference gene locus and the mutant gene locus are A or T to obtain the first locus set,

[0049] Step 2: Screen the sites in the first locus set where the minor allele frequency (MAF) of A mutating to T or T mutating to A is greater than 0.3 and less than 0.55 to obtain the second locus set,

[0050] Step 3: Screen the sites in the second locus set that are located in the low CG region of the reference genome to obtain the third locus set,

[0051] Step 4: Based on the third locus set, obtain the SNP detection region.

[0052] In Step 1, the database can be the dbSNP database (The Single Nucleotide Polymorphism database).

[0053] In Step 2, the minor allele frequency is obtained with reference to the 1000 Genomes Project database.

[0054] In Step 3, the low CG region refers to the region in the reference genome where there are no CG bases within a 200bp bin interval. Here, the bin interval is an artificially divided interval of a certain length in the reference genome. For example, the 200bp bin interval means dividing the reference genome into multiple 200bp intervals. The low CG region in this application is obtained by counting these multiple 200bp intervals. If there are no CG bases in the interval, it is considered a low CG region. The number of sites in the third locus set is more than 1M, such as 1M - 5M, 1M - 3M, 1M - 2M, etc.

[0055] In Step 4, based on the third locus set, the obtained SNP detection region is:

[0056] Extend the sites in the third site set 100 - 200 bp before and after (for example, it can be 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 200 bp), and merge the overlapping regions to obtain the SNP detection region.

[0057] In a specific embodiment, the length of the SNP detection region is more than 300 M bp, for example, it can be 300 M - 1000 M bp, 300 M - 800 M bp, 300 M - 500 M bp, 300 M - 400 M bp.

[0058] In a specific embodiment, a method for determining the SNP detection region includes the following steps: Screen the sites in the dbSNP library (The Single Nucleotide Polymorphism database) where the reference gene locus and the mutant gene locus are A or T to obtain the first site set; Screen the sites in the first site set where the minor allele frequency (MAF, Minor Allele Frequency) of A mutated to T or T mutated to A is greater than 0.3 and less than 0.55 to obtain the second site set; Screen the sites in the second site set that are located in the low CG region of the reference genome to obtain the third site set, where the number of sites in the third site set is more than 1 M; Extend the sites in the third site set 100 - 200 bp before and after, and merge the overlapping regions to obtain the SNP detection region, where the length of the SNP detection region is more than 300 M.

[0059] Compared with performing SNP detection using all SNP sites, performing SNP detection on the SNP detection region determined by the above method greatly improves the detection efficiency on the basis of ensuring accuracy, increasing the SNP detection speed from 2.5 days to 5 hours. At the same time, the determination of the SNP detection region by the method of the present application is an optimization of the SNP detection region. Compared with the whole-genome SNP detection, it is only an optimization of the detection region and does not affect the accuracy of SNP detection.

[0060] The present application also provides a method for correcting a sequencing sample, and the method includes the following steps:

[0061] Step 1: Sequencing multiple samples using ordinary sequencing and methylation sequencing respectively to obtain the ordinary sequencing result and the methylation sequencing result.

[0062] Step 2: Obtain the first SNP site set based on the ordinary sequencing result, and obtain the second SNP site set based on the methylation sequencing result.

[0063] Step 3: Use the overlapping sites of the first SNP locus set and the second SNP locus set as the third SNP locus set.

[0064] Step 4: Obtain a consistency alignment result based on the detection results of the ordinary sequencing at the third SNP locus set and the detection results of the methylation sequencing at the third SNP locus set.

[0065] Step 5: When the consistency alignment result is greater than a specified threshold, determine that the results of the ordinary sequencing and the methylation sequencing are the results of the ordinary sequencing and the methylation sequencing of the same sample.

[0066] When the consistency alignment result is less than or equal to the specified threshold, determine that the results of the ordinary sequencing and the methylation sequencing are the results of the ordinary sequencing and the methylation sequencing of different samples.

[0067] In Step 1, ordinary sequencing is in contrast to methylation sequencing, which is a sequencing method that requires special treatment of genomic DNA (such as bisulfite treatment). Ordinary sequencing does not require special treatment of DNA and is directly used for library construction and sequencing. Ordinary sequencing can include whole-genome sequencing, targeted genome sequencing, microarray sequencing, etc. Methylation sequencing can include bisulfite whole-genome sequencing, targeted genome methylation sequencing, etc.

[0068] Those skilled in the art can understand that for different sequencing methods, Step 1 may further include steps for processing the sample before sequencing. For example, the method may include the step of extracting cfDNA in the sample to obtain a DNA sample, and further sequencing the obtained sample.

[0069] In a specific embodiment, the ordinary sequencing result is the result after removing low-quality sequencing. For example, for whole-genome sequencing, the fastp quality control software can be used to view the sequencing quality, remove low-quality reads, and then the BWA alignment software is used to align the quality-controlled data to the reference genome.

[0070] In a specific embodiment, the methylation sequencing result is the result after removing low-quality sequencing. For example, for bisulfite whole-genome sequencing, the fastp quality control software can be used to view the sequencing quality, remove low-quality reads, and then the bismark alignment software is used to align the quality-controlled data to the reference genome.

[0071] In step two, obtaining the first SNP locus set based on the general sequencing results includes: performing SNP detection on the general sequencing results to obtain preliminary SNP loci; screening SNP loci with a variant locus read depth greater than 8 and the number of reads with variants greater than 3 among the preliminary SNP loci to obtain the first SNP locus set. Among them, the variant locus read depth in the preliminary SNP loci refers to the total number of reads covering the variant locus, and the reads with variants refer to the number of reads supporting the variant locus.

[0072] In a specific embodiment, performing SNP detection on the general sequencing results is to perform SNP detection on all loci.

[0073] In a specific embodiment, performing SNP detection on the general sequencing results is to perform SNP detection on the results of the general sequencing results in the SNP detection region. In a specific embodiment, the SNP detection region is the SNP detection region determined by the method for determining the SNP detection region described above.

[0074] Among them, for SNP detection, various tools and methods known in the art can be used. For example, for general sequencing data, common methods include GATK, samtools, FreeBayes, etc., and for methylation sequencing data, common methods include BisSNP, etc.

[0075] In a specific embodiment, the general sequencing is WGS. Obtaining the first SNP locus set based on the WGS results includes: the conventional WGS data call snp method (performing SNP detection on the file using GATK according to the aligned bam file. SNP locus filtering: screening SNP loci with a variant locus read depth greater than 8 (i.e., DP>8) and the number of reads with variants greater than 3 (i.e., AD>3).

[0076] Obtaining the second SNP locus set in the SNP detection region based on the methylation sequencing results includes: performing SNP detection on the results of the methylation sequencing results in the SNP detection region to obtain preliminary SNP loci; screening SNP loci with a variant locus read depth greater than 8 and the number of reads with variants greater than 3 among the preliminary SNP loci to obtain the second SNP locus set.

[0077] In a specific embodiment, the SNP detection region is the SNP detection region determined by the method for determining the SNP detection region described above.

[0078] In a specific embodiment, the methylation sequencing is WGBS. The second SNP locus set in the SNP detection region obtained based on the WGBS results includes: extracting the bam in the SNP detection region from the aligned bam file, using BisSNP (https: / / people.csail.mit.edu / dnaase / bissnp2011 / ) to detect SNPs in the file, and filtering SNP loci: screening for variant loci where the read depth is greater than 8 (i.e., DP>8) and the number of reads with mutations is greater than 3 (i.e., AD>3), and screening for SNPs in the specific locus set.

[0079] In step four, the consistency alignment result is the ratio of the loci with the same genotype in the third SNP locus set between the ordinary sequencing and the methylation sequencing to the total loci in the third SNP locus set. The consistency alignment can be performed by various software and methods known in the art, or can be run by using self-written programs or scripts. For example, a script can be written in shell.

[0080] When performing the consistency alignment, the method of the present application only performs the consistency alignment on the loci jointly detected in the results of the ordinary sequencing and the methylation sequencing, that is, the third SNP loci. Compared with detecting all SNP loci, it not only improves the accuracy of the consistency alignment, but also increases the alignment speed from 30 minutes to 10 seconds.

[0081] Step five is to judge whether the detected samples are the same samples based on the consistency alignment result. Specifically, when the consistency alignment result is greater than the specified threshold, it is determined that the results of the ordinary sequencing and the methylation sequencing are the ordinary sequencing results and the methylation sequencing results of the same sample, that is, the detected samples are the same samples. When the consistency alignment result is less than or equal to the specified threshold, it is determined that the results of the ordinary sequencing and the methylation sequencing are the ordinary sequencing results and the methylation sequencing results of different samples, that is, the detected samples are different samples.

[0082] The consistency thresholds for different omics data alignments (such as WGBS, methylation capture panel data and WGS, genomic capture panel data, etc.) are usually different. The method for setting the threshold is usually set according to the distribution of the consistency ratios of the same samples and different samples, so as to achieve the purpose of accurately identifying the same samples and different samples by setting a threshold.

[0083] In a specific embodiment, as Figure 2 shown, the specified threshold is between the maximum consistency alignment result of different samples and the minimum consistency alignment result of the same sample.

[0084] In a specific embodiment, the method for calibrating a sequencing sample in the present application includes the following steps: Sequencing a plurality of samples using WGS and WGBS respectively to obtain WGS results and WGBS results; obtaining a first set of SNP sites based on the WGS results, and obtaining a second set of SNP sites based on the WGBS results; taking the overlapping sites of the first set of SNP sites and the second set of SNP sites as the third set of SNP sites; obtaining a consistency alignment result based on the detection results of WGS in the third set of SNP sites and the detection results of WGBS in the third set of SNP sites; when the consistency alignment result is greater than a specified threshold, determining that the WGS result and the WGBS result are the WGS result and the WGBS result of the same sample, and when the consistency alignment result is less than or equal to the specified threshold, determining that the WGS result and the WGBS result are the WGS result and the WGBS result of different samples. Among them, obtaining the first set of SNP sites based on the WGS results includes: performing SNP detection on the WGS results to obtain preliminary SNP sites; screening SNP sites in the preliminary SNP sites where the read depth of the variant sites is greater than 8 and the number of reads with mutations is greater than 3 to obtain the first set of SNP sites. Obtaining the second set of SNP sites based on the WGBS results includes: performing SNP detection on the results of the WGBS results in the SNP detection region to obtain preliminary SNP sites; screening SNP sites in the preliminary SNP sites where the read depth of the variant sites is greater than 8 and the number of reads with mutations is greater than 3 to obtain the second set of SNP sites.

[0085] In a specific embodiment, the method for correcting a sequencing sample in the present application includes the following steps: Sequencing multiple samples using WGS and WGBS respectively to obtain WGS results and WGBS results; obtaining a first SNP site set based on the WGS results, and obtaining a second SNP site set based on the WGBS results; taking the overlapping sites of the first SNP site set and the second SNP site set as the third SNP site set; obtaining a consistency alignment result based on the detection results of WGS at the third SNP site set and the detection results of WGBS at the third SNP site set; when the consistency alignment result is greater than a specified threshold, determining that the WGS result and the WGBS result are the WGS result and the WGBS result of the same sample, and when the consistency alignment result is less than or equal to the specified threshold, determining that the WGS result and the WGBS result are the WGS result and the WGBS result of different samples. Among them, obtaining the first SNP site set based on the WGS results includes: performing SNP detection on the results of the WGS results in the SNP detection region to obtain preliminary SNP sites; screening the SNP sites in the preliminary SNP sites where the read depth of the variant sites is greater than 8 and the number of reads with mutations is greater than 3 to obtain the first SNP site set. Obtaining the second SNP site set in the SNP detection region based on the WGBS results includes: performing SNP detection on the results of the WGBS results in the SNP detection region to obtain preliminary SNP sites; screening the SNP sites in the preliminary SNP sites where the read depth of the variant sites is greater than 8 and the number of reads with mutations is greater than 3 to obtain the second SNP site set.

[0086] In a specific embodiment, the method for correcting a sequencing sample in the present application includes the following steps: sequencing a plurality of samples using WGS and WGBS respectively to obtain WGS results and WGBS results; obtaining a first set of SNP sites based on the WGS results, and obtaining a second set of SNP sites based on the WGBS results; taking the overlapping sites of the first set of SNP sites and the second set of SNP sites as the third set of SNP sites; obtaining a consistency alignment result based on the detection results of WGS at the third set of SNP sites and the detection results of WGBS at the third set of SNP sites; when the consistency alignment result is greater than a specified threshold, determining that the WGS result and the WGBS result are the WGS result and the WGBS result of the same sample, and when the consistency alignment result is less than or equal to the specified threshold, determining that the WGS result and the WGBS result are the WGS result and the WGBS result of different samples. Among them, obtaining the first set of SNP sites based on the WGS results includes: performing SNP detection on the WGS results to obtain preliminary SNP sites; screening SNP sites in the preliminary SNP sites where the read depth of the variant sites is greater than 8 and the number of reads with mutations is greater than 3 to obtain the first set of SNP sites. Obtaining the second set of SNP sites based on the WGBS results includes: performing SNP detection on the results of the WGBS results in the SNP detection region to obtain preliminary SNP sites; screening SNP sites in the preliminary SNP sites where the read depth of the variant sites is greater than 8 and the number of reads with mutations is greater than 3 to obtain the second set of SNP sites. The SNP detection region is determined by the following method: screening sites in the dbSNP library where the reference gene site and the mutant gene site are A or T to obtain the first set of sites; screening sites in the first set of sites where the minor allele frequency of A mutated to T or T mutated to A is greater than 0.3 and less than 0.55 to obtain the second set of sites; screening sites in the second set of sites that are located in low CG regions of the reference genome to obtain the third set of sites; extending the sites in the third set of sites by 100-200 bp before and after, and merging the overlapping regions to obtain the SNP detection region.

[0087] In a specific embodiment, the method for correcting a sequencing sample in the present application includes the following steps: Sequencing a plurality of samples using WGS and WGBS respectively to obtain WGS results and WGBS results; obtaining a first SNP locus set based on the WGS results, and obtaining a second SNP locus set based on the WGBS results; taking the overlapping loci of the first SNP locus set and the second SNP locus set as the third SNP locus set; obtaining a consistency alignment result based on the detection results of WGS at the third SNP locus set and the detection results of WGBS at the third SNP locus set; when the consistency alignment result is greater than a specified threshold, determining that the WGS result and the WGBS result are the WGS result and the WGBS result of the same sample, and when the consistency alignment result is less than or equal to the specified threshold, determining that the WGS result and the WGBS result are the WGS result and the WGBS result of different samples. Among them, obtaining the first SNP locus set based on the WGS results includes: performing SNP detection on the results of the WGS results in the SNP detection region to obtain preliminary SNP loci; screening SNP loci with a variant locus reads depth greater than 8 and the number of reads with mutations greater than 3 among the preliminary SNP loci to obtain the first SNP locus set. Obtaining the second SNP locus set based on the WGBS results includes: performing SNP detection on the results of the WGBS results in the SNP detection region to obtain preliminary SNP loci; screening SNP loci with a variant locus reads depth greater than 8 and the number of reads with mutations greater than 3 among the preliminary SNP loci to obtain the second SNP locus set. Among them, the SNP detection region is determined by the following method: screening loci in the dbSNP library where the reference gene locus and the mutant gene locus are A or T to obtain the first locus set; screening loci in the first locus set where the minor allele frequency of A mutated to T or T mutated to A is greater than 0.3 and less than 0.55 to obtain the second locus set; screening loci in the second locus set that are located in low CG regions of the reference genome to obtain the third locus set; extending the loci in the third locus set 100 - 200 bp forward and backward, and merging overlapping regions to obtain the SNP detection region.

[0088] Example

[0089] Example 1 Determining the SNP Detection Region

[0090] The flow schematic diagram for determining the SNP detection region is as Figure 1 shown.

[0091] 1) Screen the sites with ref and alt sites being A or T in the dbSNP (The Single Nucleotide Polymorphism database) library of the NCBI database, and the minimum allele frequency (MAF, Minor Allele Frequency, referring to the 1000 Genomes Project database) of the site mutation A→T or T→A is greater than 0.3 and less than 0.5, obtaining approximately 3.97M SNP sites.

[0092] 2) Retain the sites in the low CG regions of the reference genome among the SNP sites obtained in step 1) (defined as List1), and a total of approximately 1.3M SNP sites are screened.

[0093] 3) For the SNP sites obtained in step 2), extend 150bp before and after their positions to obtain the SNP detection region (defined as Bed1), with a length of approximately 316M.

[0094] Determining the SNP detection region through the above method is an optimization of the SNP detection region. Compared with the whole-genome SNP detection, it is only an optimization of the detection region and does not affect the accuracy of SNP detection. Example 2 SNP Detection and Alignment of Sequencing Data

[0095] The schematic diagram of SNP detection and alignment of sequencing data is as Figure 3 shown.

[0096] WGBS method for SNP detection: Extract the reads aligned to the Bed1 detection region from the aligned bam file to obtain bam1, use BisSNP to detect SNPs in bam1, and filter the detected SNP sites: Screen the variant sites with a reads depth greater than 8 and the number of reads with mutations greater than 3. After filtering, select the SNP sites in List1 among the SNP sites.

[0097] WGS method for SNP detection: According to the aligned bam file, use GATK4 to detect SNPs in the file, and filter the SNP sites: Screen the SNP sites with a reads depth greater than 8 and the number of reads with mutations greater than 3.

[0098] According to the SNP sites detected by WGBS and WGS, based on their positions on the reference genome, screen the sites jointly detected by the two omics data, that is, the SNPs at the overlapping sites, and the SNP sites of WGBS and WGS can be obtained respectively.

[0099] Write a script using shell to obtain the consistency alignment result based on the detection results of ordinary sequencing at overlapping sites and the detection results of methylation sequencing at overlapping sites, that is, calculate the ratio of the number of SNP sites with consistent genotypes at overlapping sites to the total number of overlapping SNP sites.

[0100] Example 3 Determine the consistency threshold for different omics data samples

[0101] Based on the SNP alignment method described in Example 2, WGS and WGBS omics studies were simultaneously conducted on 20 samples. The SNP site consistency alignment results of different omics studies of the samples are shown in Table 1, where the threshold is set between 0.6353 - 0.7125 according to the alignment results. For example, the threshold can be set to 0.65 。

[0102] Table 1

[0103]

[0104]

[0105] Among them, the alignment result of the same sample refers to the SNP consistency alignment result of the same sample.

[0106] The alignment result of different samples refers to the maximum value among the SNP consistency alignment results of this sample and non-self samples.

[0107] Calculation method for the consistency of SNP results of two types of data: Consistency = Number of SNP sites with consistent genotypes / Total number of SNP sites in the sample

[0108] Example 4 Verify the sample consistency of 1000 same samples

[0109] Based on the SNP alignment method described in Example 2 and the specified threshold determined in Example 3, the consistency verification was carried out on 1000 same samples. The results show that 100% of the samples are consistent.

[0110] Example 5 Traceability of mixed samples

[0111] Take 10 samples and scramble the label of one of the omics data. Through the SNP alignment method described in Example 2 and the specified threshold determined in Example 3, the truly corresponding consistent samples can be accurately found. The results are shown in Table 2.

[0112] Table 2

[0113]

[0114]

[0115] Among them, alignment result 1 refers to the SNP consistency alignment result of the same sample.

[0116] The comparison result 2 refers to the maximum value among the SNP consistency comparison results between the sample and non-self samples.

[0117] For a single sequencing sample, the method of the present application increases the SNP detection speed from 2.5 days to 5 hours, and the consistency comparison calculation speed from 30 minutes to 10 seconds, greatly improving the speed of the entire sequencing sample correction. When facing a large number of sequencing samples, the improvement in the correction speed will demonstrate greater advantages.

Claims

1. A method for determining SNP detection regions, the method comprising the following steps: Screening the sites in the database where the reference gene locus and the mutant gene locus are A or T to obtain a first site set; Screening the sites in the first site set where the minimum allele frequency of A mutated to T or T mutated to A is greater than 0.3 and less than 0.55 to obtain a second site set; Screening the sites in the second site set that are located in the low CG region of the reference genome to obtain a third site set; Based on the third site set, obtaining the SNP detection region; Wherein, based on the third site set, obtaining the SNP detection region is: Extending the sites in the third site set forward and backward by 100 - 200 bp, and merging the overlapping regions to obtain the SNP detection region; The low CG region refers to the region without CG bases within a 200 bp bin interval of the reference genome.

2. The method according to claim 1, wherein the number of sites in the third site set is more than 1M.

3. The method according to claim 1, wherein the length of the SNP detection region is more than 300M bp.

4. The method according to claim 1, wherein the database is the dbSNP database.

5. A method for correcting sequencing samples, the method comprising the following steps: Sequencing multiple samples respectively by ordinary sequencing and methylation sequencing to obtain the ordinary sequencing result and the methylation sequencing result; Obtaining a first SNP site set based on the ordinary sequencing result and a second SNP site set based on the methylation sequencing result; Taking the overlapping sites of the first SNP site set and the second SNP site set as the third SNP site set; Based on the detection result of the ordinary sequencing at the third SNP site set and the detection result of the methylation sequencing at the third SNP site set, obtaining a consistency alignment result; When the consistency alignment result is greater than the specified threshold, determining that the results of the ordinary sequencing and the methylation sequencing are the ordinary sequencing result and the methylation sequencing result of the same sample; When the consistency alignment result is less than or equal to the specified threshold, determining that the results of the ordinary sequencing and the methylation sequencing are the ordinary sequencing result and the methylation sequencing result of different samples; Wherein obtaining the first SNP site set based on the ordinary sequencing result includes: Performing SNP detection on the result of the ordinary sequencing in the SNP detection region to obtain preliminary SNP sites; Screening the SNP sites among the preliminary SNP sites where the reads depth of the variant sites is greater than 8 and the reads with variation are greater than 3 to obtain the first SNP site set; Obtaining the second SNP site set based on the methylation sequencing result includes: Performing SNP detection on the result of the methylation sequencing in the SNP detection region to obtain preliminary SNP sites; Screening the SNP sites among the preliminary SNP sites where the reads depth of the variant sites is greater than 8 and the reads with variation are greater than 3 to obtain the second SNP site set; Wherein the SNP detection region is determined by the method described in any one of claims 1 - 4.

6. The method according to claim 5, wherein the ordinary sequencing is whole-genome sequencing, targeted genome sequencing or microarray sequencing, and the methylation sequencing is bisulfite whole-genome sequencing or targeted genome methylation sequencing.

7. The method according to claim 5, wherein the consistency alignment result is the ratio of the number of sites with the same genotype in the third SNP site set by the ordinary sequencing and the methylation sequencing to the total number of sites in the third SNP site set.

8. The method according to claim 5, wherein the ordinary sequencing result is the sequencing result after removing low-quality sequencing.

9. The method according to claim 5, wherein the methylation sequencing result is the sequencing result after removing low-quality sequencing.

10. The method according to claim 5, wherein the specified threshold is between the maximum consistency alignment result of different samples and the minimum consistency alignment result of the same sample.

Citation Information

Patent Citations

  • Construction method of genome presumptive area nucleic acid sequencing library and device thereof

    CN105624272A

  • Homologous pseudogene variation detection method

    CN111081315A