Method, device, equipment and storage medium for evaluating contamination ratio of high-throughput sequencing data

Through genome comparison and variation detection, combined with the standard curve of AF value and pollution ratio, the homologous pollution ratio of high-throughput sequencing data is evaluated, which solves the problem of difficult detection of homologous pollution and improves the accuracy and reliability of sequencing data.

CN119132393BActive Publication Date: 2025-05-02SUZHOU SMK GENE TECH LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411590311.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-05-02
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

There are data pollution problems in high-throughput sequencing technology, especially homologous pollution is difficult to be effectively detected and corrected, which affects the accuracy and reliability of the results.

Method used

By sequencing the standards and homologous pollution sources separately, mixing different proportions of sequencing data for genome comparison and variation detection, statistically count the recall and false positive rates of variation, drawing a standard curve between AF value and pollution ratio, and then evaluating the pollution ratio and risk of the samples to be tested.

Benefits of technology

This method can effectively evaluate the proportion of homologous pollution in high-throughput sequencing data, improve the accuracy and efficiency of pollution detection, determine the pollution ratio 4% as the risk threshold, and ensure the quality of subsequent variation detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119132393B_ABST
    Figure CN119132393B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of high-throughput sequencing technology, and in particular to methods, devices, equipment and storage media for evaluating the contamination ratio of high-throughput sequencing data. The method provided by the present invention is based on the statistical quartile statistical results of the AF value, and constructs a standard curve of the contamination ratio and the IQR distribution and a standard curve of the maximum peak distribution of the contamination ratio and the density distribution between the [0.8,1] value range, and then uses quality control products to evaluate the changes in the recall rate and precision of the true set variation with different contamination ratios, and then determines that the contamination ratio of 4% can be used as the risk threshold of the impact of contamination on the accuracy of variation detection. The method is simple and efficient, has high resolution, and has high universality, and its analysis process is simple and convenient to deploy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of high-throughput sequencing, and in particular to a method, device, equipment and storage medium for evaluating the contamination ratio of high-throughput sequencing data. Background Art

[0002] With the rapid development of precision medicine, the use of high-throughput sequencing for whole-genome variation detection and assisting clinicians in exploring pathogenic mechanisms and determining the causes of diseases has become an increasingly common technology. High-throughput sequencing has many experimental links, starting from sample collection, nucleic acid extraction, library construction, and sequencing, sample contamination may occur. Among them:

[0003] During the sample collection process, different types of samples, such as blood, tissue, cells, etc., have different collection methods and storage conditions. If the collection process is not standardized, it is easy to introduce external impurities or contaminants. These external contaminants may be misjudged as components of the sample itself in the subsequent experimental process, thereby interfering with the results.

[0004] The reagents and consumables used in the nucleic acid extraction and library construction process may become sources of contamination, thereby affecting the quality of the nucleic acid or library, and further adversely affecting the subsequent sequencing results.

[0005] If the cleaning is not thorough during the sequencing process, there may be residual nucleic acid fragments from the previously sequenced samples. These residual nucleic acids may be misread during the sequencing of new samples, resulting in data contamination. At the same time, impurities or microorganisms in the sequencing reagents may also affect the accuracy and reliability of sequencing.

[0006] In addition, high-throughput sequencing technology itself has certain defects. For example, the accuracy of sequencing is not absolute, and incorrect base readings may occur. Although such errors can be corrected to a certain extent through bioinformatics analysis, if the error rate is too high or the distribution of errors has a certain regularity, it may have a serious impact on the subsequent analysis results.

[0007] It can be seen that many steps in sequencing will lead to mutual contamination between the final sequencing data. The occurrence of contamination usually leads to abnormalities in subsequent analysis results and affects the accuracy of the results. After the data are contaminated with each other, it is usually difficult to detect. If the contaminant is non-homologous sample data, such as human sequencing data contaminated by sequencing data from microorganisms or plants and animals, it can usually be determined by statistical genome alignment rates. However, if the contaminated data comes from homologous source data, such as human sequencing data contaminated by another sequencing sample from humans, the conventional detection alignment rate cannot effectively detect the contamination, especially low-proportion human contamination, which will introduce false-positive variant sites with low variation ratios, and such variant sites are easily judged as mosaic types, resulting in false positive results. At present, there is no good way to solve the data contamination caused by homologous contaminants. Summary of the invention

[0008] In view of this, the technical problem to be solved by the present invention is to provide a method, device, equipment and storage medium for evaluating the contamination ratio of high-throughput sequencing data.

[0009] The method for evaluating the homology contamination ratio of high-throughput sequencing data provided by the present invention comprises:

[0010] Sequence the standards and homologous contamination sources separately, mix the sequencing data in different proportions, and perform genome alignment and variation detection;

[0011] Perform statistical analysis on the recall rate and false positive rate of the variant detection results to determine the contamination proportional risk threshold;

[0012] Perform distribution statistics on variant AF values ​​and draw a standard curve between AF value and contamination ratio;

[0013] The samples to be tested undergo genome alignment and variation detection, and the variation AF values ​​are calculated;

[0014] According to the distribution of variant AF and the standard curve of contamination ratio, the contamination ratio of the sample to be tested is determined, and the risk of the sample to be tested is evaluated based on the contamination ratio risk threshold.

[0015] In the present invention, the specific sources of the samples to be tested, the standard substances and the pollution sources are not limited, and they can all be evaluated by the method of the present invention. The homologous pollution source refers to a sample from the same species / subspecies as the standard substance. For example, the standard substance, the sample to be tested and the pollution source are all from the same species, for example, they are all from humans, Escherichia coli, Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila, mice, rats, rabbits, dogs, cats, monkeys, tobacco or Arabidopsis. In the scheme of the present invention, for clinical applications, the sample to be tested, the standard substance and the pollution source are all from humans.

[0016] In the embodiment of the present invention, the sequencing data of the standard comes from NA12878. The homologous contamination source comes from a randomly selected non-standard human sample. The reference genome version can be hg19 or hg38.

[0017] Optionally, the mixing of the sequencing data in different proportions and performing genome alignment and variation detection includes:

[0018] The Fastq sequencing data file of the standard sample is obtained through sequencing;

[0019] Obtain the sequencing data file Fastq of the homologous contamination source through sequencing;

[0020] The Fastq sequencing data file of the contamination source is mixed with the Fastq sequencing data file of the standard sample in different proportions to generate a simulated Fastq file;

[0021] Align the simulated Fastq file to the reference genome of the species corresponding to the standard to obtain a genome alignment file;

[0022] Based on the genome comparison file, the genome variation detection of the standard and the contamination data at each ratio was performed to obtain the VCF file.

[0023] In this step, the pollution source data is proportionally added to the standard product data to obtain test data with different pollution ratios for subsequent statistical analysis.

[0024] In the present invention, the sequencing method is not limited. Preferably, it is NGS sequencing. In some embodiments, it is WES sequencing or WGS sequencing.

[0025] In the present invention, the mixing method is not limited. As a feasible case, the mixing ratio is 1% to 30%. In a specific embodiment, the mixing ratio is: 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25% or 30%.

[0026] Optionally, performing statistical analysis on the recall rate and false positive rate of the variation detection results to determine the contamination proportion risk threshold includes:

[0027] The variation in the standard VCF is regarded as a positive variation. The VCF files of the contaminated data at each ratio are compared with the VCF files of the standard. The recall rate and precision of the positive SNV variation are counted, and then the trend curve of the contamination ratio, recall rate and precision is drawn.

[0028] Taking the recall rate and precision of not less than 99% as the standard, the minimum contamination ratio is determined, and this contamination ratio is used as the risk threshold.

[0029] In this step, the VCF of the data with contamination (1%-30%) was compared with that of the data without contamination (0%), and the recall rate and precision index of the variant site detection in the standard were calculated, and the proportional risk threshold was determined.

[0030] In the present invention, the variation is SNV or InDel, and preferably, the variation is SNV.

[0031] In the present invention, the contamination ratio risk threshold is determined based on the recall rate and the false positive rate. Specifically, the contamination ratio of 4% is determined as the risk threshold based on the standard that both the recall rate and the false positive rate are not less than 99%. If the contamination ratio is higher than 4%, it indicates that the contamination will affect the data analysis. In addition, based on the different judgment methods applicable to 10% and below in the method, it is determined that the data is unqualified if the contamination ratio is higher than 10%.

[0032] In the present invention, the recall rate or precision is generated by a single test, or by taking an average value after multiple tests, which is not limited in the present invention. In order to improve the accuracy of the test, the average value of multiple test data is taken to construct a standard curve.

[0033] In the present invention, the risk threshold is a contamination ratio of 4%. If the sample contamination ratio is ≥ 4%, it is determined that there is a risk of variant detection. In the present invention, the contamination ratio of 4% is determined as the risk threshold based on the standard that both the recall rate and the false positive rate are not less than 99%. If the contamination ratio is higher than 4%, it indicates that the contamination will affect the data analysis. Due to differences in experimental operation procedures in different laboratories, the cut off value may change.

[0034] Optionally, the variant AF values ​​are distributed and statistically analyzed to draw a standard curve between the AF value and the contamination ratio. Previous experiments have shown that sample contamination will have a direct impact on the AF of the original variant site of the sample, and the higher the degree of contamination, the more significant the change in the AF value. Therefore, compared with other parameters, the AF value is a more effective indicator for evaluating contamination. Specifically, drawing a standard curve includes:

[0035] Screen the VCF files for heterozygous variants with sequencing depth ≥50X and variant type SNV, and extract the variant ratio AF value;

[0036] The AF values ​​of samples with different contamination ratios were quartile-statisticed, the interquartile range was calculated, and the standard curve A of contamination ratio and IQR distribution was summarized;

[0037] The AF values ​​of samples with different pollution ratios are plotted based on the density distribution;

[0038] According to the density distribution results, the maximum peak value between 0.8 and 1.0 is statistically analyzed, and a standard curve B between the pollution ratio and the peak distribution is drawn.

[0039] In this step, for the VCF files of samples with different contamination ratios and uncontaminated standard samples, the mutation ratio AF information of SNV variants is statistically analyzed, and a standard curve of the density distribution of contamination ratio and AF value is drawn.

[0040] Optionally, the sample to be tested undergoes genome alignment and variation detection, and the variation AF value is counted, including:

[0041] Obtain the sequencing data file Fastq of the sample to be tested through sequencing;

[0042] Align the sequencing data file Fastq of the sample to be tested to the reference genome to obtain a genome alignment file of the sample to be tested;

[0043] Based on the genome comparison file of the sample to be tested, perform genome variation detection to generate a VCF format file;

[0044] The VCF files were screened for heterozygous variants with a sequencing depth of ≥50X and a variant type of SNV, and the variant ratio AF value of each variant site was calculated.

[0045] In this step, the Fastq files of the clinical samples to be tested are subjected to whole genome alignment, and variation detection is performed to generate VCF files. The variation ratio AF values ​​of SNV type variations with a depth ≥ 50X are statistically calculated based on the VCF files.

[0046] Optionally, determining the contamination ratio of the sample to be tested according to the variant AF distribution and the contamination ratio standard curve, and evaluating the risk of the sample to be tested based on the contamination ratio risk threshold, includes:

[0047] Perform quartile statistics on the distribution of AF values ​​of the samples to be tested and calculate the interquartile range (IQR) value;

[0048] Based on the standard curve A, the contamination ratio range corresponding to the IQR value of the sample to be tested is evaluated;

[0049] If the contamination ratio is ≥10%, the feedback data is unqualified and the evaluation process is terminated;

[0050] If the contamination ratio is less than 10%, the AF value is statistically plotted to extract the maximum peak value of the density distribution within the range of [0.8, 1], and then the low contamination ratio of the sample to be tested is evaluated based on the standard curve B;

[0051] If the pollution ratio is ≥4%, the feedback data is unqualified;

[0052] If the pollution ratio is less than 4%, the feedback data is qualified.

[0053] In this step, based on the AF value of the variant of the sample to be tested, the quartile distribution is statistically analyzed to obtain the IQR value, and the density distribution is statistically analyzed to obtain the maximum peak distribution in the value range [0.8,1]. Based on the determined contamination ratio standard curve, the contamination ratio of the sample to be tested is determined; and based on the 4% contamination ratio risk threshold, it is evaluated whether the contamination ratio of the sample to be tested has a variant detection risk.

[0054] In this step, IQR statistics and density distribution statistics can be performed simultaneously or not, and the present invention does not limit this. When determining whether the data is qualified, a comprehensive determination is made by combining the pollution ratio estimated by the IQR method with the pollution ratio estimated by the density distribution.

[0055] Furthermore, the present invention provides a device for evaluating the homology contamination ratio of high-throughput sequencing data, comprising:

[0056] A sequencing module is used to obtain sequencing data of standards, contamination sources or samples to be tested;

[0057] The analysis module is used to generate simulated data files with different contamination ratios, genome alignment and mutation detection, statistical analysis of mutation results, and drawing of standard curves;

[0058] The evaluation module is used to determine the contamination ratio of the sample to be tested and evaluate the risk of the sample.

[0059] Furthermore, the present invention also provides an electronic device, comprising:

[0060] Memory, used to store computer programs;

[0061] A processor is used to execute the computer program to implement the steps of the method for evaluating the homologous contamination ratio of high-throughput sequencing data as described above.

[0062] Furthermore, the present invention also provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the method for evaluating the homologous contamination ratio of high-throughput sequencing data as described above are implemented.

[0063] The method provided by the present invention constructs a standard curve of contamination ratio and IQR distribution based on the statistical quartile statistical results of AF value, which is used to evaluate high contamination ratio data (contamination ratio>=10%), and constructs a standard curve of contamination ratio and maximum peak distribution of density distribution between [0.8,1] value range, which is used to evaluate low contamination ratio data (contamination ratio<10%), and uses quality control products to evaluate the changes in recall and precision of true set variation with different contamination ratios, and further determines that the contamination ratio of 4% can be used as the risk threshold of the impact of contamination on variation detection accuracy. The method has at least one of the following advantages:

[0064] 1) The method is simple and efficient: The method provided by the present invention can evaluate the contamination ratio and judge the risk of variant detection based on the AF distribution of the sample only, and does not introduce reference materials or information such as the population frequency of the variant. This reduces the complexity of the statistical method and improves the efficiency of contamination assessment;

[0065] 2) High resolution: The method provided by the present invention can evaluate a wide range of pollution levels, from 1% to 30% pollution ratio, and the invention can effectively identify it with high resolution.

[0066] 3) High universality: The method provided by the present invention has high universality and is applicable to a variety of high-throughput sequencing formats, such as whole genome sequencing (WGS), whole exome sequencing (WES), and customized gene panel sequencing, etc.; contamination assessment can be performed with only the variant detection VCF file of the sample to be tested;

[0067] 4) The analysis process is simple and easy to deploy: The method provided by the present invention has a simple process deployment and is easy to use and operate. The whole process analysis can be completed by deploying relevant computing nodes. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0069] Figure 1 The overall flow chart of the scheme of the present invention is shown;

[0070] Figure 2 Flow chart showing step 1);

[0071] Figure 3 Flow chart showing step 2);

[0072] Figure 4 Display recall rate and precision statistics and change trend distribution chart;

[0073] Figure 5 Flow chart showing step 3);

[0074] Figure 6 The standard curve showing the distribution of contamination ratio and IQR value;

[0075] Figure 7 Shows the density distribution of AF values ​​at different pollution ratios (range 0.8~1);

[0076] Figure 8 The standard curve showing the pollution ratio and peak distribution;

[0077] Fig. 9 4) is a flowchart showing step 4);

[0078] Fig.10 Flow chart showing step 5);

[0079] Fig.11 A schematic diagram showing the structure of a device for evaluating the homology contamination ratio of high-throughput sequencing data;

[0080] Fig.12 Figure 2 shows the structure of an electronic device for evaluating the homologous contamination ratio of high-throughput sequencing data. DETAILED DESCRIPTION

[0081] The present invention provides methods, devices, equipment and storage media for evaluating the contamination ratio of high-throughput sequencing data. Those skilled in the art can refer to the content of this article and appropriately improve the process parameters for implementation. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art, and they are all considered to be included in the present invention. The methods and applications of the present invention have been described through preferred embodiments, and relevant personnel can obviously modify or appropriately change and combine the methods and applications of this article without departing from the content, spirit and scope of the present invention to implement and apply the technology of the present invention.

[0082] Unless otherwise defined herein, scientific and technical terms related to the present invention shall have the meanings understood by those of ordinary skill in the art.

[0083] Herein, "standard" refers to samples that have been clearly determined to have all variant backgrounds and can be used as standards. Samples related to the Genome in a Bottle (GIAB) published by the National Institute of Standards and Technology (NIST) of the United States can also be used as standards. The present invention uses the NA12878 sample published by GIAB as a standard.

[0084] In this article, “contamination source” refers to human samples other than the selected standards, which can be used as simulated contamination sources;

[0085] In this article, “samples to be tested” refer to real samples for clinical testing;

[0086] In this article, "WES" refers to whole exome sequencing, a sequencing experiment that enriches genomic target sequences through probe capture and obtains whole exome sequence information through next-generation high-throughput sequencing (NGS).

[0087] In this article, "WGS" refers to whole genome sequencing, a sequencing experiment that obtains whole genome sequence information by randomly interrupting or amplifying the whole genome and performing next-generation high-throughput sequencing (NGS).

[0088] In this article, "Fastq" refers to the sequencing result file obtained after the sample is subjected to WES / WGS high-throughput sequencing. The file format is Fastq;

[0089] In this article, "BAM" refers to the genome alignment result file of sequencing data obtained after the sample Fastq data is aligned with the human reference genome. The file format is BAM (Binary Alignment Map);

[0090] In this article, "VCF" refers to the analysis of each base of the genome in the BAM file through variant detection software, and finally the specific location and type of variant in the whole genome are obtained. The result file format of recording the variant site is VCF (Variant Call Format); there are two types of variants: single nucleotide variant (SNV) and multi-base insertion or deletion (InDel).

[0091] In this article, "sequencing depth X" refers to the number of times a base on a genome is sequenced. After WES / WGS sequencing, the number of times a specified base position is covered by sequencing. For example, 50X means that the target base is covered by sequencing 50 times. Generally speaking, the higher the base sequencing depth X, the higher the reliability of the variation detection result of the base;

[0092] In this article, "AF" refers to the specific variant site described in the VCF file, the ratio of the sequencing depth X of the mutation to the total sequencing depth of the variant base, as described by the formula:

[0093] AF = sequencing depth of site mutation / total sequencing depth of site;

[0094] The problems solved by the present invention mainly include three aspects:

[0095] 1. The present invention is used to evaluate the contamination of samples from different sources. The degree of contamination can be evaluated based on the frequency distribution of sample variation itself. No control samples or reference population frequency-related information are required. The algorithm is simple and the performance is efficient.

[0096] 2. Construct a pollution assessment standard curve, based on which the sample to be tested can be assessed for contamination and the degree of contamination;

[0097] 3. The risk threshold of the contamination ratio is determined. If the contamination ratio is lower than the threshold, it indicates that the data has passed the quality control and is available, and has no impact on subsequent variant detection. If the contamination ratio is higher than the threshold, it indicates that the data has failed the quality control and there is a risk, and subsequent variant detection analysis cannot be continued;

[0098] The overall process of the technical solution of the present invention ( Figure 1 ) are summarized as follows:

[0099] 1) Select standards and pollution sources, generate pollution data simulations with different proportions, and perform genome comparison and variation detection;

[0100] 2) Perform statistical analysis on the recall rate and false positive rate of the variant detection results to determine the contamination ratio risk threshold;

[0101] 3) Perform distribution statistics on the variant AF values ​​and draw a standard curve of contamination ratio;

[0102] 4) Perform genome comparison and variation detection on the samples to be tested, and calculate the variation AF value;

[0103] 5) Calculate the distribution of AF variation and refer to the standard curve of contamination ratio to determine the contamination ratio of the sample to be tested, and evaluate the risk of the sample based on the contamination ratio risk threshold;

[0104] The present invention will be further described below in conjunction with embodiments:

[0105] Step 1). Select standards and pollution sources, generate pollution data simulations with different proportions, and perform genome comparison and variation detection ( Figure 2 );

[0106] effect:

[0107] The pollution source data is proportionally added to the standard product data to obtain test data with different pollution ratios for subsequent statistical analysis;

[0108] Build process:

[0109] A. Generation of Fastq sequencing data of standard product NA12878: Obtain Fastq sequencing data file of standard product sample through NGS sequencing experiment;

[0110] B. Generation of Fastq sequencing data of contaminated source samples: Randomly select non-standard human samples and obtain Fastq sequencing data files of standard samples through NGS sequencing experiments;

[0111] C. Simulate and generate Fastq files of data with different pollution ratios: Generate simulated Fastq files according to the pollution source Fastq file in B and the standard Fastq file in A according to the pollution ratios of 0%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, and 30%;

[0112] D. Genome alignment to generate BAM files: Align the Fastq files with different contamination ratios in C to the reference genome to obtain the genome alignment file (Binary Alignment Map, BAM). The reference genome version can be hg19 or hg38; alignment software includes but is not limited to the following categories: bwa, bowtie, MAQ, SOAP2, etc.;

[0113] E. Generate VCF file by mutation detection: Based on the BAM file in D, perform genome mutation detection to generate VCF format file; mutation detection software includes but is not limited to the following categories: GATK (The Genome Analysis Toolkit), Sentieon, Deepvariant, etc.

[0114] Input File:

[0115] Raw data Fastq file of the sample, reference gene Fasta file (hg19 or hg38).

[0116] Related Software:

[0117] Comparison software (bwa, bowtie, MAQ, SOAP2, etc.).

[0118] Variant detection software (GATK, Sentieon, Deepvariant, etc.).

[0119] Output file:

[0120] VCF files with different contamination ratios.

[0121] Step 2). Perform statistical analysis on the recall rate and false positive rate of the variant detection results to determine the contamination proportional risk threshold ( Figure 3 );

[0122] effect:

[0123] By comparing the VCF of the data with contamination (0%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%) and the data without contamination (0%), the recall rate and precision index of the variant site detection in the standard were calculated, and the proportional risk threshold was determined;

[0124] Build process:

[0125] A. Statistical recall and precision: The variants in the standard VCF are defined as positive variants. The VCF of the contaminated data is compared with the VCF of the uncontaminated standard sample, and the recall and precision of the positive SNV variants are statistically analyzed; statistical software includes but is not limited to hap.py;

[0126] B. Draw the recall rate and precision change curve: Draw the change trend curve of contamination ratio, recall rate and precision ( Figure 4 );

[0127] C. Set the contamination ratio risk threshold: With the recall rate and precision not less than 99% as the standard, determine the minimum contamination ratio, and use this contamination ratio as the risk threshold; according to statistical results, a contamination ratio of 4% can be used as the quality control threshold, and if the sample assessment contamination ratio is ≥4%, it can be determined that there is a risk of variant detection;

[0128] Input File:

[0129] VCF files of samples with different contamination ratios and VCF files of uncontaminated standards.

[0130] Output file:

[0131] The recall and precision of variant detection of samples with different contamination ratios relative to the standard, and the contamination risk threshold;

[0132] Statistical results:

[0133] Statistics and changing trends of SNV variation detection recall rate and precision of samples with different contamination ratios relative to the standard.

[0134]

[0135] Step 3). Perform distribution statistics on the variant AF values ​​and draw a standard curve of contamination ratio ( Figure 5 );

[0136] effect:

[0137] For the VCF files of samples with different contamination ratios and uncontaminated standard samples, the mutation ratio AF information of SNV variants was counted, and a standard curve of the density distribution of contamination ratio and AF value was drawn;

[0138] Build process:

[0139] A. Extract AF value from VCF file: Screen the VCF file with a sequencing depth of >= 50X and a heterozygous variant of SNV, and extract the variant ratio AF value;

[0140] B. Statistical AF distribution IQR: Perform quartile statistics on the AF values ​​of samples with different contamination ratios, calculate the interquartile range (IQR), and summarize the standard curve of contamination ratio and IQR distribution ( Figure 6 );

[0141] IQR statistics of AF values ​​at different pollution ratios.

[0142]

[0143] C. Statistical AF density distribution map: Draw a distribution map of the AF values ​​of samples with different pollution ratios based on the density distribution (KDE, Kernel Density Estimation) ( Figure 7 );

[0144] D. According to the density distribution results, the maximum peak value between 0.8 and 1.0 is counted, and a standard curve between the pollution ratio and the peak distribution is drawn;

[0145] The maximum peak value of AF value with different pollution ratios in the range [0.8,1] is:

[0146]

[0147] Input File:

[0148] VCF file of sample variation results with different contamination ratios;

[0149] Output file:

[0150] Standard curve of the distribution of contamination proportion and IQR value.

[0151] Standard curve of the contamination ratio and the maximum peak of the density distribution in the range [0.8,1].

[0152] Step 4). Perform genome alignment and variation detection on the samples to be tested, and calculate the AF density distribution ( Fig. 9 );

[0153] effect:

[0154] The Fastq files of the clinical samples to be tested are aligned to the whole genome, and the variants are detected to generate VCF files. The variant ratio AF values ​​of SNV type variants with a depth of >= 50X are calculated based on the VCF files;

[0155] Build process:

[0156] A. Fastq file acquisition: The raw data file Fastq of the clinical sample to be tested is obtained through WES or WGS sequencing;

[0157] B. Genome alignment to generate BAM files: Align the Fastq files with different contamination ratios in C to the reference genome to obtain the genome alignment file (Binary Alignment Map, BAM). The reference genome version can be hg19 or hg38; alignment software includes but is not limited to the following categories: bwa, bowtie, MAQ, SOAP2, etc.;

[0158] C. Generate VCF file by mutation detection: Based on the BAM file in B, perform genome mutation detection to generate VCF format file; mutation detection software includes but is not limited to the following categories: GATK (The Genome Analysis Toolkit), Sentieon, Deepvariant, etc.

[0159] D. SNV variation AF value statistics: Screen the VCF files with sequencing depth >= 50X and heterozygous variation of SNV type, and count the variation ratio AF value of each variation site;

[0160] Related Software:

[0161] Comparison software (bwa, bowtie, MAQ, SOAP2, etc.).

[0162] Variant detection software (GATK, Sentieon, Deepvariant, etc.).

[0163] Input File:

[0164] Fastq file of sample sequencing results.

[0165] Output file:

[0166] The variation ratio AF value of the variant site.

[0167] Step 5). Statistically calculate the distribution of variant AF, and refer to the contamination ratio standard curve to determine the contamination ratio of the sample to be tested, and evaluate the risk of the sample based on the contamination ratio risk threshold ( Fig.10 );

[0168] effect:

[0169] Based on the AF value of the variation of the sample to be tested, the quartile distribution is statistically analyzed to obtain the IQR value, and the density distribution is statistically analyzed to obtain the maximum peak distribution in the value range [0.8,1]. Based on the determined contamination ratio standard curve, the contamination ratio of the sample to be tested is determined; and based on the 4% contamination ratio risk threshold, it is evaluated whether the contamination ratio of the sample to be tested has a variation detection risk;

[0170] Build process:

[0171] A. Calculation of interquartile range (IQR): Perform quartile statistics on the distribution of AF values ​​of the samples to be tested and calculate the interquartile range (IQR);

[0172] B. Determination of high contamination ratio and risk assessment: Based on the standard curve of contamination ratio and IQR value distribution generated in 3), evaluate the contamination ratio range corresponding to the IQR value of the sample to be tested; if the contamination ratio is greater than or equal to 10%, the feedback data has a high proportion of contamination, which will seriously affect variant detection, the data is unqualified, the assessment process is terminated, and no subsequent assessment is performed;

[0173] C. Density distribution statistics and calculation of the maximum peak value in the target range: AF value statistics are drawn based on the density distribution (KDE, Kernel Density Estimation), and the maximum peak value of the density distribution in the range of [0.8, 1] is extracted;

[0174] D. Determination of low-proportion contamination and risk assessment: If the sample contamination ratio is <=10% as measured by IQR in B, the low contamination ratio of the sample to be tested needs to be evaluated based on the standard curve of contamination ratio and peak distribution generated in 3); if the contamination ratio is >=4%, the feedback data has a contamination ratio of >=4%, the degree of contamination affects variant detection, and the data is unqualified; if the contamination ratio is <4%, the feedback data is suspected to have a contamination ratio of <4%, the degree of contamination does not affect variant detection, and the data is qualified;

[0175] Input File:

[0176] List of variant AF values ​​of the samples.

[0177] Output file:

[0178] The estimated contamination ratio of the sample to be tested and whether there is a risk of variant detection;

[0179] Step 6). Replace the contaminated sample to evaluate and verify the performance of this methodology;

[0180] effect:

[0181] After replacing the pollution source sample, the new pollution source data is proportionally added to the standard data to obtain test data with different pollution ratios, and the IQR distribution of the AF value is statistically analyzed. At the same time, based on the density distribution (KDE, Kernel Density Estimation), the maximum peak value of the density distribution of the AF value in the range of [0.8,1] is extracted, and the pollution ratio is estimated through the standard curve. It is verified and checked with the simulated pollution ratio to evaluate the performance of this methodology;

[0182] Build process:

[0183] A. Generation of Fastq sequencing data of standard product NA12878: Obtain Fastq sequencing data file of standard product sample through NGS sequencing experiment;

[0184] B. Replace the contamination source sample and perform sequencing to generate Fastq data: Randomly select another non-standard human sample and obtain the Fastq sequencing data file of the standard sample through NGS sequencing experiment;

[0185] C. Simulate and generate Fastq files of data with different pollution ratios: Generate simulated Fastq files according to the pollution source Fastq file in B and the standard Fastq file in A according to the pollution ratios of 0%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, and 30%;

[0186] D. Genome alignment to generate BAM files: Align the Fastq files with different contamination ratios in C to the reference genome to obtain the genome alignment file (Binary Alignment Map, BAM). The reference genome version can be hg19 or hg38; alignment software includes but is not limited to the following categories: bwa, bowtie, MAQ, SOAP2, etc.;

[0187] E. Generate VCF file by mutation detection: Based on the BAM file in D, perform genome mutation detection to generate VCF format file; mutation detection software includes but is not limited to the following categories: GATK (The Genome Analysis Toolkit), Sentieon, Deepvariant, etc.

[0188] F. Estimate the sample contamination ratio: Screen the VCF file for heterozygous variants with sequencing depth >= 50X and variant type SNV, count the variant ratio AF value of each variant site, and calculate the maximum peak of the density distribution within the range of IQR value and AF value [0.8,1]; compare the standard curve of contamination ratio and IQR value with the standard curve of contamination ratio and peak distribution to estimate the sample contamination ratio;

[0189] Input File:

[0190] Raw data Fastq file of the sample, reference gene Fasta file (hg19 or hg38).

[0191] Related Software:

[0192] Comparison software (bwa, bowtie, MAQ, SOAP2, etc.).

[0193] Variant detection software (GATK, Sentieon, Deepvariant, etc.).

[0194] Output file:

[0195] VCF files with different contamination ratios and the estimated contamination ratio of the samples;

[0196] By analyzing the sample data of the second pollution source, relevant statistical indicators were obtained and the pollution ratio of the simulated data was estimated based on the standard curve. The data risk was evaluated and the performance of the methodology was evaluated by comparing it with the actual pollution ratio of the simulation process.

[0197] The second example shows the sample pollution simulation and pollution ratio assessment results.

[0198]

[0199] In summary: the estimated contamination ratio of the samples is basically consistent with the contamination ratio determined during the simulation, which verifies the accuracy of this methodology;

[0200] Reference Fig.11 As shown, the embodiment of the present invention discloses a device for evaluating the homology contamination ratio of high-throughput sequencing data, comprising:

[0201] A sequencing module 11 is used to obtain sequencing data of a standard, a contamination source or a sample to be tested;

[0202] Analysis module 12, used to generate simulation data files with different contamination ratios, genome alignment and variation detection, statistical analysis of variation results and drawing of standard curves;

[0203] The evaluation module 13 is used to determine the contamination ratio of the sample to be tested and evaluate the risk of the sample.

[0204] Furthermore, an embodiment of the present application also discloses an electronic device, which includes a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the computer program to implement the steps of the method for evaluating the homologous contamination ratio of high-throughput sequencing data as described above.

[0205] Fig.12 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.

[0206] Fig.12 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the method for evaluating the homologous contamination ratio of high-throughput sequencing data disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0207] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0208] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0209] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0210] Among them, the operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20, so as to realize the operation and processing of the massive data 223 in the memory 22 by the processor 21, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the method for evaluating the homologous contamination ratio of high-throughput sequencing data performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks. In addition to data transmitted from an external device received by the electronic device, the data 223 can also include data collected by its own input and output interface 25, etc.

[0211] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned method for evaluating the homologous contamination ratio of high-throughput sequencing data is implemented. The specific steps of the method can be referred to the corresponding contents disclosed in the aforementioned embodiments, and will not be repeated here.

[0212] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0213] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0214] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be directly implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable EPROM (Electrically Programmable Read Only Memory), an electrically erasable programmable EEPROM (Electric Erasable Programmable Read Only Memory), a register, a hard disk, a removable disk, a CD-ROM (Compact Disc-Read Only Memory), or any other form of storage medium known in the art.

[0215] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0216] The above are only preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for evaluating the homology contamination ratio of high-throughput sequencing data, characterized in that: include: Sequence the standards and homologous contamination sources separately, mix the sequencing data in different proportions, and perform genome alignment and variation detection; Perform statistical analysis on the recall rate and false positive rate of the variant detection results to determine the contamination proportional risk threshold; Perform distribution statistics on variant AF values ​​and draw a standard curve between AF values ​​and contamination ratios: screen heterozygous variants with sequencing depth ≥50X and variant type SNV in VCF files, extract variant ratio AF values; perform quartile statistics on AF values ​​of samples with different contamination ratios, calculate interquartile ranges, and summarize standard curve A of contamination ratio and IQR distribution; draw a distribution map of AF values ​​of samples with different contamination ratios based on density distribution; according to the results of the density distribution map, count the maximum peak value between 0.8 and 1.0, and draw a standard curve B of contamination ratio and the peak distribution; The sample to be tested is subjected to genome comparison and variation detection, and variation AF values ​​are counted: a sequencing data file Fastq of the sample to be tested is obtained by sequencing; the sequencing data file Fastq of the sample to be tested is compared to the reference genome to obtain a genome comparison file of the sample to be tested; based on the genome comparison file of the sample to be tested, genome variation detection is performed to generate a VCF format file; Screen the VCF files for heterozygous variants with a sequencing depth of ≥50X and a variant type of SNV, and calculate the variant ratio AF value of each variant site; Determine the contamination ratio of the sample to be tested based on the distribution of variant AF and the contamination ratio standard curve, and evaluate the risk of the sample to be tested based on the contamination ratio risk threshold; perform quartile statistics on the AF value distribution of the sample to be tested, and calculate the interquartile range IQR value; evaluate the contamination ratio range corresponding to the IQR value of the sample to be tested based on the standard curve A; The step of performing statistical analysis on the recall rate and false positive rate of the variation detection results to determine the contamination ratio risk threshold includes: The variation in the standard VCF is regarded as a positive variation. The VCF files of the contaminated data at each ratio are compared with the VCF files of the standard. The recall rate and precision of the positive SNV variation are counted, and then the trend curve of the contamination ratio, recall rate and precision is drawn. The minimum contamination ratio is determined with the recall rate and precision being no less than 99% as the standard, and this contamination ratio is used as the risk threshold; the risk threshold is a contamination ratio of 4%, and if the sample contamination ratio is ≥4%, it is determined that there is a risk of variant detection; If the contamination ratio is ≥10%, the feedback data is unqualified and the evaluation process is terminated; If the contamination ratio is less than 10%, the AF value is statistically plotted to extract the maximum peak value of the density distribution within the range of [0.8, 1], and then the low contamination ratio of the sample to be tested is evaluated based on the standard curve B; If the pollution ratio is ≥4%, the feedback data is unqualified; If the pollution ratio is less than 4%, the feedback data is qualified.

2. The method according to claim 1, characterized in that The step of mixing the sequencing data in different proportions and performing genome alignment and variation detection comprises: The Fastq sequencing data file of the standard sample is obtained through sequencing; Obtain the sequencing data file Fastq of the homologous contamination source through sequencing; The Fastq sequencing data file of the contamination source is mixed with the Fastq sequencing data file of the standard sample in different proportions to generate a simulated Fastq file; Align the simulated Fastq file to the reference genome of the species corresponding to the standard to obtain a genome alignment file; Based on the genome comparison file, the genome variation detection of the standard and the contamination data at each ratio was performed to obtain the VCF file.

3. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the method for evaluating the homologous contamination ratio of high-throughput sequencing data as described in claim 1 or 2.

4. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein, when the computer program is executed by a processor, the steps of the method for evaluating the homologous contamination ratio of high-throughput sequencing data as described in claim 1 or 2 are implemented.

Citation Information

Patent Citations

  • Gene sample pollution batch detection method and device based on regression model, equipment and medium

    CN116612814A

  • Method for detecting sample cross contamination based on SNP (Single Nucleotide Polymorphism) sites and application of method

    CN117059164A

  • Method and system for detecting sample contamination in high-throughput sequencing based on embryogenic mutation

    CN117253539A