Methods for identifying hot-spot prcs and sequencing errors in high-throughput sequencing data
The MIG analysis method with a two-stage correction step identifies and corrects high-frequency hotspot PCR and sequencing errors in high-throughput sequencing data, solving the problem of difficulty in distinguishing PCR and sequencing errors in existing technologies, and improving the accuracy and signal-to-noise ratio of the data, especially in antibody and T-cell receptor library analysis.
Patent Information
- Application Number
- CN202110695359.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-23
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2041-06-23
AI Technical Summary
Among the millions of reads generated by high-throughput sequencing, it is difficult to effectively distinguish between PCR and sequencing errors, which affects in-depth analysis of the diversity and composition of genomic DNA and cDNA libraries.
The MIG analysis method employs a two-stage correction process. First, dominant sequence variants within the sequencing read group are identified to filter out sequences represented by a single sequencing read group. Then, high-frequency hotspot PCR and sequencing errors are identified and corrected by counting and comparing nucleotide mismatch frequencies.
It effectively eliminates a large number of errors in the PCR amplification and high-throughput sequencing process, improving the accuracy and reliability of high-throughput sequencing data, especially in antibody and T-cell receptor library analysis, significantly improving the signal-to-noise ratio.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] This invention provides a method for analyzing high-throughput sequencing data, and a method for identifying high-frequency (hotspot) PCR and sequencing errors. In one aspect, this invention provides a method for analyzing high-throughput sequencing data of antibody and T-cell receptor libraries, specifically a method for identifying hotspot PCR and sequencing errors in high-throughput sequencing data. Background Technology
[0002] High-throughput sequencing can generate millions of reads, revealing profound information about the diversity and composition of PCR libraries based on genomic DNA and cDNA. However, the value of these analyses has been greatly diminished due to our limited ability to distinguish true sample diversity from PCR and sequencing errors.
[0003] To address this problem, Kinde et al. (Patent Application No. PCT / US2012 / 033207) proposed the Safe-Seqs algorithm using unique molecular identifiers. This algorithm requires identifying a dominant sequence variant within a set of sequencing reads from the same molecule that was initially labeled with a unique molecular identifier. To infer the correct nucleotide sequence of the original cDNA or DNA molecule, sequencing reads labeled with the same unique molecular identifier are organized into molecular identifier groups (MIGs, referred to as “families” in the reference). Since most errors accumulate in later PCR cycles or during sequencing, erroneous sequence variants are often represented by a small number of related reads when many template molecules are present. Therefore, most MIGs are either error-free or contain only a small subset, in the latter case, where errors can be reliably corrected. In some cases, a single dominant sequence is not observed in the MIG due to early errors during PCR amplification, and such ambiguous MIG cases can be distinguished and filtered out. To eliminate rare errors, those sequence variants represented by a single MIG (i.e., single molecular events not confirmed by independent molecular events) can be discarded.
[0004] However, due to the randomness of PCR amplification, this simple co-assembly cannot correct for those PCR errors that occur in the early amplification stages and subsequently dominate in the MIG. This can often occur in NGS depth analysis and cannot be resolved by filtering out rare variants or by repeat sequencing because PCR errors are highly reproducible in specific DNA contexts.
[0005] Through this invention, we introduce a second stage of MIG analysis that enables us to identify and eliminate such high-frequency (hotspot) PCR and sequencing errors specific to particular analytical libraries by using the information contained in the MIG. Summary of the Invention
[0006] The technical problem this invention aims to solve is that high-throughput sequencing can generate millions of reads, revealing in-depth information about the diversity and composition of PCR libraries based on genomic DNA and cDNA. However, the value of these analyses has been greatly diminished due to our limited ability to distinguish true sample diversity from PCR and sequencing errors.
[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical means:
[0008] A method for identifying hotspot PRCs and sequencing errors in high-throughput sequencing data, the steps of which are as follows:
[0009] Step 1: Receipt Acquisition Steps, obtain raw sequencing data;
[0010] Step 2: Data grouping step, grouping sequencing reads belonging to the same starting molecule into sequencing read groups;
[0011] Step 3: Data analysis steps, including two-level correction steps or steps to identify and correct sequencing errors.
[0012] As a preferred embodiment, a further technical solution of the present invention is:
[0013] The two-stage correction process consists of a first stage of MIG correction and a second stage of MIG correction. The first stage of MIG correction identifies dominant sequence variants within a set of sequencing reads, filtering out those reads where dominant sequence variants are represented by a single set of sequencing reads, and filtering out groups of sequencing reads in which no single dominant sequence is observed. The second stage of MIG correction counts the frequency of sequencing reads containing specific nucleotide mismatches within the sequence of readings, comparing the abundance of each specific nucleotide mismatch before and after the first and second stages of MIG correction. Specific nucleotide mismatches with a reduced relative frequency in the first stage of MIG correction are considered high-frequency, hotspot PCR, or sequencing errors. Those sequencing reads containing the identified high-frequency or hotspot errors are discarded or appropriately corrected in the second stage of MIG correction.
[0014] The two-stage correction process consists of a first stage and a second stage of MIG correction. The first stage involves identifying dominant sequence variants within a set of sequencing reads, filtering out those reads where dominant sequence variants are represented by a single set of sequencing reads, and filtering out groups of readings in which no single dominant sequence is observed. The second stage involves counting the relative number of sequencing read groups containing a small number of readings with specific nucleotide mismatches, and counting groups of readings containing the same sequence and the same nucleotide sequence. Specific nucleotide mismatches after the data grouping step are considered as the primary sequencing reads. For high-frequency, hotspot errors, the frequency of the previous set of sequencing reads should be higher than the frequency of the next set of sequencing reads. Sequencing read groups containing the identified high-frequency, hotspot errors are discarded or appropriately corrected in the second stage of MIG correction.
[0015] The step of identifying and correcting sequencing errors involves identifying and correcting high-frequency hotspot PCR and sequencing errors. This involves counting the frequency of sequencing reads containing specific nucleotide mismatches within the sequencing read array. Specific nucleotide mismatches with a reduced relative frequency are considered high-frequency, hotspot PCR, or sequencing errors. Sequencing read arrays containing identified high-frequency, hotspot errors are discarded or appropriately corrected in the step of identifying and correcting sequencing errors.
[0016] The step of identifying and correcting sequencing errors involves identifying and correcting high-frequency hotspot PCR and sequencing errors. This includes counting the relative number of sequencing read groups containing a small number of sequencing reads with specific nucleotide mismatches, and counting those groups containing the same sequence and the same nucleotide sequence. Specific nucleotide mismatches are read as the primary sequencing reads after the data grouping step. For high-frequency, hotspot errors, the frequency of the preceding sequencing read group should be higher than the frequency of the following sequencing read group. Sequencing read groups containing identified high-frequency, hotspot errors in variants will be discarded or appropriately corrected in the step of identifying and correcting sequencing errors.
[0017] The sequencing read set consists of sequencing reads, each marked with a unique molecular identifier, and the sequence contains at least one degenerate or semi-degenerate nucleic acid sequence; the sequencing reads are located at the same starting nucleotide position.
[0018] The sequencing read analysis described herein is for antibody gene library profiling, T cell receptor gene library analysis, and genomic library analysis based on high-throughput sequencing.
[0019] Through this invention, we introduce a second stage of MIG analysis that enables us to identify and eliminate high-frequency (hotspot) PCR and sequencing errors specific to particular analytical libraries by using the information contained in the MIG. This is highly effective in eliminating the large number of errors that occur during PCR amplification and high-throughput sequencing. Detailed Implementation
[0020] The present invention will be further described below with reference to the embodiments. Specific Implementation Example 1:
[0022] This invention relates to a method for identifying high-frequency, hotspot PCR, and sequencing nucleotide errors in high-throughput sequencing data by analyzing sequencing reads derived from the same starting molecule.
[0023] On one hand, the present invention relates to a method for identifying high-frequency, hotspot PCR and sequencing nucleotide errors in high-throughput sequencing data by counting the frequency of sequencing reads containing specific nucleotide mismatches within the sequencing read array.
[0024] In one embodiment of the first aspect of the method of the present invention, the following steps are included:
[0025] (1) Obtain raw sequencing data;
[0026] (2) Group sequencing reads belonging to the same starting molecule;
[0027] (3) The first stage of MIG correction;
[0028] (4) The second stage of MIG correction.
[0029] In one embodiment, step (3) of the method of the present invention includes identifying a dominant sequence variant within a set of sequencing reads. In other embodiments, step (3) further includes filtering out those sequencing read sets, wherein the dominant sequence variant is represented by a single set of sequencing reads. In other embodiments, step (3) further includes filtering out groups of sequencing reads in which no single dominant sequence was observed.
[0030] In one embodiment, step (4) of the method of the present invention includes counting the frequency of sequencing reads containing specific nucleotide mismatches within the sequencing read array.
[0031] In one embodiment, the abundance of each specific nucleotide mismatch was compared before and after step (3) in step (4) of the method of the present invention. In one embodiment, specific nucleotide mismatches with a reduced relative frequency that are typically corrected in step (3) are considered high-frequency (hotspot) PCR or sequencing errors. Therefore, those sequencing reads containing the identified high-frequency (hotspot) errors in the main step variants are discarded or appropriately corrected in step (4).
[0032] In a second aspect, the present invention relates to a method for identifying high-frequency (hotspot) PCR and sequencing nucleotide errors in high-throughput sequencing data, the method being achieved by counting the relative frequencies of sequencing read groups containing a specific number of sequencing reads. Nucleotide mismatches.
[0033] In one embodiment of the second aspect, the method of the present invention includes the following steps:
[0034] (1) Obtain raw sequencing data;
[0035] (2) Group sequencing reads belonging to the same starting molecule;
[0036] (3) The first stage of MIG correction;
[0037] (4) The second stage of MIG correction.
[0038] In one embodiment, step (3) of the method of the present invention includes identifying a dominant sequence variant within a set of sequencing reads. In other embodiments, step (3) further includes filtering out those sequencing read sets, wherein the dominant sequence variant is represented by a single set of sequencing reads. In other embodiments, step (3) further includes filtering out groups of sequencing reads in which no single dominant sequence was observed.
[0039] In one embodiment, step (4) of the method of the present invention includes counting the relative number of sequencing read groups containing a smaller number of sequencing reads with specific nucleotide mismatches, and counting those groups of sequencing reads containing the same sequence and having the same nucleotide sequence. After step (2), the specific nucleotide mismatch is taken as the primary sequencing read. For high-frequency (hotspot) errors, the frequency of the previous group of sequencing reads should be higher than the frequency of the subsequent group of sequencing reads. Therefore, those sequencing read groups containing the identified high-frequency (hotspot) errors in the primary step variant are discarded or appropriately corrected in step (4).
[0040] In a third aspect, the present invention relates to a method for identifying high-frequency (hotspot) PCR and sequencing nucleotide errors in high-throughput sequencing data by counting the frequency of sequencing reads containing specific nucleotide mismatches within a sequencing read array.
[0041] In one embodiment of the third aspect, the method of the present invention includes the following steps:
[0042] (1) Obtain raw sequencing data; (2) Group sequencing reads belonging to the same starting molecule;
[0043] (3) Identify and correct high-frequency (hotspot) PCR and sequencing errors.
[0044] In one embodiment, step (3) of the method of the present invention includes counting the frequency of sequencing reads containing specific nucleotide mismatches within the sequencing read array.
[0045] In one implementation, specific nucleotide mismatches with a reduced relative frequency are considered high-frequency (hotspot) PCR or sequencing errors. Therefore, those sequencing reads containing the identified high-frequency (hotspot) errors in the major sequence variants are discarded or appropriately corrected in step (3).
[0046] In a fourth aspect, the present invention relates to a method for identifying high-frequency (hotspot) PCR and sequencing nucleotide errors in high-throughput sequencing data, the method being achieved by counting the relative frequencies of sequencing read groups containing a small number of sequencing reads with a specific number of nucleotide mismatches.
[0047] In one embodiment of the fourth aspect, the method of the present invention includes the following steps:
[0048] (1) Obtain raw sequencing data;
[0049] (2) Group sequencing reads belonging to the same starting molecule;
[0050] (3) Identify and correct high-frequency (hotspot) PCR and sequencing errors.
[0051] In one embodiment, step (3) of the method of the present invention includes counting the relative number of sequencing read groups containing a small number of sequencing reads with specific nucleotide mismatches, and counting those groups containing the same sequence and having the same nucleotide sequence. The specific nucleotide mismatch is read as the primary sequencing read after step (2). For high-frequency (hotspot) errors, the frequency of the preceding group of sequencing reads should be higher than the frequency of the following group. Therefore, those sequencing read groups containing identified high-frequency (hotspot) errors in the primary step variant will be discarded or appropriately corrected in step (3).
[0052] In a preferred embodiment of the method for identifying high-frequency (hotspot) PCR and sequencing nucleotide errors in high-throughput sequencing data according to the first, second, third, or fourth aspect of the invention, the sequencing read set consists of sequencing reads. Each read is labeled with the same unique molecular identifier and contains at least one degenerate or semi-degenerate nucleic acid sequence.
[0053] In another preferred embodiment of the method for identifying high-frequency (hotspot) PCR and sequencing nucleotide errors in high-throughput sequencing data according to the first, second, third, or fourth aspect of the invention, the sequencing read set consists of sequencing reads characterized by having the same starting nucleotide position.
[0054] In another preferred embodiment of the method for identifying high-frequency (hotspot) PCR and sequencing nucleotide errors in high-throughput sequencing data according to the first, second, third, or fourth aspect of the invention, the sequencing read set consists of paired ends. The sequencing reads are characterized by having the same starting nucleotide position.
[0055] In another implementation, the group analyzing the sequencing reads refers to the antibody gene library profile obtained through high-throughput sequencing.
[0056] In another implementation, the group analyzing sequencing reads refers to the analysis of T-cell receptor gene libraries obtained through high-throughput sequencing.
[0057] In yet another implementation, high-throughput sequencing analysis of sequencing reads refers to genomic libraries.
[0058] The method of this invention is highly effective in eliminating a large number of errors that occur during PCR amplification and high-throughput sequencing. Specific Implementation Example 3:
[0060] Table 1. Reproducibility of PCR errors. a. One type. Percentage of erroneous readings in a given sample (row) that are also present in other samples (columns).
[0061] exp1 exp2_r1 exp2_r2 exp2_r3 exp2_r4 exp1 - 91% 90% 91% 90% exp2_r1 94% - 99% 98% 99% exp2_r2 94% 99% - 99% 99% exp2_r3 94% 95% 95% - 94% exp2_r4 94% 99% 99% 98% -
[0062] b. The percentage of erroneous variants in a given sample (row) that also exist in other samples (columns).
[0063] exp1 exp2_r1 exp2_r2 exp2_r3 exp2_r4 exp1 - 51% 48% 54% 50% exp2_rl 23% - 42% 42% 38% exp2_r2 22% 43% - 48% 41% Specific Implementation Example 4:
[0065] Table 2. Spike-in TCR and IG variants used in Spike-in Clones Experiment 1. Note that the relative abundance of the mixed T cell clones and their TCR-α and TCR-β chains is quantified with relative accuracy (compare the last two columns).
[0066]
[0067]
[0068]
[0069] b. Spike-in IG heavy chain variant used in Experiment 2. Red indicates introduced nucleotide and amino acid substitutions.
[0070]
[0071] Specific Implementation Example 5:
[0073] Table 3. Experiment Summary.
[0074]
[0075] Specific Implementation Example 6:
[0077] * - Geometric mean of the size of the molecular identifier group (MIG), the number of readings.
[0078] **- Unique molecular identifiers by size threshold (MIG of 30 readings for Experiment 1, MIG of 10 readings / MIG for Experiment 2).
[0079] MIGEC analysis effectively removed almost all erroneous spiked clonal subtypes from the pooled duplicate samples of Experiment 2, while preserving truly rare spiked subtypes. The signal-to-noise ratio (the ratio of the number of EHEB clonal reads to the most common erroneous subvariables with one or two mismatches) increased from approximately 1,000:1 and 20,000:1 in the conventionally processed data to approximately 12,000:1 and 60,000:1 in the MIGEC analysis (Table 4).
[0080] Table 4. Signal-to-noise ratio.
[0081]
[0082]
[0083] Various signal-to-noise ratios were compared between standard data processing (CDR3 region extracted from raw reads, filtered only by Phred Q>25) and Level 2 MIG correction. Read count ratios were compared between the original EHEB clone, artificial subvariants with one (EHEB-V1) and two (EHEB-V2) mismatches, and top-observed EHEB subvariants with one (Err1) and two (Err2) errors. S1–S4 – Replicas from Experiment 2.
[0084] In summary, MIGEC analysis can make the IG supermutation tree of the donor blood sample basically visible. Specific Implementation Example 7:
[0086] Two-stage MIG analysis. Vertical "barcode" lines indicate molecular identifiers used to group sequencing reads; white indicates correct sequences, and dashed lines indicate sequences containing errors. MIG - Molecular identifier grouping. ROI - Region of Interest (i.e., CDR3 for the immune library). a. Overall protocol for library preparation and analysis; b. Two-stage MIG correction. Various MIG scenarios are shown.
[0087] The gain and loss of sequencing read variants after the first stage of MIG correction. The same erroneous sequence variants that remained after the first stage of MIG correction in Experiment 1 also appeared in several successfully corrected MIGs.
[0088] After the first stage of MIG correction, the sequencing reads show the gains and losses of variants.
[0089] After the first stage of MIG correction, different variations in sequencing read counts for the correct clonal type and its incorrect subvariants resulted in the same repetitive PCR errors. Correct sequence variations were obtained (white), while repetitive PCR errors (dashed lines) resulted in lost sequencing reads. The relative number of sequencing reads for the true piercing CDR3 and its incorrect subvariants were shown after the first stage of MIG correction. Data for (c) 12 piercing clonal types from Experiment 1 and (d) three piercing clonal types and their incorrect subvariants from four replicate samples from Experiment 2 are shown.
[0090] Supersequencing histograms of unique molecular recognizers. X-axis: MIG size. Y-axis: Percentage of total reads represented by MIGs containing a given number of reads. The histograms for Experiment 1 depict B cells sequenced for IGH and T cells sequenced for TRA and TRB. b. The geometric mean per MIG for IGH, TRA, and TRB are 122, 244, and 186 reads, respectively. The histograms for Experiment 2 depict four replicates, each an independently synthesized IGH cDNA sequenced from mixed donor blood with spike clones. The geometric mean per MIG for each sample SI-4 is 15, 14, 18, and 17 reads, respectively.
[0091] A graph-based representation of MIGEC analysis of inserted clones. Central nodes represent peak clonal types, while gray represents erroneous clonal types differing from their parent clonal type by one mismatch. Node regions are scaled to indicate read sharing of clonal types. Standard analysis (extraction of CDR3 sequences with Phred > 25), low-frequency variant filtering (less than 30 reads for each clonal type in Experiments 1 and 10 reads in Experiments 2), first-stage MIG correction, and full MIG correction are displayed. IGH, TRA, and TRB clonal types in Experiment 1 and EHEB clonal type in Experiment 2 are shown.
[0092] MIG filtering in Experiment 2. Spike clonal types. MIG-guided data processing eliminates errors while preserving the correct rare subvariants. Subplots of EHEB gene clonal types are shown for standard data processing (left) and MIGEC (right), with subvariants having one (EHEB-V1) and two (EHEB-V2) mismatches and one mismatch error. Note that PCR errors cause bridges to form between IGH-CDR3 variants in the conventionally processed data, thus linking the EHEB-V2 insertion signal with two mismatches to the EHEB insertion signal through an incorrectly linked subvariant. MIGEC has successfully eliminated this type of error that complicates the analysis of hypervariable trees.
[0093] Characterization of nucleotide substitutions in MIGEC-filtered human blood IGH-CDR3 library sequencing data. Several representative MIGEC-filtered hypervariable maps of medium size are shown. CDR3 amino acid sequences for clonal variants are provided, with node size reflecting the total number of reads for a given clonal variant. Dashed and solid line edges represent a mismatched nucleotide substitution, respectively, for silencing and substitution. Specific Implementation Example 8:
[0095] The essence of this method is that high-frequency (hotspot) PCR and sequencing errors remain after the first stage of MIG analysis (e.g., "Safe-SeqS"). These errors are reproducible across MIGs and therefore exist as minor sequences. Successful correction is achieved across multiple MIGs. Therefore, such hotspot errors can be identified by analyzing MIG sets—either by counting the frequency of sequencing reads containing specific nucleotide mismatches within a MIG, or by counting the frequency of MIGs containing smaller sequence subvariants with specific nucleotide mismatches.
[0096] For example, to detect frequent PCR errors, the number of reads for each sequence in the initial data can be compared to the output of the first-stage MIG analysis. High-frequency PCR errors introduced in later PCR cycles should already be present and corrected in multiple MIGs, resulting in a reduction in the total number of such reads in the output data. Indeed, we observed erroneous subvariants of the stinging sequence variants (see Example 1), with an average presence level 50% lower than in the input. Conversely, the correct stinging sequence variants always “gain” additional reads after the first-stage analysis if erroneous subvariants corrected in their MIGs are taken into account.
[0097] Alternatively, the relative number of MIGs containing each sequence variant as a minor sequence subvariant can be calculated, along with the relative number of MIGs containing the same sequence variant as the major sequence subvariant. For hotspot errors, the frequency of the preceding MIG should be higher than that of the following MIG (at least 4 times higher, typically at least 10 times higher, and in most cases at least 50 times higher). For truly diverse variants, the frequency of the preceding MIG should be comparable to or lower than that of the following MIG.
[0098] In immunology, high-throughput sequencing of T-cell receptor (TCR) and antibody (IG) libraries can track specific clonoids in tissue samples and span the entire timeframe. However, this process particularly limits the direct and definitive analysis of IG hypermutated chains. Hypermutation naturally produces nearly identical sequence chains, resulting in clonoids with potentially drastically different concentrations, making it impossible to distinguish between minor and major clonoid-derived erroneous variants.
[0099] Our method effectively eliminates a large number of errors that occur during PCR amplification and high-throughput sequencing, as shown in the following two examples.
[0100] example
[0101] Example 1. In Example 1 (Experiment 2), three cDNA libraries were prepared and analyzed using Illumina MiSeq paired-end sequencing: 1) two known B cell clones sequenced for the IG heavy chain (IGH) complementarity-determining region 3 (CDR3) and five known T cell clones mixed in defined proportions, sequenced for the 2) TCRα (TRA) and 3) TCRβ (TRB) chains CDR3. Libraries were generated using a template switching scheme 8'17, with unique molecular identifiers introduced within the cDNA synthesis adapter 5'18. PCR amplification using low initial amounts of cDNA allowed for significant hypersequencing of the template molecules—the mean number of reads labeled with the same identifier was greater than 100 in all three samples (Table 3).
[0102] The two-stage MIG-based data processing eliminated almost all the artificial diversity contained in the raw data. On average, high-quality raw sequencing data (Phred>25 threshold for each nucleotide in CDR3) contained 5% erroneous reads, producing -100 erroneous variants for each spiked CDR3 clonal type. In contrast, the filtered data showed no erroneous variants in IGH and TRA samples, and for one of the five TRB CDR3 clonal types, only two erroneous variants with 0.05% reads were retained. Meanwhile, most sequencing information was preserved. Compared to standard analysis with Phred>25, the first-stage processing resulted in a 6% read loss (530 million and 5 million CDR3-containing reads extracted, respectively), and a <1% read loss between the first and second stages.
[0103] Importantly, we do not lose small homologous TCR or IG clonoids if they are present in the sample, because correction does not rely on blind frequency-based clustering based on homologous cDNA variants. This is particularly important for deep error-free IG analysis, as hypermutations often result in homologous IG variants that will be lost when error correction algorithms are applied based on merging / eliminating single-precision and double-mismatched clonoids (a method typically used in effective T-cell receptor library analysis).
[0104] Example 2.
[0105] To demonstrate that MIGEC has sufficient resolution to reveal deep, error-free IG analysis of hypermutant subvariants, we analyzed an IGH library from healthy human donor blood samples (Experiment 2). A plasmid encoding EHEB IGH and two artificially generated subvariants with one and two known substitutions in CDR3 (named EHEB-VI and -V2, respectively) were added as spiked controls. To achieve deep sequencing of the library, we used more cDNA molecules initially and sequenced them on a high-throughput Illumina HiSeq 2500 platform. We performed four replicates of this experiment, obtaining highly reproducible statistical data demonstrating the robustness of the protocol (Table 3). In particular,
[0106] Almost all erroneous readings and most of the artificial diversity of the inserted clone EHEB were shared between replicates in Experiment 2 and between independent experiments (Experiment 1 vs. Experiment 2). Exp2_r1-4 started from independent cDNA synthesis, corresponding to the four replicates of Experiment 2.
[0107] Since the above description is only a specific embodiment of the present invention, the protection of the present invention is not limited thereto. Any equivalent changes or substitutions of the technical features of the present invention that can be conceived by those skilled in the art are covered within the protection scope of the present invention.
Claims
1. A method for identifying hotspot PRCs and sequencing errors in high-throughput sequencing data, characterized in that: The identification method steps are as follows: Step 1: Data acquisition step, obtain raw sequencing data; Step 2: Data grouping step, grouping sequencing reads belonging to the same starting molecule into sequencing read groups; Step 3: Data analysis steps, including two-level correction steps, or: steps to identify and correct sequencing errors; The two-stage correction steps consist of a first stage of MIG correction and a second stage of MIG correction. The first stage of MIG correction involves identifying dominant sequence variants within a set of sequencing reads, filtering out those sequencing read sets where dominant sequence variants are represented by a single set of sequencing reads, and filtering out sequencing read sets where no single dominant sequence is observed. The second stage of MIG correction involves counting the frequency of sequencing reads containing specific nucleotide mismatches within the sequencing read set, and comparing the abundance of each specific nucleotide mismatch before and after the first stage of MIG correction and the second stage of MIG correction. Specific nucleotide mismatches with a reduced relative frequency in the first stage of MIG correction are considered high-frequency, hotspot PCR, or sequencing errors. Sequencing read sets containing the identified high-frequency or hotspot errors are discarded or appropriately corrected in the second stage of MIG correction. Alternatively: The two-stage correction steps are a first stage of MIG correction and a second stage of MIG correction. The first stage of MIG correction is to identify dominant sequence variants within a set of sequencing reads, filtering out those sequencing read groups where dominant sequence variants are represented by a single set of sequencing reads, and filtering out sequencing read groups where no single dominant sequence is observed. The second stage of MIG correction is to count the relative number of sequencing read groups containing a small number of sequencing reads with specific nucleotide mismatches, and to count those sequencing read groups containing the same sequence and having the same nucleotide sequence. Specific nucleotide mismatches after the data grouping step are used as the primary sequencing reads. For high-frequency, hotspot errors, the frequency of the previous set of sequencing reads should be higher than the frequency of the next set of sequencing reads. Sequencing read groups containing the identified high-frequency, hotspot errors are discarded or appropriately corrected in the second stage of MIG correction. The step of identifying and correcting sequencing errors involves identifying and correcting high-frequency hotspot PCR and sequencing errors. Identifying and correcting high-frequency hotspot PCR and sequencing errors involves counting the frequency of sequencing reads containing specific nucleotide mismatches within the sequencing read array. Specific nucleotide mismatches with a reduced relative frequency are considered high-frequency, hotspot PCR, or sequencing errors. Sequencing read arrays containing identified high-frequency, hotspot errors are discarded or appropriately corrected in the step of identifying and correcting sequencing errors. Alternatively: The step of identifying and correcting sequencing errors involves identifying and correcting high-frequency hotspot PCR and sequencing errors, counting the relative number of sequencing read groups containing a small number of sequencing reads with specific nucleotide mismatches, and counting those groups containing the same sequence and the same nucleotide sequence. Specific nucleotide mismatches are read as the primary sequencing reads after the data grouping step. For high-frequency, hotspot errors, the frequency of the previous group of sequencing reads should be higher than the frequency of the next group of sequencing reads. Sequencing read groups containing identified high-frequency, hotspot errors in variants will be discarded or appropriately corrected in the step of identifying and correcting sequencing errors.
2. The method for identifying hotspot PRCs and sequencing errors in high-throughput sequencing data according to claim 1, characterized in that: The sequencing read set consists of sequencing reads, each marked with a unique molecular identifier, and the sequence contains at least one degenerate or semi-degenerate nucleic acid sequence; the sequencing reads are located at the same starting nucleotide position.
3. The method for identifying hotspot PRCs and sequencing errors in high-throughput sequencing data according to claim 2, characterized in that: The sequencing read analysis described herein is for antibody gene library profiling, T cell receptor gene library analysis, and genomic library analysis based on high-throughput sequencing.
Citation Information
Patent Citations
Method for reducing deep sequencing errors
CN108359723A
Methods for targeted nucleic acid sequence enrichment with applications to error corrected nucleic acid sequencing
CN110520542A