Mutation site screening method based on amplicon depth sequencing and sequence counting
Through the method of amplicon deep sequencing and sequence counting, the problems of high false positive rate and insufficient detection accuracy in mutant screening in polyploid species were solved, and high-throughput and high-efficiency mutation site positioning and accurate screening were achieved.
Patent Information
- Application Number
- CN202510769337.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing mutant screening technologies have a high false positive rate in polyploid species and rely on reference genomes, resulting in insufficient detection accuracy. In particular, short-read sequencing data in polyploid species cannot be accurately compared.
The method of amplicon deep sequencing and sequence counting is adopted to simplify the experimental process, improve the detection throughput and sensitivity, and reduce the risk of cross-contamination through mixed sample pool construction, tag amplification, deep sequencing, frequency threshold screening and directional monomer verification.
It achieves high-throughput and high-efficiency mutation site positioning, shortens the screening cycle and reduces experimental costs, and improves the accuracy and sensitivity of mutation detection.
Smart Images

Figure CN120683235A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mutant screening, and in particular to a method for screening mutation sites based on amplicon deep sequencing and sequence counting. Background Art
[0002] Mutant screening technology is widely used in fields such as genetics, functional genomics, and crop breeding. Generally speaking, mutant screening technology can be divided into forward and reverse techniques. Forward screening technology is also called phenotypic screening, which is to discover mutants with specific phenotypes through morphological observation and determination of physiological and biochemical indicators. Therefore, it is often used in the screening of mutants with simple and easy-to-observe phenotypes, such as screening for dwarfing, leaf color variation, etc. in plants; reverse screening technology is also called genotypic screening, which is to discover mutants with specific genotypes by genotyping the genome or candidate genes. Reverse screening technology is more often used in the screening of mutants whose phenotypes are difficult to identify, or whose phenotypes are easily affected by the environment or difficult to detect, such as plant metabolic mutations, animal genetic diseases and other phenotypes.
[0003] With the emergence of high-throughput sequencing technology, especially the significant reduction in sequencing costs, whole genome sequencing or target gene amplicon sequencing has been increasingly used in mutant screening due to its high throughput, high sensitivity, and high accuracy. The typical experimental process is as follows:
[0004] 1) Individual genome resequencing (or target gene amplicon sequencing);
[0005] 2) Sequencing read quality control;
[0006] 3) The quality-controlled reads are then mapped back to the genome;
[0007] 4) Mutation detection.
[0008] It can be seen that whether screening mutants through genome resequencing or amplicon sequencing, the sequencing reads are back-linked to the reference genome before detecting variants, which has certain limitations or drawbacks. First, it relies on the reference genome, and the reference genome of some species is unknown. Second, existing methods rely on genome back-linking, but the similarity of multiple alleles in polyploid species often exceeds 95% (for example, the similarity of the FAD2 gene in hexaploid sea kale is the highest, exceeding 95%), resulting in the inability to accurately align short-read sequencing data to the correct allele. Traditional threshold setting methods (such as a fixed frequency threshold of 0.01) have a false positive rate of up to 15%-20% in polyploids. Summary of the Invention
[0009] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a mutation site screening method based on amplicon deep sequencing and sequence counting, which has the characteristics of high detection throughput, simple sequence analysis method, and sensitive and reliable mutation detection, and is suitable for screening gene mutants in different types of species.
[0010] To achieve the above object, the present invention provides the following solutions:
[0011] A method for screening mutation sites based on amplicon deep sequencing and sequence counting, comprising:
[0012] Mixing equal amounts of samples from multiple individuals to be screened at a ratio of 10-20 to construct a mixed sample pool, and extracting genomic DNA from each of the mixed sample pools;
[0013] Design primers based on the target gene sequence and add a tag sequence to the 5' end of the primer, perform polymerase chain reaction amplification using the genomic DNA as a template to obtain the target amplified fragment, and perform deep sequencing on the target amplified fragment using a high-throughput sequencing platform to obtain raw sequencing data;
[0014] Cleaning, removing redundancy, and screening the raw sequencing data to obtain candidate sequences;
[0015] Compare the candidate sequences whose expression numbers are within the set frequency range with the target gene sequence to identify whether there are mutation sites and obtain the comparison results;
[0016] If the comparison result shows the presence of a mutation site, the individual samples in the corresponding mixed sample pool are amplified and sequenced respectively to determine the individual samples that specifically contain the mutation.
[0017] Preferably, samples from multiple individuals to be screened are mixed in equal amounts at a ratio of 10-20 to construct a mixed sample pool, and genomic DNA is extracted from each of the mixed sample pools, comprising:
[0018] Randomly selecting a plurality of individuals to be screened from a sample population to be screened, and obtaining tissue materials of each of the individuals to be screened;
[0019] Lysing and purifying the tissue material, and extracting genomic DNA from each individual to be screened;
[0020] Weighing and mixing equal amounts of genomic DNA solutions from each individual, with ten to twenty individuals forming a group, to construct a plurality of mixed sample pools;
[0021] The mixed sample pools are numbered and saved respectively.
[0022] Preferably, primers are designed according to the target gene sequence and a tag sequence is added to the 5' end of the primer. Polymerase chain reaction amplification is performed using the genomic DNA as a template to obtain the target amplified fragment, and the target amplified fragment is deeply sequenced using a high-throughput sequencing platform to obtain raw sequencing data, including:
[0023] Designing a pair of specific primers based on the conserved region sequence of the target gene, and adding a tag sequence for distinguishing different mixed sample pools to the 5' end of each specific primer;
[0024] Using the genomic DNA in the mixed sample pool as a template, polymerase chain reaction technology is used to amplify a target gene fragment to obtain a target amplified fragment; the amplification conditions are set according to the characteristics of the primers;
[0025] The target amplified fragment is purified, and the purified fragment is subjected to single-end or double-end sequencing using the high-throughput sequencing platform to obtain the raw sequencing data.
[0026] Preferably, the properties of the primers include denaturation temperature, annealing temperature and extension time.
[0027] Preferably, the raw sequencing data is cleaned, redundancy removed, and screened to obtain candidate sequences, including:
[0028] Performing quality control on the raw sequencing data, removing sequencing adapters and primer index sequences using a preset quality control tool, and removing reads containing chimeric structures using a chimera identification algorithm to obtain valid reads after quality control;
[0029] Performing a redundancy removal operation on the valid read segments, extracting a set of non-repeated sequences to obtain a set of unique sequences, and counting the number of times each unique sequence appears in the sample pool as depth information of the unique sequence;
[0030] Calculate the expression frequency in the mixed sample pool based on the depth information of each unique sequence and the total depth information of all unique sequences;
[0031] Calculating the lower limit and upper limit of frequency screening based on the ploidy characteristics of the individuals to be screened constituting the mixed sample pool and the individual mixing multiples adopted;
[0032] The unique sequence whose expression frequency is within the range of frequency screening is screened out as the candidate sequence.
[0033] Preferably, the lower limit value of the frequency screening is a value obtained by dividing the product of the ploidy and the individual mixing multiple by one.
[0034] Preferably, the upper limit value of the frequency screening is ten times the lower limit value.
[0035] Preferably, the candidate sequences whose expression amount is within the set frequency range are compared with the target gene sequence to identify whether there is a mutation site, and the comparison results are obtained, including:
[0036] Importing the set of candidate sequences and the corresponding target gene sequence into a preset sequence alignment tool, and designating the target gene sequence as the reference sequence;
[0037] Setting alignment parameters and performing sequence alignment to generate an alignment document of the candidate sequence and the reference sequence; the alignment parameters include a match score, a mismatch penalty, and an insertion-deletion overhead;
[0038] Searching the alignment file for mismatches, insertions, or deletions at the coordinates of the candidate sequence and the reference sequence, marking the alignment coordinates corresponding to the mismatches, insertions, or deletions as differential sites, and recording the alignment results; the alignment results include: the sequence position of the differential site, the type of difference, and depth information of the candidate sequence;
[0039] Determine whether the expression frequency corresponding to the differential site is still within the frequency screening range. If so, mark the differential site as a mutation site, and output the comparison result including a mutation site list.
[0040] Preferably, if the comparison result indicates the presence of a mutation site, the individual samples in the corresponding mixed sample pool are amplified and sequenced separately to determine the specific individual samples containing the mutation, including:
[0041] Sampling each individual to be screened in the mixed sample pool where the comparison results show the presence of a mutation site, and extracting genomic DNA of the individual to be screened separately;
[0042] Using the same pair of primers with the tag sequence as a basic template, polymerase chain reaction is performed on the genomic DNA of each individual to be screened to amplify the target gene fragment to obtain an amplified product;
[0043] Purifying the amplified products and sequencing them using a first-generation sequencing method to obtain the sequencing sequence of each individual to be screened;
[0044] The sequencing sequence of each individual to be screened is compared with the reference sequence of the target gene to check whether it contains base differences consistent with the determined mutation site;
[0045] The individuals to be screened that are determined to contain the mutation site in the comparison results are marked as individual samples containing mutations, and an identification list of individual samples containing mutations is output.
[0046] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0047] The present invention integrates and simplifies the time-consuming steps of single-strain library construction and whole-genome alignment in traditional mutant screening through a closed-loop process of "mixed pool construction-label amplification-deep sequencing-frequency threshold screening-directional monomer verification": ① First, 10-20 times of individual samples are mixed in equal amounts and genomic DNA is extracted uniformly, which can simultaneously detect a large number of samples in a single sequencing run, significantly improving the screening throughput; ② After embedding the label at the 5' end of the primer, target amplification is performed, which not only ensures the specificity of the amplicon, but also can visually distinguish different mixed pools during the sequencing stage, effectively reducing the risk of cross-contamination; ③ Using the screening rules of upper and lower frequency limits, only unique sequences with expression levels in the theoretical range are retained for comparison, which greatly reduces redundant data, suppresses sequencing noise, and improves the sensitivity of mutation detection; ④ The mixed pools with positive comparison results are subjected to monomer verification again, which can lock in the individuals with true mutations at one time and avoid the waste of repeated sequencing. Overall, this method achieves high-accuracy and high-efficiency mutation site positioning with a minimum of experimental batches, shortens the screening cycle and reduces experimental costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 A flow chart of a method provided by an embodiment of the present invention;
[0050] Figure 2 A schematic diagram of the CaFAD2 amplicon reference sequence and primer positions provided in an embodiment of the present invention;
[0051] Figure 3 A schematic diagram of mutant sites in sample pool P1 provided in an embodiment of the present invention;
[0052] Figure 4 A schematic diagram of mutant sites in sample pool P6 provided in an embodiment of the present invention;
[0053] Figure 5 Schematic diagram of the rdl gene mutation sites in sample pools Rdl-P1 and Rdl-P3 provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0055] The purpose of the present invention is to provide a mutation site screening method based on amplicon deep sequencing and sequence counting, which has the characteristics of high detection throughput, simple sequence analysis method, and sensitive and reliable mutation detection, and is suitable for screening gene mutants of different types of species.
[0056] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0057] Figure 1 A flow chart of the method provided in the embodiment of the present invention is shown in FIG. Figure 1 As shown, the present invention provides a method for screening mutation sites based on amplicon deep sequencing and sequence counting, comprising:
[0058] Step 100: Mixing equal amounts of samples from multiple individuals to be screened at a ratio of 10-20 to construct a mixed sample pool, and extracting genomic DNA from each mixed sample pool;
[0059] Step 200: Design primers based on the target gene sequence and add a tag sequence to the 5' end of the primer. Perform polymerase chain reaction amplification using genomic DNA as a template to obtain the target amplified fragment. The target amplified fragment is deeply sequenced using a high-throughput sequencing platform to obtain raw sequencing data.
[0060] Step 300: Cleaning, removing redundancy, and screening the raw sequencing data to obtain candidate sequences;
[0061] Step 400: comparing candidate sequences whose expression levels are within a set frequency range with the target gene sequence to identify whether a mutation site exists and obtain a comparison result;
[0062] Step 500: If the comparison result shows the presence of a mutation site, the individual samples in the corresponding mixed sample pool are amplified and sequenced separately to determine the individual samples containing the mutation.
[0063] The present invention provides an analysis method based on deep amplicon sequencing combined with non-genomic response of sequencing reads for high-throughput screening of mutants of a target gene. The method first mixes a sample population to be screened to form a pool (sample pool) to improve the detection throughput, then obtains an amplicon by amplifying the target gene in each sample pool and performs deep sequencing, then removes redundancy from the reads sequenced in each mixed pool to obtain unique reads, and calculates the frequency of each unique read in the mixed pool sequencing by calculating the number of repetitions (depth) of each unique read. The theoretical frequency of mutant reads in the mixed pool is then used as a critical value to further screen the unique reads, and a small number of unique reads within the critical value range are screened out for comparison with the target gene sequence. If any unique read contains a mutation site, it indicates that the mixed pool contains a mutant, and each individual in the pool can be further verified by first-generation sequencing to identify the specific mutant. Otherwise, it indicates that no mutant individual exists in the mixed pool.
[0064] Furthermore, the technical process of this embodiment is as follows:
[0065] 1. Sample mixing
[0066] The sample population to be screened is mixed in a 10-20-fold ratio to construct a pool. Specifically, 10-20 individuals are mixed in equal amounts to construct a series of sample pools, and genomic DNA is extracted for subsequent mutant screening.
[0067] 2. Amplicon Sequencing
[0068] Specific amplification primers are designed based on the target gene sequence. A 6-base index sequence is added to the 5' end of each primer to facilitate subsequent differentiation between sample pools. PCR amplification is performed using the sample pool DNA as a template. PCR amplification conditions are adjusted based on primer characteristics (such as annealing temperature and product length). The resulting amplicons are deep sequenced on a next-generation sequencing platform, with a sequencing depth of at least 100 for each target gene allele.
[0069] The formula for calculating the number of target gene alleles is: Number of alleles = Sample ploidy * 2 * Pooling ratio (Formula 1). For example, after 10x pooling of a hexaploid species, the number of target gene alleles in each pool is: 6 * 2 * 10 = 120. Therefore, the sequencing depth should be at least 12,000 reads.
[0070] 3. Amplicon sequence cleaning and de-redundancy
[0071] Sequencing data was segmented into different sample pool amplicon sequences based on the tag sequence, and each mixed pool amplicon was analyzed separately. The specific analysis steps included the following: 1) Quality control: The raw reads were quality controlled using the fastp tool, the sequencing adapters and tag sequences were removed, and the sequencing reads containing chimeras were removed using the uchime2_denovo tool in the vsearch software to obtain clean reads; 2) Unique sequence acquisition: The reads were de-redundant using the fastx_uniques tool in the vsearch software to obtain unique reads and the number of each unique read (parameters were --minuniquesize 1 --strand both --sizeout); 3) Unique depth calculation and filtering: A shell script was written to calculate the total depth of unique reads in each sample pool and the frequency of each unique read. The calculation formula is:
[0072] Frequency (P) = Nu / Nt (Formula 2)
[0073] Here Nu is the depth of each unique read, and Nt is the total depth of all unique reads.
[0074] 1. Filter unique reads
[0075] Set the depth threshold. The calculation formula for the threshold is:
[0076] Pmin = 1 / (k × 2 × n) (Formula 3)
[0077] Here k is the species ploidy and n is the depth of the mixing pool.
[0078] Pmax = Pmin × 10 (Formula 4)
[0079] Unique reads were filtered, retaining only those with a depth within a critical range for subsequent sequence alignment analysis, i.e., Pmin <= P <= Pmax. For example, for a 10-fold pool of hexaploid species, Pmin = 1 / (6 × 2 × 10) = 0.008, and Pmax = 0.08, meaning only unique reads with a frequency range of 0.008-0.08 were retained for subsequent analysis.
[0080] 2. Unique read sequence alignment to screen candidate mutation sites
[0081] MEGA11 was used to align the above unique reads with the target gene reference sequence. If there was a mutation site in the unique reads, it indicated that the sample pool contained a mutant with the mutation site. Sanger sequencing was then used to sequence the target gene of each individual in the mixed pool, and the specific mutant was finally identified. If there was no mutation site in the unique reads, it indicated that all samples in the mixed pool were wild-type and no further identification was performed.
[0082] As an optional embodiment, the steps of the technical route for screening FAD2 gene mutants from the hexaploid species Crambe oleracea in this example are as follows:
[0083] 1. Sample mixing
[0084] Sixty M2 plants were randomly selected from a Crambe oleracea EMS-induced mutant population. Young leaves were harvested after sowing and genomic DNA was extracted. DNA from every ten plants was mixed in equal molar amounts to construct six 10-fold DNA pools, numbered P1-P6.
[0085] 2. Amplicon amplification and sequencing
[0086] Because crambe is a hexaploid species, theoretically there are 6 FAD2 alleles in the homozygous genome. After comparing multiple FAD2 gene sequences, the present invention designed a pair of primers in the sequence conserved region for PCR amplification (the primer specific part is GGGTCTGCCAAGGCTGTGTCC; CACGGTGCGTCCAAGAGGGT, see Figure 2 , yellow indicates primer positions). A 6-base index sequence was added to the 5' end of each primer for subsequent differentiation between sample pools. PCR amplification was performed using DNA from each of the six sample pools as templates. A 20-μl PCR reaction system included 4 μl of Phusion HF buffer (NEB, USA), 0.2 μl of 10 mM dNTPs, 0.5 μl of each forward and reverse primer (10 μM), 0.1 μl of Phusion DNA polymerase (2 U / μl), and 1 μl of template DNA (100 ng). PCR reaction conditions were: initial denaturation at 98°C for 30 seconds, followed by 30 cycles of 98°C for 10 seconds, 58°C for 10 seconds, and 72°C for 30 seconds, followed by a final extension at 72°C for 5 minutes. The amplified product, approximately 270 bp in length, was purified from an agarose gel and sequenced on an Illumina HiSeq 2000 platform (100 bp per direction).
[0087] 3. Amplicon sequence redundancy removal
[0088] Sequencing data were segmented into six different pool amplicon sequences based on the index sequence, and each pool amplicon was analyzed separately. The analysis included the following steps: 1) Quality Control: The raw reads were quality-controlled using the fastp tool, followed by excision of the sequencing adapters and index sequences. Sequencing reads containing chimeras were then removed using the uchime2_denovo tool in the vsearch software to obtain clean reads. 2) Unique Sequence Acquisition: The reads were de-redundant using the fastx_uniques tool in the vsearch software to obtain unique reads and the number of unique reads per strand (with the following parameters: --minuniquesize 1 --strand both --sizeout). 3) Unique Depth Calculation and Filtering: A shell script was used to calculate the total depth of unique reads in each pool and the frequency of each unique read. The frequency of each unique read was then calculated using Equation 2.
[0089] 4. Algorithm demonstration of theoretical minimum frequency Pmin and maximum frequency Pmax
[0090] 4.1 Algorithm demonstration of the theoretical minimum frequency Pmin
[0091] According to Formula 3, Pmin = 1 / (k × 2 × n) (Formula 3), hexaploid species are mixed 10 times, Pmin = 1 / (6 × 2 × 10) = 0.008
[0092] 4.2 Algorithm demonstration of the theoretical maximum frequency Pmax
[0093] Assume that sequencing errors follow a Poisson distribution (common in high-throughput sequencing) and that the true mutation frequency should be significantly higher than the background error rate (usually the Illumina sequencing error rate is ≤0.5%. To minimize false positives, the sequencing error rate is assumed to be 1%).
[0094] Pmax derivation process:
[0095] If the sequencing error rate is λ = 0.01, when the coverage D = 100, the expected value of error reads is λD = 1. According to the Poisson distribution theory, the upper limit of the 99.9% confidence interval is:
[0096] Pnoise=(λ+3×square root(λ)) / D=(0.01+3×1) / 100=0.0301
[0097] The frequency of true mutations must be significantly higher than the noise, Pmax = C × Pmin >> Pnoise, that is, for the hexaploid pool (Pmin = 0.008), if: Pmax = C × Pmin >> Pnoise, substituting Pmin = 0.008, when C = 10: Pmax = 0.08 >> 0.031 (that is, the signal-to-noise ratio of 0.08 / 0.031>20); when C = 5: Pmax = 0.04 (that is, the signal-to-noise ratio of 0.04 / 0.031 is approximately equal to 1). The standard signal-to-noise ratio in second-generation sequencing requires more than 10, so C = 10 is the best, and at this time Pmax = 10 × Pmin.
[0098] 5. Filter unique reads
[0099] The critical value of the unique read depth is calculated according to formulas 3 and 4, and then the unique reads are filtered according to the frequency of each unique read, and only unique reads with a depth within the critical value range are retained for subsequent sequence alignment analysis. The critical value range calculated in this example is [0.008, 0.08], and the specific calculation process is:
[0100] Pmin=1 / (6×2×10)=0.008
[0101] Pmax=0.008×10=0.08
[0102] Unique reads with a depth between 0.008 and 0.08 of the total depth were retained for sequence comparison analysis. The depth of unique sequences retained in each pool is shown in Table 1. Among them, the minimum depth of unique reads retained after filtering for P1 was 2 and the maximum was 22; the minimum depth of unique reads retained after filtering for P2 was 20 and the maximum was 208; the minimum depth of unique reads retained after filtering for P3 was 90 and the maximum was 989; the minimum depth of unique reads retained after filtering for P4 was 73 and the maximum was 736; the minimum depth of unique reads retained after filtering for P5 was 21 and the maximum was 211; and the minimum depth of unique reads retained after filtering for P6 was 62 and the maximum was 622.
[0103] Table 1 Filtering results of sequencing data of 6 sample pools
[0104]
[0105] 6. Unique reads sequence alignment to screen mutation sites
[0106] The unique reads filtered out from each sample pool were aligned with the target gene sequences of Crambe (CraFAD2-A, CraFAD2-B, CraFAD2-C1, CraFAD2-C2, CraFAD2-C3). It was found that sample pools P1 and P3 each contained a mutation site of one gene. Both mutations were located in the gene CraFAD2-C3. The unique read depth of the mutation in P1 was 2 ( Figure 3 ), the unique read depth of the mutation in P6 is 118 ( Figure 4 ), the mutation types are C>T and G>A, which are typical EMS-induced mutation types ( Figure 3 、 Figure 4 ). Figure 3 The mutant sites in sample pool P1 are shown (mutation type is C>T, unique reads depth size = 2). Figure 4 The mutant sites in sample pool P6 are shown (mutation type is G>A, unique reads depth size = 118).
[0107] As another optional embodiment, this example screens for rdl gene mutations in Cape Verde mosquitoes by targeted amplicon sequencing. The raw sequencing data used in this example comes from the literature (Collins EL, Phelan JE, HubnerM, Spadar A, Campos M, Ward D, et al. (2022) A next generation targeted amplicon sequencing method to screen for insecticide resistance mutations in Aedesaegypti populations reveals a rdl mutation in mosquitoes from Cabo Verde. PLoSNegl Trop Dis 16(12): e0010935. https: / / doi.org / 10.1371 / journal.pntd.0010935), and the data can be downloaded from the ENA database (ERS12467832-ERS12467834).
[0108] The specific implementation plan is as follows:
[0109] 1. Experimental Materials
[0110] The original study involved 152 Aedes aegypti mosquitoes from Cape Verde. This protocol selected 60 A. aegypti mosquitoes for rdl locus amplicon analysis. Genomic DNA was extracted from each mosquito individually for PCR amplification.
[0111] 2. Amplicon amplification and sequencing
[0112] A pair of primers was designed based on the rdl gene reference sequence for PCR amplification (CCAACCGATGTATCTTCTTC; CTGGTTATTTGTACAAGTAGCA, product size 498 bp). PCR amplification was performed using DNA from each individual sample as a template. A 20-μl PCR reaction system contained 4 μl of Q5 High-Fidelity PCR reagent (New England Biolabs, UK), 0.2 μl of 10 mM dNTPs, 0.5 μl of each forward and reverse primer (10 μM), and 4 μl of template DNA (100 ng). PCR reaction conditions were: 98°C for 30 s, followed by 35 cycles of 98°C for 10 s, 60.4°C for 70 s, and 72°C for 90 s, with a final extension at 72°C for 5 min. PCR products were purified using AMPure XP magnetic beads (Beckman Coulter). All PCR products were then diluted and mixed at equal concentrations to create pools with a total volume of 25 μl and a concentration of 20 ng / μl. Each pool contained 20 amplicons (1 amplicon × 20 mosquitoes / pool). A second round of PCR was then performed to add Illumina adapters and indices, allowing multiple pools to be sequenced in parallel in the same Illumina sequencing run.
[0113] 3. Amplicon sequence redundancy removal
[0114] Sequencing data were segmented into three different pool amplicon sequences based on the index sequence, and each pool amplicon was analyzed separately. The analysis included the following steps: 1) Quality Control: The raw reads were quality-controlled using the fastp tool, followed by excision of the sequencing adapters and index sequences. Sequencing reads containing chimeras were then removed using the uchime2_denovo tool in the vsearch software to obtain clean reads. The sequencing yield and theoretical depth for each pool are shown in Table 1. 2) Unique Sequence Acquisition: Reads were de-redundant using the fastx_uniques tool in the vsearch software to obtain unique reads and the number of unique reads per strand (with the parameters --minuniquesize 1 --strandboth --sizeout). 3) Unique Depth Calculation and Filtering: A shell script was used to calculate the total depth of unique reads and the frequency of each unique read in each pool. The frequency of each unique read was calculated according to Equation 2.
[0115] 4. Filter unique reads
[0116] The critical value of the unique read depth is calculated according to formulas 3 and 4, and the unique reads are filtered, and only unique reads with a depth within the critical value range are retained for subsequent sequence alignment analysis. The critical value range calculated in this embodiment is [0.025, 0.25]. The specific calculation process is:
[0117] Pmin=1 / (1×2×20)=0.025
[0118] Pmax=0.025×10=0.25
[0119] Unique reads with a depth between 0.025 and 0.25 of the total depth were retained for sequence comparison analysis. The depth of unique sequences retained in each pool is shown in Table 2. The minimum depth of unique reads retained after filtering for Rdl-P1 was 16 and the maximum was 160; the minimum depth of unique reads retained after filtering for Rdl-P2 was 4 and the maximum was 44; and the minimum depth of unique reads retained after filtering for Rdl-P3 was 1 and the maximum was 4.
[0120] Table 2 Filtering results of sequencing data of 3 sample pools
[0121]
[0122]
[0123] 5. Unique reads sequence alignment to screen mutation sites
[0124] The unique reads filtered out from each sample pool were aligned with the reference sequence (rdl_reference) of the target gene rdl. The same mutation site was found in both sample pools (Rdl-P1 and Rdl-P3). The mutation type was G>T. The mutation occurred at the 886th base of the reference sequence, resulting in the amino acid mutation A296S. This mutation site has been confirmed in the literature (Collins EL, 2022). The depth of unique reads for the representative sequence of this mutation in Rdl-P1 was 117, and the depth of unique reads for the representative sequence of this mutation in Rdl-P1 was 4 and 1 respectively. Figure 5 ). Figure 5 Shows the rdl gene mutation site (G>T) in sample pools Rdl-P1 and Rdl-P3
[0125] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0126] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.
Claims
1. A method for screening mutation sites based on amplicon deep sequencing and sequence counting, characterized in that: include: Mixing equal amounts of samples from multiple individuals to be screened at a ratio of 10-20 to construct a mixed sample pool, and extracting genomic DNA from each of the mixed sample pools; Design primers based on the target gene sequence and add a tag sequence to the 5' end of the primer, perform polymerase chain reaction amplification using the genomic DNA as a template to obtain the target amplified fragment, and perform deep sequencing on the target amplified fragment using a high-throughput sequencing platform to obtain raw sequencing data; Cleaning, removing redundancy, and screening the raw sequencing data to obtain candidate sequences; Compare the candidate sequences whose expression numbers are within the set frequency range with the target gene sequence to identify whether there are mutation sites and obtain the comparison results; If the comparison result shows the presence of a mutation site, the individual samples in the corresponding mixed sample pool are amplified and sequenced respectively to determine the individual samples that specifically contain the mutation.
2. The method for screening mutation sites based on amplicon deep sequencing and sequence counting according to claim 1, characterized in that: Mixing equal amounts of samples from multiple individuals to be screened at a ratio of 10-20 to construct a mixed sample pool, and extracting genomic DNA from each of the mixed sample pools, including: Randomly selecting a plurality of individuals to be screened from a sample population to be screened, and obtaining tissue materials of each of the individuals to be screened; Lysing and purifying the tissue material, and extracting genomic DNA from each individual to be screened; Weighing and mixing equal amounts of genomic DNA solutions from each individual, with ten to twenty individuals forming a group, to construct a plurality of mixed sample pools; The mixed sample pools are numbered and saved respectively.
3. The method for screening mutation sites based on amplicon deep sequencing and sequence counting according to claim 1, characterized in that: Primers are designed according to the target gene sequence and a tag sequence is added to the 5' end of the primer. The genomic DNA is used as a template for polymerase chain reaction amplification to obtain the target amplified fragment. The target amplified fragment is deeply sequenced using a high-throughput sequencing platform to obtain raw sequencing data, including: Designing a pair of specific primers based on the conserved region sequence of the target gene, and adding a tag sequence for distinguishing different mixed sample pools to the 5' end of each specific primer; Using the genomic DNA in the mixed sample pool as a template, polymerase chain reaction technology is used to amplify a target gene fragment to obtain a target amplified fragment; the amplification conditions are set according to the characteristics of the primers; The target amplified fragment is purified, and the purified fragment is subjected to single-end or double-end sequencing using the high-throughput sequencing platform to obtain the raw sequencing data.
4. The method for screening mutation sites based on amplicon deep sequencing and sequence counting according to claim 3, characterized in that: The properties of the primers include denaturation temperature, annealing temperature and extension time.
5. The method for screening mutation sites based on amplicon deep sequencing and sequence counting according to claim 1, characterized in that: Cleaning, removing redundancy and screening the raw sequencing data to obtain candidate sequences includes: Performing quality control on the raw sequencing data, removing sequencing adapters and primer index sequences using a preset quality control tool, and removing reads containing chimeric structures using a chimera identification algorithm to obtain valid reads after quality control; Performing a redundancy removal operation on the valid read segments, extracting a set of non-repeated sequences to obtain a set of unique sequences, and counting the number of times each unique sequence appears in the sample pool as depth information of the unique sequence; Calculate the expression frequency in the mixed sample pool based on the depth information of each unique sequence and the total depth information of all unique sequences; Calculating the lower limit and upper limit of frequency screening based on the ploidy characteristics of the individuals to be screened constituting the mixed sample pool and the individual mixing multiples adopted; The unique sequence whose expression frequency is within the range of frequency screening is screened out as the candidate sequence.
6. The method for screening mutation sites based on amplicon deep sequencing and sequence counting according to claim 1, characterized in that: The lower limit value of the frequency screening is the value obtained by dividing the product of the ploidy and the individual mixing multiple by one.
7. The method for screening mutation sites based on amplicon deep sequencing and sequence counting according to claim 5, characterized in that: The upper limit value of the frequency screening is ten times the lower limit value.
8. The method for screening mutation sites based on amplicon deep sequencing and sequence counting according to claim 1, characterized in that: The candidate sequences whose expression numbers are within the set frequency range are compared with the target gene sequence to identify whether there are mutation sites and obtain the comparison results, including: Importing the set of candidate sequences and the corresponding target gene sequence into a preset sequence alignment tool, and designating the target gene sequence as the reference sequence; Setting alignment parameters and performing sequence alignment to generate an alignment document of the candidate sequence and the reference sequence; the alignment parameters include a match score, a mismatch penalty, and an insertion-deletion overhead; Searching the alignment file for mismatches, insertions, or deletions at the coordinates of the candidate sequence and the reference sequence, marking the alignment coordinates corresponding to the mismatches, insertions, or deletions as differential sites, and recording the alignment results; the alignment results include: the sequence position of the differential site, the type of difference, and depth information of the candidate sequence; Determine whether the expression frequency corresponding to the differential site is still within the frequency screening range. If so, mark the differential site as a mutation site, and output the comparison result including a mutation site list.
9. The method for screening mutation sites based on amplicon deep sequencing and sequence counting according to claim 1, characterized in that: If the comparison result indicates the presence of a mutation site, the individual samples in the corresponding mixed sample pool are amplified and sequenced to determine the specific individual samples containing the mutation, including: Sampling each individual to be screened in the mixed sample pool where the comparison results show the presence of a mutation site, and extracting genomic DNA of the individual to be screened separately; Using the same pair of primers with the tag sequence as a basic template, polymerase chain reaction is performed on the genomic DNA of each individual to be screened to amplify the target gene fragment to obtain an amplified product; Purifying the amplified products and sequencing them using a first-generation sequencing method to obtain the sequencing sequence of each individual to be screened; The sequencing sequence of each individual to be screened is compared with the reference sequence of the target gene to check whether it contains base differences consistent with the determined mutation site; The individuals to be screened that are determined to contain the mutation site in the comparison results are marked as individual samples containing mutations, and an identification list of individual samples containing mutations is output.
Citation Information
Patent Citations
Tumor exon sequencing data analysis method
CN111223525A
Method for detecting specific mutation of liver cancer based on cfDNA base mutation frequency distribution
CN114182022A
Device and method for sample analysis
CN114269916A
Mutant rapid detection method based on sample pool mixing and multiplex PCR amplification and application
CN116287118A
High throughput method of screening a population for members comprising mutation(s) in a target sequence using alignment-free sequence analysis
US20170335388A1