Technical method system for quantitative identification and detection of gene variant allele frequency
Patent Information
- Application Number
- PCT/CN2025/078138
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2026-08-27
Smart Images

Figure 00000008_0000 
Figure 00000009_0000
Abstract
Description
A method system for quantitative identification and detection of gene allelic variation frequencies Technical Field
[0001] This invention relates to the field of biotechnology, specifically to nucleotide sequence variation type detection technology, or the frequency of different sequence variation types, or the absolute or relative percentage (gene frequency). Background Technology
[0002] Seed resource screening and identification are crucial steps in the breeding process and germplasm resource improvement. Currently, important traits in breeding are closely linked to DNA markers. By screening for specific DNA markers, superior traits can be enriched into varieties. Due to the unprecedented rapid development of biotechnology, new biotechnologies are constantly emerging. Technologies such as gene editing play a disruptive role in creating new seed resources. However, the performance of seed resources created through these technologies varies considerably. Modifications, especially in terms of edibility and metabolic pathways, are often difficult to observe with the naked eye. However, with the continuous development of synthetic biology, these changes in quality and metabolites have extremely high added value and can generate significant economic benefits. Therefore, the ability to rapidly screen these germplasm resources is a very important step.
[0003] One of the more economical and direct methods is DNA sequence identification. However, existing technologies include PCR, quantitative PCR, KASP (Kinetic Acid Spectroscopy), CAPS (Carbohydrate and Polymeric Acids), dCAPS (Dual Capsule Biomarkers), and microsatellite SSR (Silicon-Satellite Spectroscopy) marker identification. Among these technologies, PCR and quantitative PCR can only detect the concentration of DNA. CAPS, dCAPS, and microsatellite SSR marker identification are typically inefficient, consuming significant manpower, resources, and time. Sequencing technology generally has low sensitivity and accuracy, with a sensitivity of only 20%.
[0004] Currently, screening and identification of DNA sequence-corresponding markers requires determining the correspondence between genotype and phenotype through individual plant analysis. This method is extremely time-consuming and labor-intensive. Breeding typically involves identifying tens of thousands of seeds. Traditional methods for identifying individual plant genotypes are not only time-consuming and labor-intensive but also extremely costly. Furthermore, these techniques usually only identify negative or positive samples and cannot accurately determine the proportion of a specific type of DNA sequence.
[0005] Germplasm resource screening typically involves screening mixed libraries of thousands or even millions of samples. Accurate, rapid, and low-cost screening is a key technological challenge and a crucial area of expertise in the industry. Quickly identifying target markers from these large samples requires an identification technique capable of detecting the percentage (frequency) of a specific type. The method described in this invention can conveniently, rapidly, and accurately identify whether a specific type of DNA marker is present in a particular genetic population, and the proportion thereof. Summary of the Invention
[0006] This invention claims to protect a technique for the precise determination of the frequency of specific variant types in nucleotide samples.
[0007] This invention claims to protect the techniques and methods for determining and calculating the percentage of a specific nucleotide type in a nucleic acid sample.
[0008] This invention claims protection for experimental methods and techniques for identifying the frequency (percentage) of specific nucleic acid types in a sample.
[0009] This invention detects specific variant types, including but not limited to single nucleotide polymorphisms (SNPs), deletion / insertion variants, multiple copy variants, and microsatellite nucleotide fragment polymorphisms.
[0010] This invention is a technical system and computational method for amplifying a specific target variant nucleic acid type.
[0011] The calculation method protected by this invention obtains the Ct value of amplified nucleotide fragments of a specific variant type, as well as the derived data obtained by mathematical operations based on the Ct value, to determine the percentage (frequency) of any type of single nucleotide sequence variant and its combination of nucleotide variants at one or more sites in the same and / or different samples.
[0012] Compared with other methods, the present invention has advantages such as simplicity, speed, high accuracy, low cost, and high detection rate. Attached Figure Description
[0013] Figure 1 shows the methods for obtaining Ct values and data examples for different mutation types.
[0014] Figure 2 shows the method for obtaining the Ct value of the internal control sequence and a data example. Detailed Implementation
[0015] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0016] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0017] It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Unless otherwise defined or stated, the scientific terms used in this patent have the same meaning as understood by one of ordinary skill in the art.
[0018] Experimental examples
[0019] (1) Extraction of mixed DNA from plant samples. The plant tissue samples were ground into a slurry or powder using a grinder, and then CTAB extraction solution was added. The samples were digested and lysed at 60°C for 1 hour. 400 μL of chloroform was added, and the mixture was repeatedly mixed and centrifuged at 12,000 rpm for 10 minutes. 400 μL of the supernatant was collected and precipitated with an equal volume of isopropanol. The mixture was centrifuged at 12,000 rpm for 10 minutes. The sample was washed twice with 5% ethanol and then dissolved in double-distilled water.
[0020] 2) Implementation of PCR reaction system.
[0021] Table 1. Primers for detecting the frequency of allelic variants at the physical location 391158732 on wheat chromosome 4B.
[0022] ctl-F: ACCgCCAgCTCTTCCACCCTSEQ ID NO 1ctl-R: TCACTggggCATAggAggAASEQ ID NO 2391158732-Ft: TCATTTAGAGCATCTCCAACAGTTGTTSEQ ID NO 3391158732-Fc: TCATTTAGAGCATCTCCAACAGTTGTCSEQ ID NO 4391158732-R:CTAGAGTCACGCGTGCTGTATSEQ ID NO 5Block:GCATCTCCAACAGTTGTTGAACACGTCGCGAGC-NH2SEQ ID NO 6
[0023] 2×PCR buffer
[0024] Template DNA 20 μL
[0025] Internal control upstream primer 100 μM
[0026] Internal control downstream primer 100 μM
[0027] Haplotype I upstream primer 100 μM
[0028] Haplotype I downstream primer 100 μM
[0029] Haplotype II upstream primer 100 μM
[0030] Haplotype II downstream primer 100 μM
[0031] Block primers 50-500 mM
[0032] Taq enzyme 1-10 μL
[0033] dNTP 27 mM
[0034] MgCl2 27 mM
[0035] Add pure water to make up 100 μL
[0036] In this embodiment, the specific reaction system is as follows:
[0037] Reaction 1:
[0038] 2×PCR buffer
[0039] Template DNA 20 μL
[0040] Internal control upstream primer 100 μM
[0041] Internal control downstream primer 100 μM
[0042] Taq enzyme 1-10 μL
[0043] dNTP 27 mM
[0044] MgCl2 27 mM
[0045] Add pure water to make up 100 μL
[0046] The primer sequences are shown in internal control SEQ ID NO 1 to internal control SEQ ID NO 2.
[0047] Reaction 2:
[0048] 2×PCR buffer
[0049] Template DNA 20 μL
[0050] Haplotype I upstream primer 100 μM
[0051] Haplotype I downstream primer 100 μM
[0052] Block primers 50-500 mM
[0053] Taq enzyme 1-10 μL
[0054] dNTP 27 mM
[0055] MgCl2 27 mM
[0056] Add pure water to make up 100 μL
[0057] The primer sequences are shown in SEQ ID NO 3, SEQ ID NO 5, and SEQ ID NO 6.
[0058] Reaction 3:
[0059] 2×PCR buffer
[0060] Template DNA 20 μL
[0061] Haplotype II upstream primer 100 μM
[0062] Haplotype II downstream primer 100 μM
[0063] Block primer 100 mM
[0064] Taq enzyme 1μL
[0065] dNTP 27mM
[0066] MgCl2 27 mM
[0067] Add pure water to make up 100 μL
[0068] The primer sequences are shown in SEQ ID NO 4, SEQ ID NO 5, and SEQ ID NO 6.
[0069] In this example, the specific PCR reaction system is as follows: 95℃ pre-denaturation for 10 minutes, 1 cycle; 95℃ denaturation for 40 seconds, 64℃ annealing for 20 seconds, 72℃ extension for 30 seconds, 15 cycles; 95℃ denaturation for 40 seconds, 60℃ annealing for 20 seconds, 72℃ extension for 50 seconds, 45 cycles. Amplification signals are collected in the third stage.
[0070] (3) Calculation of gene frequencies based on Ct values displayed by fluorescence amplification instruments
[0071] The Ct value for reaction 1 is Ct1 (as shown in Figure 1), the Ct value for reaction 2 is Ct2 (as shown in Figure 2), and the Ct value for reaction 3 is Ct3 (as shown in Figure 2).
[0072] The formula for calculating the gene frequency of haplotype 1 is: F1 = 2^(Ct2-Ct1). The frequency results are shown in Table 2 below.
[0073] The formula for calculating the gene frequency of haplotype 2 is: F2 = 2^(Ct3-Ct1). The frequency results are shown in Table 2 below.
[0074] The theoretical distribution range formula for gene frequency is: A1-2 = 2 ^ (Ct3- Ct2) ~ 2 ^ (Ct2- Ct3).
[0075] Table 2
[0076] Haplotype I frequency F1 Haplotype I frequency F2 Gene frequency range A 10,000 DNA samples Detection times This invention 1.37082E-056.38169E-114.66E-06~2.15E+0510000 Other PCR identification Unavailable Unavailable Unavailable 1~30
[0077] Using this invention to identify gene mutation frequencies is not only simple and convenient to operate, but also highly efficient (data can be obtained in just 120 minutes), saving time and money. Most importantly, the frequency resolution can reach an accuracy of less than one part in 100,000 (the sum of positive and negative orders of magnitude can reach 10 to 11 orders of magnitude).
Claims
1. The polymerase chain reaction (PCR) reaction system used for gene frequency detection is characterized by designing an arbitrary number of primers 1, 2, ..., n near the location of a variation (nucleotide polymorphism) in the target nucleotide sequence (hereinafter referred to as gene sequence). An arbitrary number of primers I, II, III, ..., n are designed upstream and / or downstream of the variation in the target nucleotide sequence.
2. Combine two and / or three and / or more primers in the same or different PCR reactions. Detection of the PCR reaction is performed by quantitative real-time PCR (including but not limited to dye-based and probe-based methods) to detect the Ct value (hereinafter referred to as Ct value).
3. Ct values include, but are not limited to, one, two, several, or more. Ct value calculations include, but are not limited to, mathematical operations on any combination of Ct values, and mathematical operations on Ct values for different PCR reaction tubes. Ct calculation methods include, but are not limited to, absolute values and relative values. Ct-related numerical values include, but are not limited to, Ct values and / or ratios of Ct values and / or differences of Ct values, which can represent gene frequencies.
4. Gene frequencies include, but are not limited to, the absolute abundance of a nucleotide sequence type (hereinafter referred to as gene sequence), the relative abundance of one and / or multiple nucleotide sequence types, and the percentage of any one and / or nucleotide sequence type in all types.
5. Types of nucleotide sequences include, but are not limited to, the same gene (different variant types), different genes, coding and non-coding regions of genes from different species, intergenic regions, whole genome sequences, nucleotide sequence fragments of any size at any position in the whole genome sequence, complete and / or incomplete fragments of genes, and artificially synthesized nucleotide sequences.
6. PCR reaction systems include, but are not limited to, conventional PCR reaction systems, isothermal amplification, and other newly developed nucleotide amplification methods and techniques.
7. The Ct value of the PCR reaction in variant type I is Ct1, the Ct value of the PCR reaction in variant type II is Ct2, …, the Ct value of the PCR reaction in variant type n is Ctn. The Ct value of the primer in unknown sample a is Cta. The theoretical maximum or minimum range of gene frequency distribution interval is defined as including but not limited to: Ct1 / (Ct2+…+Ctn), Ct2 / (Ct1+…+Ctn), Ct1-(Ct2+…+Ctn), Ct2-(Ct1+…+Ctn). Gene frequencies of unknown gene types are defined as including, but not limited to, (Ct1+…+Ctn) / Cta, Cta / (Ct1+…+Ctn), (Ct2+…+Ctn) / Cta, Cta / (Ct2+…+Ctn), Ct1+…+Ctn-Cta, Cta-Ct1-…-Ctn, Ct2+…+Ctn-Cta, and Cta-Ct2-…-Ctn. Gene frequencies of unknown gene types are defined as extended values obtained by operations, including but not limited to, using the above values as exponents, bases, or other mathematical operators on arbitrary values.
8. Considering the existence of different allelic variations, the formula for calculating the frequency distribution of each allelic variation is the same. The frequencies of each gene variation can be freely combined for mathematical operations. The general formula for calculating the frequency of haplotype n includes, but is not limited to, Fn = 2^(Ctn-…-Ct1). The formula for the theoretical distribution range of haplotype n gene frequency includes, but is not limited to, the following formulas. A1-n: 2^ (Ct1+ Ct2..+ Ctn-1- Ctn)~ 2^ (Ctn - Ct1-.. Ctn-2- Ctn-1).