A combination of SNP loci for cotton variety identification, a gene chip, and its applications.

By designing SNP site combinations and molecular probe combinations based on capture sequencing, the problems of limited number of sites and complex operation in existing cotton variety identification technologies have been solved, enabling rapid and accurate identification and large-sample analysis of cotton varieties.

CN119177314BActive Publication Date: 2025-10-28COTTON RES INST HEBEI ACAD OF AGRI & FOREST SCI

Patent Information

Application Number
CN202411422798.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-12
Publication Date
2025-10-28
Estimated Expiration
2044-10-12

AI Technical Summary

Technical Problem

Existing cotton variety identification marker technologies, such as the SSR marker method, have limitations such as a limited number of loci, uneven distribution, low genotypic polymorphism, and are complex to operate and have long experimental cycles, making it difficult to achieve accurate identification and data integration across laboratories.

Method used

By employing a combination of SNP sites and molecular probes based on capture sequencing and utilizing high-throughput sequencing technology, a liquid-phase probe with 128 SNP marker combinations was designed. Through capture sequencing and genotyping, rapid and accurate identification of cotton varieties was achieved.

Benefits of technology

It enables efficient and rapid identification of cotton varieties with a 100% detection rate, reduces human error, is suitable for large-sample analysis, and is economical and efficient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119177314B_ABST
    Figure CN119177314B_ABST
Patent Text Reader

Abstract

This invention discloses a combination of SNP loci, a gene chip, and their applications for cotton variety identification, relating to the fields of bioinformatics and molecular breeding. From 100 cotton whole-genome resequencing data, this invention screened out 128 SNPs, which can be used as SNP fingerprints for cotton varieties. Using the SNP molecular markers and liquid-phase probes developed in this invention, cotton identification achieved a 100% detection rate, with a minimum allele frequency greater than 0.05, capable of distinguishing any two cotton varieties, enabling efficient and rapid cotton variety identification. Furthermore, this invention utilizes next-generation sequencing technology, resulting in a simple and standardized procedure that reduces human error, leading to extremely low unit costs for batch identification analysis. It is economical and efficient, particularly suitable for large-sample analysis and identification, and has promising application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bioinformatics and molecular breeding technology, and more specifically to a combination of SNP sites, a gene chip, and their applications for cotton variety identification. Background Technology

[0002] The selection and promotion of superior crop varieties are the prerequisites and foundations for achieving high-yield and high-quality agriculture, and a crucial link in ensuring my country's food security. Variety identification is an important guarantee for the selection and promotion of superior varieties, a key component of variety approval and seed market supervision, and plays a vital role in new variety breeding, seed production, processing, and marketing. Detection methods are key to accurate variety identification.

[0003] Currently, most cotton variety identification markers use the SSR marker method. However, SSR markers have drawbacks, including a limited number of loci, uneven distribution across the genome, low genotypic polymorphism, insufficient genome coverage, and difficulty in accurately integrating and comparing data across laboratories. Existing molecular markers such as ISSR, SRAP, and SSR are complex to operate, have long experimental cycles, and low throughput.

[0004] Therefore, providing a simple, efficient, and accurate method for cotton variety identification is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a combination of SNP loci, a gene chip, and their applications for cotton variety identification. SNP molecular markers are abundant throughout the genome, enabling better differentiation of cotton materials and assisting in the construction of cotton fingerprint profiles. Furthermore, SNP detection technology based on capture sequencing offers fast detection speed and short detection cycle, making it more suitable for the identification of cotton molecular markers and cotton varieties.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] The SNP locus combinations used for cotton variety identification are shown in Table 2. The physical location information of the SNP locus combinations in Table 2 is determined based on the alignment of the cotton G. hirsutum_TM-1_ICR genome sequence.

[0008] Another object of the present invention is to provide a molecular probe array for cotton variety identification, the molecular probe array being used to detect SNP site combinations as shown in Table 2.

[0009] Preferably, the nucleotide sequences of the molecular probe assembly are shown in SEQ ID NO.1-SEQ ID NO.128.

[0010] Another object of the present invention is to provide a kit comprising the above-described molecular probe combination.

[0011] Another object of the present invention is to provide a gene chip loaded with the above-described combination of molecular probes.

[0012] Another object of the present invention is to provide any one of the following applications of the above-described SNP site combination, or the above-described molecular probe combination, or the above-described kit, or the above-described gene chip:

[0013] A. Cotton variety identification;

[0014] B. Cotton variety selection;

[0015] C. Cotton variety traceability;

[0016] D. Identification of cotton variety purity;

[0017] E. Construction of SNP fingerprinting for cotton varieties;

[0018] F. Molecular marker-assisted breeding of cotton varieties;

[0019] G. Cotton genotyping;

[0020] H. Analysis of genetic diversity in cotton;

[0021] I. Functional genomics research on cotton.

[0022] Another object of the present invention is to provide a method for cotton variety identification, comprising the following steps:

[0023] (1) Genomic DNA was extracted from cotton leaf tissue using the CTAB method. After the DNA quality inspection was qualified, the genomic DNA was randomly fragmented (200-300bp), and the DNA fragment ends were repaired and ligated to the adapters for Pre-PCR amplification library.

[0024] (2) The probe hybridizes with the target region;

[0025] (3) Streptomycin affinity-labeled magnetic beads capture hybridization probes;

[0026] (4) Enrich the target fragment, elute it, and perform post-PCR amplification after capture;

[0027] (5) Next-generation sequencing of the target region; compare the sequencing results with the above SNP site combinations to determine the cotton variety.

[0028] Preferably, the probe is the molecular probe combination described above.

[0029] Beneficial effects:

[0030] This invention, based on a cotton genome SNP dataset, screened 128 SNP molecular markers and developed 128 liquid-phase probes. Using the SNP molecular markers and liquid-phase probes described in this invention, cotton identification achieved a 100% detection rate, with a minimum allele frequency greater than 0.05, capable of distinguishing any two cotton varieties. This allows for efficient and rapid cotton variety identification. Furthermore, this invention utilizes next-generation sequencing technology, resulting in a simple and standardized procedure that reduces human error, leading to extremely low unit costs for batch identification analysis. It is economical and efficient, particularly suitable for large-sample analysis and identification, and has promising application prospects. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0032] Figure 1 This is a distribution diagram of the 128 SNP marker combinations of the present invention. Detailed Implementation

[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] Example 1

[0035] The method for screening SNP marker combinations (Table 2) and designing liquid-phase probes (Table 3) based on whole-genome resequencing data from 100 cotton accessions (Table 1) includes the following steps:

[0036] (1) Genomic DNA was extracted from leaf tissues of 100 cotton varieties using the CTAB method. After the DNA passed quality inspection, a whole genome resequencing library was constructed and sequenced using BGI DNBSEQ-T7.

[0037] (2) The clean data obtained from sequencing in step (1) was compared with the genome of cotton G.hirsutum_TM-1_ICR using bwa, and GATK variant detection was performed to obtain 18,009,233 high-quality SNP sites.

[0038] (3) The SNP sites selected in step (2) were further filtered using the vcftools software according to the conditions "--minDP 3--min-alleles2--max-alleles 2--max-missing 0.8--maf0.05" to obtain 100,699 high-quality SNP sites that were detected in all samples.

[0039] (4) The SNP loci selected in step (3) were further screened using a Perl script to obtain 128 SNP marker combinations (see Appendix for their chromosome distribution map). Figure 1 The DNA fingerprint pattern was successfully constructed.

[0040] The principle is as follows: The SNP loci with the highest number of differences between any two cotton varieties are evenly distributed on the chromosomes. The SNP loci and their upstream and downstream 100 bp sequences are extracted, and the extracted 201 bp fragment is processed using Bedtools software. SNP loci with a GC content between 40-60% and no N bases are retained. The extracted sequences are then aligned to the cotton *G. hirsutum*_TM-1_ICR genome using blastn software, retaining only single-copy SNP loci.

[0041] (5) Design liquid phase probe sequences from the 128 SNPs obtained in step (4) for subsequent capture sequencing.

[0042] Table 1100 Cotton Material Information

[0043]

[0044]

[0045] Table 2 Information on 128 SNP markers used for cotton variety identification

[0046]

[0047]

[0048]

[0049]

[0050] Table 3 Information on 128 SNP-labeled liquid phase probes used for cotton variety identification

[0051]

[0052]

[0053]

[0054]

[0055]

[0056]

[0057]

[0058] The number of SNP markers that differ between any two cotton varieties for the above 128 SNP marker combinations is shown in Table 4.

[0059] Table 4. Statistics of SNP markers showing differences between any two cotton varieties.

[0060]

[0061]

[0062] Example 2

[0063] Method for detecting cotton varieties using liquid phase probes corresponding to 128 SNPs obtained in Example 1

[0064] 1. Extraction of genomic DNA

[0065] Cotton leaves were collected, preserved with ice packs, and promptly transported back to the laboratory. Total DNA was extracted from the cotton using the CTAB method.

[0066] 2. Genomic DNA fragmentation and end repair

[0067] (1) Prepare whole genome DNA fragmentation system (30μL) in ice bath: Input DNA XμL (250ng), Buffer 4.5μL, Enzymes 5μL, 1×TE Buffer Up to 30μL; where XμL represents any volume not exceeding 20μL.

[0068] (2) Use a pipette to blow or shake the mixture up and down to ensure that the system is thoroughly mixed, and then briefly centrifuge it.

[0069] (3) Immediately place the PCR tubes on a PCR instrument that has been preheated to 32°C and run the fragmentation and end repair reaction program: 32°C for 20 min; 65°C for 30 min; 4°C (after the temperature drops to 4°C, immediately place the tubes on ice for the next experiment, and do not keep them at 4°C). Set the PCR instrument's heat spreader to 75°C.

[0070] 3. Connecting connector

[0071] (1) Prepare the connection joint reaction system (55 μL) in an ice bath: End Prep Reaction Mix (reaction product from step 2)

[0072] 30μL, Ligation Buffer 15μL, ddH2O 2.5μL, Ligation Enzymes 5μL, TruncatedAdaptor 2.5μL.

[0073] (2) Place the PCR tube on the PCR instrument and the reaction program is 22℃ for 15 min. The PCR instrument is not covered with a heat cap. Then store at 4℃ (proceed to the next step immediately when the temperature drops to 4℃).

[0074] (3) DNA sample purification:

[0075] Add 44 μL (0.8×) DNA Clean Beads to each sample, mix well, and incubate at room temperature for 5 min.

[0076] Place the PCR tube on a magnetic rack and let it stand for 3 minutes until the solution becomes clear. Then remove the supernatant.

[0077] Add 180 μL of 80% ethanol to rinse the magnetic beads, incubate for 30 seconds, remove the supernatant, and repeat the operation once.

[0078] Keep the PCR tube on the magnetic rack, use a 10μL pipette to remove any residual ethanol from the bottom of the tube, and dry until no ethanol remains.

[0079] Resuspend the magnetic beads in 21 μL ddH2O and let stand at room temperature for 1 min to allow the DNA on the magnetic beads to be fully released.

[0080] Place the PCR tube on a magnetic rack for 2 minutes, then transfer 20 μL of supernatant to a new PCR tube for library amplification.

[0081] 4. Library amplification and purification

[0082] (1) PCR amplification

[0083] PCR reaction system (35μL): 20μL adapter-ligated DNA, 12.5μL 2×PCR Mix, 2.5μL PCR Index Primer, 2.5μL Universal PCR Primer.

[0084] The PCR amplification procedure is shown in Table 5.

[0085] Table 5

[0086]

[0087] (2) Add 35 μL (1×) DNA Clean Beads to the amplification product, mix well, and incubate at room temperature for 5 min;

[0088] (3) Place the sample on a magnetic rack for 2 minutes. After the solution becomes clear, remove the supernatant.

[0089] (4) Add 200 μL of 80% ethanol to rinse the magnetic beads, incubate for 30 seconds, remove the supernatant, and repeat the step once.

[0090] (5) Keep the PCR tube on the magnetic rack, remove the residual ethanol at the bottom of the tube with a 10μL pipette, and open the tube cap to dry until there is no ethanol residue.

[0091] (6) Resuspend the magnetic beads in 31 μL ddH2O and let stand at room temperature for 1 min to allow the DNA on the magnetic beads to be fully released;

[0092] (7) Place the sample on a magnetic rack for 2 min, transfer 30 μL of supernatant to a new PCR tube, and store the library at -20℃ for subsequent library quality testing and sequencing.

[0093] 5. Library and probe hybridization

[0094] (1) Take 750 ng of the library constructed in step 4 and add it to a PCR tube, and label it accordingly;

[0095] (2) Add purification magnetic beads to the library and gently mix with a pipette;

[0096] (3) Incubate at room temperature for 5 min, then place the PCR tube on a magnetic rack for 3 min to allow the solution to become clear;

[0097] (4) Remove the supernatant, keep the PCR tube on the magnetic rack, add 180 μL of 80% ethanol, and let stand for 30 seconds;

[0098] (5) Remove the supernatant, add 180 μL of 80% ethanol to the PCR tube, let stand for 30 seconds and then completely remove the supernatant.

[0099] (6) Let it stand at room temperature for 5 minutes to allow the residual ethanol to evaporate completely;

[0100] (7) Prepare the hybridization reaction system (28 μL): Add 13 μL Hyb Buffer, 5 μL Hyb Human Block, 2 μL Adapter Blocker, 5 μL Rnase Block and 3 μL Nuclease Free Water to the reaction tube after the ethanol evaporates in step (6).

[0101] (8) Gently pipette and mix well, let stand at room temperature for 3 minutes, and then briefly centrifuge. Place the PCR tube on a magnetic rack and let stand for 3 minutes.

[0102] (9) Transfer 28 μL of supernatant to a new PCR tube, add 2 μL of Target Probe, gently pipette to mix, and centrifuge briefly;

[0103] (10) Set the PCR instrument parameters as follows: hot cover temperature: 85℃; 80℃ for 5 min; 50℃ for a period of time; place the PCR tube on the PCR instrument, run the above program, and incubate overnight.

[0104] 6. Capture the target region DNA library

[0105] (1) Pretreatment of magnetic beads

[0106] Remove the captured magnetic beads from 4℃, vortex them and resuspend them, then place them at room temperature for 30 minutes to equilibrate.

[0107] Add 50 μL of magnetic beads to a new PCR tube, place it on a magnetic rack for 1 min until the solution becomes clear, and remove the supernatant.

[0108] Remove the PCR tube from the magnetic rack, add 180 μL of Binding Buffer, gently aspirate and mix several times, resuspend the magnetic beads; place on the magnetic rack for 1 min, remove the supernatant; repeat this step once.

[0109] Remove the PCR tube from the magnetic rack, add 180 μL of Binding Buffer, gently pipette to resuspend and mix the magnetic beads, and set aside for later use;

[0110] (2) Capture target region DNA library

[0111] Keep the hybridization product on the PCR instrument, add 180 μL of captured magnetic beads resuspended in step (1) to the hybridization product, mix with a pipette, and place on a rotary mixer to bind at room temperature for 30 min.

[0112] Place the PCR tube on a magnetic rack for 2 minutes to allow the solution to clarify, then remove the supernatant.

[0113] Add 150 μL of preheated Wash Buffer at 50 °C, gently aspirate and mix, then briefly centrifuge and incubate at 50 °C for 10 min on a constant temperature shaker.

[0114] After a brief centrifugation, place the PCR tube on a magnetic rack for 2 minutes to allow the solution to clarify, then remove the supernatant; repeat this step twice, washing the magnetic beads a total of 3 times.

[0115] Keep the sample on the magnetic rack, add 150 μL of 80% ethanol to the PCR tube, let it stand for 30 seconds, then completely remove the ethanol solution and air dry at room temperature.

[0116] Add 24 μL of Nuclease-free Water to the PCR tube, remove the PCR tube from the magnetic rack, and gently resuspend and mix the magnetic beads with a pipette.

[0117] 7. Post-capture PCR amplification

[0118] (1) Take out PostPCR MasterMix and PostPCR Primer from the -20℃ freezer, place them on an ice box to melt, mix them well after melting and place them on ice or at 4℃ for later use.

[0119] (2) After capture, the DNA library was amplified by PCR.

[0120] PCR reaction system (50 μL): 24 μL of DNA library capturing the target region in step 6, 25 μL of Post PCR Master Mix, and 1 μL of Post PCR Primer (select according to library type). Adjust the pipette to 40 μL, gently pipette and mix 6 times, then immediately place it on the PCR instrument;

[0121] PCR instrument program: Hot lid temperature: 105℃; Program: 95℃ 1min; 98℃ 20s; 60℃ 30s N cycles; 72℃ 30s; 72℃ 5min; Hold at 4℃;

[0122] (3) After PCR, add 55 μL of purified magnetic beads to the sample, vortex or pipette to mix, and let stand at room temperature for 5 min.

[0123] (4) Briefly centrifuge and place the PCR tube on a magnetic rack for 3 minutes to allow the solution to clarify;

[0124] (5) Keep the PCR tube on the magnetic rack, remove the supernatant, add 180 μL of 80% ethanol solution to the PCR tube, and let stand for 30 seconds.

[0125] (6) Keep the PCR tube on the magnetic rack, remove the supernatant, add 180 μL of 80% ethanol solution to the PCR tube again, let it stand for 30 seconds and then completely remove the supernatant; let it stand at room temperature for 5 minutes to allow the residual ethanol to evaporate completely.

[0126] (7) Add 25 μL of Nuclease-free water, remove the PCR tube from the magnetic rack, vortex or pipette to mix 10 times, and let stand at room temperature for 2 min.

[0127] (8) Briefly centrifuge, place the PCR tube on a magnetic rack for 2 minutes to allow the solution to clarify;

[0128] (9) Use a pipette to transfer 23 μL of supernatant to a 1.5 mL centrifuge tube and label the sample information;

[0129] (10) Take 1 μL of the library and quantify it using the Qubit dsDNAHS Assay Kit. Record the library concentration, which is approximately 1-20 ng / μL.

[0130] (11) Take 1 μL of sample and use Agilent 2100 Bioanalyzer system (Agilent DNA 1000 Kit) to determine fragment length.

[0131] 8. High-throughput sequencing

[0132] The library obtained after capture and amplification in step 7 was subjected to high-throughput sequencing using the BGI Genomics DNBSEQ-T7 sequencing platform to obtain the sequencing results of cotton genomic DNA, and the obtained data underwent basic cleaning.

[0133] 9. Evaluation of liquid phase probe capture rate

[0134] (1) Use bwa software to align the cleaned sequencing data to the cotton reference sequence, and use GATK software to obtain SNP genotyping data of cotton varieties;

[0135] (2) Use SNP genotyping data to complete the fingerprint analysis of cotton varieties.

[0136] Example 3

[0137] Liquid phase probe capture efficiency verification

[0138] Using the method described in Example 2, 128 SNP liquid phase probes were used to hybridize and genotype individual plants of the 12 cotton varieties in Table 6 to obtain the genotyping data for each individual plant, as shown in Table 7.

[0139] Table 6 12 cotton samples

[0140] Serial Number variety Serial Number variety Serial Number variety 1 Handan 5158 2 Handan Cotton 802 3 Hebei 668 4 Ji Mian 20 5 Lu Mianyan 21 6 Ji 228 7 Ji Mian 616 8 Guoxin Cotton No. 3 9 Quick Breeding 66 10 Handan Wu 23 11 Jifeng 914 12 Ji Mian 298

[0141] Table 7 Genotyping results of 712 cotton samples

[0142]

[0143]

[0144]

[0145]

[0146] Twelve cotton samples were tested using a liquid-phase probe designed based on 128 SNP marker combinations. The number of SNP genotypes was obtained, and the detection rate for the cotton materials was 100%. Using these 128 SNP genotyping results, SNP fingerprint profiles for the 12 tested varieties can be constructed for variety authenticity identification. Furthermore, genetic diversity analysis of these 12 varieties can be performed. This demonstrates that the liquid-phase probe set developed using the 128 SNP marker combinations of this invention can efficiently and rapidly identify cotton varieties.

[0147] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0148] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

The application of 1,128 SNP site combinations is characterized by, The application is any of the following: A. Cotton variety identification; B. Cotton variety screening; C. Cotton variety tracing; D. Cotton variety purity identification; E. Construction of SNP fingerprinting of cotton varieties; F. Molecular marker-assisted breeding of cotton varieties; G. Cotton genotyping; H. Cotton genetic diversity analysis; I. Cotton functional genomics research; The cotton varieties are shown in Tables 1 and 6, and Table 1 is as follows: Table 6 is as follows: ; The physical location information of the SNP loci combinations is shown in Table 2. The physical location information of the loci combinations in Table 2 was determined based on the genome sequence alignment of cotton G. hirsutum_TM-1_ICR. Table 2 is as follows:

2. A molecular probe assembly for cotton variety identification, characterized in that, The molecular probe assembly is used to detect the SNP site combination of claim 1, and the nucleotide sequences of the molecular probe assembly are shown in SEQ ID NO.1-SEQ ID NO.

128.

3. A reagent kit, characterized in that, The kit comprises the molecular probe combination as described in claim 2.

4. A gene chip, characterized in that, The gene chip is loaded with the molecular probe combination as described in claim 2.

5. Any of the following applications of the molecular probe assembly of claim 2, the kit of claim 3, or the gene chip of claim 4: A. Cotton variety identification; B. Cotton variety selection; C. Cotton variety traceability; D. Identification of cotton variety purity; E. Construction of SNP fingerprinting for cotton varieties; F. Molecular marker-assisted breeding of cotton varieties; G. Cotton genotyping; H. Analysis of genetic diversity in cotton; I. Functional genomics research on cotton; The cotton varieties are shown in Tables 1 and 6, and Table 1 is as follows: Table 6 is as follows: 。 6. A method for identifying cotton varieties, characterized in that, Includes the following steps: (1) Genomic DNA was extracted from cotton leaf tissue using the CTAB method. After the DNA quality inspection was qualified, the genomic DNA was randomly fragmented (200-300bp), and the DNA fragment ends were repaired and ligated to the adapters for Pre-PCR amplification library. (2) The probe hybridizes with the target region; (3) Streptomycin affinity-labeled magnetic beads capture hybridization probes; (4) Enrich the target fragment, elute it, and perform post-PCR amplification after capture; (5) Next-generation sequencing of the target region; the sequencing results are compared with the SNP site combination described in claim 1 to determine the cotton variety; The cotton varieties are shown in Tables 1 and 6, and Table 1 is as follows: Table 6 is as follows: 。 7. The cotton variety identification method according to claim 6, characterized in that, The probe is the molecular probe combination described in claim 2.

Citation Information

Patent Citations

  • Upland cotton SNP marker and application thereof

    CN105349537A

  • Cotton whole genome SNP chip and application thereof

    CN108779459A

Cited By

  • SNP (Single Nucleotide Polymorphism) marker set for identifying upland cotton variety and application of SNP marker set

    CN122279098A