SNP molecular markers associated with linolenic acid in upland cotton and their applications
Through genome-wide association analysis, SNP molecular markers of linolenic acid-associated on land cotton have been discovered, which solved the problem of insufficient research on linolenic acid in cotton breeding, achieved early prediction and screening of cotton linolenic acid traits, and promoted cotton breeding process and quality improvement.
Patent Information
- Application Number
- CN202210769028.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-06-30
AI Technical Summary
The prior art neglects the research on linolenic acid in cotton breeding, resulting in low yield and quality of cottonseed oil, low market recognition, and lack of effective molecular marker-assisted selection methods to improve the nutritional composition and quality of cottonseed oil.
Through genome-wide association analysis, SNP molecular markers associated with upland cotton linolenic acid were found, primers and kits were designed for early prediction and screening of cotton linolenic acid content, assisted breeding, and improved cotton germplasm resources.
Early prediction and screening of cotton linolenic acid traits is achieved, suitable for all tissues and development stages, and is not affected by seasons and the environment, supporting rapid and large-scale screening, and promoting the cotton breeding process.
Smart Images

Figure BDA0003723250090000021 
Figure BDA0003723250090000041 
Figure BDA0003723250090000051
Abstract
Description
Technical Field
[0001] The present invention relates to a SNP molecular marker associated with linolenic acid in upland cotton and an application thereof, and belongs to the fields of molecular biology and bioinformatics. Background Art
[0002] Vegetable oil is a valuable resource, serving as a source of dietary nutrition and a key raw material for products such as biofuels, lubricants, and animal feed. Cotton (Gossypium spp.) is not only a significant fiber crop but also a crucial edible oilseed crop, ranking sixth globally as the world's largest source of vegetable oil and playing a crucial role in ensuring edible oil safety in my country. In recent years, linolenic acid has been found to have promising therapeutic effects on cardiovascular and cerebrovascular diseases. However, since linolenic acid is found in certain aquatic organisms and plants and is not consumed in high quantities by humans, its development has become a new area of research. However, research and application of cotton often focuses solely on improving yield and quality, while neglecting research on linolenic acid in cottonseed. This has resulted in low cottonseed oil yield and quality, and limited market acceptance. Therefore, research on linolenic acid in cottonseed oil is crucial for improving the nutritional profile and shelf life of cottonseed oil.
[0003] Genome-wide association studies (GWAS) combine genome-wide genetic variation and polymorphisms with the phenotypic diversity of a target trait across multiple individuals. This approach directly identifies gene loci or markers that are closely associated with phenotypic variation and possess specific functions. Conducting a comprehensive genome-wide study provides a comprehensive overview of desirable traits, making it suitable for research such as identifying desirable traits.
[0004] Molecular marker technology can quickly identify markers closely linked to the QTL for linolenic acid content. Using molecular markers to assist selection for linolenic acid content can accelerate cottonseed oil quality breeding. Researchers used molecular markers such as SLAF and SSR to initially locate candidate genes for linolenic acid in cotton and found that the genetic distance between the candidate genes and the molecular markers was generally large, indicating that the molecular mechanism of linolenic acid formation is very complex and requires further research and exploration. Fully exploring and utilizing these genes that control linolenic acid will enrich the genetic resources for improving fiber quality and provide an important foundation for breeding new cotton varieties that meet various needs.
[0005] In recent years, with the rapid development of high-throughput DNA sequencing technology, the inventors have successfully resequenced 1,279 cotton NAM population resources. Through bioinformatics data analysis and comparison, a large number of high-quality single-nucleotide polymorphisms (SNPs) were obtained. These SNPs can be used to construct haplotype maps, genetic maps, association maps, and fingerprint maps, providing important support for molecular breeding, phylogenetic evolution, and germplasm resource identification. Using genome-wide association analysis, the present invention has discovered a group of SNP molecular markers associated with linolenic acid in upland cotton, laying the foundation for marker-assisted selection and polygenic breeding to improve cottonseed oil quality. Summary of the Invention
[0006] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a group of SNP molecular markers associated with linolenic acid in upland cotton and their applications.
[0007] In order to achieve the above object, the technical solution of the present invention is:
[0008] A SNP molecular marker associated with linolenic acid in upland cotton, wherein the SNP molecular marker is at least one of the nucleotide sequences shown in SEQ ID NO.1-SEQ ID NO.28.
[0009] The SNP molecular site mutated at the 51 bp position of the sequence, and the mutation form of the SNP molecular marker is as follows:
[0010]
[0011] An application of the SNP molecular marker in early prediction and screening of cotton linolenic acid content.
[0012] An application of the SNP molecular marker in molecular marker-assisted breeding of high-yield cotton.
[0013] An application of the SNP molecular marker in improving cotton germplasm resources.
[0014] An application of the SNP molecular marker in identifying high-yield cotton varieties.
[0015] A primer or reagent for detecting the SNP molecular marker.
[0016] A kit for detecting the SNP molecular marker.
[0017] A gene chip containing the SNP molecular marker.
[0018] A method for analyzing linolenic acid in upland cotton using the SNP molecular marker comprises the following steps:
[0019] (1) Extracting genomic DNA from the sample to be tested;
[0020] (2) Using the extracted DNA as a template, primers were designed based on the SNP molecular markers and PCR amplification was performed respectively;
[0021] (3) Analyze the linolenic acid content in upland cotton based on the PCR amplification products.
[0022] Beneficial effects of the present invention:
[0023] The SNP molecular markers associated with linolenic acid in upland cotton provided by this invention can be used for early prediction and screening of linolenic acid traits in cotton. They are directly expressed in DNA and can be detected in all tissues and developmental stages of cotton, unaffected by seasonal or environmental constraints or expression issues. They are neutral and do not affect the expression of target traits. These SNPs are suitable for rapid, large-scale screening. Genomic screening of SNPs often requires only a plus / minus analysis, rather than fragment length analysis, facilitating the development of automated technologies for SNP screening or detection. DETAILED DESCRIPTION
[0024] The specific embodiments of the present invention are further described in detail below with reference to the examples.
[0025] Example 1: Acquisition of SNP molecular markers
[0026] (1) Determination of linolenic acid:
[0027] The population underwent a 1-point 2-replication experiment in 2017 and 2018. 1,271 progeny and 8 parents (3 maternal parents and 5 paternal parents) were randomly arranged within and between subpopulations. The parents of the subpopulation were randomly added to the subpopulation. Three controls were set up for the entire population, namely the maternal parents of the population, Zhongzhimian No. 3, Lumianyan 28, and Jinke 178. The three controls appeared in the population every 15 materials and eventually evenly covered the entire population. The experimental site is: Anyang, Henan (AY). The experimental site was planted in a single-row area with a row length of 2m. The number of plants in each row was between 10-30 plants (depending on the local cultivation pattern). The sampling time varied from September 20 to October 20 (depending on the local frost period and cultivation pattern). Except for the 2 plants at the two ends, the bolls in the middle of the remaining plants close to the main stem were sampled in each plot. 1-2 bolls were taken from each plant, for a total of 20 bolls. The linolenic acid content of cottonseed harvested from each plot was determined using our laboratory's gas chromatograph (GC-2030, Shimadzu, Japan). Linolenic acid content was determined by extraction with boiling petroleum ether and methylation with a potassium hydroxide-methanol solution. The chromatographic column used was an SH-Rtx-65 (30 m × 0.25 mm × 0.50 μm) and the detector was a flame ionization detector (FID). High-purity nitrogen (99.999%) was used as the carrier gas, and the injection volume was 1 μl (containing n-hexane solvent and seed fatty acid methyl ester sample) with a split ratio of 39:1. The temperature program was set as follows: (1) maintain column temperature at 100°C, retention time for 1 min, (2) increase the temperature from 100°C to 210°C at a rate of 4°C / min, (3) maintain column temperature at 210°C, retention time for 4 min, (4) increase the temperature from 210°C to 230°C at a rate of 4°C / min, and (5) maintain column temperature at 230°C, retention time for 5 min. The maximum measurement temperature during the experiment was set at 280°C, and the injection port temperature was set at 250°C.
[0028] (2) SNP detection:
[0029] A total of 1,279 upland cotton samples were collected for genome resequencing. When collecting samples, seeds of each variety were sown in an incubator and young leaves of the cotton plants were collected. 5 μg of high-quality cotton genomic DNA was extracted from all samples using the CTAB method. The extracted genomic DNA was sent to Shenzhen BGI Genomics Technology Co., Ltd. for genome resequencing. High-quality clean data was obtained by sequencing, with a data volume of 20.47 Tb, an average sequencing depth of 35X for parents, and an average sequencing depth of more than 4X for offspring. Sequence positioning was performed using the genome of high-quality tetraploid cotton (G. hirsutum'Texas Marker 1') as the reference genome. Before mapping, all unassembled contigs were connected to a pseudo-chromosome (named "ChrUN"). BWA (v.0.7.12) software was used to map the short sequences of the 1,279 samples to the reference genome, and all unaligned reads and low-quality reads (mapping quality less than 20) were removed. GATK UnifiedGenotyper (v.3.8.0) was then used to identify variants in each sample, and the variant files for all samples (n=1279) were merged into a single VCF file. Finally, 11,856,129 and 4,543,742 high-quality SNPs and indels were identified, respectively. VCFtools was used to further filter variant sites based on minor allele frequencies greater than 0.05 and deletion rates less than 0.2, resulting in 1,855,955 high-quality SNPs and 1,309,084 high-quality indels for subsequent genome-wide association analysis. The effects of all variants were annotated using ANNOVAR.
[0030] (3) Genome-wide association analysis of linolenic acid traits in upland cotton:
[0031] The genome-wide scanning (GWAS) of linolenic acid traits in upland cotton was used to perform statistical analysis on the linolenic acid trait results obtained in step (1) and the genotype data obtained in step (2) using a mixed linear model using the Efficient Mixed-Model Association Expedited (EMMAX) statistical analysis software. For details, please refer to (http: / / csg.sph.umich.edu / kang / emmax / download / index.html). The statistical model is:
[0032] y=Xα+Zβ+Wμ+e
[0033] y is the phenotypic trait, X is the indicator matrix of fixed effects, α is the estimated parameter of fixed effects; Z is the indicator matrix of SNPs, β is the effect of SNPs; W is the indicator matrix of random effects, μ is the predicted random individual, e is the random residual, and it obeys e~(0,δe 2 ). In this model, the population analysis was corrected by adding a kinship matrix to μ. The analysis found that a total of 28 SNPs were significantly associated with the linolenic acid trait of upland cotton. The allele loci of the SNP markers are shown in Table 1. The reference sequence is the upland cotton cultivar TM-1, reference genome version number G.hirsutum_TM-1_ICR (http: / / grand.cricaas.com.cn / page / download / download). The nucleotide sequences 50 bp upstream and downstream of these SNP sites are shown in SEQ ID NO.1-SEQ ID NO.28.
[0034] Table 1 SNP molecular markers associated with linolenic acid in upland cotton
[0035]
[0036]
[0037] (4) Verification: The effects of the above SNPs were verified using the BLUP values (best linear unbiased prediction values) of linolenic acid content in 1279 cotton multi-parent populations under 10 environments with 5 points over 2 years. The results showed that 100% of the SNPs showed a significant effect on the variation of the linolenic acid content trait in upland cotton. SEQUENCE LISTING <110> Cotton Research Institute, Chinese Academy of Agricultural Sciences <120> SNP molecular markers associated with linolenic acid in upland cotton and their applications <160> 28 <170> PatentIn version 3.5 <210> 1 <211> 101 <212> DNA <213> Cotton (Gossypium spp.) <400> 1 tttccggcga attcacccac ttctccatgg aatttttgct ttctaataaa tagcttcgaa 60 cttcggcctc ttcgatgtag cgctcattgg ccatgatctg c 101 <210> 2 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 2 gtttgaaggt aatcatggat gatttcttta atttcatgcc gatgctaagg gttctggact 60 tgtctagaaa tatgaatttg gaagaactgc cagtaggaat t 101 <210> 3 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 3 ttttggacga gcataccaat ttgagcaatg ccagaaacca gggacatttc attataggtt 60 ctcttattga cgcatgctta ttggagaaac aaggtagtca t 101 <210> 4 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 4 aaacccttcg aatgcatcta gatatttgca agttagctga aacggtggta gaagagtgtg 60 ctgggctacc tctcgctctc attacatttg ggcgggccat g 101 <210> 5 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 5 cccgactacg ttccagttca atcaacagaa tgttgctgat gaaatattta acgagttatt 60 tggtgctttt ggtggtggtc tgagaggtac gagattctct a 101 <210> 6 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 6 gaggattgga ggagatgatt ggtgagaact tgggttcgtt ggaaagctcg ttccattgtg 60 gaatctgggt tttgtttttc tcttttctat tggaatttgc a 101 <210> 7 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 7 ccctgttcaa catcctggct gtttctctag gggagtcctc tttctcggac gcttctaaat 60 cccacttacg gccggatact cctccttcat cgttcttcga t 101 <210> 8 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 8 tgtttctcta ggggagtcct ctttctcgga cgcttctaaa tcccacttac ggccggatac 60 tcctccttca tcgttcttcg atacccttcc tcctctgtcg t 101 <210> 9 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 9 gctcaatgaa tttgatgttc caattgcgtt gtctgctatc tgcatctgcc aatgatatat 60 ctaggaccct attatgataa gcatataagg tacgacgatg a 101 <210> 10 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 10 gaccggccgc atacccgata gtcttattgg aagaccgggg ctggaagttc agtaagtata 60 aattcacgaa agttgtatga acttccaact ttaaacctaa c 101 <210> 11 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 11 agccgctgtc ttcgcaacta gatggcctct tggtggtctt acaaagataa ccttaagtcc 60 tgctgctaat tcaaatggta gtcctttgat caatgctgga g 101 <210> 12 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 12 catgttttct atgatgtgtg aagacctttt tcctcacaca ggttagacaa ttagttttaa 60 ctgctattgc tatgaaagga cttggtgcta ttctctttgt a 101 <210> 13 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 13 tttcttctgg ggaatgaaga actcgataac aaagaggcgg aagaagaaag ccccgaagtc 60 gaaaactgct tagtttatat ccaagaggta gctggtgaag a 101 <210> 14 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 14 aaagaaaaaa acaatgggtt ttgcttcatt tgttggacga gtttgctttg gttccatttt 60 cattctatca gcttggcaaa tgtaatcctc tctttctttt a 101 <210> 15 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 15 aagatgttac aataaaggaa ggggaaggga gcaagaaagg ggacacaatg cttgtggatt 60 caactagtga ttgcgttgat cattacattc gatcttcttt g 101 <210> 16 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 16 attaggtaca aatctgtgat caatgatgga ttcaagttgg atctccatca tgctctcgcc 60 ctcgagaagg tatattttct tcatttcgtt cgagaagaat a 101 <210> 17 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 17 acaaatctgt gatcaatgat ggattcaagt tggatctcca tcatgctctc gccctcgaga 60 aggtatattt tcttcatttc gttcgagaag aataataata t 101 <210> 18 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 18 cgttgaagct ttgagtcatc tcagtctgga aaacggttcc actccttctc tttcggctcc 60 acgcggtggt cgagctgtac gtgaagctac tgataaagct t 101 <210> 19 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 19 acgatgtggc gacggtggcg actaatgtca atgcgtgcga tgagcagttt tctgagaaaa 60 cgccttttag tgataggaat agacttgttc atgatctttc t 101 <210> 20 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 20 acaggcttcc caattcttgt aagcgtagtt tctgattgat tcgatagaac tatgtatctg 60 caggaaaaat aaataaaatt gaagttttca tcagaaacaa t 101 <210> 21 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 21 gtagtactgt gaatcaccat ctgatgacag atttagacgg aaatgaagca cctacaactc 60 ataatgcata ccaatcaggc caaaaattcc attcaaaaaa t 101 <210> 22 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 22 tttcatgggt tgggctgata gatgtatggg ttttcttaat cgtcttgaaa gaacttctaa 60 agaatttgat accttctttc aagaactcat tgatgaacat c 101 <210> 23 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 23 tcagagcaag aggacatagt tgatgtgtta ctacggataa ggacggatca gatattttca 60 tttgatctta ccatagatca cataaaagct attcttatgg t 101 <210> 24 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 24 agagcaagag gacatagttg atgtgttact acggataagg acggatcaga tattttcatt 60 tgatcttacc atagatcaca taaaagctat tcttatggtc a 101 <210> 25 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 25 tggctgaaaa gttgataaaa gaggggaaag cctatgtgga tgatacccca tgtgagcaaa 60 tgcagaaaga aaggatggat ggcattgaat caaaatgtag g 101 <210> 26 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 26 tccaaattct gaaaagtttg gcgaagccgt acaaagagca gaaaatatct gcatacaatc 60 accttcaaga atcacttgca aatagtcctc cgatttgagc c 101 <210> 27 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 27 tgggttcaag tatagaatgt caatggcata ttttggggcg ccaccacccc gtgcatggct 60 cagtgtccca ccagaactgg tcacacccca taggccattg c 101 <210> 28 <211> 101 <212> DNA <213> Cotton (Gossypium spp) <400> 28 aatcagattg cctctataaa caactccata cccaccatca ccgataatat tatcctttga 60 aaactggtta gtggcaagtt gtaggtctct tagagtgaac c 101
Claims
1. Application of a SNP molecular marker in early prediction and screening of linolenic acid content in upland cotton, characterized in that: The SNP molecular marker is at least one of the nucleotide sequences shown in SEQ ID NO.1 to SEQ ID NO.28; The SNP molecular site mutated at the 51 bp position of the sequence, and the mutation form of the SNP molecular marker is as follows:
2. A method for analyzing linolenic acid in upland cotton using SNP molecular markers, characterized in that: The following steps are involved: (1) Extracting genomic DNA from the sample to be tested; (2) Using the extracted DNA as a template, primers were designed based on the SNP molecular markers and PCR amplification was performed respectively; (3) Analyze the linolenic acid content of upland cotton based on PCR amplification products; The SNP molecular marker is at least one of the nucleotide sequences shown in SEQ ID NO.1 to SEQ ID NO.28; The SNP molecular site mutated at the 51 bp position of the sequence, and the mutation form of the SNP molecular marker is as follows: