A method for identifying rubber tree cultivar B based on chloroplast SNP markers of rubber tree leaves
The identification of high-product varieties through the combination of rubber chloroplast SNP markings has solved the problem of screening high-yield and stress-resistant varieties in the existing technology, and achieved more efficient breeding process and yield improvement.
Patent Information
- Application Number
- CN202510157297.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-02-13
AI Technical Summary
It is difficult for the prior art to effectively screen out high yield and stress-resistant rubber tree cultivars, which affects the breeding process.
A combination of SNP tags based on rubber chloroplast SNP tags is provided, including multiple loci such as s8165-T/G, s9952-G/A, etc., for identification of high-quality rubber tree species and for variety identification in combination with genotyping technology.
It can carry out more precise variety improvement and new variety breeding, improve the yield of rubber trees and improve planting efficiency.
Smart Images

Figure CN119614744B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of molecular markers, and particularly relates to a method for identifying Hevea brasiliensis B-type cultivated varieties based on chloroplast SNP markers of rubber trees. Background Art
[0002] Natural rubber mainly comes from Hevea brasiliensis, which is an important industrial raw material and is widely used in fields such as transportation, medical and health. Due to its excellent properties in elasticity, abrasion resistance, impact resistance, etc., it still cannot be replaced by synthetic rubber so far. Hevea brasiliensis is native to the Amazon Basin in Brazil. After multiple introductions, it is now widely cultivated in tropical regions of Asia. Historically, there were two relatively famous introductions of rubber trees. In 1876, the British Wickham collected rubber tree germplasm in Brazil, then brought the seeds and seedlings to the Kew Garden in the UK for breeding, and then planted them in Sri Lanka, and subsequently introduced them to Southeast Asian countries such as Malaysia, Indonesia, and Thailand for planting. The second time was in 1981 when the International Rubber Research and Development Board (IRRDB) organized the collection and introduction of rubber tree germplasm. In addition, it also includes other rubber tree germplasm available in the main rubber-producing countries. After years of introduction, cultivation, and screening of crossbred and selected varieties, China has bred multiple high-yield varieties with cold tolerance and wind resistance, including high-yield varieties such as Reyan 879 and Reyan 73397, wind-resistant high-yield varieties such as Reyan 917, and cold-resistant high-yield varieties such as Yunyan 774. Among them, Reyan 879 is an ultra-high-yield variety, and its yield per mu is more than 50% higher than the control, which is the variety with the highest yield per unit area in the world currently.
[0003] With the continuous progress of sequencing technology, the genomes of more and more rubber tree varieties, germplasms, and rubber tree species have been analyzed, laying a solid foundation for further carrying out genomic selection and molecular-assisted breeding to accelerate the rubber breeding process. Chloroplast is the place where green plants carry out photosynthesis, and it has an independent genetic system. Its genetic structure is relatively stable and gene recombination occurs less frequently. Therefore, chloroplast genome sequences are widely used in plant phylogeny and species identification. Plant genome resequencing mostly uses leaves as materials, and mesophyll cells contain a large number of chloroplasts. In addition, the size of the chloroplast genome is very small, only more than 100 kb. Therefore, compared with the nuclear genome, only a small amount of leaf resequencing data can be used for chloroplast genome assembly and SNP analysis. Rubber tree cultivated varieties all originate from the screening of wild germplasms in the Amazon Basin of Brazil, and generally have characteristics such as high yield and stress resistance. By comparing and analyzing the chloroplast genomes of cultivated varieties and wild germplasms, effective SNP marker sites are screened, which will accelerate the high-yield and stress-resistant breeding process of rubber trees. Summary of the Invention
[0004] The object of the present invention is to provide an SNP marker combination that can be used for high-yielding varieties of rubber trees, enrich molecular markers available for breeding excellent rubber tree varieties, and provide new ideas for high-yield and stress-resistant breeding of rubber trees.
[0005] The present invention provides an SNP marker combination for identifying high-yielding varieties of rubber trees. The SNP marker combination includes s8165-T / G, s9952-G / A, s23507-G / A, s26511-A / G, s28467-C / A, s29302-T / G, s29303-C / A, s29894-G / T, s43759-T / C, s46966-A / C, s69276-G / T, s69443-C / A, s74109-G / A, s79616-G / T, s86288-C / T, s119902-C / A, s122256-C / T, s129310-G / A, and s133745-T / C; s8165-T / G is located at the 65th position from the 5'-end of the sequence shown in SEQ ID No.1, and there is a natural variation of T / G at this site, specifically: 5'-AGTCAAAATACAATTATTTTCAAAAAAAAAATGCTTGTTATGCTTAATATTTTTAGTTTAATTTTTATCTGTTTTAATTCTGCCATTTTTTCAAGCAATTTTTTCTTTACAAAATTGCCCGAAGCCTACGCCTTTTTGAA-3'; its corresponding wild-type sequence is as shown in SEQ ID No.2, specifically: 5'-AGTCAAAATACAATTATTTTCAAAAAAAAAATGCTTGTTATGCTTAATATTTTTAGTTTAATTTGTATCTGTTTTAATTCTGCCATTTTTTCAAGCAATTTTTTCTTTACAAAATTGCCCGAAGCCTACGCCTTTTTGAA-3'; s9952-G / A is located at the 72nd position from the 5'-end of the sequence shown in SEQ ID No.3, and there is a natural variation of G / A at this site, specifically: 5'-TAATTATAATGGATAATGACAATTAGAATGAATCTTTCATGAATATAAGAATCAAAAAATTCTATGGAATCGTGAAAGACAGAAAGATTCTTCGAATTTAGTCATAACATTTCGTTATATTGACAATTTCAAAAACTGTTCATACTATGAGCATAGTATG-3'; its corresponding wild-type sequence is as shown in SEQ ID No.As shown in Figure 4, specifically: 5'-TAATTATAATGGATAATGACAATTAGAATGAATCTTTCATGAATATAAGAATCAAAAAATTCTATGGAATCATGAAAGACAGAAAGATTCTTCGAATTTAGTCATAACATTTCGTTATATTGACAATTTCAAAAACTGTTCATACTATGAGCATAGTATG-3'; The s23507-G / A is located at the 67th position from the 5' end of the sequence shown in SEQ ID No. 5. There is a natural G / A variation at this site, specifically: 5'-GTAGATACACAAATCGCTTTAAATACAAGAAGTCGGGTGGGCGGATTGGTCCGAGTGGAGAGAAAAGAAAAAAAAATGGAACTTAAAATCTTTTCTGGAGATATCCATTTTCCGGGAGAGACAGATAAAATATCCCGACA-3'; Its corresponding wild-type sequence is as shown in SEQ ID No. 6, specifically: 5'-GTAGATACACAAATCGCTTTAAATACAAGAAGTCGGGTGGGCGGATTGGTCCGAGTGGAGAGAAAAAAAAAAAAAATGGAACTTAAAATCTTTTCTGGAGATATCCATTTTCCGGGAGAGACAGATAAAATATCCCGACA-3'; The s26511-A / G is located at the 81st position from the 5' end of the sequence shown in SEQ ID No. 7. There is an A / G natural variation at this site, specifically: 5'-TTCGATTCCAACGAATGATGACGCTATAGCTTCAATCCGATTAATTCTTAATAAATTAGTATTTGCAATTTGTGAGGGTCATTCTAGCTATATACGAAATCCCTGATTAGGAATAAGTAAGATAGATTAAATAAGATAACTAAATTATTTTGTGAACTCC-3'; Its corresponding wild-type sequence is as shown in SEQ ID No. 8, specifically: 5'-TTCGATTCCAACGAATGATGACGCTATAGCTTCAATCCGATTAATTCTTAATAAATTAGTATTTGCAATTTGTGAGGGTCGTTCTAGCTATATACGAAATCCCTGATTAGGAATAAGTAAGATAGATTAAATAAGATAACTAAATTATTTTGTGAACTCC-3'; The s28467-C / A is located at SEQ ID No.At the 87th position from the 5'-end of the sequence shown in Figure 9, there is a natural C / A variation at this site, specifically: 5'-TCCTTTCTTTTCGTTTGTAACTGTGAATTGAATATTGAATAAACAGATTAAATCAAGAAAAAGAACGAAGAGTTCGAATTCAAAAACTGATTCGTAACAAAGAAAGAAATGGAAAAACAAAAAAGTTGTATGAATTCACAAAAACTTTGTGAATAGGAAC-3'; The corresponding wild-type sequence is as shown in SEQ ID No. 10, specifically: 5'-TCCTTTCTTTTCGTTTGTAACTGTGAATTGAATATTGAATAAACAGATTAAATCAAGAAAAAGAACGAAGAGTTCGAATTCAAAAAATGATTCGTAACAAAGAAAGAAATGGAAAAACAAAAAAGTTGTATGAATTCACAAAAACTTTGTGAATAGGAAC-3'; The s29302-T / G is located at the 142nd position from the 5'-end of the sequence shown in SEQ ID No. 11, there is a natural T / G variation at this site, the s29303-C / A is located at the 143rd position from the 5'-end of the sequence shown in SEQ ID No. 11, there is a natural C / A variation at this site, specifically: 5'-CCAATCGAAACCTTTTTTTTAAGGAAATGTTGCAACAGGCATTTTTGTTTTGCAATTAATACGAGACCTGTTACCGAATTCCCGAATTCTAAACTGAATTTTAATAAGAATAATATTTATTAGTATTAGAGATTTCAAATTTCAAATTGAAAAAAAAAATGCATTTATAATTCGAATATAAATCGAATATTTTTTTTACTCTATTTTTAAATAGAATAAAAAAAAAAAAAAACAAGAAAAAAACAAGAACAATAATAAACAATAACATAT-3'; The corresponding wild-type sequence is as shown in SEQ ID No.As shown in SEQ ID No. 12, specifically: 5'-CCAATCGAAACCTTTTTTTTAAGGAAATGTTGCAACAGGCATTTTTGTTTTGCAATTAATACGAGACCTGTTACCGAATTCCCGAATTCTAAACTGAATTTTAATAAGAATAATATTTATTAGTATTAGAGATTTCAAATTGAAAATTGAAAAAAAAAATGCATTTATAATTCGAATATAAATCGAATATTTTTTTTACTCTATTTTTAAATAGAATAAAAAAAAAAAAAAACAAGAAAAAAACAAGAACAATAATAAACAATAACATAT-3'; the s29894-G / T is located at the 134th position from the 5'-end of the sequence shown in SEQ ID No. 13. There is a natural G / T variation at this site, specifically: 5'-ATTCGAAATTCAGAAAAACTACGCGAGGGGGCTATTGAACAGCTGGAAAAAGCCCGGGCCCGCTTACGGAAAGTGGAAATAGAAGCAGATCAGTTTCGAACGAATGGATATTCTGAGATAGAACGAGAAAAATGGAATTTGATTAATTCAACTTATAAGACTTTGGAACAATTAGAAAATTACAAAAATGAAACCATTCATTTTGAACAACAACGAACGATTAATCAAGTCCGACAACGGGTTTTCCAAC-3'; the corresponding wild-type sequence is as shown in SEQ ID No. 14, specifically: 5'-ATTCGAAATTCAGAAAAACTACGCGAGGGGGCTATTGAACAGCTGGAAAAAGCCCGGGCCCGCTTACGGAAAGTGGAAATAGAAGCAGATCAGTTTCGAACGAATGGATATTCTGAGATAGAACGAGAAAAATTGAATTTGATTAATTCAACTTATAAGACTTTGGAACAATTAGAAAATTACAAAAATGAAACCATTCATTTTGAACAACAACGAACGATTAATCAAGTCCGACAACGGGTTTTCCAAC-3'; the s43759-T / C is located in SEQ ID No.At the 139th position from the 5'-end of the sequence shown in Figure 15, there is a natural T / C variation at this site, specifically: 5'-AAGATTTGCTTTATCCGGTATCAAACGAGAGCTGCGGGCAAATAGAACTCCTTTCAGAAGTATCAATACCGTCACATGAATCGTGAATGCATGAATGTGATGTACCAAAAAATCCGCGGTTCCTAATGGAATCGGTAATAAAGCAACCTTGCCACCAACTGCCACTAAATCACCACCCCCCCAAGTTAAACTGGTGCTTGCTGTTGCACCAGGAGCCGTTGCACCAGGTGCTAAAGCATGGGTGTTTTGTATCCATTGAGCAAAGACAGG-3'; The corresponding wild-type sequence is shown in SEQ ID No. 16, specifically: 5'-AAGATTTGCTTTATCCGGTATCAAACGAGAGCTGCGGGCAAATAGAACTCCTTTCAGAAGTATCAATACCGTCACATGAATCGTGAATGCATGAATGTGATGTACCAAAAAATCCGCGGTTCCTAATGGAATCGGTAACAAAGCAACCTTGCCACCAACTGCCACTAAATCACCACCCCCCCAAGTTAAACTGGTGCTTGCTGTTGCACCAGGAGCCGTTGCACCAGGTGCTAAAGCATGGGTGTTTTGTATCCATTGAGCAAAGACAGG-3'; The s46966-A / C is located at the 136th position from the 5'-end of the sequence shown in SEQ ID No. 17, and there is a natural A / C variation at this site, specifically: 5'-AAGCCTATTTCTAGTAAATTTTTTTCTTTTTTTTTTTTTTTCTTTTTTCTTTCTATAGTGGAGATAGTCGCACGTAATGACAGATCACGGCCATATTATTTAAAGCTTGTGGTAAGAAGGGGTTTCGTTCTAGTGACCGAAAATAATATTCCAAAGCTTTTGTGTGTTCTCCATTACTTGTGTGAATAAGGCCTATATTATAGAGTATATAACTTCGATCATAGGGATCAATTTCTAGCC-3'; The corresponding wild-type sequence is as shown in SEQ ID No.As shown in Figure 18, specifically: 5'-AAGCCTATTTCTAGTAAATTTTTTTCTTTTTTTTTTTTTTTCTTTTTTCTTTCTATAGTGGAGATAGTCGCACGTAATGACAGATCACGGCCATATTATTTAAAGCTTGTGGTAAGAAGGGGTTTCGTTCTAGTGCCCGAAAATAATATTCCAAAGCTTTTGTGTGTTCTCCATTACTTGTGTGAATAAGGCCTATATTATAGAGTATATAACTTCGATCATAGGGATCAATTTCTAGCC-3'; The s69276-G / T is located at the 196th position from the 5'-end of the sequence shown in SEQ ID No. 19. There is a natural G / T variation at this site, specifically: 5'-CAAAAAAATAATTGGAATAATCAATATTCGCGATGCACGTATTCGTGTATTAGATCCAAAAGGTCTTTCTTGACACTTAACTACAAAGAAAAGAAGTTTTTTTTAATATGGAAAAAATGTAAAATCTAGAAAAGAAAAAGATAATCGGAATTACTAGTCTTAGAATCTATTTGTACAAAAACAAAAGAAAGATTTGATTTTTTCGAGTGATCTGATCATGCGATACTTTTTTTTCTTATTCGAGATATTGTGTGAATGAATCTACTAATAAATCTAATAA-3'; The corresponding wild-type sequence is as shown in SEQ ID No. 20, specifically: 5'-CAAAAAAATAATTGGAATAATCAATATTCGCGATGCACGTATTCGTGTATTAGATCCAAAAGGTCTTTCTTGACACTTAACTACAAAGAAAAGAAGTTTTTTTTAATATGGAAAAAATGTAAAATCTAGAAAAGAAAAAGATAATCGGAATTACTAGTCTTAGAATCTATTTGTACAAAAACAAAAGAAAGATTTTATTTTTTCGAGTGATCTGATCATGCGATACTTTTTTTTCTTATTCGAGATATTGTGTGAATGAATCTACTAATAAATCTAATAA-3'; The s69443-C / A is located at SEQ ID No.At the 143rd position from the 5'-end of the sequence shown in Figure 21, there is a natural C / A variation at this site, specifically: 5'-CGATACTTTTTTTTCTTATTCGAGATATTGTGTGAATGAATCTACTAATAAATCTAATAAGAGTTAAATTCAATTTTAAAATGAAAAATGACACAAATTCGATACAAATATACAAATAAATAGTAAATAAATATTAAAAATCCATTTTAAAATAAACTAAAAGAAACTATCAAAATAGAAATGATTTTCCAATTTTATTTTGTAGTGATTTCCGTAGTAATTCAAAGAAAGTTTTTTGAATAAAGTTATACAACACAGTTTTTATTACTT-3'; the corresponding wild-type sequence is as shown in SEQ ID No. 22, specifically: 5'-CGATACTTTTTTTTCTTATTCGAGATATTGTGTGAATGAATCTACTAATAAATCTAATAAGAGTTAAATTCAATTTTAAAATGAAAAATGACACAAATTCGATACAAATATACAAATAAATAGTAAATAAATATTAAAAATCAATTTTAAAATAAACTAAAAGAAACTATCAAAATAGAAATGATTTTCCAATTTTATTTTGTAGTGATTTCCGTAGTAATTCAAAGAAAGTTTTTTGAATAAAGTTATACAACACAGTTTTTATTACTT-3'; the s74109-G / A is located at the 129th position from the 5'-end of the sequence shown in SEQ ID No. 23, and there is a natural G / A variation at this site, specifically: 5'-ATTTCCCTGCTTAATGGATAATCATTTCCTACCAATGGGAAATTCTTTCTCGTTTTAAATTGAAAATTGAGGTGATTGAATTTGCACCAATGAAACCATAAATTTGATACACAATAGACGAATATGATGAATCTTTTTTTTTAGATAATGAAGGGGATTCTTCCATTCTATTCTATTCATCCGTACTGATCACAGATAGTGGAAAATTTT-3'; the corresponding wild-type sequence is as shown in SEQ ID No.As shown in 24, specifically: 5'-ATTTCCCTGCTTAATGGATAATCATTTCCTACCAATGGGAAATTCTTTCTCGTTTTAAATTGAAAATTGAGGTGATTGAATTTGCACCAATGAAACCATAAATTTGATACACAATAGACGAATATGATAAATCTTTTTTTTTAGATAATGAAGGGGATTCTTCCATTCTATTCTATTCATCCGTACTGATCACAGATAGTGGAAAATTTT-3'; The s79616-G / T is located at the 116th position from the 5' end of the sequence shown in SEQ ID No. 25. There is a natural G / T variation at this site, specifically: 5'-CGTGCAATTTCTTTTTTTTTTTCTGTATTTCTGGAATATGAGTGTGCGACTTGTTATAATTTATCCTATTGATAGTACAGAGAGTGGGTCTGTCATCTTGATAGAGATGATTTTTGACCTCGTCAGGTATTTATTCTAGTATCTTTTATTCTAGTATCTGGAGCACGGAATATATGAAATAGATTACGAAGTATTAGAAC-3'; Its corresponding wild-type sequence is shown in SEQ ID No. 26, specifically: 5'-CGTGCAATTTCTTTTTTTTTTTCTGTATTTCTGGAATATGAGTGTGCGACTTGTTATAATTTATCCTATTGATAGTACAGAGAGTGGGTCTGTCATCTTGATAGAGATGATTTTTTACCTCGTCAGGTATTTATTCTAGTATCTTTTATTCTAGTATCTGGAGCACGGAATATATGAAATAGATTACGAAGTATTAGAAC-3'; The s86288-C / T is located at SEQ ID No.At the 128th position from the 5'-end of the sequence shown in Figure 27, there is a natural C / T variation at this site, specifically: 5'-TTCTCGCTATATTTTCTGCTACTCCGCCCATTTCATAAAGTATTCTACCTGGTTTAACGACAGCTACCCAATATTCGGGAGATCCTTTCCCCGAACCCATACGTGTTTCCGTAGGTCTTAAAGTAACCGGTTTGTCGGGAAATATGCGTACCCATATTTTTCCACCGCGGCGTGCATTTCGTGTCATTGCTCGTCGCCCCGCTTCTATTT-3'; The corresponding wild-type sequence is as shown in SEQ ID No. 28, specifically: 5'-TTCTCGCTATATTTTCTGCTACTCCGCCCATTTCATAAAGTATTCTACCTGGTTTAACGACAGCTACCCAATATTCGGGAGATCCTTTCCCCGAACCCATACGTGTTTCCGTAGGTCTTAAAGTAACTGGTTTGTCGGGAAATATGCGTACCCATATTTTTCCACCGCGGCGTGCATTTCGTGTCATTGCTCGTCGCCCCGCTTCTATTT-3'; The s119902-C / A is located at the 132nd position from the 5'-end of the sequence shown in SEQ ID No. 29, and there is a natural C / A variation at this site, specifically: 5'-CAAAAAAAGAAAAAAGAGAAGACACAAAGTTTTATCTTTCTTTTTATTATTAGGATATTTTATCATTTTCAGGATAGGGATTATTATTTTCCCCATCGATCCATTTGTCACAATCACAATAATAAATATTCCTAAGTTTTGTCTTTTTTATATTTATACTATTTGAATATAAAAGAGCAAATCTTTGAATGATCTACTTTAATCATATGCTTGGGCTAGAAAATGGATCA-3'; The corresponding wild-type sequence is as shown in SEQ ID No.As shown in 30, specifically: 5'-CAAAAAAAGAAAAAAGAGAAGACACAAAGTTTTATCTTTCTTTTTATTATTAGGATATTTTATCATTTTCAGGATAGGGATTATTATTTTCCCCATCGATCCATTTGTCACAATCACAATAATAAATATTCATAAGTTTTGTCTTTTTTATATTTATACTATTTGAATATAAAAGAGCAAATCTTTGAATGATCTACTTTAATCATATGCTTGGGCTAGAAAATGGATCA-3'; The s122256-C / T is located at the 146th position from the 5' end of the sequence shown in SEQ ID No. 31. There is a natural C / T variation at this site, specifically: 5'-GGGAAGCTAGTGATAAGATACTGAATGTCGTGAATATTTTTGGCATTGAGGTAGCCATTCCACCCATTTCATCAAGATAAACACGACGTATTCTATCATAACCCGTTCCTGCCAAGAAAAAAAGTGCGGCACCAATAAATCCATGCGATATTATTTGTAAAATGGCTCCATTGAGTCCCATATCACTTATAGAGCAAATTCCTATAATTATGAAACCCATATGAGATACAGAAGAATAGGCTATTCTTTTTTTTAAATTTCGTTGACCAGGAGATGTTGA-3'; Its corresponding wild-type sequence is shown in SEQ ID No. 32, specifically: 5'-GGGAAGCTAGTGATAAGATACTGAATGTCGTGAATATTTTTGGCATTGAGGTAGCCATTCCACCCATTTCATCAAGATAAACACGACGTATTCTATCATAACCCGTTCCTGCCAAGAAAAAAAGTGCGGCACCAATAAATCCATGTGATATTATTTGTAAAATGGCTCCATTGAGTCCCATATCACTTATAGAGCAAATTCCTATAATTATGAAACCCATATGAGATACAGAAGAATAGGCTATTCTTTTTTTTAAATTTCGTTGACCAGGAGATGTTGA-3'; The s129310-G / A is located at SEQ ID No.At the 170th position from the 5'-end of the sequence shown in SEQ ID No. 33, there is a natural G / A variation, specifically: 5'-AATAATCCCAACGTGTTACATAGGGCAAATATTGTATAATTGTTCGATTTTCCGCAATTTTTTCCATTCCTCTGTGTAAATAACCTAATATTGGTTCACAGTCAATAACATCTTCCCCGTCTAGAGTAACGATGAGGCGAAGAACACCATGCATTGATGGGTGGTGGGGGCCCATATTAACTATCATAAGGTCTTTTCGTGTAGCTGGTACATTCATAGGGTTTCCTTGATCCATTCGTCCATGAATTACTGAAAACAAAAAGAAGTTCATCTCAAGATCTAATAAATCAAATAATTCAA-3'; The corresponding wild-type sequence is shown in SEQ ID No. 34, specifically: 5'-AATAATCCCAACGTGTTACATAGGGCAAATATTGTATAATTGTTCGATTTTCCGCAATTTTTTCCATTCCTCTGTGTAAATAACCTAATATTGGTTCACAGTCAATAACATCTTCCCCGTCTAGAGTAACGATGAGGCGAAGAACACCATGCATTGATGGGTGGTGGGGACCCATATTAACTATCATAAGGTCTTTTCGTGTAGCTGGTACATTCATAGGGTTTCCTTGATCCATTCGTCCATGAATTACTGAAAACAAAAAGAAGTTCATCTCAAGATCTAATAAATCAAATAATTCAA-3'; The s133745-T / C is located in SEQ ID No.At the 115th position from the 5'-end of the sequence shown in Figure 35, there is a natural T / C variation at this site, specifically: 5'-AAAAGAGGAGAATGCACATTTGCTTGAAACAGTTCCCAAATAGCTATTTTGCGTCTTTGTGCGCGCACGGATCCTTTTATTATGTCTCGGCGAAAATCCGATTGTTGTGAATAATGTATCAAAGTCACTTCTTCTATTTGATCAGAATTCGTTGTATCTTTGGTATTATTATAAGTATCTGTATTTTTTTTATTATCAGTAAAAATCACTACCCGTTTGGCTTTTCGTGAACGAATTTCATAATCTTCCG-3'; The corresponding wild-type sequence is as shown in SEQ ID No.As shown in 36, specifically: 5'-AAAAGAGGAGAATGCACATTTGCTTGAAACAGTTCCCAAATAGCTATTTTGCGTCTTTGTGCGCGCACGGATCCTTTTATTATGTCTCGGCGAAAATCCGATTGTTGTGAATAACGTATCAAAGTCACTTCTTCTATTTGATCAGAATTCGTTGTATCTTTGGTATTATTATAAGTATCTGTATTTTTTTTATTATCAGTAAAAATCACTACCCGTTTGGCTTTTCGTGAACGAATTTCATAATCTTCCG-3'; the bolded bases represent the SNP marker sites; the SNP marker combination is related to the latex yield of rubber trees; when the base at the s8165-T / G site is T, the base at the s9952-G / A site is G, the base at the s23507-G / A site is G, the base at the s26511-A / G site is A, the base at the s28467-C / A site is C, the base at the s29302-T / G site is T, the base at the s29303-C / A site is C, the base at the s29894-G / T site is G, the base at the s43759-T / C site is T, the base at the s46966-A / C site is A, the base at the s69276-G / T site is G, the base at the s69443-C / A site is C, the base at the s74109-G / A site is G, the base at the s79616-G / T site is G, the base at the s86288-C / T site is C, the base at the s119902-C / A site is C, the base at the s122256-C / T site is C, the base at the s129310-G / A site is G, and the base at the s133745-T / C site is T, it indicates that the rubber tree is a high-yield variety.
[0006] Preferably, the high yield means the latex yield > 50 mL per cut.
[0007] The present invention also provides the application of the SNP marker combination described in the above technical solution in the breeding of high-yield rubber tree varieties.
[0008] The present invention also provides a method for identifying high-yielding rubber tree varieties based on SNP markers in rubber tree chloroplasts, comprising the following steps: amplifying the DNA of the sample to be tested, detecting the polymorphism of SNP sites using the SNP marker combination described in the above technical solution, performing genotyping on the sample to be tested according to the detection results, when the base at the s8165-T / G site is T, the base at the s9952-G / A site is G, the base at the s23507-G / A site is G, the base at the s26511-A / G site is A, the base at the s28467-C / A site is C, the base at the s29302-T / G site is T, the base at the s29303-C / A site is C, the base at the s29894-G / T site is G, the base at the s43759-T / C site is T, the base at the s46966-A / C site is A, the base at the s69276-G / T site is G, the base at the s69443-C / A site is C, the base at the s74109-G / A site is G, the base at the s79616-G / T site is G, the base at the s86288-C / T site is C, the base at the s119902-C / A site is C, the base at the s122256-C / T site is C, the base at the s129310-G / A site is G, and the base at the s133745-T / C site is T, it indicates that the sample is a high-yielding rubber tree variety.
[0009] The present invention also provides the application of the SNP marker combination described in the above technical solution in the preparation of a product for identifying high-yielding rubber tree varieties.
[0010] Preferably, the product includes a kit and / or a gene chip.
[0011] The present invention also provides a kit for identifying high-yielding rubber tree varieties, and the kit includes the SNP marker combination described in the above technical solution.
[0012] The present invention also provides the application of the SNP marker combination described in the above technical solution in constructing a rubber tree gene library.
[0013] Advantages of the present invention:
[0014] By searching the NCBI and the National Genomics Data Center of China, the present invention found that there were hundreds of genome re-sequencing data of different varieties and wild germplasms. Among them, the data with the project number PRJCA004986 released in September 2023 had the largest data volume, including more than 100 cultivated varieties and more than 200 wild germplasm materials. By mining and analyzing these data, it was beneficial to screen effective SNP marker loci, laying a foundation for further assisting the high-yield and disease-resistant breeding of rubber trees. Using the SNP marker combination provided by the present invention, variety improvement and new variety breeding can be carried out more accurately, and then excellent varieties with higher yields can be selected to improve the planting efficiency of rubber trees. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments.
[0016] Figure 1 It is the phylogenetic tree of 670 chloroplast genome sequences of 335 samples in Example 1;
[0017] Figure 2 It is the heat map of the clustering analysis of 299 core SNP loci of 335 samples in Example 2;
[0018] Figure 3 It is the heat map of the clustering analysis of 25 SNP loci in Example 2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to further illustrate the present invention, the following will describe in detail a method for identifying B-type cultivated varieties of rubber trees based on chloroplast SNP markers provided by the present invention in combination with the drawings and embodiments, but they cannot be understood as limiting the protection scope of the present invention.
[0020] The present invention downloaded the genome re-sequencing data of 335 rubber tree cultivated varieties and wild germplasms from the National Genomics Data Center, used the GetOrganelle software to assemble the chloroplasts of 335 materials including 127 cultivated varieties and 208 wild germplasms, then took the chloroplast genome of the main cultivated variety Reyan 73397 of rubber tree as a reference, used the snippy software to screen and analyze the core SNPs, and finally used the Heatmap analysis tool of the TBtools software to perform clustering analysis on the core SNPs to screen out the SNP loci between cultivated varieties and wild germplasms. The methods described in the following embodiments are all conventional methods unless otherwise specified.
[0021] Example 1
[0022] Chloroplast assembly of 335 materials
[0023] Downloaded the genome re-sequencing data of 335 rubber tree cultivars and wild germplasms from the National Genomics Data Center (https: / / ngdc.cncb.ac.cn / ) to the local server (project number: PRJCA004986). The GetOrganelle software was used to assemble the chloroplast genomes of each sample. The specific command was: get_organelle_from_reads.py -1 CRR286406_f1.fastq.gz -2 CRR286406_r2.fastq.gz -F embplant_pt -o CRR286406 -R 15 -t 20 --reverse-lsc. The chloroplast genomes of 335 samples were obtained. The chloroplast genome of each sample included two configurations, namely embplant_pt.K85.complete.graph1.1.path_sequence.fasta and embplant_pt.K85.complete.graph1.2.path_sequence.fasta, which were named CRR286406_1 and CRR286406_2 respectively (other samples were named in the same way). The specific commands were: ll|grep CRR|awk '{print"cat "$NF" / *complete.graph1.1.path_sequence.fasta>Result / "$NF"_1"}'|sh and ll|grep CRR|awk '{print"cat "$NF" / *complete.graph1.2.path_sequence.fasta>Result / "$NF"_2"}'|sh. Then, the cat command was used to merge 670 sequences of 335 samples into a fasta file named CRR286-cp.fasta. The mafft software was used for multiple sequence alignment. The command was: mafft CRR286-cp.fasta>CRR286-cp.fasta.mafft. Then, the fasttree software was used to construct the phylogenetic tree. The command was: fasttree CRR286-cp.fasta.mafft>CRR286-cp.fasta.mafft.tree. The obtained phylogenetic tree file CRR286-cp.fasta.mafft.tree was transferred to a personal computer, and the Figtree software was used to view and display the phylogenetic tree. As Figure 1 shown, a total of 335 chloroplast genome sequences in the branch where CRR286406_2 was located were selected for subsequent SNP analysis.
[0024] Example 2
[0025] SNP Analysis of Chloroplast Genomes of 335 Materials
[0026] Based on the phylogenetic tree of the chloroplast genome sequences constructed in Example 1, a total of 335 chloroplast genome sequences in the branch where CRR286406_2 is located were selected, and the software snippy was used for core SNP analysis. The specific command is: snippy --outdir CRR286364_1a --ref. / CRR286406_2 --ctgs CRR286364_1 --cpus 10. Taking the command with CRR286364_1 as an example, the chloroplast genome sequence CRR286406_2 of the cultivated variety Reyan 73397 was used as the reference genome, and the above command was used to analyze the 335 chloroplast genome sequences respectively; then core SNP analysis was carried out. The specific command is: snippy-core CRR286348_1a CRR286348_2a CRR286349_1a CRR286349_2a CRR286350_1a... (a total of 335 folder names). The analysis result obtained the core.tab file, which is the core SNP matrix file, including a total of 299 core SNP loci. Using the sed command, A, T, C, and G were replaced with the numbers 1, 2, 3, and 4 respectively. The specific command is: more core.tab|sed -e's / A / 1 / g' -e's / T / 2 / g' -e's / C / 3 / g' -e's / G / 4 / g'>core.tab.xlsx. Transfer the core.tab.xlsx to a personal desktop computer, open it with Excel, select all and perform selective paste in a new sheet, check Transpose, and delete the first row to obtain the SNP locus matrix. Using the Heatmap analysis tool of TBtools software, a heatmap was drawn for the core.tab.xlsx file after transposition processing, and column clustering analysis was carried out. As Figure 2 shown, where Cul is a cultivated variety material obtained by large-scale cultivation of high-yield and resistant materials selected and cultivated from wild germplasms; WAC, WRO, and WMG represent wild germplasm materials collected from three regions in Brazil. Through comparative analysis, it was found that 25 of these SNP loci could divide the high-yield varieties into two complementary types, as Figure 3As shown in the figure, two complementary types are named as type A (A class) and type B (B class) respectively. An SNP locus covering both type A and type B is named as an AB type (AB class) SNP locus. Among them, type A includes 5 SNP loci, type B includes 19 SNP loci, and AB type includes 1 SNP locus. Among the 19 SNP loci of type B, the structure of s-phyical location of the chloroplast genome of Reyan 73397 - genotype of cultivated variety / wild germplasm genotype is adopted for representation. The variation information is as follows: s8165 - T / G, s9952 - G / A, s23507 - G / A, s26511 - A / G, s28467 - C / A, s29302 - T / G, s29303 - C / A, s29894 - G / T, s43759 - T / C, s46966 - A / C, s69276 - G / T, s69443 - C / A, s74109 - G / A, s79616 - G / T, s86288 - C / T, s119902 - C / A, s122256 - C / T, s129310 - G / A and s133745 - T / C.
[0027] Using the 19 SNP loci of type B respectively to identify wild germplasm numbered WAC - 1~WAC - 33, WRO - 1~WRO - 53, WMG - 1~WMG - 29 and cultivated variety samples numbered Cul - 1~Cul - 127. When the bases of the s8165 - T / G locus are T, the bases of the s9952 - G / A locus are G, the bases of the s23507 - G / A locus are G, the bases of the s26511 - A / G locus are A, the bases of the s28467 - C / A locus are C, the bases of the s29302 - T / G locus are T, the bases of the s29303 - C / A locus are C, the bases of the s29894 - G / T locus are G, the bases of the s43759 - T / C locus are T, the bases of the s46966 - A / C locus are A, the bases of the s69276 - G / T locus are G, the bases of the s69443 - C / A locus are C, the bases of the s74109 - G / A locus are G, the bases of the s79616 - G / T locus are G, the bases of the s86288 - C / T locus are C, the bases of the s119902 - C / A locus are C, the bases of the s122256 - C / T locus are C, the bases of the s129310 - G / A locus are G, and the bases of the s133745 - T / C locus are T, it indicates that the sample is the high - yielding variety type B of rubber tree; the rest are type b. The latex yield of type B rubber tree is higher than that of type b. The identification results are shown in Tables 1~2, where Cul is the cultivated variety material, and WAC, WRO and WMG represent wild germplasm materials collected from three regions in Brazil.
[0028] Table 1 Identification Results of WAC-1~WAC-33, WRO-1~WRO-53, WMG-1~WMG-29
[0029]
[0030] Table 2 Identification Results of Cul-1~Cul-127
[0031]
[0032] As can be seen from Table 1, among 115 wild-type samples, a total of 2 samples were detected as type B. As can be seen from Table 2, among the cultivated variety samples numbered Cul-1~Cul-127, 87 samples were type B, and the results were Figure 3 matched. It can be seen that by using the SNP marker combination provided by the present invention, the cultivated varieties of rubber trees of type B can be identified, and potential high-yield and disease-resistant germplasm materials can also be identified from wild germplasms, which can greatly improve the breeding process.
[0033] Test Example 1
[0034] Part of the varieties were selected, and their latex yields were collected, expressed by the volume of latex that can be collected per cut, and then the screening results of the present invention were verified, as shown in Tables 3~4 below.
[0035] Table 3 Latex Yields of Samples of WAC-1~WAC-33, WRO-1~WRO-53, WMG-1~WMG-29
[0036]
[0037] Table 4 Latex Yields of Samples of Cul-1~Cul-93
[0038]
[0039] Combining Tables 1~4, it can be seen that among 242 samples, 89 cultivated varieties of type B were screened out by using the SNP loci of the present invention, and among them, a total of 57 samples had a latex yield > 50 mL / cut, belonging to high-yield varieties.
[0040] There are certain differences between the screening results and the yield results. This is because there are many influencing factors for the yield. For example, climate conditions, pests and diseases, planting management methods, and tapping techniques will all cause deviations in the yield results. Among the germplasms screened by using the SNP markers of the present invention, 64.1% are high-yield varieties, which can show that the SNP marker combination provided by the present invention is related to the latex yield of rubber trees, and the SNP marker combination provided by the present invention can be used to screen high-yield germplasms of rubber trees.
[0041] As can be seen from the above embodiments, the present invention obtains a chloroplast genome SNP marker combination that can effectively distinguish cultivated varieties and wild germplasms through resequencing data collection, downloading, chloroplast genome assembly, and SNP screening and analysis. By performing cluster analysis on the screened SNPs, the cultivated varieties can be divided into three types: type A, type B, and type AB. The SNP marker combination provided by the present invention can be used to identify the cultivated rubber tree varieties of type B. Combining Tables 1 to 4, it can be seen that among the rubber trees with cultivated varieties of type B, 64% are high-yield varieties. The SNP marker combination provided by the present invention can be used to screen high-yield germplasms of rubber trees and can assist in the breeding of high-yield rubber tree varieties.
[0042] Although the above embodiments have described the present invention in detail, they are only a part of the embodiments of the present invention, rather than all embodiments. People can also obtain other embodiments based on this embodiment without creative efforts, and these embodiments all fall within the protection scope of the present invention.
Claims
1. Use of a reagent for detecting SNP marker combinations in identifying high-yielding rubber tree varieties, characterized in that, The SNP marker combination is s8165-T / G, s9952-G / A, s23507-G / A, s26511-A / G, s28467-C / A, s29302-T / G, s29303-C / A, s29894-G / T, s43759-T / C, s46966-A / C, s69276-G / T, s69443-C / A, s74109-G / A, s79616-G / T, s86288-C / T, s119902-C / A, s122256-C / T, s129310-G / A, and s133745-T / C; The s8165-T / G is located at the 65th position from the 5'-end of the sequence shown in SEQ ID No.1, and there is a natural variation of T / G at this site; The s9952-G / A is located at the 72nd position from the 5'-end of the sequence shown in SEQ ID No.3, and there is a natural variation of G / A at this site; The s23507-G / A is located at the 67th position from the 5'-end of the sequence shown in SEQ ID No.5, and there is a natural variation of G / A at this site; The s26511-A / G is located at the 81st position from the 5'-end of the sequence shown in SEQ ID No.7, and there is a natural variation of A / G at this site; The s28467-C / A is located at the 87th position from the 5'-end of the sequence shown in SEQ ID No.9, and there is a natural variation of C / A at this site; The s29302-T / G is located at the 142nd position from the 5'-end of the sequence shown in SEQ ID No.11, and there is a natural variation of T / G at this site; The s29303-C / A is located at the 143rd position from the 5'-end of the sequence shown in SEQ ID No.11, and there is a natural variation of C / A at this site; The s29894-G / T is located at the 134th position from the 5'-end of the sequence shown in SEQ ID No.13, and there is a natural variation of G / T at this site; The s43759-T / C is located at the 139th position from the 5'-end of the sequence shown in SEQ ID No.15, and there is a natural variation of T / C at this site; The s46966-A / C is located at the 136th position from the 5'-end of the sequence shown in SEQ ID No.17, and there is a natural variation of A / C at this site; The s69276-G / T is located at the 196th position from the 5'-end of the sequence shown in SEQ ID No.19, and there is a natural variation of G / T at this site; The s69443-C / A is located at the 143rd position from the 5'-end of the sequence shown in SEQ ID No.21, and there is a natural variation of C / A at this site; The s74109-G / A is located at the 129th position from the 5'-end of the sequence shown in SEQ ID No.23, and there is a natural variation of G / A at this site; The s79616-G / T is located at the 116th position from the 5'-end of the sequence shown in SEQ ID No.25, and there is a natural variation of G / T at this site; The s86288-C / T is located at the 128th position from the 5'-end of the sequence shown in SEQ ID No. 27, and there is a natural C / T variation at this site; The s119902-C / A is located at the 132nd position from the 5'-end of the sequence shown in SEQ ID No. 29, and there is a natural C / A variation at this site; The s122256-C / T is located at the 146th position from the 5'-end of the sequence shown in SEQ ID No. 31, and there is a natural C / T variation at this site; The s129310-G / A is located at the 170th position from the 5'-end of the sequence shown in SEQ ID No. 33, and there is a natural G / A variation at this site; The s133745-T / C is located at the 115th position from the 5'-end of the sequence shown in SEQ ID No. 35, and there is a natural T / C variation at this site; The SNP marker combination is related to the latex yield of rubber trees; when the base at the s8165-T / G site is T, the base at the s9952-G / A site is G, the base at the s23507-G / A site is G, the base at the s26511-A / G site is A, the base at the s28467-C / A site is C, the base at the s29302-T / G site is T, the base at the s29303-C / A site is C, the base at the s29894-G / T site is G, the base at the s43759-T / C site is T, the base at the s46966-A / C site is A, the base at the s69276-G / T site is G, the base at the s69443-C / A site is C, the base at the s74109-G / A site is G, the base at the s79616-G / T site is G, the base at the s86288-C / T site is C, the base at the s119902-C / A site is C, the base at the s122256-C / T site is C, the base at the s129310-G / A site is G, and the base at the s133745-T / C site is T, it indicates that the rubber tree is a high-yield variety, and the high yield means that the latex yield > 50 mL per cut.
2. A method for identifying high-yield rubber tree varieties based on chloroplast SNP markers of rubber tree leaves, characterized in that, Comprising the following steps: The DNA of the sample to be tested is amplified, the SNP site polymorphisms of the SNP marker combination described in claim 1 are detected, and the sample to be tested is genotyped according to the detection results. When the base at the s8165-T / G site is T, the base at the s9952-G / A site is G, the base at the s23507-G / A site is G, the base at the s26511-A / G site is A, the base at the s28467-C / A site is C, the base at the s29302-T / G site is T, the base at the s29303-C / A site is C, the base at the s29894-G / T site is G, the base at the s43759-T / C site is T, the base at the s46966-A / C site is A, the base at the s69276-G / T site is G, the base at the s69443-C / A site is C, the base at the s74109-G / A site is G, the base at the s79616-G / T site is G, the base at the s86288-C / T site is C, the base at the s119902-C / A site is C, the base at the s122256-C / T site is C, the base at the s129310-G / A site is G, and the base at the s133745-T / C site is T, it indicates that the sample is a high-yielding variety of rubber tree, and the high yield means that the latex yield > 50 mL per cut.
3. Use of a reagent for detecting the SNP marker combination according to claim 1 in the preparation of a product for identifying high-yielding varieties of rubber trees, characterized in that, The application includes: amplifying the DNA of the sample to be tested, detecting the SNP site polymorphisms of the SNP marker combination described in claim 1, and genotyping the sample to be tested according to the detection results. When the base at the s8165-T / G site is T, the base at the s9952-G / A site is G, the base at the s23507-G / A site is G, the base at the s26511-A / G site is A, the base at the s28467-C / A site is C, the base at the s29302-T / G site is T, the base at the s29303-C / A site is C, the base at the s29894-G / T site is G, the base at the s43759-T / C site is T, the base at the s46966-A / C site is A, the base at the s69276-G / T site is G, the base at the s69443-C / A site is C, the base at the s74109-G / A site is G, the base at the s79616-G / T site is G, the base at the s86288-C / T site is C, the base at the s119902-C / A site is C, the base at the s122256-C / T site is C, the base at the s129310-G / A site is G, and the base at the s133745-T / C site is T, it indicates that the sample is a high-yielding variety of rubber tree, and the high yield means that the latex yield > 50 mL per cut.
4. The application according to claim 3, wherein The product includes a kit and / or a gene chip.
Citation Information
Patent Citations
SNP marker related to rubber yield of rubber tree trunk and application of SNP marker
CN105861498A
SNP (Single Nucleotide Polymorphism) site combination related to breeding traits of rubber trees and application of SNP site combination
CN118813855A