SCAR (sequence characterized amplified region) molecular marker, primer group and kit for identifying sex of cannabis sativa and application
Patent Information
- Application Number
- CN202511613725.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-02-24
AI Technical Summary
[0007]本发明的目的在于提供用于鉴定大麻性别的SCAR分子标记、引物组、试剂盒及应用,本发明针对大麻早期性别鉴定困难及现有分子标记稳定性不足的问题,基于大麻性染色体结构特征,开发稳定、准确的大麻早期性别鉴定分子标记,为生产中的不同时期大麻的性别鉴定提供可靠技术手段
本发明基于大麻性染色体的结构特征,综合利用多份染色体级别大麻单倍型基因组数据,利用生物信息学方法从基因组层面挖掘Y染色体雄性特异序列片段(MSY),并在此基础上,以36份已知性别大麻样品(包括15份雄株和21份雌株)为材料,采用PCR扩增及琼脂糖凝胶电泳技术,对分子标记进行特异性与稳定性验证,并与已报道标记(MADC5、MADC6)进行对比分析。结果显示,本发明共筛选获得15,648个Y染色体特异片段,从中开发出12个SCAR标记。经验证,其中8个标记(MSY99M-2、MSY99M-3、MSY99M-4、MSY99M-5、MSY100M-1、MSY100M-4、MSY100M-5、MSY100M-7)在雄株中均能扩增出清晰且稳定的特异条带,而在雌株中未出现条带,鉴定准确率达100%,展现出稳定的雄性特异性,优于部分已报道标记,而在这之中的5个标记(MSY99M-3、MSY99M-4、MSY99M-5、MSY100M-1、MSY100M-7)条带明亮程度更优,可作为优选。这一结果显示,本发明开发的标记在大麻幼苗期的性别鉴定中具有较高的稳定性和应用价值。与已有的OPV-08、SCAR1以及InDel等分子标记的研究结果一致(赵铭森,方书生,陈瑶,等. 籽用大麻性别连锁标记的验证及SCAR标记开发[J]. 热带作物学报,2019.;孙玉婷,丁美云,王璐瑶,等. 工业大麻种子及苗期性别鉴定方法的研究[J]. 甘肃农业大学学报,2023.;孙哲,王金娥,乔永刚. 工业大麻发育早期雌雄株鉴定方法的研究[J]. 农业与技术,2021.;陶杰,潘根,黄思齐,等. 工业大麻性别连锁Indel标记的筛选与鉴定[J]. 中国麻叶科学,2022.),本发明进一步证明了基于性染色体序列比对分析的分子标记开发方法是实现大麻早期性别精准鉴定的有效策略,本发明为大麻性别鉴定分子标记的开发提供新的思路。
Smart Images

Figure CN121555672A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of plant molecular genetics, and in particular to SCAR molecular markers, primer sets, kits, and applications for identifying the sex of cannabis. Background Technology
[0002] marijuana( Cannabis sativa Cannabis (L.) is a dioecious crop with a long history of domestication, with female plants having higher economic value. In medicinal use, the traditional Chinese medicine hemp seed is derived from the mature fruit of the female plant, while the important active ingredient in cannabis, cannabidiol (CBD), is mainly concentrated in the unpollinated female inflorescences. Cannabis has been cultivated and utilized globally for thousands of years due to its medicinal value. The roots, leaves, flowers, and seeds of cannabis can all be used medicinally. Its dried, mature fruit, hemp seed, has a laxative effect and is used for symptoms such as blood deficiency, fluid depletion, and constipation. Besides its traditional Chinese medicine applications, the active ingredient CBD has become a current research hotspot due to its potential therapeutic effects in diseases such as Alzheimer's disease, epilepsy, and schizophrenia. Epidiolex, a drug with CBD as a single component, has been approved by the US Food and Drug Administration, the European Commission, and the UK National Health Service for the treatment of rare congenital epilepsy in children, such as Lenox-gastuat or Dravet syndrome.
[0003] In actual production, cannabis, as a plant source of hemp seeds and CBD, primarily targets the harvesting of fruits and unpollinated female flowers, respectively. Therefore, rationally controlling the male-to-female ratio is crucial for increasing yield. Because cannabis is a wind-pollinated plant with strong pollen dispersal capabilities, in cultivation aimed at CBD extraction, even a small number of male plants can lead to widespread pollination of female plants, resulting in smaller inflorescences and a CBD yield reduction of up to 60%, causing severe economic losses. Therefore, sex identification during the cannabis seedling stage is of great significance. However, the sex dimorphism of cannabis is not obvious in the early growth stages, and accurate identification through morphological characteristics is usually only possible during the flowering period, thus limiting its application in early production management.
[0004] With the development of molecular biology, genetics, and bioinformatics, molecular markers have provided a new approach for sex identification in cannabis seedlings. Currently, commonly used molecular markers for plant sex identification mainly include Random Amplified Polymorphic of DNA (RAPD), Amplified Fragment Length Polymorphism (AFLP), Sequence Characterized Amplified Regions (SCAR), Simple Sequence Repeats (SSR), and Single Nucleotide Polymorphism (SNP), and have been widely applied to dioecious plants such as sea buckthorn, ginkgo, sweet flag, hops, and papaya. Compared with morphological or chemical methods, molecular markers are abundant, highly accurate, and unaffected by plant growth and development stages or tissue locations. The detection process is simple and quick, making them particularly suitable for sex identification in seedlings. In cannabis research, several male-linked molecular markers (such as MADC1-MADC6) have been developed, and based on these, SCAR markers with high stability and reproducibility have been constructed. Furthermore, Zhao Mingsen et al. (Zhao Mingsen, Fang Shusheng, Chen Yao, et al. Validation of sex-linked markers in seed cannabis and development of SCAR markers [J]. Journal of Tropical Crops, 2019.) achieved a 98.34% accuracy rate in identifying male and female plants in seed cannabis through primer redevelopment of existing markers such as MADC2. Sun Yuting (Sun Yuting, Ding Meiyun, Wang Luyao, et al. Research on sex identification methods for industrial hemp seeds and seedlings [J]. Journal of Gansu Agricultural University, 2023.) and Sun Zhe (Sun Zhe, Wang Jin'e, Qiao Yonggang. Research on identification methods for male and female plants in the early development stage of industrial hemp [J]. Agriculture and Technology, 2021.) have successively demonstrated that OPV-08 and SCAR1 can accurately distinguish between male and female plants in the seed and seedling stages. However, these markers were initially developed based on random amplification techniques such as RAPD and AFLP, and their screening process still requires a significant amount of work and has certain limitations in terms of specificity and stability.
[0005] In recent years, with the development of molecular biology techniques, more stable and efficient molecular strategies for sex identification have been gradually established. For example, Tao Jie et al. (Tao Jie, Pan Gen, Huang Siqi, et al. Screening and identification of sex-linked Indel markers in industrial hemp [J]. Chinese Journal of Cannabis Science, 2022.) constructed InDel markers Is-02 and Is-08 based on whole-genome information, achieving a validation accuracy of up to 100% in various samples. These advances have promoted the transformation of early sex identification in cannabis from random amplification markers to highly specific molecular marker systems, providing strong support for sex identification research and agricultural production.
[0006] Sex in cannabis is primarily controlled by the X / Y sex chromosome system. Sex chromosomes typically consist of one or two pseudoautosomal regions (PARs) that can still recombine during meiosis, and an unpaired sex-determining region (SDR). The SDR includes a male-specific Y chromosome region (MSY) and an X chromosome-specific region (XSR). Sequence-wise, the PAR region is highly homologous, while the MSY, due to recombination repression, has evolved independently of its corresponding XSR and gradually accumulated sequence divergence. In recent years, with the continuous improvement in the quantity and quality of cannabis genome assembly, haplotype genomes and sex chromosome sequences from numerous different varieties of male and female plants have been successfully assembled and published, providing a solid foundation for identifying the sex-determining region and developing male-specific molecular markers based on the MSY region. Summary of the Invention
[0007] The purpose of this invention is to provide SCAR molecular markers, primer sets, kits, and applications for identifying the sex of cannabis. This invention addresses the difficulties in early sex identification of cannabis and the insufficient stability of existing molecular markers. Based on the structural characteristics of cannabis sex chromosomes, it develops stable and accurate molecular markers for early sex identification of cannabis, providing a reliable technical means for sex identification of cannabis at different stages of production.
[0008] To achieve the above-mentioned objectives, the present invention provides the following technical solution: This invention provides SCAR molecular markers for identifying the sex of cannabis, including MSY99M-2, MSY99M-3, MSY99M-4, MSY99M-5, MSY100M-1, MSY100M-4, MSY100M-5 or MSY100M-7; The nucleotide sequence of MSY99M-2 is shown in SEQ ID No. 1; The nucleotide sequence of MSY99M-3 is shown in SEQ ID No. 2; The nucleotide sequence of MSY99M-4 is shown in SEQ ID No. 3; The nucleotide sequence of MSY99M-5 is shown in SEQ ID No. 4; The nucleotide sequence of MSY100M-1 is shown in SEQ ID No. 5; The nucleotide sequence of the MSY100M-4 is shown in SEQ ID No. 6; The nucleotide sequence of the MSY100M-5 is shown in SEQ ID No. 7; The nucleotide sequence of the MSY100M-7 is shown in SEQ ID No. 8.
[0009] This invention also provides the application of the aforementioned SCAR molecular marker in identifying the sex of cannabis.
[0010] This invention also provides the application of the aforementioned SCAR molecular marker in cannabis breeding.
[0011] This invention also provides a primer set for amplifying the SCAR molecular marker to identify the sex of cannabis, the primer set comprising the nucleotide sequence MSY99M-2-F as shown in SEQ ID No. 9 and the nucleotide sequence MSY99M-2-R as shown in SEQ ID No. 10; or Including nucleotide sequences such as MSY99M-3-F as shown in SEQ ID No. 11 and MSY99M-3-R as shown in SEQ ID No. 12; or Including nucleotide sequences such as MSY99M-4-F as shown in SEQ ID No. 13 and MSY99M-4-R as shown in SEQ ID No. 14; or Including nucleotide sequences such as MSY99M-5-F as shown in SEQ ID No. 15 and MSY99M-5-R as shown in SEQ ID No. 16; or Including nucleotide sequences such as MSY100M-1-F as shown in SEQ ID No. 17 and MSY100M-1-R as shown in SEQ ID No. 18; or Including nucleotide sequences such as MSY100M-4-F as shown in SEQ ID No. 19 and MSY100M-4-R as shown in SEQ ID No. 20; or Including nucleotide sequences such as MSY100M-5-F as shown in SEQ ID No. 21 and MSY100M-5-R as shown in SEQ ID No. 22; or This includes MSY100M-7-F, as shown in SEQ ID No. 23, and MSY100M-7-R, as shown in SEQ ID No. 24.
[0012] The present invention also provides a kit for identifying the sex of cannabis, comprising the aforementioned primer set.
[0013] This invention also provides the application of the primer set described above in the preparation of products for identifying the sex of cannabis.
[0014] The beneficial effects of this invention compared to the prior art are as follows: This invention, based on the structural characteristics of cannabis sex chromosomes, comprehensively utilizes multiple chromosome-level cannabis haplotype genome data. Bioinformatics methods are used to mine male-specific Y-chromosome (MSY) sequence fragments at the genomic level. Based on this, using 36 cannabis samples of known sexes (including 15 male and 21 female plants), PCR amplification and agarose gel electrophoresis techniques are employed to verify the specificity and stability of molecular markers, and comparative analysis is performed with previously reported markers (MADC5, MADC6). The results show that this invention screened and obtained 15,648 Y-chromosome-specific fragments, from which 12 SCAR markers were developed. Verification showed that eight markers (MSY99M-2, MSY99M-3, MSY99M-4, MSY99M-5, MSY100M-1, MSY100M-4, MSY100M-5, and MSY100M-7) amplified clear and stable specific bands in male plants, while no bands appeared in female plants, achieving a 100% accuracy rate and demonstrating stable male specificity, superior to some previously reported markers. Among these, five markers (MSY99M-3, MSY99M-4, MSY99M-5, MSY100M-1, and MSY100M-7) exhibited superior band brightness and can be considered preferred. This result demonstrates that the markers developed in this invention have high stability and application value in sex identification during the cannabis seedling stage. Consistent with existing research results on molecular markers such as OPV-08, SCAR1, and InDel (Zhao Mingsen, Fang Shusheng, Chen Yao, et al. Validation of sex-linked markers in seed hemp and development of SCAR markers [J]. Journal of Tropical Crops, 2019.; Sun Yuting, Ding Meiyun, Wang Luyao, et al. Study on sex identification methods for industrial hemp seeds and seedlings [J]. Journal of Gansu Agricultural University, 2023.; Sun Zhe, Wang Jin'e, Qiao Yonggang. Study on identification methods for male and female plants in the early development stage of industrial hemp [J]. Agriculture and Technology, 2021.; Tao Jie, Pan Gen, Huang Siqi, et al. Screening and identification of sex-linked Indel markers in industrial hemp [J]. Chinese Journal of Cannabis Science, 2022.), this invention further demonstrates that the molecular marker development method based on sex chromosome sequence alignment analysis is an effective strategy for achieving accurate early sex identification of hemp. This invention provides a new approach for the development of molecular markers for sex identification in hemp.
[0015] Compared to previous sex markers developed using random amplification techniques such as RAPD and AFLP (e.g., the MADC series), this invention directly utilizes Y chromosome-specific regions to develop SCAR markers, significantly improving development efficiency, reliability, and reproducibility. For example, validation results show that MADC5 and MADC6 exhibit weak positive bands or non-specific amplification in females, suggesting that some early-developed markers may have insufficient stability under different genetic backgrounds. The method employed in this invention effectively avoids the false positives and polymorphism loss problems common with random amplification markers, thereby improving the accuracy and applicability of sex identification and significantly reducing experimental screening work, thus increasing marker development efficiency. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 The results are the initial screening and validation results of the candidate male-specific molecular markers; where M represents DL2000 plusMarker, 1~6 represent varieties - numbers: 1 represents Longma 1-2, 2 represents Fangke 1-4, 3 represents Longma 2-1, 4 represents Longma 2-5, 5 represents Qingma 1-6, and 6 represents Tatanka. Figure 2 The validation results of the candidate molecular markers in different samples are shown. In the M-DL2000 plus Marker, 1~12 represent primers: 1 represents primer SCAR119, 2 represents primer SCAR323, 3 represents primer MSY99M-2, 4 represents primer MSY99M-3, 5 represents primer MSY99M-5, 6 represents primer MSY99M-4, 7 represents primer MSY100M-1, 8 represents primer MSY100M-4, 9 represents primer MSY100M-5, 10 represents primer MSY100M-7, 11 represents primer ITS2, and 12 represents primer ITS2 (blank template). Figure 3The experimental results obtained in Example 5 are shown below; where M-DL2000 plus Marker, 1~12 represent primers: 1 represents primer SCAR119, 2 represents primer SCAR323, 3 represents primer MSY99M-2, 4 represents primer MSY99M-3, 5 represents primer MSY99M-5, 6 represents primer MSY99M-4, 7 represents primer MSY100M-1, 8 represents primer MSY100M-4, 9 represents primer MSY100M-5, 10 represents primer MSY100M-7, 11 represents primer ITS2, and 12 represents primer ITS2 (blank template). Detailed Implementation
[0018] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.
[0019] It should be understood that the terminology used in this invention is merely for describing particular embodiments and is not intended to limit the invention. Furthermore, with respect to numerical ranges in this invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Every smaller range between any stated value or intermediate value within a stated range, and any other stated value or intermediate value within said range, is also included in this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.
[0020] Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. While only preferred methods and materials have been described herein, any methods and materials similar or equivalent to those described herein may be used in the implementation or testing of this invention. All references to this specification are incorporated by way of citation to disclose and describe methods and / or materials associated with those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.
[0021] Various modifications and variations can be made to the specific embodiments described in this specification without departing from the scope or spirit of the invention, as will be apparent to those skilled in the art. Other embodiments derived from this specification will also be apparent to those skilled in the art. This specification and embodiments are merely exemplary.
[0022] The terms “include,” “including,” “have,” “contain,” etc., used in this article are all open-ended terms, meaning that they include but are not limited to.
[0023] This invention provides SCAR molecular markers for identifying the sex of cannabis, including MSY99M-2, MSY99M-3, MSY99M-4, MSY99M-5, MSY100M-1, MSY100M-4, MSY100M-5 or MSY100M-7; The nucleotide sequence of MSY99M-2 is shown in SEQ ID No. 1; The nucleotide sequence of MSY99M-3 is shown in SEQ ID No. 2; The nucleotide sequence of MSY99M-4 is shown in SEQ ID No. 3; The nucleotide sequence of MSY99M-5 is shown in SEQ ID No. 4; The nucleotide sequence of MSY100M-1 is shown in SEQ ID No. 5; The nucleotide sequence of the MSY100M-4 is shown in SEQ ID No. 6; The nucleotide sequence of the MSY100M-5 is shown in SEQ ID No. 7; The nucleotide sequence of the MSY100M-7 is shown in SEQ ID No. 8.
[0024] In the present invention, the nucleotide sequence of MSY99M-2 is TGTGTCTGTACTGTGCTTGTATAATCTTGAAGCCATTTGACCCATATATGATGTTTTAAGCTATAGGTCGTTATAGATTACACTAACAGGTGTTTTCCTTCTCCCTCCTTGTAGATACAAGTTTCCTTGAATTTTGCTTGAAACAATCAAGATCAGAGGCTGAGGTAATTTTGACGGAGAACCTTGGTTCATTTGATCCTGATCACGAGTTTATTGATAAATTTCTTAACTACAAAGAGTTGTTACCTGCTGATGTCCTTGAAATGGCTTTCCAAAGGCTAAATGATCAAATTGTTACTGGGTTCAGTGCTGTAGACGCTAAGTCCAACC (SEQ ID No.1); the nucleotide sequence of MSY99M-3 is CAACTATCAGCACCGTACCTACTATATTTGTTCTTTAGTTAAGTTTGTTGACAAAACTGTTGTATGTTTTCTCCTTGTATAAATTTTTTCGTTTTCCAAAGGACGGGCTCTTGTTGTGATTAAATTTATTTTATTATCATTAATTTATGTCTGATGGATGTATATACAAATGGTAAATAGTGAGAGGAAGTCTCAGCTTGTTGTTGTCATACCTACTTGGTATTTCCTGTTCATAAGTGGATTTACTTATTATGAGAACAAAGATTATAAGAATATATAAGTATATTCTAATTACTGCTTATGAATTTATTGCACGTGTCTTTGGTGATAGTCCTAAGGCTGCACATTAATGCTCCCAATAGTGAATGCTGTTTTTGGAGATGCATTGCTCTTGACTTGAATTTACTTATGAGCTGTGAACTCTCATGCTTAATCATTGTTGTTTTTGTTGAACGAGAGGGTATATCTAACTTTTTGCTTGTTCATTTTGGTAGCAGAACATATTTTCTTAGTCTCTCAGACCAGGGTAGTAGGGGAGTGGAAAACAACC (SEQ ID No.2); The nucleotide sequence of MSY99M-4 is GTAGAAGCCCCATCCCTTACTGTGGATGAGCCACTAGAGCCTATGGCATTGTCTTCACGAATTCCTGAGGAAATAGTAAGTTACATTTAATACTGTGTCCTTAATGCTTTATACTCTTGAGGAGGATATTGTAAGCAGTGGTGTTACTCCACAGATATCTAAAGAGCAACGAGATCAGGTGGGCTTTATGCAGTCAAGAACTAAGATTGGTATTTTACCTATTCATATAGTCGATTGAAAAACATTCATTTTTAGTTAATTCTCATGCTTTCAACCTGTGATGAATGGATTTGTTGTGCCTTTGCATTGCAGGTCGTAGAGAAGAC (SEQ ID No.3); The nucleotide sequence of MSY99M-5 is CCAGTTGACCAAACAAAGCAAGAAACTAAACAGTAGGATTTCTCTTTTTTCGTGTTTGCAGTTTGATAATTTAGCTTAGATTGTGCATACAATTTTTGAAGTTTTCTTCTTACCTTTTAAATTGAGTAAGGGTCGGTCTGTACTTTTTTATTTTTCATGTGATTAAGCAATTGAACTGGTTTTATGTAGATCCTGTTTAGTGCTGGCTGGGTGGTTGCAAAGTTTAGATAGATGTTTGTTTGAATGTGTTTGACTTATGTTTGTGCATGTTGTGTTTCCTAATTCTACCATGTGGTACCTTGACAGTGCTTACCAATAAATAACAGGAACAAATGTTAAATGGAGGTACCTGGTCTAAGACAACTATCAGCACCGTACCT (SEQ ID No.4); The nucleotide sequence of MSY100M-1 is CCCCCAATAGTTAAGTCGCTTAGGTAATATTAGCATGAGAGGGCAATGTGAAGATTCTCAGGCGAAAAAACCCTTAAGGTGCTAAATGAGGTTAGCGGATGGGAGGGTTGGTGGAGGTACCATGGAAAAGGTACCTTACGACTTATTAGAAAGGCAAAGAGATACCTCGTCCAACCATGTTTGATTCAAGGCTGAGACTCTCTTTGTATTCATAAGACTCTGGGTTTTATTGCAAAAAGTTTTAGGCAAGTTATTGTGTGTATTTTCTCTAGAACTTTTTACTTTTAATCTTCCATAGGGTATTAGTAACTTTCTAAAAAAATTTATTAGAGGTGGTGCCCATAAATTTTGGGAGAGCGATGCATGAAGTGAAGTCCTCAACCTTGGTGCCCTAAAAGGTTCGACACAGGTTTCT (SEQ ID No.5); The nucleotide sequence of MSY100M-4 is GGTGCTTCCTTTTGTGTGTTGAGATGGTCACTAATGTGGTGATGCTAAGATTGTGGTGCTGAGATGTCACCAACACAATCTCATGGTAAATGTCGAGATGGTGGTGTTGAGATGGTGGACTAGTGCATGTCGAGATTGTGGTGGTGAGATGGTCACCGACACCATCAGATTGTGCATGCCAAAATTGTAGTGTTGAGATTGTTAACAATGCCATCTCATGGTGCATGGCGAGATGGTGGTACCGAGATGATCCCCAACACCATATCCTGGTGCATGTCGAGATGGCACCAACTCCATCTCATGATGCATGCTAAGATGGCCTTCAAGCCTGGCTTGAACCTTGAGTGTTGATGCC (SEQ ID No.6); The nucleotide sequence of MSY100M-5 is AGTGCCAATGTAGACAACCACGATGTACACATGACAAATTTTCATTGATTCACTTAATATTAATAATTTATATTTTAAAAAAGTAATTTAAAATATTTAAAAAATAAAAAAAATTAAAAATTATTTAATATTTTTTTTCTATTTTCTTTCTTTTTCTTCTTCATCTTTCTTCTTCGTTTTTCTTTTTTTTTTTTTTTTTT CAGAAATCAAACCACCGCCATTATCATCTTCACGAGAAAATCGAACCACAACAACAACTACCACACCACCACCATTATCACCTTCACACTACAATCAACAAATTTTAACAGATTTAAAGAAAAAACACATTCTTTTGCCAAAAATAAGGCAAAAACACCCCAAATCAACATCTTTGTTCCAAGCTTTACAAAAATATTTTGAAGAGGAGAGGAGGTGAA (SEQ ID No.7); The nucleotide sequence of MSY100M-7 is CGGGTGGAACTTTGATCCTATTCATTATTTGCATCGGAGGTGCGTGAGGTACCTCTCTTGTTCCATCAATCTTTGATTTCTAAACAATGTTTTGGGGAAAGGACCTTATGTTTCTTCCCACTTAGGAGGGGGCCATAGGCGACGTGTTTGCTAGCCTAAAACATTTTGGATGAGTTTCCTCCCATGTTAATGGGG TGTTCCTTCAGTGTGACATTTAGAGATAGGATCGACTTATGTTGCCTCTGCTGACCTGAAGGATGTTTGACTAGCAACATGTCTTTTCTACATTTATCTTCATGGATATATTTAATAGAGAGCGTGGAGTTCTCTGCATTATCAAGTCTGATGGATTCTCGTTTCTCTTATTGTTTGCTTTGTATTTGTAGAGTGATGCTTGGGTTTTTCCCTC (SEQ ID No.8). .
[0025] This invention also provides the application of the aforementioned SCAR molecular marker in identifying the sex of cannabis.
[0026] In this invention, the cannabis is cannabis at any stage of growth.
[0027] This invention also provides the application of the aforementioned SCAR molecular marker in cannabis breeding.
[0028] This invention also provides a primer set for amplifying the SCAR molecular marker to identify the sex of cannabis, the primer set comprising the nucleotide sequence MSY99M-2-F as shown in SEQ ID No. 9 and the nucleotide sequence MSY99M-2-R as shown in SEQ ID No. 10; or Including nucleotide sequences such as MSY99M-3-F as shown in SEQ ID No. 11 and MSY99M-3-R as shown in SEQ ID No. 12; or Including nucleotide sequences such as MSY99M-4-F as shown in SEQ ID No. 13 and MSY99M-4-R as shown in SEQ ID No. 14; or Including nucleotide sequences such as MSY99M-5-F as shown in SEQ ID No. 15 and MSY99M-5-R as shown in SEQ ID No. 16; or Including nucleotide sequences such as MSY100M-1-F as shown in SEQ ID No. 17 and MSY100M-1-R as shown in SEQ ID No. 18; or Including nucleotide sequences such as MSY100M-4-F as shown in SEQ ID No. 19 and MSY100M-4-R as shown in SEQ ID No. 20; or Including nucleotide sequences such as MSY100M-5-F as shown in SEQ ID No. 21 and MSY100M-5-R as shown in SEQ ID No. 22; or This includes the nucleotide sequences MSY100M-7-F as shown in SEQ ID No. 23 and MSY100M-7-R as shown in SEQ ID No. 24.
[0029] In this invention, the nucleotide sequence of MSY99M-2-F is TGTGTCTGTACTGTGCTTGT (SEQ ID No. 9); the nucleotide sequence of MSY99M-2-R is GGTTGGACTTAGCGTCTACA (SEQ ID No. 10); the nucleotide sequence of MSY99M-3-F is CAACTATCAGCACCGTACCT (SEQ ID No. 11); the nucleotide sequence of MSY99M-3-R is GGTTGTTTTCCACTCCCCTA (SEQ ID No. 12); the nucleotide sequence of MSY99M-4-F is GTAGAAGCCCCATCCCTTAC (SEQ ID No. 13); the nucleotide sequence of MSY99M-4-R is GTCTTCTCTACGACCTGCAA (SEQ ID No. 14); and the nucleotide sequence of MSY99M-5-F is CCAGTTGACCAAACAAAGCA (SEQ ID No. 9). No. 15); the nucleotide sequence of MSY99M-5-R is AGGTACGGTGCTGATAGTTG (SEQ ID No. 16); the nucleotide sequence of MSY100M-1-F is CCCCCAATAGTTAAGTCGCT (SEQ ID No. 17); the nucleotide sequence of MYS100M-1-R is AGAAACCTGTGTCGAACCTT (SEQ ID No. 18); the nucleotide sequence of MSY100M-4-F is GGTGCTTCCTTTTGTGTGTT (SEQ ID No. 19); the nucleotide sequence of MYS100M-4-R is GGCATCAACACTCAAGGTTC (SEQ ID No. 20); the nucleotide sequence of MSY100M-5-F is AGTGCCAATGTAGACAACCA (SEQ ID No. 21); the nucleotide sequence of MYS100M-5-R is TTCACCTCCTCTCCTCTTCA (SEQ ID No. 15). No. 22); the nucleotide sequence of MSY100M-7-F is CGGGTGGAACTTTGATCCTA (SEQ ID No. 23); the nucleotide sequence of MSY100M-7-R is GAGGGAAAAACCCAAGCATC (SEQ ID No. 24).
[0030] The present invention also provides a kit for identifying the sex of cannabis, comprising the aforementioned primer set.
[0031] This invention also provides the application of the primer set described above in the preparation of products for identifying the sex of cannabis.
[0032] In this invention, the cannabis is cannabis at any stage of growth.
[0033] Example 1 Sample Processing Seeds from eight cannabis varieties—Longdama No. 1, Qingdama No. 1, Muma No. 1, Jingma No. 1, Fangke No. 1, Longdama No. 2, Dinamed KUSH, and Tatanka—were used for cannabis cultivation and phenotypic observation. Detailed sample information is shown in Table 1. After observing the inflorescence phenotype, the sex of the samples was recorded. Leaves were collected and immediately dried in silica gel for preservation and total DNA extraction experiments.
[0034] Genomic data were obtained from Lynch (Lynch RC, Padgitt-Cobb LK, Garfinkel AR, et al. Domesticated cannabinoid synthases amid a wild mosaic). cannabis The genome sequences of 24 chromosome-level haplotypes from 4 males and 8 females published by Pangenome [J]. Nature, 2025, 643(8073):1001-1010. are available at https: / / resources.michael.salk.edu / root / tools.html?tool=%2Fresources%2Fcannabis_genomes%2Findex.html. Detailed genome information is shown in Table 2.
[0035] Table 1 Sample Information
[0036] Table 2. Detailed Genome Information
[0037] Example 2: Sex chromosome alignment and screening for specific regions Winnowmap v2.03 was used to perform cross-alignment of the genomic sequences in Table 2. Bedtools v2.26.0 was used to process the alignment intervals, including merging, intersecting, and complementing operations, and GetFastA was used to extract the corresponding sequences. First, the Y chromosome sequence in the BCMb genome was aligned to the female haplotype genome of BCMa to obtain candidate Y chromosome-specific regions that could not be aligned (dataset 1). Then, the X chromosome sequences in dataset 1 were aligned separately with the X chromosome sequences in the remaining 19 genomes (Table 2 numbers: 1, 5-8, 10-13, 15-24). The union was then taken, and the complement was taken with dataset 1 to filter out sequence fragments in dataset 1 that could be aligned to the female cannabis genome, resulting in further filtered Y chromosome-specific regions (dataset 2). Simultaneously, Dataset 1 was compared with the Y chromosomes of the other three genomes (Table 2 numbers: 2, 9, 14), and the intersection was taken to screen out the Y chromosome-specific segments common to different varieties (Dataset 3). Finally, the intersection of Dataset 2 and Dataset 3 was taken to obtain the final Y chromosome-specific sequence (Dataset 4).
[0038] Comparative analysis revealed 16,120 candidate Y chromosome-specific regions (dataset 1), with an average length of approximately 1,003 bp. After filtering through 19 X chromosome sequences, dataset 2 contained 29,438 regions, with an average length of approximately 180 bp. Dataset 3, obtained by intersecting with three other Y chromosome sequences, contained 13,496 regions, with an average length of approximately 1,095 bp. Finally, combining datasets 2 and 3, 15,648 Y chromosome-specific sequences were selected (dataset 4), with an average length of approximately 275 bp. All regions were distributed across the Y chromosome in the BCMb genome, ranging from 28.6 Mb to 107.8 Mb, similar to those found by Lynch et al. (Lynch RC, Padgitt-Cobb LK, Garfinkel AR, et al. Domesticated cannabinoid synthases amid a wild mosaic). cannabis The Y chromosome SDR region reported in pangenome[J]. Nature, 2025, 643(8073):1001-1010. is consistent with that reported in pangenome[J]. Nature, 2025, 643(8073):1001-1010.
[0039] Example 3: Design of male-specific molecular markers Dataset 4 was sorted by length, and the longer regions ChrY:99.94Mb–99.95Mb and ChrY:100.78Mb–100.79Mb were selected as candidate regions for marker development. Specific primers (200-600bp) were designed using Primer-Blast (https: / / www.ncbi.nlm.nih.gov / tools / primer-blast) and synthesized by Beijing Liuhe BGI Genomics Co., Ltd. Based on the two longer regions ChrY:99.94Mb–99.95Mb and ChrY:100.78Mb–100.79Mb (12,025bp and 10,252bp respectively), a total of 12 primer pairs were designed. The primer sequences are shown in Table 3.
[0040] Table 3 Primer Sequences
[0041] Example 4: Extraction of Genomic DNA Take an appropriate amount of the sample from Table 1, wipe the surface with 75% ethanol, blot dry, and place 20 mg in a 1.5 ml centrifuge tube. Grind the sample finely using a high-throughput tissue homogenizer. Extract total DNA from the sample using a plant genomic DNA extraction kit (purchased from Nanjing Novizan Biotechnology Co., Ltd., DC104-01 FastPure Plant DNA Isolation Mini Kit) following the instructions. The concentration and purity of the extracted DNA were detected using a micro-spectrophotometer: OSE-260-03 TGem Pro spectrophotometer (purchased from Tiangen Biotech (Beijing) Co., Ltd.). The specific genomic DNA extraction results are shown in Table 4.
[0042] Table 4 Genomic DNA Extraction Quality
[0043] Example 5: Screening and Validation of Molecular Markers Three male and three female plants were selected as initial screening materials, namely serial numbers 2, 27, 29, 33, 11 and 36 in Table 1. PCR amplification was performed using 12 pairs of designed primers. ITS2 (Internal Transcribed Spacer, GenBank ID: GQ434337.1) and the reported male-specific markers MADC5 and MADC6 were set as positive controls, and a blank template was used as a negative control (the PCR system included ITS2 primers but no genomic DNA). The PCR reaction system was: 13 μL of 2×Taq PCR Mix (purchased from Nanjing Novizan Biotechnology Co., Ltd.), 1 μL of forward primer, 1 μL of reverse primer, 9 μL of ddH2O, and 1 μL of genomic DNA; the amplification program was: 95℃ for 4 min, [95℃ for 30 s, 55℃ (… SCAR 119 F / R 52℃ 30s, 72℃ 10s] 35 cycles, 72℃ 10min. The amplified products were detected by 1% agarose gel electrophoresis (TSINGKE TSJ001 high-concentration low-electroosmotic agarose, purchased from Beijing Qingke Biotechnology Co., Ltd.). A DL2000 Plus DNA Marker (purchased from Nanjing Novizan Biotechnology Co., Ltd.) was used as a DNA band control, and the results were recorded using a GenoSens 2150 gel imaging system (purchased from Shanghai Qinxiang Scientific Instruments Co., Ltd.). Candidate primers with good specificity were screened. The results are as follows: Figure 1 As shown.
[0044] Subsequently, using the candidate primers obtained from the screening, 36 sex-identified samples (15 males and 21 females) from Table 1 were further selected for amplification and validation. The PCR reaction system and amplification procedure were the same as those used during the screening. ITS2 (Internal Transcribed Spacer, GenBank ID: GQ434337.1) and the reported male-specific markers MADC5 and MADC6 were used as positive controls, and a blank template was used as a negative control (the PCR system included ITS2 primers but not genomic DNA). This was to assess the effectiveness and stability of the markers. The results are shown below. Figure 2 , 3 As shown.
[0045] Table 5 ITS2 primers and control group primers
[0046] Depend on Figure 1The initial PCR screening results showed that 8 out of 12 primer pairs amplified clear bands in male plants but no amplification products were found in female plants, demonstrating male specificity. The remaining 4 pairs showed non-specific amplification or weak positive bands in some female samples. The positive control ITS2 amplified in all samples, while the negative control showed no band, indicating the reliability of the experimental system. The previously reported MADC6 marker showed non-specific bands in the 750-1000 bp range in some female plants, while the MADC5 marker showed weak positive bands in some female plants, suggesting insufficient specificity.
[0047] Depend on Figure 2 and Figure 3 It was found that all eight candidate markers exhibited 100% identification accuracy in 36 samples. Among them, MSY99M-3, MSY99M-4, MSY99M-5, MSY100M-1, and MSY100M-7 showed bright amplification bands in male samples, but no obvious target bands or non-specific amplification bands in female samples, demonstrating the best specificity and stability, and can be used as preferred molecular markers for sex identification of cannabis seedlings.
[0048] As can be seen from the above embodiments, the present invention provides SCAR molecular markers, primer sets, kits and applications for identifying the sex of cannabis. The SCAR markers developed by the present invention based on Y chromosome-specific segments have high specificity and stability, and can be used for rapid and accurate identification of early sex in cannabis, which has important application value for improving the production efficiency of medicinal cannabis-related products.
[0049] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A SCAR molecular marker for identifying the sex of cannabis, characterized in that, Including MSY99M-2, MSY99M-3, MSY99M-4, MSY99M-5, MSY100M-1, MSY100M-4, MSY100M-5 or MSY100M-7; The nucleotide sequence of MSY99M-2 is shown in SEQ ID No. 1; The nucleotide sequence of MSY99M-3 is shown in SEQ ID No. 2; The nucleotide sequence of MSY99M-4 is shown in SEQ ID No. 3; The nucleotide sequence of MSY99M-5 is shown in SEQ ID No. 4; The nucleotide sequence of MSY100M-1 is shown in SEQ ID No. 5; The nucleotide sequence of the MSY100M-4 is shown in SEQ ID No. 6; The nucleotide sequence of the MSY100M-5 is shown in SEQ ID No. 7; The nucleotide sequence of the MSY100M-7 is shown in SEQ ID No.
8.
2. The application of the SCAR molecular marker as described in claim 1 in identifying the sex of cannabis.
3. The application of the SCAR molecular marker as described in claim 1 in cannabis breeding.
4. A primer set for amplifying the SCAR molecular marker of claim 1 for identifying the sex of cannabis, characterized in that, The primer set includes the nucleotide sequence MSY99M-2-F as shown in SEQ ID No. 9 and the nucleotide sequence MSY99M-2-R as shown in SEQ ID No. 10; or Including nucleotide sequences such as MSY99M-3-F as shown in SEQ ID No. 11 and MSY99M-3-R as shown in SEQ ID No. 12; or Including nucleotide sequences such as MSY99M-4-F as shown in SEQ ID No. 13 and MSY99M-4-R as shown in SEQ ID No. 14; or Including nucleotide sequences such as MSY99M-5-F as shown in SEQ ID No. 15 and MSY99M-5-R as shown in SEQ ID No. 16; or Including nucleotide sequences such as MSY100M-1-F as shown in SEQ ID No. 17 and MSY100M-1-R as shown in SEQ ID No. 18; or Including nucleotide sequences such as MSY100M-4-F as shown in SEQ ID No. 19 and MSY100M-4-R as shown in SEQ ID No. 20; or Including nucleotide sequences such as MSY100M-5-F as shown in SEQ ID No. 21 and MSY100M-5-R as shown in SEQ ID No. 22; or This includes MSY100M-7-F, as shown in SEQ ID No. 23, and MSY100M-7-R, as shown in SEQ ID No.
24.
5. A reagent kit for identifying the sex of cannabis, characterized in that, Includes the primer set as described in claim 4.
6. The use of the primer set according to claim 4 in the preparation of products for identifying the sex of cannabis.