A cauliflower 15k snp liquid chip based on targeted capture sequencing and application thereof

By developing a cauliflower 15K SNP liquid-phase chip based on targeted capture sequencing, utilizing machine learning to optimize sites, and combining probe mixtures and hybridization capture reagents, the technical gap of high precision and low cost in cauliflower breeding has been bridged, achieving high-throughput genotyping of the whole genome and improving breeding efficiency and adaptability.

CN121472478BActive Publication Date: 2026-04-10CHINA AGRI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA AGRI UNIV
Filing Date
2026-01-12
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

There is a technological gap in cauliflower breeding, namely, high precision, low cost, and strong generalization. Linear models and randomized label combinations cannot resolve non-additive effects. The economic barriers of high-density chips hinder the widespread adoption of the technology. Static design strategies are difficult to adapt to breeding scenarios in multiple environments.

Method used

We developed a liquid-phase chip for cauliflower 15K SNPs based on targeted capture sequencing. It employs a machine learning-driven site optimization strategy, combined with probe mixtures and hybridization capture reagents, to achieve low-cost, high-throughput genotyping of the whole genome. The modular design supports dynamic functional expansion and is adaptable to complex breeding scenarios.

Benefits of technology

It enables low-cost, high-throughput genotyping of the entire cauliflower genome, improves breeding efficiency, supports flexible adaptation to breeding scenarios in multiple environments, and provides efficient tools for genetic diversity analysis, variety identification, and molecular-assisted breeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121472478B_ABST
    Figure CN121472478B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of cauliflower 15K SNP liquid chip based on targeted capture sequencing and its application.The present application discloses a kind of cauliflower high-density liquid chip based on 14432 specific SNP sites and its breeding application.The core site of the chip is derived from the whole genome resequencing data analysis of multiple cauliflower germplasm resources, after the resequencing data is compared to cauliflower reference genome "C-8 (V2)", it is filtered after strict quality control, polymorphism evaluation and genome distribution optimization, finally 15K high-value SNP marker set covering whole genome is screened out.High-throughput genotyping is carried out on cauliflower germplasm or breeding material using this chip, which can efficiently support germplasm resource genetic background analysis, variety authenticity and purity rapid identification, important trait gene / QTL fine mapping and whole genome selection breeding and other key links.Application of the technology will significantly improve cauliflower breeding efficiency, shorten the breeding cycle and accelerate the process of new variety breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of genome breeding technology, specifically to a cauliflower 15KSNP liquid phase chip based on targeted capture sequencing and its application. Background Technology

[0002] Cauliflower ( Brassica oleracea var. botrytis As a globally important cruciferous vegetable crop, cauliflower's planting area exceeded 1.4 million hectares in 2023, with an annual output of 26 million tons. Its nutritional value is outstanding, rich in vitamin C, dietary fiber, and glucosinolates with anti-cancer activity, making it a promising candidate for functional food development. However, cauliflower yield-related traits exhibit typical polygenic control characteristics and are significantly influenced by genotype-environment interactions (G×E), resulting in low breeding efficiency. Although marker-assisted selection (MAS) technology was once highly anticipated, the limited number of locatable quantitative trait loci (QTLs) limits its explanatory power for phenotypic variations in complex traits, making it difficult to support the breeding needs for high-yielding and stable varieties.

[0003] The rise of Genomic Selection (GS) technology has brought revolutionary breakthroughs to crop breeding. This method uses whole-genome molecular markers to calculate the Genomic Predicted Breeding Value (GEBV), avoiding the limitations of single QTL models. However, the application of GS in cauliflower has lagged behind for a long time. Ultra-high density marker coverage is required to capture effective genetic signals. At the same time, mainstream linear statistical models (such as GBLUP and BayesB) assume that traits are normally distributed and cannot analyze non-additive genetic effects, resulting in insufficient prediction accuracy for key yield traits. The high cost of single-sample detection using commercial high-density SNP chips also leads to serious waste of resources.

[0004] Existing genotyping technologies face significant challenges in cauliflower breeding. Low-throughput PCR-based markers (such as SSR and CAPS) typically cover only a limited number of sites, making it difficult to analyze multi-gene regulatory networks. Simplified genome sequencing technologies (GBS and RAD-seq) can perform medium-throughput genotyping, but their random site detection and poor data compatibility between platforms hinder collaboration across breeding projects. While high-density SNP chips provide whole-genome coverage, their high cost and noise interference severely limit their large-scale application.

[0005] In recent years, machine learning (ML) algorithms (such as Gradient Boosting Tree (GBDT) and Random Forest (RF)) have shown breakthrough potential in crop GS (Gross Genesis). These models can capture nonlinear genetic interactions, theoretically improving the accuracy of complex trait predictions. Practice in wheat breeding has shown that ML models significantly improve yield prediction accuracy compared to traditional methods. However, the application of ML technology in cauliflower has not yet been realized.

[0006] Liquid phase capture chip is based on DNA capture, followed by sequencing, which performs multiple sequencing on the same position, has higher accuracy, low cost, and flexible SNP chip density development; new sites can be supplemented at any time, and upgrading is flexible; in the future, different combinations of chips can be used for different use scenarios and use requirements, that is, detection. However, this technology has not been applied in cauliflower, and the development of the cauliflower chip is helpful to the development of the cauliflower breeding industry.

[0007] In summary, there is a "high precision-low cost-strong generalization" triple technical gap in the current cauliflower breeding field: the combination of linear models and random markers cannot break through the non-additive effect analysis bottleneck; the economic barrier of high-density chips hinders the technology popularization; the static design strategy is difficult to adapt to the multi-environment breeding scene. Therefore, a cauliflower-specific SNP chip that integrates machine learning site optimization, targeted capture technology, and multi-environment verification system is needed to drive the paradigm upgrade of breeding efficiency. SUMMARY

[0008] In view of the defects in the prior art, the purpose of the present application is to provide a cauliflower 15K SNP liquid phase chip based on targeted capture sequencing and its application. The use of the liquid phase chip of the present application can realize the rapid typing of cauliflower germplasm resources, and solve the problem that a large amount of genetic information cannot be applied to actual breeding.

[0009] To achieve the above purpose, the technical scheme adopted by the present application is:

[0010] A cauliflower 15K liquid phase chip, characterized in that the genotyping object of the chip includes 14432 SNP sites on the whole genome of cauliflower; the position information of the SNP sites is shown in the last attached table of the specification.

[0011] The reference genome of the position of the SNP site is cauliflower genome C-8 (v2) (https: / / www.ncbi.nlm.nih.gov / nuccore / JAMKOK000000000); the chromosome genome described in the following of the present application is the above genome.

[0012] The above-mentioned cauliflower 15K liquid phase chip is a cauliflower 15K liquid phase gene chip based on targeted capture sequencing.

[0013] On the basis of the above-mentioned scheme, characterized in that the cauliflower 15K liquid phase chip is composed of a probe mixed solution and a hybrid capture reagent;

[0014] The probe mixed solution includes cauliflower SNP targeted capture 15K probe, universal blocking solution I, and universal blocking solution II.

[0015] The hybridization capture reagent comprises TCGBS 2x hybridization Buffer, TCGBS repeat sequence blocking solution, TCGBS 2x Beads Wash Buffer, TCGBS Wash Buffer I / Wash Buffer II / Wash Buffer III or TCGBS Stringent Wash Buffer.

[0016] On the basis of the above scheme, the Cauliflower SNP targeting capture 15K probe in the probe mixture is a DNA double-stranded probe.

[0017] The application of a Cauliflower 15K liquid chip in genetic diversity analysis of Cauliflower germplasm resources, variety identification, gene positioning or molecular assisted breeding.

[0018] On the basis of the above scheme, the method of the application comprises the following steps: genotyping a Cauliflower sample by using the Cauliflower 15K liquid chip.

[0019] On the basis of the above scheme, the application of the genotyping is specifically: through the genotyping, molecular assisted breeding of Cauliflower is carried out; the genotyping is based on a mutation site at Chromosome 6 Chr6-40570621 on Cauliflower genome C-8 (v2), the alleles of the mutation site are T / G, and different alleles are selected for molecular assisted breeding of tight type (GG and GT type) or loose (TT type) Cauliflower.

[0020] The primer for identifying the genotype of the above mutation site is characterized in that the sequence of the primer is shown as SEQ ID NO. 7-8.

[0021] The Cauliflower 15K SNP liquid chip based on targeted capture sequencing and the application thereof have the following beneficial effects:

[0022] High-density Cauliflower liquid chip technology system is created.

[0023] (1) The 15K SNP liquid chip (containing 14432 sites) developed by the application realizes low-cost and high-throughput genotyping of the whole Cauliflower genome for the first time through a site optimization strategy driven by machine learning.

[0024] (2) Modular design supports dynamic expansion of functions

[0025] The chip adopts an open probe architecture, and users can flexibly supplement new functional sites according to breeding targets. After the new site is verified by a standard biological information process, it can be directly integrated into the existing chip system without the need to redesign the platform, which significantly improves the technical iteration efficiency.

[0026] (3) Full process adaptation to complex breeding scene

[0027] The cauliflower 15K liquid phase gene chip provides an efficient genome typing tool for genetic diversity analysis, variety identification, gene positioning, background screening and molecular marker assisted breeding of cauliflower; meanwhile, according to existing research, the application can design new probes according to new sites at any time, and has no higher requirement for sample quantity, so as to meet different use scenes and use requirements of users, and has high application value in cauliflower molecular breeding. BRIEF DESCRIPTION OF DRAWINGS

[0028] The application has the following drawings:

[0029] Figure 1 Phylogenetic tree based on whole genome markers of cauliflower;

[0030] Figure 2 Phylogenetic tree based on 15K markers of cauliflower;

[0031] Figure 3 Hybrid Jinpin 75 female parent;

[0032] Figure 4 Hybrid Jinpin 75 male parent;

[0033] Figure 5 Hybrid Jinpin 75;

[0034] Figure 6 Difference SNP distribution map between female parent PN_0601 and male parent QB-4;

[0035] Figure 7 SNP typing map for detecting Jin 75 hybrid;

[0036] Figure 8 Tight and loose cauliflower in QU-26 and QU-27 mixed pool F2;

[0037] Figure 9 BSA analysis based on whole genome;

[0038] Figure 10 BSA analysis based on 15K markers;

[0039] Figure 11 SNP typing map for detecting mixed pool F2 hybrid;

[0040] Figure 12 DNA fingerprint. DETAILED DESCRIPTION

[0041] The application will be further described in detail below in combination with the drawings.

[0042] Example 1 Design of Cauliflower 15K Liquid Chip

[0043] The design of Cauliflower 15K Liquid Chip is based on whole genome resequencing and machine learning driven site optimization strategy. 394 representative cauliflower inbred lines were used as materials, and their resequencing data (Illumina NovaSeq 6000 platform, PE150) were aligned to the cauliflower reference genome "C-8(V2)", and variant detection was performed by GATK HaplotypeCaller. High-quality SNP sites were screened through multi-level quality control: secondary allele sites were retained, the minor allele frequency (MAF) was required to be ≥0.05, the detection rate was required to be ≥80%, and linkage disequilibrium (LD) pruning was performed by PLINK (window 100 kb, step 50 bp, r²<0.9), and finally 550K background SNP set was obtained.

[0044] Site screening combines feature importance ranking and saturation analysis for dual optimization. Four machine learning algorithms (GBDT, GBM, RF, XGBoost) were used for genomic prediction of 12 flower ball yield-related traits (flower ball diameter, flower ball height, flower ball single ball weight, whole plant weight, second node branch length, first node branch length, middle leaf length, plant high foot degree, maximum leaf length, maximum leaf width, plant type angle, and total leaf number). The model performance was evaluated by 10-fold cross-validation. Based on the feature importance ranking of the GBM model, SNP with high contribution to trait prediction was selected preferentially; further optimization by chromosome distribution ensured that the sites were evenly distributed on 9 chromosomes (average density 1,700 SNP / chromosome), and redundant sites in high LD regions were removed, and finally a high-representative 15K SNP core set was screened.

[0045] Example 2 Preparation of Cauliflower 15K Liquid Chip

[0046] 1. Extraction of Cauliflower Leaf Genomic DNA Sample: After the cauliflower grows true leaves, CTAB method is used for nucleic acid extraction, then 1% agarose gel electrophoresis is used to detect the integrity of DNA, and Agilent 2100 instrument is used to determine the concentration. The qualified sample is stored in -20℃ refrigerator for standby.

[0047] 2. Data determination of liquid chip: the qualified sample is randomly broken by ultrasonic, the breaking is carried out by ultrasonic fragmentation by using Covaris S2 ultrasonic instrument, the value of Covaris system is set according to the standard, 3 cycles x 60 s, water bath temperature: 4℃, duty cycle: 20%, intensity: 5, mode: Frequency sweeping, the fragments after breaking are recovered by gel electrophoresis to obtain DNA fragments with a length of 300-500 bp, the ends of the fragments are added with adaptors to form a whole genome library. The library is amplified by using LM-PCR to form a sequencing library, the library is freeze-dried in liquid nitrogen, 1.5 μL of 15K probe of Cauliflower SNP targeted capture and related hybridization reagents (universal blocking solution I 2 μL, universal blocking solution II 2 μL; TCGBS 2x hybridization buffer 2 μL, TCGBS repeat sequence blocking solution Block 2 μL, TCGBS 2x bead wash buffer 8.5 μL, TCGBS wash buffer I / wash buffer II / wash buffer III 2.7 μL, TCGBS stringent wash buffer 1.8 μL) are added, and hybridization and denaturation are carried out at 65℃ for 12-16 h, the unhybridized DNA is washed away by repeatedly washing with washing solution for 2-3 times, and 5 cycles of PCR reaction (PCR mix 25 μL, hybridization DNA 10 μL, primer mix 2.5 μL 95℃ 5 s, 65℃ 20 s, 72℃ 10 s) are carried out, the hybridization and capture sequencing library is formed, and sequencing is carried out, and 1 Gb of data is output for each sample.

[0048] The design and synthesis process of the 15K probe is as follows: the sequences of 50-60 bp of the upstream and downstream of the site information described in synthesis example 1, that is, the probe length is about 120 bp, the number of homologous regions is ≤3, a biotin group is labeled at the 5' end of each probe, and each probe in the probe group is mixed to form a 15K probe chip in an equimolar ratio.

[0049] Example 3 Application of 15K liquid chip in Cauliflower genetic diversity analysis

[0050] The application specifically utilizes the 15K liquid chip of Cauliflower to group the core germplasm materials.

[0051] The accurate evaluation of Cauliflower germplasm resources is a key basis for breeding improvement. Traditional phenotypic identification is easily disturbed by the environment, while whole genome resequencing can provide comprehensive information, but it needs high cost to reach 10x depth in the face of about 568 Mb genome. The 15K liquid chip of Cauliflower developed in the application realizes high-throughput genotyping at a cost of less than 90% of resequencing by covering specific sites of the whole genome, and provides an efficient solution for germplasm resource evaluation.

[0052] Population genetic analysis was conducted on 394 core germplasm accessions (i.e., the 394 representative cauliflower inbred lines mentioned above, including 364 early-maturing and 30 late-maturing germplasm accessions) from the Tianjin Academy of Agricultural Sciences using this chip. Figure 2 The early-maturing germplasm includes typical early-maturing materials PN_0526, PN_0538, and PN_0157, and the late-maturing germplasm includes late-maturing materials PN_0013, PN_0026, and PN_0052. A maximum likelihood phylogenetic tree based on 15K SNPs and a phylogenetic tree constructed based on the whole genome (…) Figure 1 The similarity between the two groups, along with their mutual embedding, indicates that there is no significant genetic differentiation between them. This confirms the historical background of the introduction of Chinese cauliflower germplasm homology and demonstrates that the chip can accurately analyze subtle variations within the subgroup.

[0053] Example 4: Application of 15K liquid phase chip in cauliflower variety identification

[0054] Ensuring the purity of cauliflower seeds is a core requirement for industrialization. Traditional methods rely on phenotypic observation, which takes up to 70 days and is prone to significant errors. However, the 15K SNP liquid-phase chip provided by this invention achieves accurate and efficient identification by detecting the frequency deviation of heterozygous sites in the population. Combined with high-throughput genotyping capabilities, seed testing can be completed in a short time, completely solving the timeliness problem of seed quality control for breeding companies. This case study focuses on Jinpin 75 cauliflower (… Figure 5 Parents are respectively as follows Figure 3 and Figure 4 As shown, 15K chip detection was performed, and KASP markers were applied based on the differential sites between the parents to achieve seed purity identification in a low-cost and efficient manner. The specific implementation is as follows:

[0055] 1. Obtaining homozygous differential loci in parents

[0056] (1) Perform quality control on the raw sequencing data to obtain clean data for analysis; the raw data is RAW data to obtain high-quality clean reads for subsequent analysis. The sequencing data filtering steps are as follows: 1) Remove reads containing adapters; 2) Remove reads with more than 3 N; 3) Remove low-quality reads (the number of bases with quality value Q < 5 accounts for more than 20% of the entire read);

[0057] (2) Align the Clean Data with the Reference Genome; After obtaining the clean reads, use BWA software to align the clean reads with the reference genome. The initial alignment results are in SAM format, and then use SAMtools software to convert the results to BAM format and sort them. If the results of a sample contain multiple libraries, use SAMtools to merge the BAM results of multiple libraries, use picard to mark repetitive sequences, and perform basic data information statistics;

[0058] (3) SNP mutation detection was performed; GATK software (v3.8 second-generation chip sequencing mutation detection software https: / / software.broadinstitute.org / gatk / ) was used to detect SNPs;

[0059] (4) SNP screening: filter out sites with QUAL value (base quality value) less than 30, MQ value less than 30, and DP value less than 2, and select SNP sites that are homozygous and differential between the parents as candidate sites.

[0060] 2. Molecular markers developed based on differential analysis

[0061] According to the parents of Jinpin 75 ( Figure 3 and Figure 4 Differences between sequencing sequences () Figure 6 Candidate SNP sites were selected, and KASP primers (as shown in Table 1) were designed using the online primer design software SNP Primer (www.snpway.com). These primers consisted of a pair of specific primers for SNP alleles containing different fluorescent adapters (FAM and HEX) and a reverse common primer. The FAM fluorescent adapter sequence was GAAGGTGACCAAGTTCATGCT, and the HEX fluorescent adapter sequence was GAAGGTCGGAGTCAACGGATT. Different differentially expressed sites were revealed by the fluorescence of either FAM or HEX. Primer pairs with appropriate specificity and annealing temperatures were selected, and the primers were synthesized by a third-party company.

[0062] Table 1 KASP primer design

[0063] Chr5-49428221-F GAAGGTGACCAAGTTCATGCTAGAGAATTAACATTTCGGCATCAA (SEQ ID NO. 1) Chr5-49428221-R GAAGGTCGGAGTCAACGGATTAGAGAATTAACATTTCGGCATCAC (SEQ ID NO. 2) Chr5-49428221-C CACAGAAGTGGTCAAGGATAGTCA (SEQ ID NO. 3) Chr7-56503584-F GAAGGTGACCAAGTTCATGCTTGAACGAAGACTAACCTTTTTGCTA (SEQ ID NO. 4) Chr7-56503584-R GAAGGTCGGAGTCAACGGATTTGAACGAAGACTAACCTTTTTGCTG (SEQ ID NO. 5) Chr7-56503584-C TGGAGATGGTTTGAACGAAGACTA (SEQ ID NO. 6)

[0064] The application of molecular markers specifically includes the following steps:

[0065] (1) Using the genomic DNA of the sample to be tested as a template, Touchdown PCR amplification is performed using the amplification primers of the molecular marker to obtain the amplification product; the Touchdown PCR amplification procedure is as follows: 94°C for 15 min; 95°C for 20 s; 65°C-56°C for 60 s, 10 cycles, with the annealing and extension temperature decreasing by 0.8°C for each cycle; 94°C for 20 s; 57°C for 60 s, 26 cycles;

[0066] (2) The amplification product is detected and analyzed; if the sample PCR product detects the HEX fluorescent signal corresponding to the primers Chr5-49428221 and Chr7-56503584, the corresponding detection site is C:C, G:G genotype, and it is determined to be the paternal type false hybrid; if the sample PCR product detects the FAM fluorescent signal corresponding to the primers Chr5-49428221 and Chr7-56503584, the corresponding detection site is A:A, A:A genotype, and it is determined to be the maternal type false hybrid; if both FAM and HEX fluorescent signals are detected, the detection site is A:C, A:G genotype, and it is determined to be a true hybrid; if no fluorescent signal is detected, it is other genotype combination, and it is determined to be other type hybrid (in actual detection, no fluorescent signal is produced, which may be caused by powder mixing in actual production);

[0067] (3) The ratio of true hybrids to total detection samples is calculated, and the flower cauliflower hybrid seed purity is calculated (Table 2).

[0068] Table 2 Seed purity identification results

[0069] Detection index Number of strains Number of effective data strains 188 Number of true hybrid strains determined 185 Number of female type hybrid strains 3 Number of male type hybrid strains 0 Number of other type hybrid strains 0 Purity 98.4%

[0070] In breeding, by molecular marker identification and screening, retaining the detection of the FAM fluorescent signal corresponding to the primers Chr5-49428221 and Chr7-56503584 can determine the maternal false hybrid; retaining the detection of the HEX fluorescent signal corresponding to the primers Chr5-49428221 and Chr7-56503584 can determine the paternal false hybrid; if both fluorescent signals are detected, the detection site is A:C, A:G genotype, and it is determined to be a true hybrid (true hybrid) Figure 7 ), and the purity of the commodity can be identified through the previous molecular marker screening.

[0071] 3. Low-cost and low-computing-power linkage trait marker development based on 15K chips

[0072] The loose-type flower cauliflower advanced inbred material 'QU-27' is used as the paternal parent, the compact-type flower cauliflower advanced inbred material 'QU-26' is used as the maternal parent, 'QU-26' and 'QU-27' are crossed to obtain F1 generation, and F2 population is obtained after selfing of the F1 generation.

[0073] 10 loose and 10 tight zucchini plants were randomly selected from F2 population and their leaves were mixed respectively (Mix 1 and Mix 2) Figure 8 ), and the total DNA of the two mixed pools was extracted by CTAB method. The DNA of the two mixed pools was constructed into library by TruSeq DNA LT Sample Prep Kit, and the genome resequencing was performed by Illumina Novaseq 6000. The data obtained was mapped by bwa and SNP calling was performed by samtools. The filtering standard was that the base quality was greater than or equal to 30, the sequencing depth was greater than or equal to 2, and the mapping quality value was greater than or equal to 30. 79743 high-quality homozygous difference SNP sites were screened out. Subsequently, a SNP-index distribution map was drawn with 1 Mb as the window and 0.1 Mb as the step (as shown in Figure 9 ), the candidate interval was on chromosome 6, and the highest peak was Chr6:40570621. At the same time, based on the mixed pool analysis of 15k chip markers, the results were consistent with the previous ones, indicating that the same results as the whole genome could be achieved based on the 15k chip (as shown in Figure 10 ).

[0074] 4. Development of linkage markers

[0075] By using the BSA population of 15K chip combined with whole genome BSA positioning, a marker linked to the ball tightness trait was located, and the mutation site was located at Chr6-40570621 on chromosome 6. After developing KASP molecular markers associated with the ball tightness trait, the molecular markers could be directly used for identification of the ball tightness trait phenotype and corresponding genotype, and then assisted breeding was carried out relying on the molecular markers to speed up the breeding efficiency. By using the molecular markers early, the target plants can be quickly screened, the planting scale is effectively reduced, the workload of later field identification is reduced, and the selection efficiency and accuracy are improved.

[0076] Embodiment:

[0077] KASP primers were designed using online primer design software SNP Primer (www.snpway.com), and the actual operation was carried out according to the operation process described in the seed purity identification part of the above text. The marker is:

[0078] Chr6-40570621-F: GAAGGTGACCAAGTTCATGCTGCACCAAACATTGCTCCTAGATG (SEQ ID NO. 7)

[0079] Chr6-40570621-R: GAAGGTCGGAGTCAACGGATTGCACCAAACATTGCTCCTAGATT (SEQ ID NO. 8)

[0080] Chr6-40570621-C: ATCTGGGACACCTTGAAAAGTTTG (SEQ ID NO. 9)

[0081] The F2 population of 359 individuals of QU-26 and QU-27 was genotyped with the marker. There were three fluorescence signals, among which the fluorescence signal of G:G had 90 individuals, the fluorescence signal of G:T had 181 individuals, and the fluorescence signal of T:T had 88 individuals, which was consistent with the segregation of 1:2:1. Combined with the phenotype investigation data, it was found that the genotype was completely consistent with the tightness phenotype, and the coincidence rate reached 100% ( Figure 11 ). The above results fully demonstrate that the Chr6-40570621 marker has universality and accuracy, and can be applied to the prediction, identification and screening of cauliflower ball tightness.

[0082] 5. Precise identification of materials based on 15K chip and construction of fingerprint map

[0083] In view of the chaos of infringement and germplasm plagiarism of cauliflower varieties, the chip innovatively constructs a whole genome machine learning authentication system, pre-stores the fingerprint database of 48 SNP markers of the previously described 394 core germplasms ( Figure 12 ), and calculates the variety similarity of the test sample by GBM algorithm, which can accurately distinguish materials with similar genetic backgrounds, and can also trace the parent plagiarism behavior, provide a judicial level evidence chain for variety right protection, and promote the upgrading of seed industry intellectual property protection technology.

[0084] The contents not described in detail in the specification belong to the prior art known to those skilled in the art.

[0085] Position information of SNP sites in the attached table

[0086] ;

[0087] ;

[0088] ;

[0089]

[0090] ;

[0091] ;

[0092] ;

[0093] ;

[0094]

[0095] ;

[0096]

[0097] ;

[0098] ;

[0099] ;

[0100]

Claims

1. A cauliflower 15K liquid phase chip, characterized in that, The genotyping target of this chip includes 14,432 SNP sites on the entire cauliflower genome; The reference genome for the location of the SNP site is the cauliflower genome C-8 (v2). The cauliflower 15K liquid phase chip consists of a probe mixture and a hybridization capture reagent; The probe mixture includes a cauliflower SNP-targeted capture 15K probe, universal blocking solution I, and universal blocking solution II; The hybridization capture reagents include TCGBS 2× Hybridization Buffer, TCGBS Repeat Sequence Block, TCGBS 2× Beads Wash Buffer, TCGBS Wash Buffer I / Wash Buffer II / Wash Buffer III or TCGBS Stringent Wash Buffer; The location information of the SNP sites is shown in the table below: ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; 。 2. The cauliflower 15K liquid phase chip as described in claim 1, characterized in that, The cauliflower SNP-targeting 15K probe in the probe mixture is a DNA double-stranded probe.

3. The application of the cauliflower 15K liquid phase chip as described in any one of claims 1-2 in the genetic diversity analysis, variety identification, gene mapping, or molecular-assisted breeding of cauliflower germplasm resources.

4. The application as described in claim 3, characterized in that, The application includes: using the cauliflower 15K liquid phase chip to perform genotyping on cauliflower samples.

5. The application as described in claim 4, characterized in that, This genotyping method is used for molecular-assisted breeding of cauliflower. The genotyping is based on a mutation site at Chr6-40570621 on chromosome 6 of the cauliflower genome C-8 (v2). The genotyping of this mutation site is T / G. Different genotypes are selected for molecular-assisted breeding of compact or loose cauliflower. The compact cauliflower is of the GG and GT types, and the loose cauliflower is of the TT type.

6. Primers for identifying the genotype of the mutation site as described in claim 5, characterized in that, The sequences of the primers are shown in SEQ ID NO.7-8.

Citation Information

Patent Citations

  • KASP marker primer group linked with broccoli ball-flower yield character and application of KASP marker primer group

    CN117660689A

  • Brassica campestris whole genome liquid phase chip and application thereof

    CN120310946A