A grape liquid-phase capture probe combination based on a pan-genome, a detection kit and application thereof

By screening grape SNP sites based on multi-source resequencing data and designing liquid-phase microarrays, the deterministic bias problem of existing grape SNP arrays in varieties with large differences in genetic background has been solved, enabling high-throughput and flexible molecular marker applications and improving the accuracy and efficiency of grape breeding.

CN122105000APending Publication Date: 2026-05-29NANJING AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing grape SNP arrays rely on a single or limited number of reference genomes during the design process, resulting in significant differences in genetic background between Eastern cultivars and wild grape germplasm, leading to deterministic bias, making it difficult to fully reflect genetic diversity, and thus having insufficient application value in diversified breeding.

Method used

Based on resequencing data from 466 grape germplasm resources, combined with GWAS analysis and literature mining, 8576 core SNP sites were screened, biotin-labeled single-stranded DNA or RNA probes were designed, and liquid-phase microarrays were constructed to cover low-frequency variations and complex genomic regions in different varietal populations, forming a high-throughput and flexible molecular marker system.

Benefits of technology

It improves the accuracy and efficiency of grape breeding, especially showing high prediction accuracy in the early identification of seedless traits and skin color traits, and shortens the breeding cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The application discloses a grape whole genome SNP site combination, a capture probe, a liquid phase chip and application thereof. Based on the grape PN40024 T2T complete reference genome, 20094 SNP sites covering the whole genome and uniformly distributed are screened through large-scale resequencing and whole genome association analysis (GWAS), including 8576 core sites and 11518 flanking sites. The site combination covers functional markers closely related to important agronomic traits such as seedlessness and fruit peel color. The application further provides a liquid phase hybridization capture probe and a kit designed based on the above-mentioned site. Compared with a traditional solid phase chip, the liquid phase chip has the advantages of high site coverage, high detection sensitivity, accurate genotyping and flexible cost, and can be widely applied to grape variety identification, molecular assisted breeding and whole genome selection breeding, and significantly improves the new variety breeding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of molecular marker-assisted breeding technology, specifically to a grape SNP molecular marker combinatorial system, probe group, liquid-phase gene chip constructed based on pan-genome variation information, and its application in grape breeding. Background Technology

[0002] Grapes are one of the most widely cultivated, largest-area, and oldest fruit trees in the world. Their fruit is widely used for winemaking, fresh consumption, juicing, and drying. Rich in nutrients and with high economic benefits, grapes have attracted widespread consumer attention and have promising development prospects. As breeding objectives shift from single traits to the synergistic improvement of multiple traits such as seedlessness, skin color, flavor and aroma, and stress adaptability, the demand for high-throughput, high-accuracy molecular marker tools is becoming increasingly urgent.

[0003] Single nucleotide polymorphisms (SNPs) are the most abundant form of DNA polymorphism in eukaryotic genomes. Since most SNPs have only two alleles, developing high-throughput, easily automated methods can simplify the genotyping process. Currently, more than 50 SNP arrays and 15 different types of genotyping platforms have been developed in more than 25 crops and perennial trees, including maize, rice, wheat, potato, sunflower, peanut, and cotton.

[0004] Compared to traditional solid-phase gene chips, liquid-phase gene chips rely on the core principle of covalently coupling oligonucleotide probes to the surface of fluorescently encoded microspheres, enabling hybridization with target genomic regions in a liquid environment. Because hybridization occurs in a liquid environment, they offer higher detection sensitivity, higher throughput, faster speed, and better reproducibility. Secondly, traditional solid-phase chips can only detect fixed SNP sites, which are limited in number and cannot be changed. Liquid-phase chips, on the other hand, allow for flexible combinations of microspheres with different probes to meet research needs, offering greater flexibility. Compared to traditional solid-phase gene chips, liquid-phase gene chips have significant value in research and applications requiring high throughput, high sensitivity, and flexibility, such as multiplex gene expression analysis, genotyping, and pathogen detection. However, the full realization of the advantages of the liquid-phase chip platform is highly dependent on the scientific validity and representativeness of the front-end SNP site system. If the selected SNP sites themselves have deterministic biases or insufficient population coverage, it will be difficult to fully demonstrate its technological advantages in breeding applications.

[0005] Currently, multiple SNP-related loci have been obtained through technologies such as whole-genome sequencing and association analysis, and grape gene chips have been developed. However, most of these are applied to genotyping and principal component analysis, and a systematic application scheme for practical breeding screening and early trait identification has not yet been formed. Therefore, it is necessary to design and generate liquid-phase chips based on the key SNP loci obtained from the analysis to improve the efficiency of grape breeding. Currently published grape SNP arrays (such as Vitis18kSNP) are generally constructed based on a single reference genome or a limited number of cultivar variation information during the design process. This type of design strategy has certain applicability in major European and American varieties, but in Eastern cultivar populations and wild grape germplasm, due to significant differences in genetic background, it is prone to serious ascertainment bias, resulting in a large number of real variation loci not being included in the chip design, manifested as low polymorphism detection rate and insufficient number of effective genotyping loci. The above-mentioned defects make it difficult for existing grape SNP arrays to comprehensively reflect the genetic diversity of different germplasm resources and to meet the needs of diversified breeding groups and complex breeding objectives for highly reliable molecular marker tools.

[0006] With advancements in sequencing technology, the grape reference genome has been updated from the earlier 8X and 12X versions to the PN40024 T2T (Telomere-to-Telomere) version. The T2T genome fills gaps in complex regions such as centromeres and telomeres, providing a foundation for developing more accurate and comprehensive genotyping tools. However, relying solely on a single or a few reference genomes is still insufficient to fully characterize the complex genetic variation features of grape populations. To overcome deterministic bias caused by limited sample sources, it is crucial to systematically introduce large-scale, multi-source resequencing data during the microarray design phase. The purpose of introducing multiple resequencing datasets is not simply to increase the number of samples, but to cover low-frequency variations, rare allelic variations, and structural variations in complex genomic regions specific to different varietal populations, thereby improving the universality and effectiveness of SNP loci in different genetic backgrounds. Based on these technical requirements, constructing an SNP molecular marker system that integrates large-scale resequencing data is a key technical approach to solving the deterministic bias problem of existing grape microarrays and enhancing their application value in diverse breeding populations. Therefore, the development of high-throughput liquid phase chips based on T2T genomes is of great significance for modern grape breeding. Summary of the Invention

[0007] One objective of this invention is to provide a combination of SNP molecular markers for grapes. Based on resequencing data from 466 grape germplasm resources (including 324 proprietary sequencing data and 142 public data), and using PN40024 T2T as a reference genome, this invention, combined with GWAS analysis and literature mining, screened out 8576 core SNP loci. These SNP loci are evenly distributed on chromosomes and include functional loci significantly associated with traits such as seedlessness, color, and aroma in grapes. Based on these technical requirements, constructing an SNP molecular marker system integrating large-scale resequencing data is a key technical approach to solving the deterministic bias problem of existing grape microarrays and enhancing their application value in diversified breeding populations. This technical approach differs from traditional microarray design methods based on a single reference genome or screening of variations in a small number of samples, constituting the fundamental improvement direction of this invention compared to existing technologies.

[0008] A second objective of this invention is to provide a probe array for detecting the aforementioned SNP combinations. The probes are biotin-labeled single-stranded DNA or RNA probes designed based on the aforementioned SNP sites and their flanking sequences, preferably with a length of 120 bp.

[0009] A third objective of this invention is to provide a grape liquid phase chip and kit comprising the above-described probe group.

[0010] The fourth objective of this invention is to provide the application of the above-mentioned technical solutions in grape breeding, particularly in the early identification of seedless traits and skin color traits. Detailed Implementation

[0011] Example 1: Screening of core SNP loci in the whole genome of grape

[0012] 1. Data source: Whole genome resequencing was performed on 466 representative grape germplasm resources with an average sequencing depth of 15X.

[0013] 2. Variance detection: Sequencing data were aligned to the PN40024 T2T reference genome. Variance detection was performed using the GATK pipeline, initially yielding approximately 8.59 million SNPs.

[0014] 3. Core site filtering: Set screening criteria: (1) Minimum allele frequency (MAF) > 0.1; (2) Deletion rate < 10%; (3) Uniform distribution principle: Use the sliding window method (e.g., every 200Kb or 1Mb window) to preferentially retain sites with high polymorphism information content (PIC).

[0015] 4. Functional site integration: GWAS analysis based on 2 years of field phenotypic data yielded 136 significantly associated SNPs, of which 103 were successfully designed; based on literature reports, 82 known functional SNPs were included, of which 43 were successfully designed.

[0016] 5. Final Set: After redundancy removal and probe design feasibility assessment, 20,094 SNP sites were finally identified (see Table 1), including 8,576 core SNP sites and 11,518 flanking sites. The specific physical locations, reference genomic coordinates, and variant types of the 8,576 SNP sites are detailed in Appendix Table 1 of the instruction manual. The average spacing of the 8,576 core target sites was 7 kb. Genomic characteristics of the core target sites were statistically analyzed using a 7 kb sliding window, showing that the core target sites are evenly distributed throughout the genome (see Appendix Table 1). Figure 1 Loci with a minimum allele frequency >0.1 accounted for 85.11% of the core target sites, as shown in the specific distribution below. Figure 2 As shown in the figure. Combining genome functional annotation and gene structure information, the characteristics of the core target sites are annotated, and the annotation results are as follows. Figure 3 As shown, the main annotation type is intergenic.

[0017] Table 1. Location information of all SNP sites in the grape liquid phase gene chip.

[0018]

[0019]

[0020]

[0021]

[0022]

[0023]

[0024]

[0025]

[0026]

[0027]

[0028]

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046]

[0047]

[0048]

[0049]

[0050]

[0051]

[0052]

[0053]

[0054] Example 2 Design and fabrication of liquid-phase trapping probe

[0055] Capture probes were designed for the SNP sites identified in Example 1. The design principles included:

[0056] 1. Length and GC content: The probe length is 120nt, and the GC content is controlled between 30% and 60%.

[0057] 2. Specificity assessment: The sequence was unique in the T2T genome and contained no microsatellite repeats.

[0058] 3. Thermodynamic stability: The Tm value is controlled within a certain range of 65-80°C.

[0059] 4. Model scoring: Use deep learning models to predict probe capture efficiency, remove low-scoring probes, and finally synthesize a probe library (Panel).

[0060] Example 3: Chip Performance Verification

[0061] Twenty grape varieties with known phenotypes (covering different subgroups) were selected for performance testing.

[0062] 1. Method:

[0063] (a) Extraction and detection of DNA samples: Extract sample DNA using conventional plant or animal genomic DNA extraction methods. It is preferred to use commercial genomic DNA extraction kits, such as the Suzhou Cretaceous Biological Genomic DNA Extraction Kit (CNA01990F64), but not limited to this kit.

[0064] DNA was extracted from the sample using conventional plant or animal genomic DNA extraction methods, preferably using a commercially available genomic DNA extraction kit, such as the Suzhou Cretaceous Biological Genomic DNA Extraction Kit (CNA01990F64), but not limited to this kit.

[0065] The extracted DNA was analyzed for integrity and degradation using 1.0%–1.5% agarose gel electrophoresis, and its concentration was determined using quantitative real-time fluorescence (qRF) assays, preferably using a Qubit series qRF system. The DNA sample concentration was ideally controlled at 10–50 ng / μL to meet the requirements for subsequent library construction.

[0066] (II) DNA Library Construction

[0067] DNA library construction can employ transposase-based random fragmentation methods, preferably using the Tn5 transposase library construction system, such as the YZSeq™ Tn5 Library Preparation Kit (Wuhan Yingzi, T1012-008), but is not limited to this commercial kit. Take 20–100 ng of genomic DNA for fragmentation. The transposition reaction temperature is preferably 50–60°C, and the reaction time is 5–15 min, to obtain DNA fragments with an average fragment length of approximately 150–300 bp.

[0068] The fragmented DNA is ligated with sequencing adapter sequences, which can be universal adapters for the Illumina platform or functionally equivalent adapter sequences. PCR amplification is then performed using high-fidelity DNA polymerase. Preferred PCR conditions include: pre-denaturation: 94–98°C, 1–3 min; denaturation: 94–98°C, 10–30 s; annealing: 55–65°C, 20–40 s; extension: 68–72°C, 20–60 s; cycle number: 8–16 cycles; final extension: 68–72°C, 2–5 min.

[0069] After PCR amplification, the library is purified using magnetic beads to screen for fragments. The preferred method is a two-sided magnetic bead screening ratio of 0.4×–0.6× and 0.1×–0.2× to remove fragments that are too short or too long and obtain the target library.

[0070] (III) Library hybridization and enrichment

[0071] The constructed library was hybridized with a liquid-phase microarray probe array containing probes for the target SNP sites. The hybridization reaction can refer to commercial liquid-phase microarray hybridization systems, such as the functional site liquid-phase microarray system from Wuhan Yingzi Company, but is not limited to this system. The hybridization temperature is preferably controlled at 60–65°C, and the hybridization time is 12–24 h. After hybridization, the target fragment was enriched by magnetic bead capture and eluted at 55–65°C for a preferred time of 5–15 min. The eluted products were then enriched by PCR amplification, with a preferred number of PCR cycles of 10–16.

[0072] After enrichment, the library was purified by magnetic beads, and the library concentration was determined by fluorescence quantitative PCR. The library fragment length was detected by high-sensitivity electrophoresis. The average fragment length of the library was preferably 300–400 bp.

[0073] (iv) Sequencing

[0074] Libraries that pass quality control can be sequenced on high-throughput sequencing platforms, preferably platforms that support paired-end sequencing, such as DNBSEQ™, Illumina NovaSeq, or equivalent sequencing systems. The preferred sequencing strategy is paired-end sequencing (PE100–PE150).

[0075] 2. Results: Average target region coverage >98%, genotyping consistency of replicate samples >99% (see...) Figure 4 This proves that the chip has extremely high accuracy and stability.

[0076] Example 4: Molecular marker-assisted selection for nucleate-free traits

[0077] The chip of this invention was used to test 30 hybrid offspring with unknown traits. The analysis focused on key loci located on chromosome 18 (corresponding to the VviAGL11 gene region).

[0078] 1. Results Analysis: The judgment criteria were set as follows: when the genotypes of loci SNPA20002 and SNPA20004 were heterozygous (0 / 1) or homozygous for mutation (1 / 1), they were judged as nucleus-free; when they were reference homozygous (0 / 0), they were judged as nucleus-containing.

[0079] 2. Validation: Compared with actual traits in the field, the prediction accuracy of this group of markers reached over 73% (see Table 2). This indicates that the chip can effectively replace conventional field screening and significantly shorten the breeding cycle.

[0080] Table 2. Prediction results of nucleate-free traits using liquid-phase microarray of SNPs in the grape genome. Test sample code Nucleated / Nucleated Genotype prediction accuracy (%) Test sample code Nucleated / Nucleated Genotype prediction accuracy (%) 1-1 Nucleus 100 2-1 Nucleus 100 1-2 Nucleus 100 2-2 Nucleus-free 0 1-3 Nucleus-free 0 2-3 Nucleus-free 0 1-4 Nucleus-free 0 2-4 Nucleus 0 1-5 Nucleus 100 2-5 Nucleus 100 1-6 Nucleus 100 2-6 Nucleus 100 1-7 Nucleus 100 2-7 Nucleus-free 100 1-8 Nucleus 100 2-8 Nucleus-free 100 1-9 Nucleus 100 2-9 Nucleus 100 1-10 Nucleus 100 2-10 Nucleus-free 0 1-11 Nucleus 100 2-11 Nucleus-free 100 1-12 Nucleus 100 2-12 Nucleus-free 0 1-13 Nucleus 100 2-13 Nucleus 100 1-14 Nucleus 100 2-14 Nucleus 100 1-15 Nucleus-free 0 2-15 Nucleus 100

[0081] Example 5: Application of the Identification of Fruit Peel Color Characteristics

[0082] The MybA gene cluster-related sites on Chr02 were analyzed using the chip of this invention.

[0083] 1. Results Analysis: The focus was on loci SNPA01012 and SNPA01015. When a colored haplotype was detected, the predicted pericarp was colored (red / purple / black); when a homozygous recessive haplotype was detected, the predicted pericarp was colorless (yellow / green).

[0084] 2. Validation: In 46 test samples, the color trait prediction accuracy was as high as 100% (see Table 3).

[0085] Table 3. Prediction results of color traits from grape genome-wide SNP liquid phase microarray. Test sample code Colored / Colorless Genotype prediction accuracy (%) Test sample code Colored / Colorless Genotype prediction accuracy (%) 1-1 colorless 100 2-1 colored 100 1-2 colorless 100 2-2 colorless 100 1-3 colored 100 2-3 colorless 100 1-4 colorless 100 2-4 colored 100 1-5 colorless 100 2-5 colored 100 1-6 colorless 100 2-6 colored 100 1-7 colored 100 2-7 colored 100 1-8 colored 100 2-8 colored 100 1-9 colored 100 2-9 colored 100 1-10 colored 100 2-10 colored 100 1-11 colored 100 2-11 colorless 100 1-12 colored 100 2-12 colorless 100 1-13 colored 100 2-13 colored 100 1-14 colorless 100 2-14 colored 100 1-15 colored 100 2-15 colorless 100 Attached Figure Description

[0087] Appendix Figure 1 This is a uniform distribution map of SNP loci in this invention (probe coverage density was calculated on the chromosome with an observation window of 0.1 Mb length). This map is based on the grape reference genome PN40024 T2T.

[0088] Appendix Figure 2 This shows the MAF distribution of the SNP sites in this invention.

[0089] Appendix Figure 3 This is the annotation result of the core target site features of the present invention.

[0090] Appendix Figure 4 This is a statistical chart showing the detection rate and consistency distribution of duplicate samples in this invention.

Claims

1. A genome-wide SNP molecular marker combinatorial system for grapes, characterized in that, The molecular marker combination contains 20,094 SNP sites, the physical location and variation information of which are selected from Table 1 of the specification based on the grape reference genome PN40024 T2T version.

2. A grape whole genome liquid-phase chip, characterized in that, It includes the molecular marker combination of claim 1 and the probe designed according to it.

3. A kit for grape genotyping, characterized in that, It includes the grape whole genome liquid phase chip as described in claim 2, and reagents for library construction, hybridization capture and elution.

4. The application of the SNP molecular marker combination of claim 1, the liquid phase chip of claim 2, or the kit of claim 3 in grape variety identification, genetic diversity analysis of germplasm resources, or kinship identification.

5. The application of the SNP molecular marker combination of claim 1, the liquid phase chip of claim 2, or the kit of claim 3 in marker-assisted breeding of grapes.

6. The application according to claim 5, characterized in that, The specific application is to screen for grape skin color traits or seedless traits; wherein, the SNP markers used to screen for seedless traits include at least SNPA20002 and SNPA20004 located on Chr18; and the SNP markers used to screen for skin color traits include at least SNPA01012 and SNPA01015 located on Chr02.