Pea whole genome 20K SNP liquid phase breeding chip and application thereof

By developing the pea genome 20K SNP liquid-phase breeding chip, the problems of medium and high cost and low density of pea breeding are solved, efficient genotype detection is achieved, breeding efficiency and variety innovation are improved, and it is suitable for a variety of breeding analysis.

CN120536614AActive Publication Date: 2025-08-26XIANGHU LABORATORY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510528625.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-17
Filing Date
2025-04-25
Publication Date
2025-08-26
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The prior art has high resequencing costs and low site density of traditional solid-phase chips in pea breeding, which limits the application of large genomic species and leads to inefficient breeding.

Method used

A pea whole genome 20K SNP liquid phase breeding chip is developed, which contains 21,659 high-quality SNP sites. It uses high-throughput sequencing technology and liquid phase hybridization capture technology, and is suitable for platforms such as illumina and MGI to achieve efficient and low-cost genotype detection.

Benefits of technology

It has improved the efficiency of pea breeding, promoted variety innovation and improvement, and is suitable for important trait analysis, genetic diversity analysis of germplasm resources, variety identification, population structure analysis, genetic and evolutionary analysis and kinship identification, shortening the breeding gap between pea and staple food crops.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120536614A_ABST
    Figure CN120536614A_ABST
Patent Text Reader

Abstract

The invention discloses a pea whole genome 20K SNP (Single Nucleotide Polymorphism) liquid phase breeding chip and application thereof, the liquid phase breeding chip comprises 21 and 659 SNP loci on a pea whole genome, and is specifically based on agronomic trait correlation analysis and whole genome re-sequencing data of a pea sample, re-sequencing data of a variety is added, and the pea whole genome 20K SNP liquid phase breeding chip is obtained. And 21,659 high-quality and representative SNP loci are screened from the high-quality and representative SNP loci. The chip can be widely applied to pea genotyping, and is beneficial to improvement of pea breeding efficiency and innovation of pea varieties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of biotechnology applications and relates to a pea whole-genome 20K SNP liquid-phase breeding chip and an application thereof. Background Art

[0002] Pea (Pisum sativum L.) is an important food source, animal feed, and industrial raw material, with outstanding nutritional and health benefits. In recent years, significant progress has been made in breeding and improving pea for key agronomic traits. However, due to its large and complex genome, genomics-based breeding with complex phenotyping faces numerous challenges, severely restricting the rapid development of basic research and breeding. With the rise of computational breeding and biological breeding, the breeding gap between pea and staple crops has further widened. To narrow this gap, the development of breeding chips suitable for pea is urgently needed. These will help scientists more effectively identify and utilize promising genes, improve breeding efficiency, promote the innovation and improvement of pea varieties, and bring new growth points to agricultural production.

[0003] The development and application of gene chips have played a vital role in accelerating the analysis of crop genetic characteristics, variety identification, resource conservation, and breeding efforts. Using SNP microarrays, researchers can analyze population genetic structure, identify candidate genes associated with economic traits, and implement various breeding studies, including genomic selection. Therefore, it is essential to screen for molecular markers associated with important economic traits and accelerate the selection of pea varieties using cutting-edge molecular breeding techniques.

[0004] Genome-wide association studies and genome-wide selective breeding are the foundations of modern breeding. However, the high cost of resequencing and the low locus density and flexibility of traditional solid-phase microarrays limit their application in species with large genomes, such as pea and wheat. Therefore, developing an efficient, rapid, and low-cost large-scale genotyping tool will promote the development of pea molecular breeding in my country.

[0005] Site information is the foundation of gene chip probe design, and site selection is the core and key step in chip design, directly impacting the practical value of the chip. For a long time, site selection was primarily based on criteria such as allele frequency, deletion rate, and heterozygosity. With the advancement of pea genomics research, more reference information is available for site selection and evaluation of new chip designs.

[0006] Therefore, it is necessary to develop a new breeding chip to provide new tools for pea breeding. Summary of the Invention

[0007] The present invention aims to provide a pea whole-genome 20K SNP liquid-phase breeding chip and its application to address the aforementioned existing technical problems. This pea 20K SNP liquid-phase breeding chip can be widely used in pea important trait analysis, germplasm genetic diversity analysis, variety identification analysis, assisted selection breeding, population structure analysis, genetic and evolutionary analysis, kinship identification, genome-wide association analysis, or genomic selection. It can improve pea breeding efficiency, promote innovation and improvement of pea varieties, and bring new growth points to agricultural production.

[0008] The purpose of the present invention can be achieved through the following technical solutions:

[0009] A pea whole-genome 20K SNP liquid-phase breeding chip, characterized in that the genotyping objects of the chip include 21,659 SNP sites on the pea whole genome, as shown in Table 1 below.

[0010] The reference genome of the SNP site is the pea genome "Pisum sativum cultivar Zhewan1 (PeaZW1)".

[0011] The present invention also provides the use of the liquid phase chip in genotyping pea samples.

[0012] The beneficial effects of the present invention are:

[0013] 1. The pea 20K SNP liquid-phase breeding chip of the present invention is based on association analysis of 57 agronomic traits and whole-genome resequencing data from 237 pea samples. In addition, resequencing data from 77 previously studied varieties were incorporated. From these data, 21,659 high-quality, representative SNP sites were screened and obtained, including 235 published reliable functional sites significantly associated with the 57 traits.

[0014] 2. Based on high-throughput sequencing technology, the one-time output of the result information is large, and nearly a thousand samples can be tested simultaneously. It is suitable for mainstream second-generation sequencing platforms such as Illumina and MGI, and has wide platform adaptability.

[0015] 3. The pea 20K SNP liquid-phase breeding chip of the present invention can be widely used in important pea trait analysis, germplasm genetic diversity analysis, variety identification analysis, assisted selection breeding, population structure analysis, genetic and evolutionary analysis, and kinship identification. It can improve pea breeding efficiency and promote the innovation and improvement of pea varieties. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 The figure shows the distribution of the SNP sites of the present invention on the pea genome;

[0017] Figure 2 The figure shows the number and spacing of the SNP sites on each chromosome of the present invention;

[0018] Figure 3 This is a graph showing the probe capture rate of the pea 20K SNP liquid-phase breeding chip sites in 14 representative samples of the present invention;

[0019] Figure 4 To construct a phylogenetic tree using the chip genotyping results of 14 pea samples;

[0020] Figure 5 The diagram is a pea 20K SNP liquid-phase breeding chip combined with population segregation analysis (BSA) of the present invention; wherein diagram (a) is a gene location map; and diagram (b) is a resequencing gene location map. DETAILED DESCRIPTION

[0021] The present invention will be further illustrated below with reference to the detailed description of specific embodiments. However, the following examples are merely illustrative of the technical solutions of the present invention and do not limit the technical solutions of the present invention. The experimental methods in the following examples, unless otherwise specified, are conventional methods and are performed according to the techniques or conditions described in the literature in the art or according to the product instructions. The materials, reagents, etc. used in the following examples, unless otherwise specified, can all be obtained from commercial sources.

[0022] The SNP referred to in the present invention refers to single-nucleotide polymorphism (SNP), which mainly refers to DNA sequence polymorphism at the genomic level caused by single nucleotide variation in the nucleotide sequence, including single base deletion, insertion, transversion and conversion.

[0023] Example 1

[0024] Design of pea 20K SNP liquid-phase breeding array:

[0025] The present invention provides a SNP site combination for pea variety identification, comprising 21,659 SNP sites on the whole pea genome, wherein the reference genome of the SNP site combination is the pea genome "Pisum sativum cultivar Zhewan 1 (Pea ZW1)", and the reference genome source is shown in the document Liu N, Lyu X, Zhang X, et al. Reference genome sequence and population genomic analysis of peas provide insights into the genetic basis of Mendelian and other agronomic traits [J]. Nature Genetics, 2024, 56(9): 1964-1974.

[0026] The location and variation information of the SNP sites are shown in Table 1.

[0027] Table 1

[0028]

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046]

[0047]

[0048]

[0049]

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056]

[0057]

[0058]

[0059]

[0060]

[0061]

[0062]

[0063]

[0064]

[0065]

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081]

[0082]

[0083]

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172]

[0173]

[0174]

[0175]

[0176]

[0177]

[0178]

[0179]

[0180]

[0181]

[0182]

[0183] Among them, Chr1, Chr2, Chr3, Chr4, Chr5, Chr6, and Chr7 represent the chromosome IDs where the SNP sites are located, and the number following the chromosome ID indicates the position of the SNP site on the chromosome.

[0184] The specific preparation process of the 21,659 SNP molecular marker site combination liquid phase breeding chip is as follows:

[0185] 1. Germplasm resource collection

[0186] A diverse panel of 237 varieties was cultivated at the Yangdu Research Base of the Zhejiang Academy of Agricultural Sciences (120.21551°E, 30.25308°N). These varieties were planted on November 12, 2020, and 2021. Resequencing data for 77 previously studied varieties were also included. The cultivar categories are shown in Table 2. Young leaves were collected two weeks after planting and quickly frozen in liquid nitrogen for DNA extraction. A total of 57 traits were investigated. Two replicates were performed in 2020 and 2021. Each trait was measured in five independent experiments, and the average was calculated. The pea varieties used are listed below.

[0187] Table 2

[0188]

[0189]

[0190]

[0191] 2. Data quality control

[0192] Genomic DNA was extracted from fresh leaves using the CTAB method. DNA libraries for Illumina sequencing were constructed using NEBNext Ultra DNA LibraryPrep Kits. DNA was randomly fragmented using a Covaris ultrasonic disruptor, end-repaired, and A-tail adapters were added. Purified DNA was amplified by PCR, and the library was evaluated and sequenced using DNBSEQ-T7 in PE150 mode. The resulting raw image data files were converted to raw sequencing sequences through base recognition analysis, referred to as Raw Reads. The results were stored in the FASTQ file format. To ensure the quality of information analysis, Raw Reads were filtered using fastp (v.0.12.2) with default parameters to remove sequencing adapters and low-quality bases, generating Clean Reads for subsequent information analysis.

[0193] 3. Design of 20K SNP Liquid-Phase Breeding Microarray for Pea

[0194] Clean reads were aligned to the reference genome PeaZW1 using the mem algorithm in BWA (v.0.7.17-r1188) software. After alignment, PCR duplicates were removed using the Picard tool (v.1.123). A single BAM file was generated for each sample for SNP detection. Subsequently, the SAM file format was converted and sorted using Samtools (v.1.9), and the alignment quality was filtered using the 'q30' parameter. The HaplotypeCaller module of the Genome Analysis Toolkit (GATK, v.4.1.9.0) was used to extract raw SNPs. Individual SNP extraction used hard filtering with mapping quality ≤ 20.0, minimum sequencing coverage ≤ 5, and maximum sequencing coverage ≥ 200. Then, the variants were merged into a single variant call file using GATK Combine Variants. Possible false SNPs were further filtered using VCFtools (v.0.1.16), and the criteria included:

[0195] (1) If a marker was not detected in more than 20% of the samples, the marker was filtered out, that is, the genotype missing rate was ≤ 20%, and the parameter was vcftools --max-missing 0.8;

[0196] (2) If the minimum allele frequency of a marker is lower than 0.01, it is filtered out, that is, the minimum allele frequency (MAF) is < 0.01, and the parameter is vcftools --maf 0.01;

[0197] (3) Extract the 100 bp sequence upstream and downstream of the SNP marker and perform blast alignment of the whole genome sequence. If the similarity is greater than 60% and the alignment length is greater than 60% (about 120 bp), it is considered to be a copy, and only single-copy sites are retained;

[0198] (4) Extract the 100 bp sequence upstream and downstream of the SNP marker, calculate the GC content, and retain sites with a GC content greater than 40% and less than 60%;

[0199] (5) Extract the 100 bp sequence upstream and downstream of the SNP marker, analyze whether it contains SSR sequences, and retain sites without SSR sequences;

[0200] (6) Analyze the InDels within 100 bp upstream and downstream of the SNP marker, and retain the SNPs without InDels of more than 5 bp within 100 bp upstream and downstream;

[0201] (7) Extract the 100 bp sequence upstream and downstream of the SNP marker, analyze whether the sequence contains N, and retain the sequence without N within 100 bp upstream and downstream of the SNP site.

[0202] (8) Addition of chip functional sites

[0203] A large-scale GWAS was conducted using the filtered SNPs. Association analysis was performed using the Efficient Mixed-Model Association eXpedited (EMMAx) (v.2012-021-0) program with default parameters. The first three principal components were used as covariates. The effective number of independent SNPs and the suggested P value were estimated using the Genetic type 1 Error Calculator (gec, v.0.2) software. In the mixed model, the suggested P value was P < 4.63 × 10^-7. Ultimately, 235 reliable SNP functional loci significantly associated with 57 traits were identified.

[0204] The above methods ultimately screened 21,659 high-quality, representative SNP sites, including 235 published reliable functional sites significantly associated with 57 traits. The distribution of sites on chromosomes is shown in Figure 2. Figure 1 The number and spacing of the chromosomes are shown in Figure 2 shown.

[0205] 4. Development of a 20K SNP Liquid-Phase Breeding Chip for Pea

[0206] The 21,659 SNPs ultimately identified were used by Tianjin Jizhi Gene Technology Co., Ltd. to design and synthesize liquid-phase capture probes for the development of a pea 20K SNP liquid-phase breeding array using TargetSeq liquid-phase hybridization capture technology. TargetSeq liquid-phase hybridization capture technology uses a multi-factor algorithm to design capture probes for the target genomic region, synthesizes effective and specific probes at high throughput, and then performs liquid-phase hybridization with fragmented genomic DNA to capture and enrich the target region sequence. Finally, high-throughput sequencing is performed using leading sequencing platforms such as Illumina and Life Technology.

[0207] Example 2

[0208] The method for using the pea 20K SNP liquid-phase breeding chip of the present invention to perform genotyping on pea samples is as follows:

[0209] First, the genomic DNA is fragmented to form small fragments of 200-300bp. The DNA fragments are repaired at the ends and connected to the pre-PCR amplification library. The probe is hybridized to the target region. Then, the hybridization probe is captured with streptavidin affinity-labeled magnetic beads to enrich the target fragments, elute them, and perform post-PCR amplification after capture. Finally, the target region is subjected to second-generation sequencing. The raw image data files obtained by the high-throughput sequencing DNBSEQ-T7 sequencing platform are converted into raw sequencing sequences through base recognition analysis, called RawData or Raw Reads. The results are stored in the FASTQ file format, which contains the sequence information of the sequencing sequence and its corresponding sequencing quality information. The raw sequencing sequence or Raw Reads obtained by sequencing contain low-quality reads with adapters. In order to ensure the quality of information analysis, the Raw Reads need to be filtered. The filtered Clean Reads are used for subsequent information analysis.

[0210] The main steps of data filtering are as follows:

[0211] (1) Remove reads with adapters.

[0212] (2) Filter reads with N content exceeding 10%.

[0213] (3) Remove reads with bases with a quality value lower than 10 accounting for more than 50%.

[0214] The sequencing reads obtained by resequencing need to be relocated to the reference genome before subsequent variation analysis can be performed. The bwa software is mainly used to align the short sequences obtained by the second-generation high-throughput sequencing DNBSEQ-T7 sequencing platform with the reference genome. By aligning and locating the position of Clean Reads on the reference genome, the sequencing depth, genome coverage and other information of each sample are counted, and variation detection is performed. Subsequently, Samtools (v.1.9) is used to convert the SAM file format, sort the BAM file, and filter the alignment quality with the 'q30' parameter. The HaplotypeCaller module of the Genome Analysis Toolkit (GATK, v.4.1.9.0) is used to extract the original SNPs. Individual SNP extraction uses hard filtering of mapping quality ≤20.0, minimum sequencing coverage ≤5, and maximum sequencing coverage ≥200. Then, GATK Combine Variants is used to merge the variants into a single variant call file. VCFtools (v.0.1.16) is used to further filter possible false SNPs. The criteria include:

[0215] (1) If a marker was not detected in more than 20% of the samples, the marker was filtered out, that is, the genotype missing rate was ≤ 20%, and the parameter was vcftools --max-missing 0.8;

[0216] (2) If the minimum allele frequency of a marker is lower than 0.01, it is filtered out, that is, the minimum allele frequency (MAF) is < 0.01, and the parameter is vcftools --maf 0.01;

[0217] Finally, the genotyping results of each target SNP in a specific pea material are obtained, achieving high-throughput SNP genotyping.

[0218] Example 3

[0219] To verify the genotyping effect of the pea 20K SNP liquid-phase breeding chip, genotyping of 14 pea samples was performed using the pea 20K SNP liquid-phase breeding chip designed in Example 1. The specific method is as described in Example 2.

[0220] like Figure 3 As shown in the figure, the probe capture efficiency analysis of the sequencing data of 14 representative materials showed that 98.46% of the reads were aligned to the reference genome, of which the average capture rate of the probes was 44.47%, and the off-target reads accounted for 53.98%. This shows that when the pea 20K SNP liquid-phase breeding chip was used to genotype the test materials, the target site detection rate was high, the typing results were accurate and reliable, and it can be used for SNP typing detection of different pea samples.

[0221] The genetic distance matrix was calculated using VCF2Dis software to perform cluster analysis on the 14 materials, and the phylogenetic tree was constructed using fastme software. The results were consistent with the actual grouping. Figure 4 As shown, the site has high representativeness and can be used in combination with the evolutionary tree for genetic diversity analysis of pea germplasm resources, variety identification analysis, population structure analysis, genetic and evolutionary analysis, and kinship identification.

[0222] Example 4

[0223] In order to verify the effect of the pea 20K SNP liquid phase breeding chip in gene mapping, important trait analysis, and assisted selection breeding, a mixed segregation population analysis (BSA) for pod color traits was carried out using the pea 20K SNP liquid phase breeding chip designed in Example 1. In this analysis, ZWR12 with green pod characteristics was used as the female parent, and ZWNJ004 with yellow pod characteristics was used as the male parent. 30 plants were selected from the segregation population to form a yellow pod pool (WD_Y1) and a green pod pool (WD_G1), respectively, and genotyping analysis was performed. The specific method is as described in Example 2, and resequencing was performed at the same time. The clean reads obtained by the two methods were aligned to the PeaZW1 reference genome using BWA software. After sorting using SAMtools, SNPs were identified using GATK v4.0 and filtered according to predefined criteria. The R package QTLseqr was used to calculate and plot the ΔSNP index and G value.

[0224] The results showed that the polymorphic probes of these populations were mainly located on chromosome 3, specifically between positions 225,461,652-367,063,164 and 386,708,349-552,070,089. Figure 5 Figure 1 shows the gene mapping (a) and the resequencing gene mapping (b). These locations are consistent with the genomic coordinates from next-generation sequencing. This demonstrates the accuracy of the chip in identifying genes associated with specific traits, highlighting its potential for genetic improvement research and development, such as analyzing important pea traits and assisted selection breeding.

[0225] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications derived therefrom remain within the scope of protection of the present invention.

Claims

1. A pea whole-genome 20K SNP liquid-phase breeding chip, characterized in that: The genotyping targets of this chip include 21,659 SNP sites in the whole pea genome as shown in Table 1.

2. The pea whole genome 20K SNP liquid phase breeding chip according to claim 1, characterized in that: The reference genome of the SNP site is the pea genome "Pisum sativum cultivar Zhewan1 (PeaZW1)".

3. Use of the pea 20K SNP liquid-phase breeding chip according to claim 1 in genotyping pea samples.

Citation Information

Patent Citations

  • Pea heat resistance related SNP (Single Nucleotide Polymorphism) marker developed based on SnaPshot technology and application

    CN114574626A

  • KASP molecular marker for detecting pea leaf configuration and application of KASP molecular marker

    CN114686614A

  • Pea core SNP molecular marker set developed based on KASP technology and application thereof

    CN116219048A

  • Hot pepper 5K liquid phase chip and application thereof

    CN117051151A

  • Oat 2K SNP liquid phase chip and application thereof

    CN118879916A