Citrus hybrid offspring authenticity identification method based on whole genome SNP (Single Nucleotide Polymorphism) analysis
By using whole-genome SNP analysis and the IBD algorithm, the problems of nucellar embryo interference and identification uncertainty in citrus breeding have been solved, achieving efficient and accurate identification of hybrid offspring, reducing costs and shortening the cycle, and is applicable to a variety of citrus varieties.
Patent Information
- Application Number
- CN202610026536.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2046-01-09
AI Technical Summary
Citrus breeding suffers from problems such as prolonged breeding cycles, resource waste, and identification uncertainty due to nucellar embryo interference and errors in artificial hybridization. Existing technologies are difficult to efficiently and accurately distinguish between hybrid offspring and nucellar seedlings, and are also costly.
A whole-genome SNP analysis method was adopted, which combines low-depth sequencing with high-density SNP site analysis. The IBD algorithm was used to establish a discrimination system to achieve accurate differentiation between hybrid offspring and nucellar embryos. Genotype filling technology and multi-parameter judgment criteria were used to reduce sequencing depth and data volume and optimize the bioinformatics analysis process.
It achieves 99.2% accuracy in identifying hybrid offspring and nucellar embryos, shortens the identification cycle to 12-15 working days, reduces costs to 10-15% of traditional methods, is applicable to a variety of citrus varieties, and has broad application prospects.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of agricultural biotechnology and molecular breeding, and particularly relates to a method for identifying the authenticity of citrus hybrid offspring based on whole genome SNP analysis. BACKGROUND
[0002] As one of the most important economic fruit trees in the world, citrus breeding has long been faced with two major technical challenges: nucellar embryo interference and manual hybridization operation errors. Under the characteristics of polyembryony, in addition to zygotic embryos formed by sexual hybridization, nucellar embryos also exist in seeds, which are developed from nucellar cells. These nucellar embryos are genetically identical to the maternal parent and belong to asexual clonal offspring, which makes the probability of obtaining nucellar seedlings for some varieties still as high as 80% or more even if the hybridization process is strictly followed. This directly leads to the forced extension of the breeding cycle, often requiring years of field phenotypic observation to preliminarily screen possible hybrids, while causing a large number of non-target nucellar seedlings to occupy valuable land and management resources, severely restricting the breeding process of new varieties. On the other hand, the complex manual operation steps involved in citrus hybrid breeding, such as detasseling, pollination, and bagging, are prone to technical errors, such as incomplete detasseling or improper timing leading to self-pollination, loose bagging causing external pollen contamination, pollen mixing during pollination, and unexpected hybridization caused by wind or insect pollination, which further increases the uncertainty of breeding results.
[0003] Current citrus hybrid offspring identification mainly relies on morphological identification and molecular marker technology, both of which have obvious limitations. Morphological identification preliminarily judges hybrids by observing the leaf characteristics, growth potential, and distribution of thorns of seedlings. This method is not only difficult to distinguish between nucellar seedlings and hybrid seedlings due to their similar early-stage appearance, but also susceptible to environmental factors, leading to reduced accuracy. More critically, it requires a long observation period of 2-3 years, with extremely low efficiency.
[0004] With the development of molecular biology techniques, DNA marker-based identification methods have been gradually applied to citrus breeding. Among them, the SSR (microsatellite) marker technology has been widely used due to its low cost and simple operation. Studies have shown that using 12 pairs of SSR primers for citrus hybrid offspring identification can achieve an accuracy rate of about 85%. However, this technology has obvious limitations: its detection throughput is low and resolution is limited, usually only analyzing less than 20 loci, which is difficult to fully reflect the true situation of the genome and difficult to represent the genomic polymorphism of a large range of citrus resource populations. Therefore, it is difficult to effectively distinguish such hybrid offspring from nucellar seedlings based on SSR markers alone.
[0005] To overcome the limitations of SSR markers, researchers began to use SNP chip technology. Researchers developed a citrus-specific chip (Citrus SNP Array) containing 20,000 SNP sites. Although this method significantly improves detection throughput, there are still some fundamental problems: on the one hand, chip technology relies on a reference genome of a certain citrus species and pre-designed known SNP sites, and the detection results are easily affected by the uneven distribution of chip design sites in the genome. Some key areas may lack effective SNP markers, and the identification efficiency of hybrid offspring created from distant citrus species to the reference genome species is limited. On the other hand, the citrus-specific chip is not mass-produced and is relatively expensive, which is not suitable for large-scale identification of hybrid offspring authenticity.
[0006] In recent years, simplified genome sequencing technologies such as RAD-Seq and GBS have begun to be applied to citrus breeding research. Researchers have successfully identified citrus hybrid offspring using RAD-Seq technology, detecting about 5,000 SNP sites. Researchers have also used GBS technology to analyze the genetic diversity of citrus. These technologies can discover new SNP sites, but due to sequencing only specific regions of the genome, a large amount of genetic information is lost, and the data analysis process is complex and requires high bioinformatics capabilities.
[0007] Whole genome sequencing (WGS) technology can theoretically provide the most comprehensive genetic information. Researchers used 30x depth whole genome sequencing to identify citrus hybrids, with an accuracy rate of over 99% (Wang et al 2021). However, this method requires several hundred yuan for sequencing costs per sample and strong computing resources, making it difficult to apply on a large scale in actual breeding work.
[0008] A comprehensive analysis of existing technologies reveals that the field of citrus hybrid offspring identification faces several key challenges: first, the contradiction between throughput and cost, low-throughput methods are low-cost but lack accuracy, while high-throughput methods are accurate but cost too much; second, the problem of nucellar interference has not been effectively solved, and existing marker technologies cannot distinguish between true hybrid offspring and asexually propagated nucellar seedlings; third, the problem of incomplete genome coverage, most methods can only detect part of the genome region; finally, the dependence on known variations, many technologies cannot discover new genetic variations. These technical bottlenecks severely restrict the efficiency and quality of citrus breeding work. SUMMARY
[0009] The present application aims to overcome the deficiencies of the prior art, and provides a citrus hybrid offspring authenticity identification method based on whole genome SNP analysis.
[0010] To achieve the above-mentioned object, the technical scheme designed by the present application is as follows: The present application provides a citrus hybrid offspring authenticity identification method based on whole genome SNP analysis, comprising the following steps: (1) Extracting DNA of the female parent, the male parent and each hybrid offspring respectively; (2) Constructing a library for the DNA of the female parent, the male parent and each hybrid offspring respectively to obtain the DNA library of the female parent, the male parent and each hybrid offspring respectively; (3) Performing high-depth sequencing on the DNA library of the female parent and the male parent respectively to obtain the DNA sequencing data of the female parent and the male parent respectively, and performing low-depth sequencing on the DNA library of each hybrid offspring to obtain the DNA sequencing data of each hybrid offspring; (4) Performing quality control processing on the DNA sequencing data of the female parent, the male parent and each hybrid offspring respectively to obtain clean reads of the female parent, the male parent and each hybrid offspring respectively; (5) Aligning the clean reads of the female parent, the male parent and each hybrid offspring to the citrus reference genome through the BWA-MEM algorithm to obtain the aligned BAM files of the female parent, the male parent and each hybrid offspring respectively, then removing the duplicates of the BAM files and correcting the quality thereof to obtain the BAM files of the female parent, the male parent and each hybrid offspring respectively; (6) Performing SNP calling on the BAM files of the female parent and the male parent to generate the whole genome genotype data sets of the female parent and the male parent respectively, and combining the whole genome genotype data sets of the female parent and the male parent into a unified preliminary VCF file; (7) Obtaining a high-quality SNP site set through quality filtering and screening of the preliminary VCF file; (8) Using the R language STITCH software to fill the BAM files of each hybrid offspring according to the high-quality SNP site set to obtain a population integrated VCF file containing the site genotype information of the female parent, the male parent and all hybrid offspring; (9) After filtering the population integrated VCF file, the paired IBD results between each offspring and the paternal parent and the maternal parent are calculated using the genome command of the PLINK software to perform whole genome kinship analysis, and the shared allele probability is obtained to judge the hybrid offspring, and the judgment standard is as follows: When Z1 of the hybrid offspring and the maternal parent is greater than 0.9, and Z1 of the hybrid offspring and the paternal parent is greater than 0.9, it is indicated that the hybrid offspring is the true hybrid offspring of the maternal parent and the paternal parent; On the contrary, the rest are false hybrids; If Z2 of the false hybrid and the maternal parent is greater than 0.9, it is indicated that the false hybrid is a complete nucellar embryo; Wherein, Z1 and Z2 are two key output values in the IBD result, Z1 is the probability of sharing 1 allele between the hybrid offspring and the maternal parent / paternal parent; and Z2 is the probability of sharing 2 alleles between the hybrid offspring and the maternal parent / paternal parent.
[0011] Further, the maternal parent is ponkan, and the paternal parent is navel orange; In the step (1), the concentration of DNA is greater than or equal to 20 ng / μL, the total amount of DNA is greater than or equal to 500 ng, the A 260 / A 280 of DNA is 1.8-2.0, and the A 260 / A 230 of DNA is greater than 2.0; In the step (3), the high-depth sequencing is 30x and above, and the low-depth sequencing is 5-10x; In the step (4), the quality control processing specifically includes: The data quality index is comprehensively detected by using FastQC software, and then fastp software is used for data filtering to remove low-quality sequences and adapter sequences; In the step (5), the citrus reference genome is Citrus sinensis Valencia genome v2.0, and GATK tool is used for removing repetition; In the step (7), the specific conditions for screening are: Low-quality SNP sites, low-depth SNP sites and abnormal SNP sites not meeting Mendelian genetic rules are removed, and the GQ values of the SNP sites in the maternal parent and the paternal parent are both greater than 70.
[0012] The application also provides a method for screening SNP sites based on the true hybrid offspring obtained by the identification method, comprising the following steps: Using the population integrated VCF file containing the maternal parent, the paternal parent and the true hybrid offspring site genotype information corresponding to the true hybrid offspring, 5-10 SNP sites are screened.
[0013] Further, the screening criteria of the SNP site are: ① The SNP site is uniformly distributed on the whole genome level; ② The SNP site is two different homozygous genotypes in the maternal and paternal, respectively; The number of the true hybrid offspring is greater than 20.
[0014] The application further provides a SNP site screened by the method, and the SNP site includes chrl:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714 and chr9:10463074. The SNP site chrl:16715056 is located at the 16715056th base of the first chromosome of citrus, and the polymorphic site is C or T. The SNP site chr2:20043474 is located at the 20043474th base of the second chromosome of citrus, and the polymorphic site is C or G. The SNP site chr3:31884372 is located at the 31884372th base of the third chromosome of citrus, and the polymorphic site is G or C. The SNP site chr4:6258181 is located at the 6258181th base of the fourth chromosome of citrus, and the polymorphic site is G or A. The SNP site chr5:46639852 is located at the 46639852th base of the fifth chromosome of citrus, and the polymorphic site is C or G. The SNP site chr6:15804801 is located at the 15804801th base of the sixth chromosome of citrus, and the polymorphic site is C or G. The SNP site chr7:21054402 is located at the 21054402th base of the seventh chromosome of citrus, and the polymorphic site is C or T. The SNP site chr8:2865714 is located at the 2865714th base of the eighth chromosome of citrus, and the polymorphic site is A or G. The SNP site chr9:10463074 is located at the 10463074th base of the ninth chromosome of citrus, and the polymorphic site is A or T.
[0015] The application further provides a primer pair for obtaining the sequence containing the SNP site, and the nucleotide sequence of the primer pair for obtaining the sequence xl-1 containing the SNP site chrl:16715056 is as follows: F-1 : AGGCGTAACATTTACCTCCA (SEQ ID NO: 10), R-1 : AAAGAGAAATTTTGTTGCAAAAAAAATTGT (SEQ ID NO: 11); The nucleotide sequences of the primer pairs obtained for sequence xl-2 containing SNP site chr2:20043474 are as follows: F-2: GTAAAAAATATAAACAAACGCAAGATGAGG (SEQ ID NO: 12), R-2: TTACCCCTTCGACCCTTTTT (SEQ ID NO: 13); The nucleotide sequences of the primer pairs obtained for sequence xl-3 containing SNP site chr3:31884372 are as follows: F-3: TCGAAAACTTGGTGTTTAAATAGATTTGC (SEQ ID NO: 14), R-3: TATATATTGTAGGATCTGAATCATTAATCATCTTG (SEQ ID NO: 15); The nucleotide sequences of the primer pairs obtained for sequence xl-4 containing SNP site chr4:6258181 are as follows: F-4: GCACAGGGAAAGAAAGAGGAA (SEQ ID NO: 16), R-4: GCTACTCAGCTAATTGATTTGTGG (SEQ ID NO: 17); The nucleotide sequences of the primer pairs obtained for sequence xl-5 containing SNP site chr5:46639852 are as follows: F-5: ATCTTGATTTGCCACGTGTC (SEQ ID NO: 18), R-5: TCAGAACCTGAGACAAAACTAATG (SEQ ID NO: 19); The nucleotide sequences of the primer pairs obtained for sequence xl-6 containing SNP site chr6:15804801 are as follows: F-6: AGATCTGATCTTACTTTTTTTTTATTTTTTCCC (SEQ ID NO: 20), R-6: TAATTTTTGCTATTCCCCAAAAACATTC (SEQ ID NO: 21); The nucleotide sequence of the primer pair of sequence xl-7 containing SNP site chr7:21054402 is as follows: F-7: AAGGTCAGATGGAGCAACAC (SEQ ID NO: 22), R-7: AGTTGCAATTTCAACTTAAGGGAA (SEQ ID NO: 23); The nucleotide sequence of the primer pair of sequence xl-8 containing SNP site chr8:2865714 is as follows: F-8: TGCTGACGATAGCTCTAAAACAG (SEQ ID NO: 24), R-8: ATTATAGGAATTATTTAAGCTCTCAAAAGTTTTA (SEQ ID NO: 25); The nucleotide sequence of the primer pair of sequence xl-9 containing SNP site chr9:10463074 is as follows: F-9: TCCCAAGGAAATGATCTCAACT (SEQ ID NO: 26), R-9: GAGAACTCCCGTAATTCGAAAGAAA (SEQ ID NO: 27).
[0016] Further, the nucleotide sequence of the sequence xl-1 is shown as SEQ ID NO: 1, and the SNP site chr1:16715056 is located at the 296th base of the sequence xl-1; the nucleotide sequence of the sequence xl-2 is shown as SEQ ID NO: 2, and the SNP site chr2:20043474 is located at the 312th base of the sequence xl-2; the nucleotide sequence of the sequence xl-3 is shown as SEQ ID NO: 3, and the SNP site chr3:31884372 is located at the 264th base of the sequence xl-3; the nucleotide sequence of the sequence xl-4 is shown as SEQ ID NO: 4, and the SNP site chr4:6258181 is located at the 295th base of the sequence xl-4; the nucleotide sequence of the sequence xl-5 is shown as SEQ ID NO: 5, and the SNP site chr5:46639852 is located at the 272th base of the sequence xl-5; the nucleotide sequence of the sequence xl-6 is shown as SEQ ID NO: 6, and the SNP site chr6:15804801 is located at the 385th base of the sequence xl-6; The nucleotide sequence of the sequence xl-7 is shown as SEQ ID NO: 7, and the SNP site chr7:21054402 is located at the 281st base of the sequence xl-7; The nucleotide sequence of the sequence xl-8 is shown as SEQ ID NO: 8, and the SNP site chr8:2865714 is located at the 314th base of the sequence xl-8; The nucleotide sequence of the sequence xl-9 is shown as SEQ ID NO: 9, and the SNP site chr9:10463074 is located at the 299th base of the sequence xl-9.
[0017] The application further provides a kit for identifying the authenticity of a citrus hybrid offspring, which comprises the primer pair.
[0018] The application further provides a method for identifying the authenticity of a citrus hybrid offspring by using the kit, comprising the following steps: 1) extracting DNA of the citrus hybrid offspring to be detected; 2) amplifying the extracted DNA by using the primer pair in the kit; 3) performing sequencing analysis on the PCR amplification product to obtain sequencing results; 4) obtaining genotypes based on the sequencing results, and when the SNP sites chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714 and chr9:10463074 are all heterozygous genotypes, the citrus hybrid offspring to be detected is a true hybrid offspring.
[0019] The application further provides application of the SNP site, the primer pair or the kit in identifying the authenticity of a citrus hybrid offspring or cultivating a citrus hybrid offspring.
[0020] The application has the following beneficial effects: The citrus hybrid offspring identification method provided by the application has significant progress and outstanding substantial features compared with the prior art, and the technical effects mainly include the following aspects: 1. In terms of accuracy, the invention has made a breakthrough. Through the technical route of whole genome low-depth sequencing combined with high-density SNP analysis, accurate differentiation of nucellar embryos and true hybrid offspring is achieved. Experimental data show that the multi-parameter determination system based on Z value (Z1 and Z2) has an accuracy of 99.2%±0.5% in identifying nucellar embryos and 98.7%±0.8% in determining true hybrid offspring, which is significantly better than traditional SSR marker technology (accuracy about 85%) and SNP chip technology (accuracy about 92%). Through large-scale field verification test, the hybrid offspring population screened by the method of the invention shows better stability and consistency in subsequent trait performance.
[0021] 2. In terms of detection efficiency, traditional morphological identification requires 2-3 years of field observation to obtain preliminary results, while the invention shortens the identification period to 12-15 working days. This breakthrough is mainly due to three technical innovations: first, the application of low-depth sequencing technology (5-10x) greatly reduces the detection cost of a single sample; second, the optimized bioinformatics analysis process compresses the data analysis time to 12-18 hours using parallel computing architecture; third, the establishment of high-confidence determination criteria, which can accurately distinguish between true hybrid offspring and false hybrids through multi-level parameter analysis. The system takes advantage of the IBD (identity by descent) analysis method, considering both haplotype fragment genetic characteristics and whole genome similarity patterns. Based on a large amount of experimental data, the determination system shows high accuracy and reliability, providing a scientific basis for offspring identification in citrus breeding. The entire analysis process is realized through an automated process, ensuring the consistency and repeatability of the results, while providing intuitive visual reports to assist researchers in interpreting the results.
[0022] 3. Economic efficiency is another significant advantage of the invention. Through technical innovation, the invention controls the detection cost of a single sample at 10-15% of the traditional whole genome deep sequencing (30x). The cost reduction is mainly due to: a significant optimization of sequencing depth, reducing data volume by 1 / 30; the application of genotype filling algorithm, effectively improving the utilization rate of low-depth data; the implementation of two-stage screening strategy, which requires targeted detection of 85% of the samples in the population. Specific economic indicators show that for a population of 1000 samples, the total cost of the invention is 5%-10% of the deep sequencing scheme, with incomparable advantages in data quality and information quantity.
[0023] 4. Because this invention employs a whole-genome analysis strategy rather than pre-defined marker sites, it is applicable to the identification of various citrus species, including mandarin oranges, sweet oranges, and pomelos. Validation experiments show stable detection performance in 15 major cultivated varieties (including Satsuma mandarins, navel oranges, and grapefruits), with an accuracy fluctuation range of less than 1%. The excellent scalability of the technology platform is also reflected in its compatibility with different sequencing platforms (Illumina, MGI, etc.); the analysis workflow supports flexible replacement of the reference genome; and the judgment criteria can be adjusted according to the specific characteristics of each species. These features give this invention broad application prospects.
[0024] 5. In terms of operational process optimization, this invention achieves significant simplification and improvement. By establishing a standardized experimental operation manual, it ensures that different operators can complete key steps such as DNA extraction and library construction according to unified specifications. The bioinformatics analysis adopts a modular design, integrating raw data input, sequence alignment, genotype identification, and other steps into an automated process, directly outputting structured result files containing key indicators such as genotype data and Z1 and Z2 values. Experimental verification shows that this standardized process can maintain stable detection performance under different laboratory conditions, with a batch-to-batch coefficient of variation of less than 5%. The entire technical solution can be implemented with only conventional molecular biology experimental equipment and mid-range computing configuration, significantly reducing the technical threshold and equipment investment requirements.
[0025] 6. From the perspective of technological innovation, the implementation effects of this invention are mainly reflected in: 1) The successful application of low-depth whole-genome sequencing to the identification of citrus hybrid offspring for the first time has opened up a new technical route; 2) The STITCH-based genotype filling algorithm has increased the utilization rate of low-depth data to over 85%; 3) The established Z-value multi-parameter determination system provides a new standard for the analysis of kinship under the interference of plant asexual reproduction; 4) The innovative two-stage screening strategy has achieved the best balance between high-throughput detection and economy.
[0026] The implementation of this invention not only solves key technical problems in citrus breeding, but its technical principles and methods can also be extended to the breeding of other crops with asexual reproduction interference, such as mangoes and apples. With the continuous improvement and widespread application of the technology, it is expected to have a profound impact on the field of fruit tree breeding, providing strong technical support for the development of modern agriculture. Detailed Implementation
[0027] The present invention will now be described in further detail with reference to specific embodiments, so that those skilled in the art can understand it.
[0028] Explanation of technical terms in this invention 1. SNP site (Single Nucleotide Polymorphism) refers to genetic polymorphism caused by single nucleotide (A, T, C or G) variation in genomic DNA sequence. The specific definition is as follows: If a single nucleotide at a specific site in the genomic sequence of a population of individuals of the same species is replaced, deleted or inserted by a single base, and the frequency of the variation in the population is greater than 1%, the site is called SNP site.
[0029] 2. High-depth sequencing: generally refers to sequencing coverage of 30x and above; 3. Low-depth sequencing: generally refers to sequencing coverage of 5-10x.
[0030] Here, "x" represents coverage.
[0031] A method for identifying the authenticity of citrus hybrid offspring based on whole genome SNP analysis, comprising the following steps: (1) Extracting DNA from the mother, father and each hybrid offspring, respectively; (2) Constructing DNA libraries from the mother, father and each hybrid offspring, respectively, to obtain DNA libraries of the mother, father and each hybrid offspring, respectively; (3) High-depth sequencing of the DNA libraries of the mother and father, respectively, to obtain DNA sequencing data of the mother and father, respectively, and low-depth sequencing of the DNA libraries of each hybrid offspring to obtain DNA sequencing data of each hybrid offspring; (4) Quality control processing of the DNA sequencing data of the mother, father and each hybrid offspring, respectively, to obtain clean reads of the mother, father and each hybrid offspring, respectively; (5) Aligning the clean reads of the mother, father and each hybrid offspring to the citrus reference genome by BWA-MEM algorithm, respectively, to obtain aligned BAM files of the mother, father and each hybrid offspring, respectively, then removing duplicates and correcting the quality of the BAM files, to obtain BAM files of the mother, father and each hybrid offspring, respectively; (6) SNP calling of the BAM files of the mother and father to generate whole genome genotype data sets of the mother and father, respectively, and combining the whole genome genotype data sets of the mother and father into a unified preliminary VCF file; (7) Obtaining a high-quality SNP site set by quality filtering and screening the preliminary VCF file; (8) According to the high-quality SNP site set, the BAM file of each hybrid offspring is filled using the R language STITCH software to obtain a population integrated VCF file containing the genotype information of the maternal parent, the paternal parent and all hybrid offspring sites; (9) After filtering the population integrated VCF file, the paired IBD results between each offspring and the paternal parent and the maternal parent are calculated using the genome command of the PLINK software to perform whole genome kinship analysis to obtain the shared allele probability and judge the hybrid offspring, and the judgment criteria are as follows: When Z1 of the hybrid offspring and the maternal parent is greater than 0.9, and Z1 of the hybrid offspring and the paternal parent is greater than 0.9, it is indicated that the hybrid offspring is the true hybrid offspring of the maternal parent and the paternal parent; On the contrary, the rest are false hybrids; If Z2 of the false hybrid and the maternal parent is greater than 0.9, it is indicated that the false hybrid is a complete nucellar embryo. Wherein, Z1 and Z2 are two key output values in the IBD result, Z1 is the probability of sharing 1 allele between the hybrid offspring and the maternal parent / paternal parent; and Z2 is the probability of sharing 2 alleles between the hybrid offspring and the maternal parent / paternal parent.
[0032] In the embodiment, in step (1), the concentration of DNA is ≥20 ng / μL, the total amount of DNA is ≥500 ng, the A 260 / A 280 of DNA is 1.8-2.0, and the A 260 / A 230 of DNA is >2.0. In step (3), the high-depth sequencing is 30x and above, and the low-depth sequencing is 5-10x. In step (4), the steps of quality control processing specifically include: The data quality indicators are comprehensively detected using the FastQC software, and then the fastp software is used for data filtering to remove low-quality sequences and adapter sequences. In step (5), the citrus reference genome is Citrus sinensis Valencia genome v2.0, and the GATK tool is used for removing duplicates. In step (7), the specific conditions for screening are that low-quality SNP sites, low-depth SNP sites and abnormal SNP sites not meeting the Mendelian inheritance rule are removed, and the GQ values of the SNP sites in the maternal parent and the paternal parent are both greater than 70.
[0033] The method for identifying the authenticity of citrus hybrid offspring based on whole genome SNP analysis provided by the present application will be described in detail in combination with the embodiments below, but they should not be understood as limiting the scope of protection of the present application.
[0034] Example 1 This example takes the citrus hybrid combination "ponkan" (female parent) x "navel orange" (male parent) as an example, and the technical solution of the present application is also applicable to the hybrid identification of other citrus varieties. In this example, the female parent is numbered as A, the male parent is numbered as B, and 300 hybrid offspring are numbered as AB-1 to AB-300.
[0035] Since the hybrid offspring samples of this example belong to a large population, 50 samples are randomly selected for whole genome analysis: I. Sample collection and processing Healthy and pest-free young leaves are selected as sample materials. The specific operation is as follows: 3-5 young leaves that have not fully expanded are collected from the top of the plant to be tested and immediately placed in pre-cooled sample preservation tubes. The samples are stored in a -80°C ultra-low temperature freezer within 24 hours after collection. For large population testing, a 96-well plate sample collection system can be used, and each sample well is labeled with a unique number and strictly corresponds to the field number. During transportation, the samples need to be kept at 4°C to avoid DNA degradation.
[0036] Fresh young leaf samples of the female parent "ponkan", the male parent "navel orange", and the hybrid offspring AB-1 to AB-50 are collected.
[0037] II. DNA extraction and quality control 1. DNA extraction is performed using a plant genomic DNA extraction kit (HCSCI Plant Genomic DNA Extraction Kit (Magnetic Bead Method) V3 (Pre-packaged) from Chengdu Hanchen Guangyi Technology Co., Ltd.). The specific implementation steps are as follows: (1) First, take 50-100 mg of fresh young leaves and place them in a 2 mL grinding tube, add 450 μL of lysis buffer and 3-5 sterilized steel balls, and homogenize in a tissue grinder at a frequency of 60 Hz for 60 seconds until the tissue is completely broken.
[0038] (2) Incubate the lysate obtained in step (1) at 65°C for 10 minutes, and mix well for 2-3 times during the incubation. After cooling, add 130 μL of binding buffer, mix well, and then transfer to a DNA purification column and centrifuge at 12000 rpm for 1 minute.
[0039] (3) Discard the filtrate, add 500 μL of wash buffer I, and centrifuge at 12000 rpm for 1 minute.
[0040] (4) Discard the filtrate and add 700 μL of wash buffer II twice, and finally empty for 2 minutes to completely remove residual ethanol.
[0041] (5) Transfer the purification column to a new 1.5 mL centrifuge tube, add 50-100 μL of elution buffer, stand for 2 minutes, and then centrifuge to collect the DNA.
[0042] 2. DNA quality detection adopts the following standardized procedures: (1) Use Nanodrop One microspectrophotometer to measure the A 260 / A 280 ratio (1.8-2.0 required) and A 260 / A 230 ratio (≥2.0 required) of the DNA to preliminarily evaluate the purity and concentration of the DNA.
[0043] (2) Further use Qubit 4.0 fluorometer with dsDNA HS detection kit for accurate quantification, and the DNA concentration should be ≥20 ng / μL.
[0044] (3) Integrity detection is completed by 1.2% agarose gel electrophoresis (120V, 20 min), and the qualified sample should show a clear main band and no obvious degradation signs.
[0045] For the DNA samples meeting the standards (DNA concentration ≥20 ng / μL, total DNA amount ≥500 ng, A 260 / A 280 =1.8-2.0, A 260 / A 230 >2.0), aliquot and store at -20℃ for standby, and avoid repeated freeze-thawing.
[0046] The DNA samples of the female parent "ponkan", the male parent "navel orange", and the hybrid offspring AB-1 to AB-50 were obtained by step two respectively.
[0047] III. Library construction and sequencing (1) The verified DNA fragmentation library kit (NDB627 kit from Nanjing Novogene Bio- tech Co., Ltd.) was used for library construction. 200 ng of high-quality DNA of the female parent, the male parent, and the hybrid offspring AB-1 to AB-50 obtained in step two was taken for fragmentation to obtain 450-550 bp DNA fragments, which were subjected to end repair, A tailing, and ligation of Illumina compatible adapters, followed by 6-8 cycles of PCR amplification (98℃ 45s; [98℃ 15s, 60℃ 30s, 72℃ 30s] x 6-8; 72℃ 1 min).
[0048] (2) The amplified products were purified by 1.2x beads, and the fragment distribution (main peak 350-450 bp) was detected on an Agilent 2100 Bioanalyzer to ensure that the insert size met the expected size. The qualified libraries were mixed at equal molar concentrations, and the DNA libraries of the maternal, paternal and hybrid offspring AB-1 to AB-50 were obtained, and PE150 sequencing was performed on the Illumina NovaSeq 6000 platform. The target data volume of the hybrid offspring samples was 5-10x genome coverage (citrus genome about 360 Mb, i.e. about 2G reads per sample).
[0049] The DNA libraries of the maternal and paternal were sequenced at high depth (30x), and the DNA libraries of the hybrid offspring AB-1 to AB-50 were sequenced at low depth (5-10x), and after sequencing, the fastq format DNA sequencing data of the maternal, paternal and hybrid offspring AB-1 to AB-50 were obtained.
[0050] Four, bioinformatics analysis (1) The sequencing data was first subjected to quality control processing. The raw data quality of the maternal, paternal and hybrid offspring AB-1 to AB-50 was evaluated using FastQC v0.11.9, and the data quality indicators were comprehensively detected, including base quality distribution, sequence repeat rate, and key parameters of GC content.
[0051] (2) Then fastp v0.20.1 was used for data filtering (parameters: -q 20 -u 30 -l 50 -n 5), and low-quality reads and adapter sequences were removed, and high-quality clean reads of the maternal, paternal and hybrid offspring AB-1 to AB-50 were obtained.
[0052] (3) The high-quality clean reads were aligned to the citrus reference genome (BDZ.gapless.genome.fasta (HZAU)) by BWA-MEM v0.7.17, with the parameters -M -t 8, and the aligned BAM files of the maternal, paternal and hybrid offspring AB-1 to AB-50 were obtained.
[0053] (4) The aligned BAM files were subjected to MarkDuplicates by GATK v4.2 to remove duplicates, and the quality was corrected (BaseRecalibrator), and the BAM files of the maternal, paternal and hybrid offspring AB-1 to AB-50 were obtained.
[0054] (5) Then, the differential analysis strategy was used in the processing of the maternal, paternal and hybrid offspring samples. The BAM files of the maternal and paternal samples were first subjected to SNP calling, respectively. The parameters of the parent samples were set as follows: --min-base-quality-score 20 --min-mapping-quality 30 --min-pruning 3, and the whole genome genotype datasets (VCF files) of the maternal and paternal samples were generated, respectively. Then, the whole genome genotype datasets of the maternal and paternal samples were combined into a unified preliminary VCF file.
[0055] (6) The preliminary VCF file obtained was filtered by the following filtering standards: QD < 2.0 || FS > 60.0 || MQ < 40.0 || MQRankSum < -12.5 || ReadPosRankSum < -8.0, to remove low-quality SNP sites, low-depth SNP sites and abnormal SNP sites that do not conform to Mendelian genetic rules, and to obtain a screened SNP site set.
[0056] (7) The screened SNP site set was systematically divided into three types of key sites: segregation sites, hybrid sites and homozygous fidelity sites, and a high-quality SNP site set was screened from them as a core marker set. The high-quality SNP site screening standard is as follows: the GQ value of the SNP site in the two parents is greater than 70.
[0057] V. Genotype filling and determination (1) According to the high-quality SNP site set, the R language STITCH v1.6.8 was used to fill the BAM files of AB-1 to AB-50 of the hybrid offspring, and the key parameters were set as follows: K = 4, nGen = 100, nCores = 8. The input files include: ① the BAM files of AB-1 to AB-50 of the hybrid offspring; ② the VCF file of the maternal sample and the VCF file of the paternal sample. The population integrated VCF file containing the site genotype information of samples A, B, AB-1 to AB-50 was obtained.
[0058] (2) After filtering the population integrated VCF file containing sample A, B, AB-1 to AB-50 site genotype information (QD <2.0 || FS >60.0 || MQ <40.0 || MQRankSum < -12.5 || ReadPosRankSum < -8.0), the paired IBD results between each offspring and the paternal and maternal parents were calculated using the genome command of PLINK v1.9 software, the whole genome kinship analysis was performed, and the shared allele probability was calculated: ① Z0 (the probability of hybrid offspring and maternal / paternal not sharing alleles); ② Z1 (the probability of hybrid offspring and maternal / paternal sharing one allele); ③ Z2 (the probability of hybrid offspring and maternal / paternal sharing two alleles). Z0, Z1 and Z2 are the three key output values in the IBD result, and the judgment criteria are: When Z1 of the hybrid offspring and the maternal parent is greater than 0.9, and Z1 of the hybrid offspring and the paternal parent is greater than 0.9, it indicates that the hybrid offspring is the true hybrid offspring of the maternal parent and the paternal parent; On the contrary, the rest are false hybrids; If the Z2 of the false hybrid and the maternal parent is greater than 0.9, it indicates that the false hybrid is a complete nucellar embryo.
[0059] Since citrus is diploid, it has two sets of chromosomes, so a SNP site can at most have two different alleles, so for a sample, the genotype at a SNP site can be A or a, and for two samples, the genotype at the same site can have the following different situations, such as: The first situation: sample 1 is Aa, and sample 2 is Bb, at this time there is no same allele at the SNP site of the two samples, and Z0 represents the proportion of this situation in the whole genome; The second situation: sample 1 is Aa, and sample 2 is Ab, at this time there is one same allele (i.e. A), and Z1 represents the proportion of this situation in the whole genome; The third situation: sample 1 is Aa, and sample 2 is Aa, at this time there are two same alleles (i.e. A), and Z2 represents the proportion of this situation in the whole genome; Since there is no third set of chromosomes, there will be no three same alleles, so there are only the above three situations, and the sum of the proportions of the three situations must be equal to 1, that is, Z0+Z1+Z2=1.
[0060] The results of the genetic relationship of the present example are shown in Table 1 (only the genetic relationship results of hybrid offspring AB-1 to AB-9 are shown), and according to Table 1 and the determination criteria, it can be obtained that hybrid offspring AB-1 to AB-5 are all true hybrid offspring, and hybrid offspring AB-6 to AB-9 are false hybrids, wherein hybrid offspring AB-8 and AB-9 are complete nucellar embryos.
[0061] Table 1 Genetic relationship results (hybrid offspring AB-1 to AB-9) Note, in Table 1, FID1: Family ID of the first sample; IID1: Individual ID of the first sample; FID2: Family ID of the second sample; IID2: Individual ID of the second sample; RT: Relationship Type, the most likely relationship inferred by PLINK from the calculated IBD values, where FS: Full Siblings; HS: Half Siblings; PO: Parent-Offspring; OT: Other Relationship; UN: Unrelated Individuals.
[0062] EZ: Expected Relatedness.
[0063] PI_HAT: Overall IBD Proportion Estimate, which is a key comprehensive indicator for measuring the closeness of genetic relationship, the calculation formula is: PI_HAT = (Z1 / 2) + Z2.
[0064] PHE: Phenotype Consistency, if there is phenotype information in the data (such as case-control status, represented by 1 / 2 or 0 / 1), this row shows whether the phenotypes of the two individuals are consistent, -1: lack of phenotype information, cannot be compared; 0: inconsistent phenotypes (such as one is a case and the other is a control); 1: consistent phenotypes (such as both are cases or both are controls).
[0065] DST: IBS (Identity-by-State) Distance, which is a measure based on IBS (Identity-by-State), unlike IBD, IBS only looks at whether the genotypes are the same, but does not consider whether they come from a common ancestor.
[0066] PPC: Phenotype Prior Probability, a relatively complex statistical quantity, used to evaluate the degree of support of the observed genotype data for the genetic relationship hypothesis in the case of known phenotype, less directly used in actual analysis.
[0067] RATIO: Ratio. Another statistical quantity used to assist relationship inference, the specific formula varies with the version of PLINK, and is usually based on some ratio of IBD statistics (such as used to distinguish full siblings and parent-offspring relationships).
[0068] After judging the hybrid offspring AB-1 to AB-50 samples by the judging standard, 50 hybrid offspring true hybrid information can be obtained, and among the 50 hybrid offspring, there are 35 true hybrid offspring and 15 false hybrids.
[0069] Example 2 Based on the identification method of Example 1, the SNP sites for screening true hybrid offspring are obtained, including the following steps: The above 35 true hybrid offspring are used in the steps of Example 1, and in step (1) of step five, the BAM files of the 35 true hybrid offspring are replaced by the BAM files of the hybrid offspring AB-1 to AB-50 to obtain a population integrated VCF file containing the genotype information of the maternal parent A, the paternal parent B and the true hybrid offspring sites, and 5-10 SNP sites are screened out. These sites need to meet the following conditions: (1) Uniformly distributed on the whole genome level, not concentrated in a certain area; (2) Two different homozygous genotypes in parents respectively (such as AA in maternal parent and aa in paternal parent, so that the true hybrid offspring must be Aa).
[0070] Table 2 is the qualified SNP site screened out. After selecting 9 qualified SNP sites, primers are designed for the fragments where the SNP sites are located, and the remaining 250 hybrid offspring in Example 1 are analyzed by HiTOM sequencing to analyze the true situation of the sites in the hybrid offspring. For example, the first site in Table 2 is chr1:16715056. The base situation of this site in the maternal parent is homozygous CC, and that in the paternal parent is homozygous TT. According to the sequencing results of HiTOM, the base situation of this site in the true hybrid offspring should be 40%-60% C base and 40%-60% T base; and the complete pearl embryo offspring will have more than 90% C base.
[0071] Table 2 Qualified SNP sites After sequencing the hybrid offspring AB-51 to AB-300, among the 250 hybrid offspring, there are 171 true hybrid offspring and 29 false hybrids.
[0072] In order to ensure the reliability of the method, a three-level verification system is established: (1) Internal control: 5% known samples (previously identified true hybrids and pearl embryos) are included in each batch; (2) Technical repetition: 10% samples are randomly selected for repeated detection; (3) Field verification: the determination results are tracked for 2 years of phenotype.
[0073] The results of the embodiment are verified, and the result shows that the result judgment accuracy of the embodiment is >98%, which indicates that the identification method of the embodiment has high accuracy and strong practicability.
[0074] Those skilled in the art should understand that the above embodiments are only illustrative, and the specific parameters and steps can be properly adjusted without departing from the core idea of the present application. For example, other commercial kits can be used for DNA extraction, sequencing platforms can be replaced by MGI or Ion Torrent, and analysis software can use similar tools with similar functions. The second stage can use Sanger sequencing and other methods. These adjustments should be within the protection scope of the present patent.
[0075] According to the SNP site obtained in the embodiment, the remaining hybrid offspring samples are subjected to HiTOM sequencing, and the genotype mode of each sample at a specific site is directly obtained, and the authenticity is determined by comparing with the expected hybridization mode. This strategy can reduce the total cost of population detection by 90%.
[0076] Example 3 The SNP site in Example 2 is obtained according to the citrus reference genome (BDZ.gapless.genome.fasta (HZAU)).
[0077] 1. Design primer xl-1 to amplify the sequence containing SNP site chr1:16715056, and the nucleotide sequences of the primer pair are as follows: F-1: AGGCGTAACATTTACCTCCA, R-1: AAAGAGAAATTTTGTTGCAAAAAAAATTGT; PCR amplification of citrus genomic DNA using primer pair F-1 and R-1 obtains sequence xl-1 containing SNP site chr1:16715056, and the nucleotide sequence is shown in SEQ ID NO: 1, Y is C / T, and SNP site chr1:16715056 is located at the 296th base of sequence xl-1.
[0078] 2. Design primer xl-2 to amplify the sequence containing SNP site chr2:20043474, and the nucleotide sequences of the primer pair are as follows: F-2: GTAAAAAATATAAACAAACGCAAGATGAGG, R-2: TTACCCCTTCGACCCTTTTT; The sequence xl-2 containing the SNP site chr2:20043474 is obtained by PCR amplification of the citrus genomic DNA using primer pair F-2 and R-2, the nucleotide sequence of which is shown in SEQ ID NO: 2, S is C / G, and the SNP site chr2:20043474 is located at the 312th base of the sequence xl-2.
[0079] 3. The sequence xl-3 containing the SNP site chr3:31884372 is amplified by designing primers, the nucleotide sequences of which are as follows: F-3: TCGAAAACTTGGTGTTTAAATAGATTTGC, R-3: TATATATTGTAGGATCTGAATCATTAATCATCTTG; The sequence xl-3 containing the SNP site chr3:31884372 is obtained by PCR amplification of the citrus genomic DNA using primer pair F-3 and R-3, the nucleotide sequence of which is shown in SEQ ID NO: 3, S is G / C, and the SNP site chr3:31884372 is located at the 264th base of the sequence xl-3.
[0080] 4. The sequence xl-4 containing the SNP site chr4:6258181 is amplified by designing primers, the nucleotide sequences of which are as follows: F-4: GCACAGGGAAAGAAAGAGGAA, R-4: GCTACTCAGCTAATTGATTTGTGG; The sequence xl-4 containing the SNP site chr4:6258181 is obtained by PCR amplification of the citrus genomic DNA using primer pair F-4 and R-4, the nucleotide sequence of which is shown in SEQ ID NO: 4, R is G / A, and the SNP site chr4:6258181 is located at the 295th base of the sequence xl-4.
[0081] 5. The sequence xl-5 containing the SNP site chr5:46639852 is amplified by designing primers, the nucleotide sequences of which are as follows: F-5: ATCTTGATTTGCCACGTGTC, R-5: TCAGAACCTGAGACAAAACTAATG; The sequence xl-5 containing the SNP site chr5:46639852 is amplified by PCR using the primer pair F-5 and R-5, the nucleotide sequence of which is shown in SEQ ID NO: 5, S is C / G, and the SNP site chr5:46639852 is located at the 272nd base of the sequence xl-5.
[0082] 6. The sequence xl-6 containing the SNP site chr6:15804801 is amplified by designing primers, the nucleotide sequences of which are as follows: F-6: AGATCTGATCTTACTTTTTTTTTATTTTTTCCC, R-6: TAATTTTTGCTATTCCCCAAAAACATTC; The sequence xl-6 containing the SNP site chr6:15804801 is amplified by PCR using the primer pair F-6 and R-6, the nucleotide sequence of which is shown in SEQ ID NO: 6, S is C / G, and the SNP site chr6:15804801 is located at the 385th base of the sequence xl-6.
[0083] 7. The sequence xl-7 containing the SNP site chr7:21054402 is amplified by designing primers, the nucleotide sequences of which are as follows: F-7: AAGGTCAGATGGAGCAACAC, R-7: AGTTGCAATTTCAACTTAAGGGAA; The sequence xl-7 containing the SNP site chr7:21054402 is amplified by PCR using the primer pair F-7 and R-7, the nucleotide sequence of which is shown in SEQ ID NO: 7, Y is C / T, and the SNP site chr7:21054402 is located at the 281st base of the sequence xl-7.
[0084] 8. The sequence xl-8 containing the SNP site chr8:2865714 is amplified by designing primers, the nucleotide sequences of which are as follows: F-8: TGCTGACGATAGCTCTAAAACAG, R-8: ATTATAGGAATTATTTAAGCTCTCAAAAGTTTTA; The sequence xl-8 containing the SNP site chr8:2865714 is obtained by PCR amplification of the citrus genomic DNA using the primer pair F-8 and R-8, the nucleotide sequence of which is shown in SEQ ID NO: 8, R is A / G, and the SNP site chr8:2865714 is located at the 314th base of the sequence xl-8.
[0085] 9. The sequence xl-9 containing the SNP site chr9:10463074 is designed by primers, the nucleotide sequences of which are as follows: F-9: TCCCAAGGAAATGATCTCAACT, R-9: GAGAACTCCCGTAATTCGAAAGAAA; The sequence xl-9 containing the SNP site chr9:10463074 is obtained by PCR amplification of the citrus genomic DNA using the primer pair F-9 and R-9, the nucleotide sequence of which is shown in SEQ ID NO: 9, W is A / T, and the SNP site chr9:10463074 is located at the 299th base of the sequence xl-9.
[0086] Example 4 The present embodiment provides a kit for identifying the authenticity of a citrus hybrid offspring, which comprises the primer pair F-1 and R-1, the primer pair F-2 and R-2, the primer pair F-3 and R-3, the primer pair F-4 and R-4, the primer pair F-5 and R-5, the primer pair F-6 and R-6, the primer pair F-7 and R-7, the primer pair F-8 and R-8, and the primer pair F-9 and R-9 in Example 4.
[0087] Example 5 The present embodiment provides a method for identifying the authenticity of a citrus hybrid offspring using the kit of Example 4, which comprises the following steps: 1) Extracting the DNA of the citrus hybrid offspring to be detected; 2) Amplifying the extracted DNA using the primer pairs in the kit; 3) Sequencing and analyzing the PCR amplification product to obtain sequencing results; 4) Based on the sequencing results, obtaining the genotype, and when the SNP sites chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714 and chr9:10463074 are all heterozygous genotypes, the citrus hybrid offspring to be detected is a true hybrid offspring.
[0088] The kit of the present embodiment is used to identify the hybridization offspring authenticity of the hybridization offspring AB-51 to AB-300 in Embodiment 1, among the 250 hybridization offspring, 171 are true hybridization offspring and 29 are false hybrids. The kit of the present embodiment is used to identify the hybridization offspring of citrus, and the information of the hybridization offspring can be accurately obtained with high accuracy.
[0089] Other parts not described in detail are prior art. Although the above embodiment describes the present application in detail, it is only a part of the embodiment of the present application, not all the embodiments, and other embodiments can be obtained according to the present embodiment without creativity, which belongs to the protection scope of the present application.
Claims
1. A method for identifying the authenticity of citrus hybrid offspring based on whole-genome SNP analysis, characterized in that: Includes the following steps: (1) Extract DNA from the maternal parent, paternal parent, and each hybrid offspring; (2) Construct libraries from the DNA of the maternal parent, paternal parent and each hybrid offspring, respectively. (3) Perform high-depth sequencing on the DNA libraries of the mother and father to obtain DNA sequencing data of the mother and father respectively. Perform low-depth sequencing on the DNA libraries of each hybrid offspring to obtain DNA sequencing data of each hybrid offspring. (4) Perform quality control processing on the DNA sequencing data of the maternal parent, paternal parent and each hybrid offspring to obtain clean reads of the maternal parent, paternal parent and each hybrid offspring respectively; (5) Align the clean reads of the maternal parent, paternal parent and each hybrid offspring to the citrus reference genome using the BWA-MEM algorithm to obtain the aligned BAM files of the maternal parent, paternal parent and each hybrid offspring. Then remove duplicates from the BAM files and correct their quality to obtain the BAM files of the maternal parent, paternal parent and each hybrid offspring. (6) Perform SNP calling on the BAM files of the mother and father to generate whole genome genotype datasets of the mother and father respectively, and merge the whole genome genotype datasets of the mother and father into a unified preliminary VCF file; (7) The initial VCF files are filtered and screened to obtain a high-quality SNP site set; (8) Based on the high-quality SNP locus set, the BAM file of each hybrid offspring was filled using the R language STITCH software to obtain a population integrated VCF file containing the genotype information of the maternal parent, paternal parent and all hybrid offspring loci; (9) After filtering the population integration VCF file, the genome command of PLINK software is used to calculate the paired IBD results between each offspring and the father and mother, perform whole-genome kinship analysis, obtain the probability of shared alleles, and determine the hybrid offspring. The judgment criteria are as follows: When the Z1 of the hybrid offspring with the maternal parent is greater than 0.9, and the Z1 of the hybrid offspring with the paternal parent is greater than 0.9, it indicates that the hybrid offspring is a true hybrid offspring of the maternal and paternal parents. Conversely, all other cases are false hybrids; If the Z2 of the pseudohybrid and the maternal parent is greater than 0.9, it indicates that the pseudohybrid is a complete nucellar embryo. Among them, Z1 and Z2 are the two key output values in the IBD results. Z1 is the probability that the hybrid offspring and the maternal / paternal parent share 1 allele; Z2 is the probability that the hybrid offspring and the maternal / paternal parent share 2 alleles.
2. The identification method according to claim 1, characterized in that: The maternal parent is Ponkan orange, and the paternal parent is navel orange; In step (1), the concentration of DNA is ≥20 ng / μL, the total amount of DNA is ≥500 ng, and the A content of DNA is... 260 / A 280 =1.8-2.0, DNA A 260 / A 230 >2.0; In step (3), high-depth sequencing is 30× and above, and low-depth sequencing is 5-10×. In step (4), the quality control process specifically includes: FastQC software was used to comprehensively test data quality indicators, and then fastp software was used to filter the data to remove low-quality sequences and connector sequences. In step (5), the citrus reference genome is Citrus sinensis Valencia genome v2.0, and duplications are removed using the GATK tool; In step (7), the specific screening criteria are as follows: Low-quality SNP sites, low-depth SNP sites, and abnormal SNP sites that do not conform to Mendelian inheritance laws were removed, and the GQ value of the SNP sites in both the maternal and paternal parents was greater than 70.
3. A method for obtaining true hybrid offspring screening SNP loci based on the identification method described in claim 1, characterized in that: Includes the following steps: Using population integration VCF files containing genotype information of maternal, paternal, and real hybrid offspring loci corresponding to the actual hybrid offspring, 5-10 SNP loci were obtained through screening.
4. The method according to claim 3, characterized in that: The screening criteria for the SNP sites are as follows: ①SNP sites are evenly distributed across the entire genome; ②The SNP loci are two different homozygous genotypes in the maternal and paternal parents.
5. An SNP site obtained by screening using the method of claim 3, characterized in that: The SNP sites include chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714 and chr9:10463074; The SNP site chr1:16715056 is located at base 16715056 on chromosome 1 of citrus, and the polymorphic site is C or T. The SNP site chr2:20043474 is located at base 20043474 on chromosome 2 of citrus, and the polymorphic site is C or G; The SNP site chr3:31884372 is located at base 31884372 on chromosome 3 of citrus, and the polymorphic site is either G or C. The SNP site chr4:6258181 is located at base 6258181 on chromosome 4 of citrus, and the polymorphic site is either G or A. The SNP site chr5:46639852 is located at base 46639852 on chromosome 5 of citrus, and the polymorphic site is C or G; The SNP site chr6:15804801 is located at base 15804801 on chromosome 6 of citrus, and the polymorphic site is C or G; The SNP site chr7:21054402 is located at base 21054402 on chromosome 7 of citrus, and the polymorphic site is C or T; The SNP site chr8:2865714 is located at base 2865714 on chromosome 8 of citrus, and the polymorphic site is either A or G. The SNP site chr9:10463074 is located at base 10463074 on chromosome 9 of citrus, and the polymorphic site is either A or T.
6. A primer pair for obtaining a sequence containing the SNP site as described in claim 5, characterized in that: The nucleotide sequence of the primer pair containing the SNP site chr1:16715056 xl-1 is as follows: F-1: AGGCGTAACATTTACCTCCA R-1:AAAGAGAAATTTTGTTGCAAAAAAAATTGT; The nucleotide sequences of the primer pair containing the SNP site chr2:20043474 xl-2 are as follows: F-2: GTAAAAAATATAAACAAAACCGCAAGATGAGG, R-2: TTACCCCTTCGACCCTTTTT; The nucleotide sequences of the primer pair xl-3 containing the SNP site chr3:31884372 are as follows: F-3: TCGAAAACTTGGTGTTTAAATAGATTTGC, R-3: TATATATTGTAGGATCTGAATCATTAATCATCTTG; The nucleotide sequences of the primer pair xl-4 containing the SNP site chr4:6258181 are as follows: F-4: GCACAGGGAAAGAAAGAGGAA, R-4: GCTACTCAGCTAATTGATTTGTGG; The nucleotide sequences of the primer pair xl-5 containing the SNP site chr5:46639852 are as follows: F-5: ATCTTGATTTGCCACGTGTC, R-5:TCAGAACCTGAGACAAAACTAATG; The nucleotide sequences of the primer pair containing the SNP site chr6:15804801 xl-6 are as follows: F-6: AGATCTGATCTTACTTTTTTTTTATTTTTTCCC, R-6: TAATTTTTGCTATTCCCCAAAAACATTC; The nucleotide sequences of the primer pair containing the SNP site chr7:21054402 xl-7 are as follows: F-7: AAGGTCAGATGGAGCAACAC R-7: AGTTGCAATTTCAACTTAAGGGAA; The nucleotide sequences of the primer pair xl-8 containing the SNP site chr8:2865714 are as follows: F-8: TGCTGACGATAGCTCTAAAACAG, R-8: ATTATAGGAATTATTTAAGCTCTCAAAAGTTTTA; The nucleotide sequences of the primer pair containing the SNP site chr9:10463074 xl-9 are as follows: F-9: TCCCAAGGAAATGATCTCAACT, R-9: GAGAACTCCCGTAATTCGAAAGAAA.
7. The primer pair according to claim 6, characterized in that: The nucleotide sequence of the sequence xl-1 is shown in SEQ ID NO: 1, and the SNP site chr1:16715056 is located at the 296th base of the sequence xl-1; The nucleotide sequence of the sequence xl-2 is shown in SEQ ID NO: 2, and the SNP site chr2:20043474 is located at the 312th base of the sequence xl-2. The nucleotide sequence of the sequence xl-3 is shown in SEQ ID NO: 3, and the SNP site chr3:31884372 is located at the 264th base of the sequence xl-3. The nucleotide sequence of the sequence xl-4 is shown in SEQ ID NO: 4, and the SNP site chr4:6258181 is located at the 295th base of the sequence xl-4. The nucleotide sequence of the sequence xl-5 is shown in SEQ ID NO: 5, and the SNP site chr5:46639852 is located at the 272nd base of the sequence xl-5. The nucleotide sequence of the sequence xl-6 is shown in SEQ ID NO: 6, and the SNP site chr6:15804801 is located at the 385th base of the sequence xl-6. The nucleotide sequence of the sequence xl-7 is shown in SEQ ID NO: 7, and the SNP site chr7:21054402 is located at the 281st base of the sequence xl-7. The nucleotide sequence of the sequence xl-8 is shown in SEQ ID NO: 8, and the SNP site chr8:2865714 is located at the 314th base of the sequence xl-8. The nucleotide sequence of the sequence xl-9 is shown in SEQ ID NO: 9, and the SNP site chr9:10463074 is located at the 299th base of the sequence xl-9.
8. A reagent kit for identifying the authenticity of citrus hybrid offspring, characterized in that: The kit includes the primer pair as described in claim 6.
9. A method for identifying the authenticity of citrus hybrid offspring using the kit described in claim 8, characterized in that: Includes the following steps: 1) Extract DNA from the hybrid offspring of the citrus trees to be tested; 2) Amplify the extracted DNA using the primer pair in the kit described in claim 8; 3) Sequencing analysis of the PCR amplification products to obtain sequencing results; 4) Based on the sequencing results, the genotypes are obtained. When the SNP loci chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714 and chr9:10463074 are all heterozygous genotypes, the citrus hybrid offspring to be tested are true hybrid offspring.
10. The application of the SNP site of claim 5, the primer pair of claim 6, or the kit of claim 8 in identifying the authenticity of citrus hybrid offspring or in breeding citrus hybrid offspring.
Citation Information
Patent Citations
Rapid identification primers of citrus hybrids based on SNP markers as well as method
CN110042172A
20K liquid phase chip for citrus genotype identification and application of 20K liquid phase chip
CN117305503A
Primer and method for identifying authenticity of distant hybridization offspring of clausena lansium and citrus
CN120624696A
SSR molecular markers for identifying citrus zygotic and nucellar individuals and uses thereof
KR101784244B1
Citrus whole-genome 40k liquid chip and use
WO2024197985A1
Cited By
Identification method for meiosis recombination of citrus F1-generation hybrid population based on next-generation sequencing data
CN121931230A