A method for identifying the authenticity of citrus hybrid offspring based on genome-wide SNP analysis

By using whole-genome SNP analysis and the IBD algorithm, the problems of nucellar embryo interference and identification uncertainty in citrus breeding were solved, achieving efficient and accurate identification of hybrid offspring, reducing costs and time, and applicable to a variety of citrus species.

CN121472476BActive Publication Date: 2026-04-03HUAZHONG AGRI UNIV +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Citrus breeding suffers from problems such as prolonged breeding cycles, resource waste, and identification uncertainty caused by nucellar embryo interference and errors in artificial hybridization. Existing technologies, such as SSR markers and SNP chip methods, have contradictions between throughput and cost, incomplete genome coverage, and dependence on known variations, making it difficult to accurately distinguish between hybrid offspring and nucellar seedlings.

Method used

A whole-genome SNP analysis method was adopted, which combines low-depth sequencing with high-density SNP site analysis. The IBD algorithm and genotype filling technology were used to establish a discrimination system to achieve accurate identification of true hybrid offspring.

Benefits of technology

It achieves accurate differentiation between nucellar embryos and true hybrid offspring, with an identification accuracy rate of 99.2% ± 0.5%, reduces costs to 10-15% of traditional methods, and shortens the testing cycle to 12-15 working days. It is applicable to the identification of various citrus tree species.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

This invention discloses a method for identifying the authenticity of citrus hybrid offspring based on genome-wide SNP analysis. This invention employs genome-wide genotyping technology, high-density SNP marker analysis, and the IBD (identical by descent) algorithm to establish a discriminant system capable of accurately identifying true hybrid offspring. This method overcomes the limitations of traditional methods in identifying hybrid offspring and provides reliable technical support for breeding practices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of agricultural biotechnology and molecular breeding, specifically to a method for identifying the authenticity of citrus hybrid offspring based on whole-genome SNP analysis. Background Technology

[0002] Citrus, as one of the world's most important economic fruit trees, has long faced two major technical challenges in its breeding: nucellar embryo interference and errors in artificial hybridization. Due to the polyembryonic nature of citrus, in addition to the zygote formed by sexual hybridization, seeds commonly contain nucellar embryos developed from nucellar cells. These nucellar embryos are genetically identical to the maternal parent and are asexual clones, meaning that even when strictly following hybridization procedures, the probability of obtaining nucellar seedlings in some varieties can still be as high as 80% or more. This phenomenon directly leads to a forced extension of the breeding cycle, often requiring several years of field phenotypic observation to initially screen for potential hybrids. Simultaneously, it results in a large number of non-target nucellar seedlings occupying valuable land and management resources, severely hindering the breeding process of new varieties. On the other hand, the complex manual operations involved in citrus hybridization breeding, such as emasculation, pollination, and bagging, are prone to technical errors. For example, incomplete emasculation or improper timing can lead to self-pollination, improper bagging can cause foreign pollen contamination, pollen mixing during pollination, and unexpected hybridization caused by wind or insects. These factors further increase the uncertainty of the breeding results.

[0003] The current methods for identifying citrus hybrid offspring, which mainly rely on morphological identification and molecular marker technology, have significant limitations. Morphological identification relies on observing phenotypic characteristics such as leaf features, growth vigor, and thorn distribution in seedlings to make a preliminary judgment on hybrids. This method is not only difficult to distinguish because nucellar seedlings and hybrid seedlings may look similar in the early stages, but it is also easily affected by environmental factors, leading to reduced accuracy. More importantly, it requires an observation period of up to 2-3 years, making it extremely inefficient.

[0004] With the development of molecular biology techniques, DNA marker-based identification methods have been increasingly applied to citrus breeding. Among these, SSR (microsatellite) marker technology has gained widespread use due to its low cost and ease of operation. Researchers have shown that using 12 pairs of SSR primers for identifying citrus hybrid offspring can achieve an accuracy rate of approximately 85%. However, this technology has significant limitations: its detection throughput is low and its resolution is limited, typically analyzing fewer than 20 loci, making it difficult to comprehensively reflect the true state of the genome and characterize the genomic polymorphism of large-scale citrus resource populations. Therefore, SSR markers alone cannot effectively distinguish these hybrid offspring from bulbils.

[0005] To overcome the limitations of SSR markers, researchers began using SNP microarray technology. They developed a citrus-specific microarray (CitrusSNP array) containing 20,000 SNP loci. While this method significantly improved throughput, several fundamental problems remain: First, the microarray technology relies on a reference genome of a specific citrus species and pre-designed known SNP loci, and the results are easily affected by the uneven distribution of the designed loci across the genome. Certain key regions may lack effective SNP markers, limiting its effectiveness in identifying hybrid offspring from citrus species distantly related to the reference genome. Second, citrus-specific microarrays are expensive due to lack of mass production, making them unsuitable for identifying the authenticity of offspring in large-scale hybrid populations.

[0006] In recent years, simplified genome sequencing technologies such as RAD-Seq and GBS have begun to be applied to citrus breeding research. Researchers have successfully identified citrus hybrids using RAD-Seq, detecting approximately 5,000 SNP loci. Other researchers have used GBS to analyze the genetic diversity of citrus. While these technologies can discover new SNP loci, they result in the loss of significant genetic information because they only sequence specific regions of the genome. Furthermore, the data analysis process is complex and requires a high level of bioinformatics expertise.

[0007] Whole-genome sequencing (WGS) technology can theoretically provide the most comprehensive genetic information. Researchers have used 30× depth whole-genome sequencing for citrus hybrid identification with an accuracy rate exceeding 99% (Wang et al 2021). However, this method results in a sequencing cost of several hundred yuan per sample and requires substantial computing resources, making it difficult to apply on a large scale in practical breeding work.

[0008] A comprehensive analysis of existing technologies reveals several key challenges in the identification of citrus hybrid offspring: First, there is a trade-off between throughput and cost; low-throughput methods are inexpensive but lack accuracy, while high-throughput methods are accurate but prohibitively expensive. Second, the problem of nucellar embryo interference has not been effectively resolved, and existing marker technologies struggle to distinguish true hybrid offspring from asexually propagated nucellar seedlings. Third, there is the issue of incomplete genome coverage, with most methods only detecting partial genomic regions. Finally, there is a dependence on known variations, with many technologies failing to detect new genetic variations. These technological bottlenecks severely restrict the efficiency and quality of citrus breeding efforts. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a method for identifying the authenticity of citrus hybrid offspring based on whole-genome SNP analysis. This invention employs whole-genome genotyping technology, relying on the whole-genome genotype information of the hybrid parents (obtained through deep DNA sequencing). Genotype imputation technology is used to obtain high-density SNP locus genotype information of the hybrid offspring population after low-depth DNA sequencing. Through the IBD (identical by descent) algorithm, a discrimination system capable of accurately identifying true hybrid offspring is established. This method overcomes the limitations of traditional methods in identifying hybrid offspring and provides reliable technical support for breeding practices.

[0010] To achieve the above objectives, the technical solution designed by the present invention is as follows:

[0011] This invention provides a method for identifying the authenticity of citrus hybrid offspring based on whole-genome SNP analysis, comprising the following steps:

[0012] (1) Extract DNA from the maternal parent, paternal parent, and each hybrid offspring;

[0013] (2) Construct libraries from the DNA of the maternal parent, paternal parent and each hybrid offspring, respectively.

[0014] (3) Perform high-depth sequencing on the DNA libraries of the mother and father to obtain DNA sequencing data of the mother and father respectively. Perform low-depth sequencing on the DNA libraries of each hybrid offspring to obtain DNA sequencing data of each hybrid offspring.

[0015] (4) Perform quality control processing on the DNA sequencing data of the maternal parent, paternal parent and each hybrid offspring to obtain clean reads of the maternal parent, paternal parent and each hybrid offspring respectively;

[0016] (5) Align the clean reads of the maternal parent, paternal parent and each hybrid offspring to the citrus reference genome using the BWA-MEM algorithm to obtain the aligned BAM files of the maternal parent, paternal parent and each hybrid offspring. Then remove duplicates from the BAM files and correct their quality to obtain the BAM files of the maternal parent, paternal parent and each hybrid offspring.

[0017] (6) Perform SNP calling on the BAM files of the mother and father to generate whole genome genotype datasets of the mother and father respectively, and merge the whole genome genotype datasets of the mother and father into a unified preliminary VCF file;

[0018] (7) The initial VCF files are filtered and screened to obtain a high-quality SNP site set;

[0019] (8) Based on the high-quality SNP locus set, the BAM file of each hybrid offspring was filled using the R language STITCH software to obtain a population integrated VCF file containing the genotype information of the maternal parent, paternal parent and all hybrid offspring loci;

[0020] (9) After filtering the population integration VCF file, the genome command of PLINK software is used to calculate the paired IBD results between each offspring and the father and mother, perform whole-genome kinship analysis, obtain the probability of shared alleles, and determine the hybrid offspring. The judgment criteria are as follows:

[0021] When the Z1 of the hybrid offspring with the maternal parent is greater than 0.9, and the Z1 of the hybrid offspring with the paternal parent is greater than 0.9, it indicates that the hybrid offspring is a true hybrid offspring of the maternal and paternal parents.

[0022] Conversely, all other cases are false hybrids;

[0023] If the Z2 of the pseudohybrid and the maternal parent is greater than 0.9, it indicates that the pseudohybrid is a complete nucellar embryo.

[0024] Among them, Z1 and Z2 are the two key output values ​​in the IBD results. Z1 is the probability that the hybrid offspring and the maternal / paternal parent share 1 allele; Z2 is the probability that the hybrid offspring and the maternal / paternal parent share 2 alleles.

[0025] Furthermore, the maternal parent is Ponkan orange and the paternal parent is navel orange;

[0026] In step (1), the concentration of DNA is ≥20 ng / μL, the total amount of DNA is ≥500 ng, and the A content of DNA is... 260 / A 280 =1.8-2.0, DNA A 260 / A 230 >2.0;

[0027] In step (3), high-depth sequencing is 30× and above, and low-depth sequencing is 5-10×.

[0028] In step (4), the quality control process specifically includes:

[0029] FastQC software was used to comprehensively test data quality indicators, and then fastp software was used to filter the data to remove low-quality sequences and connector sequences.

[0030] In step (5), the citrus reference genome is Citrus sinensis Valencia genome v2.0, and duplications are removed using the GATK tool;

[0031] In step (7), the specific screening criteria are as follows:

[0032] Low-quality SNP sites, low-depth SNP sites, and abnormal SNP sites that do not conform to Mendelian inheritance laws were removed, and the GQ value of the SNP sites in both the maternal and paternal parents was greater than 70.

[0033] This invention also provides a method for obtaining true hybrid offspring screening SNP sites based on the aforementioned identification method, comprising the following steps:

[0034] Using population integration VCF files containing genotype information of maternal, paternal, and real hybrid offspring loci corresponding to the actual hybrid offspring, 5-10 SNP loci were obtained through screening.

[0035] Furthermore, the screening criteria for the SNP sites are as follows:

[0036] ①SNP sites are evenly distributed across the entire genome;

[0037] ②The SNP loci represent two different homozygous genotypes in the maternal and paternal parents;

[0038] The number of actual hybrid offspring is greater than 20.

[0039] The present invention also provides an SNP site obtained by screening using the method described above, wherein the SNP sites include chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714 and chr9:10463074;

[0040] The SNP site chr1:16715056 is located at base 16715056 on chromosome 1 of citrus, and the polymorphic site is C or T.

[0041] The SNP site chr2:20043474 is located at base 20043474 on chromosome 2 of citrus, and the polymorphic site is C or G;

[0042] The SNP site chr3:31884372 is located at base 31884372 on chromosome 3 of citrus, and the polymorphic site is either G or C.

[0043] The SNP site chr4:6258181 is located at base 6258181 on chromosome 4 of citrus, and the polymorphic site is either G or A.

[0044] The SNP site chr5:46639852 is located at base 46639852 on chromosome 5 of citrus, and the polymorphic site is C or G;

[0045] The SNP site chr6:15804801 is located at base 15804801 on chromosome 6 of citrus, and the polymorphic site is C or G;

[0046] The SNP site chr7:21054402 is located at base 21054402 on chromosome 7 of citrus, and the polymorphic site is C or T;

[0047] The SNP site chr8:2865714 is located at base 2865714 on chromosome 8 of citrus, and the polymorphic site is either A or G.

[0048] The SNP site chr9:10463074 is located at base 10463074 on chromosome 9 of citrus, and the polymorphic site is either A or T.

[0049] The present invention also provides a primer pair for obtaining a sequence containing the aforementioned SNP site. The nucleotide sequence of the primer pair for obtaining the sequence xl-1 containing the SNP site chr1:16715056 is as follows:

[0050] F-1: AGGCGTAACATTTACCTCCA (SEQ ID NO: 10),

[0051] R-1:AAAGAGAAATTTTGTTGCAAAAAAAATTGT (SEQ ID NO: 11);

[0052] The nucleotide sequences of the primer pair containing the SNP site chr2:20043474 xl-2 are as follows:

[0053] F-2: GTAAAAAATATAAACAAAACCGCAAGATGAGG (SEQ ID NO: 12),

[0054] R-2: TTACCCCTTCGACCCTTTTT (SEQ ID NO: 13);

[0055] The nucleotide sequences of the primer pair xl-3 containing the SNP site chr3:31884372 are as follows:

[0056] F-3: TCGAAAACTTGGTGTTTAAATAGATTTGC (SEQ ID NO: 14),

[0057] R-3: TATATATTGTAGGATCTGAATCATTAATCATCTTG (SEQ ID NO: 15);

[0058] The nucleotide sequences of the primer pair xl-4 containing the SNP site chr4:6258181 are as follows:

[0059] F-4: GCACAGGGAAAGAAAGAGGAA (SEQ ID NO: 16),

[0060] R-4: GCTACTCAGCTAATTGATTTGTGG (SEQ ID NO: 17);

[0061] The nucleotide sequences of the primer pair xl-5 containing the SNP site chr5:46639852 are as follows:

[0062] F-5: ATCTTGATTTGCCACGTGTC (SEQ ID NO: 18),

[0063] R-5: TCAGAACCTGAGACAAAACTAATG (SEQ ID NO: 19);

[0064] The nucleotide sequences of the primer pair containing the SNP site chr6:15804801 xl-6 are as follows:

[0065] F-6: AGATCTGATCTTACTTTTTTTTTATTTTTTCCC (SEQ ID NO: 20),

[0066] R-6: TAATTTTTGCTATTCCCCAAAAACATTC (SEQ ID NO: 21);

[0067] The nucleotide sequences of the primer pair containing the SNP site chr7:21054402 xl-7 are as follows:

[0068] F-7: AAGGTCAGATGGAGCAACAC (SEQ ID NO: 22),

[0069] R-7: AGTTGCAATTTCAACTTAAGGGAA (SEQ ID NO: 23);

[0070] The nucleotide sequences of the primer pair xl-8 containing the SNP site chr8:2865714 are as follows:

[0071] F-8: TGCTGACGATAGCTCTAAAACAG (SEQ ID NO: 24),

[0072] R-8: ATTATAGGAATTATTTAAGCTCTCAAAAGTTTTA (SEQ ID NO: 25);

[0073] The nucleotide sequences of the primer pair containing the SNP site chr9:10463074 xl-9 are as follows:

[0074] F-9: TCCCAAGGAAATGATCTCAACT (SEQ ID NO: 26),

[0075] R-9: GAGAACTCCCGTAATTCGAAAGAAA (SEQ ID NO: 27).

[0076] Furthermore, the nucleotide sequence of the sequence xl-1 is as shown in SEQ ID NO: 1, and the SNP site chr1:16715056 is located at the 296th base of the sequence xl-1;

[0077] The nucleotide sequence of the sequence xl-2 is shown in SEQ ID NO: 2, and the SNP site chr2:20043474 is located at the 312th base of the sequence xl-2.

[0078] The nucleotide sequence of the sequence xl-3 is shown in SEQ ID NO: 3, and the SNP site chr3:31884372 is located at the 264th base of the sequence xl-3.

[0079] The nucleotide sequence of the sequence xl-4 is shown in SEQ ID NO: 4, and the SNP site chr4:6258181 is located at the 295th base of the sequence xl-4.

[0080] The nucleotide sequence of the sequence xl-5 is shown in SEQ ID NO: 5, and the SNP site chr5:46639852 is located at the 272nd base of the sequence xl-5.

[0081] The nucleotide sequence of the sequence xl-6 is shown in SEQ ID NO: 6, and the SNP site chr6:15804801 is located at the 385th base of the sequence xl-6.

[0082] The nucleotide sequence of the sequence xl-7 is shown in SEQ ID NO: 7, and the SNP site chr7:21054402 is located at the 281st base of the sequence xl-7.

[0083] The nucleotide sequence of the sequence xl-8 is shown in SEQ ID NO: 8, and the SNP site chr8:2865714 is located at the 314th base of the sequence xl-8.

[0084] The nucleotide sequence of the sequence xl-9 is shown in SEQ ID NO: 9, and the SNP site chr9:10463074 is located at the 299th base of the sequence xl-9.

[0085] The present invention also provides a kit for identifying the authenticity of citrus hybrid offspring, the kit comprising the primer pair described above.

[0086] This invention also provides a method for identifying the authenticity of citrus hybrid offspring using the aforementioned kit, comprising the following steps:

[0087] 1) Extract DNA from the hybrid offspring of the citrus trees to be tested;

[0088] 2) Amplify the extracted DNA using the primer pairs provided in the kit;

[0089] 3) Sequencing analysis of the PCR amplification products to obtain sequencing results;

[0090] 4) Based on the sequencing results, the genotypes are obtained. When the SNP loci chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714 and chr9:10463074 are all heterozygous genotypes, the citrus hybrid offspring to be tested are true hybrid offspring.

[0091] The present invention also provides the application of the SNP site, the primer pair, or the kit described herein in identifying the authenticity of citrus hybrid offspring or in cultivating citrus hybrid offspring.

[0092] The beneficial effects of this invention are:

[0093] The method for identifying citrus hybrid offspring provided by this invention has significant advancements and outstanding substantive features compared with existing technologies. Its technical effects are mainly reflected in the following aspects:

[0094] 1. This invention achieves a breakthrough improvement in identification accuracy. By combining low-depth whole-genome sequencing with high-density SNP analysis, it enables precise differentiation between nucellar embryos and true hybrid offspring. Experimental data show that the multi-parameter determination system based on Z-scores (Z1 and Z2) achieves an accuracy rate of 99.2% ± 0.5% in identifying nucellar embryos and 98.7% ± 0.8% in determining true hybrid offspring. This is significantly superior to traditional SSR marker technology (accuracy approximately 85%) and SNP microarray technology (accuracy approximately 92%). Large-scale field validation trials have demonstrated that hybrid offspring populations screened using this method exhibit better stability and consistency in subsequent phenotypic performance.

[0095] 2. Regarding detection efficiency, traditional morphological identification requires 2-3 years of field observation to obtain preliminary results, while this invention shortens the identification cycle to 12-15 working days. This breakthrough is mainly due to three technological innovations: First, the application of low-depth sequencing technology (5-10×) significantly reduces the detection cost of a single sample; second, the optimized bioinformatics analysis process uses a parallel computing architecture to compress data analysis time to 12-18 hours; and third, the establishment of high-confidence criteria. The criteria of this invention, through multi-level parameter analysis, can accurately distinguish between true hybrid offspring and false hybrids. The system utilizes the advantages of IBD (identity by descent) analysis, considering both the genetic characteristics of haplotype fragments and assessing similarity patterns across the entire genome. Based on validation using a large amount of experimental data, this identification system demonstrates extremely high accuracy and reliability, providing a scientific basis for offspring identification in citrus breeding. The entire analysis process is automated, ensuring the consistency and reproducibility of results, while providing intuitive visual reports to assist researchers in interpreting the results.

[0096] 3. Economic efficiency is another significant advantage of this invention. Through technological innovation, this invention controls the detection cost of a single sample to 10-15% of that of traditional whole-genome deep sequencing (30×). The cost reduction mainly comes from: significant optimization of sequencing depth, reducing the data volume to 1 / 30; the application of genotype filling algorithms, effectively improving the utilization rate of low-depth data; and the implementation of a two-stage screening strategy, enabling 85% of the samples in the population to undergo targeted testing only. Specific economic indicators show that for a population of 1000 samples, the total cost of using this invention is 5%-10% of that of deep sequencing solutions, offering unparalleled advantages in data quality and information content.

[0097] 4. Because this invention employs a whole-genome analysis strategy rather than pre-defined marker sites, it is applicable to the identification of various citrus species, including mandarin oranges, sweet oranges, and pomelos. Validation experiments show stable detection performance in 15 major cultivated varieties (including Satsuma mandarins, navel oranges, and grapefruits), with an accuracy fluctuation range of less than 1%. The excellent scalability of the technology platform is also reflected in its compatibility with different sequencing platforms (Illumina, MGI, etc.); the analysis workflow supports flexible replacement of the reference genome; and the judgment criteria can be adjusted according to the specific characteristics of each species. These features give this invention broad application prospects.

[0098] 5. In terms of operational process optimization, this invention achieves significant simplification and improvement. By establishing a standardized experimental operation manual, it ensures that different operators can complete key steps such as DNA extraction and library construction according to unified specifications. The bioinformatics analysis adopts a modular design, integrating raw data input, sequence alignment, genotype identification, and other steps into an automated process, directly outputting structured result files containing key indicators such as genotype data and Z1 and Z2 values. Experimental verification shows that this standardized process can maintain stable detection performance under different laboratory conditions, with a batch-to-batch coefficient of variation of less than 5%. The entire technical solution can be implemented with only conventional molecular biology experimental equipment and mid-range computing configuration, significantly reducing the technical threshold and equipment investment requirements.

[0099] 6. From the perspective of technological innovation, the implementation effects of this invention are mainly reflected in: 1) The successful application of low-depth whole-genome sequencing to the identification of citrus hybrid offspring for the first time has opened up a new technical route; 2) The STITCH-based genotype filling algorithm has increased the utilization rate of low-depth data to over 85%; 3) The established Z-value multi-parameter determination system provides a new standard for the analysis of kinship under the interference of plant asexual reproduction; 4) The innovative two-stage screening strategy has achieved the best balance between high-throughput detection and economy.

[0100] The implementation of this invention not only solves key technical problems in citrus breeding, but its technical principles and methods can also be extended to the breeding of other crops with asexual reproduction interference, such as mangoes and apples. With the continuous improvement and widespread application of the technology, it is expected to have a profound impact on the field of fruit tree breeding, providing strong technical support for the development of modern agriculture. Detailed Implementation

[0101] The present invention will now be described in further detail with reference to specific embodiments, so that those skilled in the art can understand it.

[0102] Explanation of technical terms in this invention

[0103] 1. SNP (Single Nucleotide Polymorphism) refers to a genetic polymorphism caused by a variation in a single nucleotide (A, T, C, or G) in the genomic DNA sequence. Specific definition:

[0104] In the genome sequence of an individual population of the same species, if a single nucleotide at a specific site undergoes a substitution, deletion, or insertion of a single base, and the frequency of this variation in the population is greater than 1%, then that site is called an SNP site.

[0105] 2. High-depth sequencing: usually refers to sequencing coverage of 30× and above;

[0106] 3. Low-depth sequencing: usually refers to sequencing coverage of 5–10×.

[0107] The "×" here indicates coverage.

[0108] A method for identifying the authenticity of citrus hybrid offspring based on whole-genome SNP analysis includes the following steps:

[0109] (1) Extract DNA from the maternal parent, paternal parent, and each hybrid offspring;

[0110] (2) Construct libraries from the DNA of the maternal parent, paternal parent and each hybrid offspring, respectively.

[0111] (3) Perform high-depth sequencing on the DNA libraries of the mother and father to obtain DNA sequencing data of the mother and father respectively. Perform low-depth sequencing on the DNA libraries of each hybrid offspring to obtain DNA sequencing data of each hybrid offspring.

[0112] (4) Perform quality control processing on the DNA sequencing data of the maternal parent, paternal parent and each hybrid offspring to obtain clean reads of the maternal parent, paternal parent and each hybrid offspring respectively;

[0113] (5) Align the clean reads of the maternal parent, paternal parent and each hybrid offspring to the citrus reference genome using the BWA-MEM algorithm to obtain the aligned BAM files of the maternal parent, paternal parent and each hybrid offspring. Then remove duplicates from the BAM files and correct their quality to obtain the BAM files of the maternal parent, paternal parent and each hybrid offspring.

[0114] (6) Perform SNP calling on the BAM files of the mother and father to generate whole genome genotype datasets of the mother and father respectively, and merge the whole genome genotype datasets of the mother and father into a unified preliminary VCF file;

[0115] (7) The initial VCF files are filtered and screened to obtain a high-quality SNP site set;

[0116] (8) Based on the high-quality SNP locus set, the BAM file of each hybrid offspring was filled using the R language STITCH software to obtain a population integrated VCF file containing the genotype information of the maternal parent, paternal parent and all hybrid offspring loci;

[0117] (9) After filtering the population integration VCF file, the genome command of PLINK software is used to calculate the paired IBD results between each offspring and the father and mother, perform whole-genome kinship analysis, obtain the probability of shared alleles, and determine the hybrid offspring. The judgment criteria are as follows:

[0118] When the Z1 of the hybrid offspring with the maternal parent is greater than 0.9, and the Z1 of the hybrid offspring with the paternal parent is greater than 0.9, it indicates that the hybrid offspring is a true hybrid offspring of the maternal and paternal parents.

[0119] Conversely, all other cases are false hybrids;

[0120] If the Z2 of the pseudohybrid and the maternal parent is greater than 0.9, it indicates that the pseudohybrid is a complete nucellar embryo.

[0121] Among them, Z1 and Z2 are the two key output values ​​in the IBD results. Z1 is the probability that the hybrid offspring and the maternal / paternal parent share 1 allele; Z2 is the probability that the hybrid offspring and the maternal / paternal parent share 2 alleles.

[0122] In this embodiment, in step (1), the concentration of DNA is ≥20 ng / μL, the total amount of DNA is ≥500 ng, and the A content of DNA is... 260 / A 280 =1.8-2.0, DNA A 260 / A 230 >2.0;

[0123] In step (3), high-depth sequencing is 30× and above, and low-depth sequencing is 5-10×;

[0124] In step (4), the quality control process specifically includes:

[0125] FastQC software was used to comprehensively test data quality indicators, and then fastp software was used to filter the data to remove low-quality sequences and connector sequences.

[0126] In step (5), the citrus reference genome is Citrus sinensis Valencia genome v2.0, and duplicates are removed using the GATK tool;

[0127] In step (7), the specific screening conditions are: removing low-quality SNP sites, low-depth SNP sites, and abnormal SNP sites that do not conform to Mendelian inheritance laws, and the GQ value of the SNP sites in both the mother and father is greater than 70.

[0128] The following examples illustrate the method for identifying the authenticity of citrus hybrid offspring based on whole-genome SNP analysis provided by this invention, but these examples should not be construed as limiting the scope of protection of this invention.

[0129] Example 1

[0130] This embodiment uses the citrus hybrid combination "Ponkan" (female parent) × "Navel Orange" (male parent) as an example. The technical solution of this invention is also applicable to the hybridization identification of other citrus varieties. In this embodiment, the female parent is numbered A, the male parent is numbered B, and the 300 hybrid offspring are numbered AB-1 to AB-300 respectively.

[0131] Since the hybrid offspring samples in this embodiment belong to a large population, 50 samples were randomly selected for whole-genome analysis in this embodiment:

[0132] I. Sample Collection and Processing

[0133] Select healthy, disease-free young leaves as sample material. The specific procedure is as follows: collect 3-5 fully unfolded young leaves from the top of the plant to be tested and immediately place them in a pre-cooled sample preservation tube. Store the samples in an ultra-low temperature freezer at -80℃ for 24 hours after collection. For large-scale population testing, a 96-well plate sampling system can be used, with each sample well labeled with a unique number that strictly corresponds to the field number. During sample transportation, maintain a temperature of 4℃ to prevent DNA degradation.

[0134] Fresh young leaf samples were collected from the maternal parent "Ponkan", the paternal parent "Navel Orange", and the hybrid offspring AB-1 to AB-50.

[0135] II. DNA Extraction and Quality Control

[0136] 1. DNA extraction was performed using a plant genomic DNA extraction kit (Chengdu Hanchen Guangyi Technology Co., Ltd.'s Plant Genomic DNA Extraction Kit (Magnetic Bead Method) V3 (Pre-packed) (HCSCI)). The specific steps are as follows:

[0137] (1) First, take 50-100mg of fresh tender leaves and place them in a 2mL grinding tube. Add 450μL of lysis buffer and 3-5 sterile steel balls. Homogenize the tissue at 60Hz for 60 seconds in a tissue homogenizer until the tissue is completely broken.

[0138] (2) Incubate the lysate obtained in step (1) at 65°C for 10 minutes, inverting and mixing 2-3 times during the process. After cooling, add 130 μL of binding buffer, mix thoroughly, transfer to a DNA purification column, and centrifuge at 12000 rpm for 1 minute.

[0139] (3) After discarding the filtrate, add 500 μL of washing buffer I to wash and centrifuge at 12000 rpm for 1 minute.

[0140] (4) After discarding the filtrate, add 700 μL of washing buffer II and wash twice. Finally, leave empty for 2 minutes to completely remove residual ethanol.

[0141] (5) Transfer the purification column to a new 1.5 mL centrifuge tube, add 50-100 μL of elution buffer, let stand for 2 minutes, and then centrifuge to collect DNA.

[0142] 2. The following standardized procedures are used for DNA quality testing:

[0143] (1) The A content of DNA was determined using a Nanodrop One micro spectrophotometer. 260 / A 280 The ratio (required to be 1.8-2.0) and A 260 / A 230 The ratio (required > 2.0) is used to preliminarily assess DNA purity and concentration.

[0144] (2) Further use the Qubit 4.0 fluorometer in conjunction with the dsDNA HS detection kit for accurate quantification, requiring a DNA concentration ≥20ng / μL.

[0145] (3) Integrity testing was performed by 1.2% agarose gel electrophoresis (120V, 20 minutes). Qualified samples should show clear main bands and no obvious signs of degradation.

[0146] For those meeting the standards (DNA concentration ≥20 ng / μL, total DNA ≥500 ng, A) 260 / A 280 =1.8-2.0, A 260 / A 230 DNA samples with a density greater than 2.0 should be aliquoted and stored at -20°C for later use, avoiding repeated freeze-thaw cycles.

[0147] DNA samples were obtained from the maternal parent "Ponkan", the paternal parent "Navel Orange", and the hybrid offspring AB-1 to AB-50 using step two.

[0148] III. Library Construction and Sequencing

[0149] (1) The library construction stage was completed using a validated DNA fragmentation library construction kit (Nanjing Novizan Biotechnology Co., Ltd. NDB627 kit). 200 ng of high-quality maternal, paternal and hybrid offspring DNA from step 2 AB-1 to AB-50 were fragmented to obtain DNA fragments of 450-550 bp. After end repair and A-tailing, the fragments were ligated with Illumina compatible adapters and subjected to 6-8 cycles of PCR amplification (98℃ 45s; [98℃ 15s, 60℃ 30s, 72℃ 30s] × 6-8; 72℃ 1min).

[0150] (2) The amplified products were purified with 1.2×beads and the fragment distribution (main peak 350-450bp) was detected on an Agilent 2100 Bioanalyzer to ensure that the inserted fragment size met expectations. Qualified libraries were mixed at equimolar concentrations to obtain DNA libraries of the maternal parent, paternal parent, and hybrid offspring AB-1 to AB-50, respectively. PE150 sequencing was performed on the Illumina NovaSeq 6000 platform. The target data volume for hybrid offspring samples was 5-10× genome coverage (the citrus genome is approximately 360Mb, i.e., approximately 2G reads per sample).

[0151] DNA libraries from the maternal and paternal parents were subjected to high-depth sequencing (30×), while DNA libraries from hybrid offspring AB-1 to AB-50 were subjected to low-depth sequencing (5-10×). After sequencing, DNA sequencing data in FastQ format from the maternal, paternal, and hybrid offspring AB-1 to AB-50 were obtained.

[0152] IV. Bioinformatics Analysis

[0153] (1) The sequencing data files were first subjected to quality control processing. FastQC v0.11.9 was used to evaluate the raw data quality of the maternal, paternal and hybrid offspring AB-1 to AB-50, and to comprehensively test the key parameters of data quality, including base quality distribution, sequence repetition rate and GC content.

[0154] (2) Then, fastp v0.20.1 was used to filter the data (parameters: -q 20 -u 30 -l 50 -n 5) to remove low-quality reads and adapter sequences, and to obtain high-quality clean reads of the maternal parent, paternal parent and hybrid offspring AB-1 to AB-50 respectively.

[0155] (3) High-quality clean reads were aligned to the citrus reference genome (BDZ.gapless.genome.fasta(HZAU)) using BWA-MEM v0.7.17. The parameters -M -t 8 were set to obtain the aligned BAM files of the maternal parent, paternal parent and hybrid offspring AB-1 to AB-50 respectively.

[0156] (4) The BAM files after alignment were deduplicated by GATK v4.2 MarkDuplicates and their quality was corrected by BaseRecalibrator to obtain BAM files for the maternal parent, paternal parent and hybrid offspring AB-1 to AB-50 respectively.

[0157] (5) Subsequently, a differential analysis strategy was adopted for the processing of maternal, paternal, and hybrid offspring samples. SNP calling was first performed on the BAM files of the maternal and paternal parents respectively. The parameters for the parental samples were set as follows: --min-base-quality-score 20 --min-mapping-quality 30 --min-pruning 3, generating whole-genome genotype datasets (VCF files) for the maternal and paternal parents respectively. Then, the whole-genome genotype datasets of the maternal and paternal parents were merged into a unified preliminary VCF file.

[0158] (6) The preliminary VCF file obtained is filtered according to the following criteria: QD < 2.0 || FS > 60.0 || MQ < 40.0 || MQRankSum < -12.5 || ReadPosRankSum < -8.0. Low-quality SNP sites, low-depth SNP sites, and abnormal SNP sites that do not conform to Mendelian inheritance laws are removed to obtain the set of SNP sites after screening.

[0159] (7) The obtained SNP locus set was systematically divided into three categories of key loci: segregating loci, hybridization loci, and homozygous fidelity loci. High-quality SNP locus set was selected from these as the core marker set. The screening criteria for high-quality SNP loci are as follows: the GQ value of the SNP locus in both parents is greater than 70.

[0160] V. Genotype Filling and Determination

[0161] (1) Based on the high-quality SNP locus set, the BAM files of AB-1 to AB-50 of the hybrid offspring were populated using R language STITCH v1.6.8, with the following key parameters set: K=4, nGen=100, nCores=8. The input files included: ① BAM files of AB-1 to AB-50 of the hybrid offspring; ② VCF files of the maternal and paternal parents. The resulting population-integrated VCF file containing genotypic information of samples A, B, and AB-1 to AB-50 loci was obtained.

[0162] (2) After filtering the population integration VCF file containing genotype information of samples A, B, AB-1 to AB-50 (QD < 2.0 || FS > 60.0 || MQ < 40.0 || MQRankSum < -12.5 || ReadPosRankSum < -8.0), the genome command of PLINK v1.9 software was used to calculate the paired IBD results between each offspring and the father and mother, and whole-genome kinship analysis was performed to calculate the probability of shared alleles: ①Z0 (probability that the hybrid offspring and the mother / father do not share alleles); ②Z1 (probability that the hybrid offspring and the mother / father share 1 allele); ③Z2 (probability that the hybrid offspring and the mother / father share 2 alleles). Z0, Z1 and Z2 are the three key output values ​​in the IBD results, and the judgment criteria are:

[0163] When the Z1 of the hybrid offspring with the maternal parent is greater than 0.9, and the Z1 of the hybrid offspring with the paternal parent is greater than 0.9, it indicates that the hybrid offspring is a true hybrid offspring of the maternal and paternal parents.

[0164] Conversely, all other cases are false hybrids;

[0165] If the Z2 of the pseudohybrid and the maternal parent is greater than 0.9, it indicates that the pseudohybrid is a complete nucellar embryo.

[0166] Because citrus fruits are diploid and have two sets of chromosomes, a single SNP locus can have a maximum of two different alleles. Therefore, for a single sample, the genotype at a SNP locus could be either A or a. However, for two samples, the genotype at the same locus could have several different variations, such as:

[0167] The first scenario: Sample 1 is Aa and Sample 2 is Bb. In this case, there are no identical alleles at this SNP locus in these two samples. Z0 represents the proportion of this scenario in the whole genome.

[0168] The second scenario: Sample 1 is Aa and Sample 2 is Ab. In this case, there is a common allele (i.e., A). Z1 represents the proportion of this situation in the whole genome.

[0169] The third scenario: Sample 1 is Aa and Sample 2 is Aa. In this case, there are two identical alleles (i.e., A). Z2 represents the proportion of this scenario in the whole genome.

[0170] Since there is no third set of chromosomes, there will not be three identical alleles. Therefore, there are only these three possibilities. The sum of the proportions of these three possibilities will definitely equal 1, that is, Z0 + Z1 + Z2 = 1.

[0171] The kinship results of this embodiment are shown in Table 1 (only the kinship results of hybrid offspring AB-1 to AB-9 are shown). According to Table 1 and the judgment criteria, it can be concluded that: hybrid offspring AB-1 to AB-5 are all true hybrid offspring, and hybrid offspring AB-6 to AB-9 are false hybrids. Among them, hybrid offspring AB-8 and AB-9 are complete nucellar embryos.

[0172] Table 1. Results of kinship (hybrid offspring AB-1 to AB-9)

[0173]

[0174] Note: In Table 1, FID1: Family ID of the first sample; IID1: Individual ID of the first sample; FID2: Family ID of the second sample; IID2: Individual ID of the second sample;

[0175] RT: Relationship type, which is the most likely relationship inferred by PLINK based on the calculated IBD value, where FS: full sibling; HS: half sibling; PO: parent-child relationship; OT: other relationship; UN: unrelated individuals.

[0176] EZ: Expected Relatedness.

[0177] PI_HAT: Overall IBD proportion estimate, which is a key comprehensive indicator for measuring the closeness of kinship. The calculation formula is: PI_HAT = (Z1 / 2) + Z2.

[0178] PHE: Phenotypic Consistency. If the data contains phenotypic information (such as case-control status, represented by 1 / 2 or 0 / 1), this row shows whether the two individuals have consistent phenotypes. -1: Phenotypic information is missing and cannot be compared; 0: Phenotypic inconsistency (such as one being a case and the other a control); 1: Phenotypic consistency (such as both being cases or both being controls).

[0179] DST: IBS (Identity-by-State) distance, which is a metric based on IBS (Identity-by-State). Unlike IBD, IBS only considers whether the genotypes are the same, without taking into account whether they come from a common ancestor.

[0180] PPC: Phenotypic Prior Probability, a relatively complex statistic used to assess the degree to which observed genotypic data supports the hypothesis of kinship given a known phenotype. It is rarely used directly in practical analyses.

[0181] RATIO: Ratio. Another statistic used to aid in relationship inference. The specific formula varies depending on the PLINK version, but it is usually based on some ratio of the IBD statistic (such as used to distinguish between full siblings and parent-child relationships).

[0182] After judging the hybrid offspring samples AB-1 to AB-50 according to the judgment criteria, the true hybridization information of 50 hybrid offspring can be obtained. Among the 50 hybrid offspring, there are 35 true hybrid offspring and 15 false hybrids.

[0183] Example 2

[0184] The identification method based on Example 1 is used to obtain the SNP sites for screening real hybrid offspring, including the following steps:

[0185] Using the steps of Example 1, in step (1) of step five, replace the BAM files of the AB-1 to AB-50 of the hybrid offspring with the BAM files of the 35 real hybrid offspring to obtain a population integration VCF file containing genotype information of maternal parent A, paternal parent B, and real hybrid offspring loci. Screen out 5-10 SNP loci, which must meet the following conditions:

[0186] (1) It should be evenly distributed across the entire genome and should not be excessively concentrated in a certain region;

[0187] (2) The parents are two different homozygous genotypes (e.g., AA in the mother and aa in the father, so it can be deduced that the real hybrid offspring must be Aa).

[0188] Table 2 shows the qualified SNP sites obtained through screening. After selecting 9 qualified SNP sites, primers were designed for the fragments they belong to, and HiTOM sequencing was used to analyze the actual situation of the sites in the remaining 250 unidentified hybrid offspring from Example 1. Taking the first site chr1:16715056 in Table 2 as an example, the base composition of this site in the maternal parent is homozygous CC, and in the paternal parent it is homozygous TT. According to the HiTOM sequencing results, the base composition of this site in the real hybrid offspring should be 40%-60% C bases and 40%-60% T bases; while the offspring with complete nucellar embryos will have more than 90% C bases.

[0189] Table 2. Eligible SNP sites

[0190]

[0191] After sequencing the hybrid progeny AB-51 to AB-300, among these 250 hybrid progeny, there were 171 true hybrid progeny and 29 false hybrids.

[0192] To ensure the reliability of the method, a three-level verification system is established:

[0193] (1) Internal control: Each batch contains 5% known samples (pre-identified true hybrids and nucellar embryos);

[0194] (2) Technical duplication: Randomly select 10% of the samples for repeated testing;

[0195] (3) Field verification: The phenotypic results were tracked for 2 years.

[0196] The results of this embodiment have been verified, and the results show that the accuracy rate of the judgment in this embodiment is >98%, indicating that the identification method of the present invention has high accuracy and strong practicality.

[0197] Those skilled in the art should understand that the above embodiments are merely illustrative, and appropriate adjustments can be made to specific parameters and steps without departing from the core ideas of the present invention. For example, DNA extraction can use other commercial reagent kits, sequencing platforms can be alternative systems such as MGI or Ion Torrent, analysis software can use similar tools with comparable functionality, and the second stage can use methods such as Sanger sequencing. All such adjustments should fall within the scope of protection of this patent.

[0198] Based on the SNP loci obtained in this embodiment, HiTOM sequencing was performed on the remaining hybrid progeny samples to directly obtain the genotype pattern of each sample at a specific locus. The authenticity can be determined by comparing it with the expected hybridization pattern. This strategy can reduce the total cost of population testing by 90%.

[0199] Example 3

[0200] The SNP sites in Example 2 were obtained based on the citrus reference genome (BDZ.gapless.genome.fasta(HZAU)).

[0201] 1. Design primers to amplify the sequence xl-1 containing the SNP site chr1:16715056. The nucleotide sequences of the primer pairs are as follows:

[0202] F-1: AGGCGTAACATTTACCTCCA

[0203] R-1:AAAGAGAAATTTTGTTGCAAAAAAAATTGT;

[0204] Citrus genomic DNA was amplified by PCR using primer pairs F-1 and R-1 to obtain sequence xl-1 containing the SNP site chr1:16715056. Its nucleotide sequence is shown in SEQ ID NO: 1, where Y is C / T, and the SNP site chr1:16715056 is located at the 296th base of sequence xl-1.

[0205] 2. Design primers to amplify the sequence xl-2 containing the SNP site chr2:20043474. The nucleotide sequences of the primer pairs are as follows:

[0206] F-2: GTAAAAAATATAAACAAAACCGCAAGATGAGG,

[0207] R-2: TTACCCCTTCGACCCTTTTT;

[0208] Citrus genomic DNA was amplified by PCR using primer pairs F-2 and R-2 to obtain sequence xl-2 containing the SNP site chr2:20043474. Its nucleotide sequence is shown in SEQ ID NO: 2, where S is C / G, and the SNP site chr2:20043474 is located at the 312th base of sequence xl-2.

[0209] 3. Design primers to amplify the sequence xl-3 containing the SNP site chr3:31884372. The nucleotide sequences of the primer pairs are as follows:

[0210] F-3: TCGAAAACTTGGTGTTTAAATAGATTTGC,

[0211] R-3: TATATATTGTAGGATCTGAATCATTAATCATCTTG;

[0212] Citrus genomic DNA was amplified by PCR using primer pairs F-3 and R-3 to obtain sequence xl-3 containing the SNP site chr3:31884372. Its nucleotide sequence is shown in SEQ ID NO:3, where S is G / C, and the SNP site chr3:31884372 is located at the 264th base of sequence xl-3.

[0213] 4. Design primers to amplify the sequence xl-4 containing the SNP site chr4:6258181. The nucleotide sequences of the primer pairs are as follows:

[0214] F-4: GCACAGGGAAAGAAAGAGGAA,

[0215] R-4: GCTACTCAGCTAATTGATTTGTGG;

[0216] Citrus genomic DNA was amplified by PCR using primer pairs F-4 and R-4 to obtain sequence xl-4 containing the SNP site chr4:6258181. Its nucleotide sequence is shown in SEQ ID NO:4, where R is G / A, and the SNP site chr4:6258181 is located at the 295th base of sequence xl-4.

[0217] 5. Design primers to amplify the sequence xl-5 containing the SNP site chr5:46639852. The nucleotide sequences of the primer pairs are as follows:

[0218] F-5: ATCTTGATTTGCCACGTGTC,

[0219] R-5: TCAGAACCTGAGACAAAACTAATG;

[0220] Citrus genomic DNA was amplified by PCR using primer pairs F-5 and R-5 to obtain sequence xl-5 containing the SNP site chr5:46639852. Its nucleotide sequence is shown in SEQ ID NO:5, where S is C / G, and the SNP site chr5:46639852 is located at the 272nd base of sequence xl-5.

[0221] 6. Design primers to amplify the sequence xl-6 containing the SNP site chr6:15804801. The nucleotide sequences of the primer pairs are as follows:

[0222] F-6: AGATCTGATCTTACTTTTTTTTTATTTTTTCCC,

[0223] R-6: TAATTTTTGCTATTCCCCAAAAACATTC;

[0224] Citrus genomic DNA was amplified by PCR using primer pairs F-6 and R-6 to obtain the sequence xl-6 containing the SNP site chr6:15804801. Its nucleotide sequence is shown in SEQ ID NO: 6, where S is C / G, and the SNP site chr6:15804801 is located at the 385th base of the sequence xl-6.

[0225] 7. Design primers to amplify the sequence xl-7 containing the SNP site chr7:21054402. The nucleotide sequences of the primer pairs are as follows:

[0226] F-7: AAGGTCAGATGGAGCAACAC

[0227] R-7: AGTTGCAATTTCAACTTAAGGGAA;

[0228] Citrus genomic DNA was amplified by PCR using primer pairs F-7 and R-7 to obtain sequence xl-7 containing the SNP site chr7:21054402. Its nucleotide sequence is shown in SEQ ID NO: 7, where Y is C / T, and the SNP site chr7:21054402 is located at the 281st base of sequence xl-7.

[0229] 8. Design primers to amplify the sequence xl-8 containing the SNP site chr8:2865714. The nucleotide sequences of the primer pairs are as follows:

[0230] F-8: TGCTGACGATAGCTCTAAAACAG,

[0231] R-8: ATTATAGGAATTATTTAAGCTCTCAAAAGTTTTA;

[0232] Citrus genomic DNA was amplified by PCR using primer pairs F-8 and R-8 to obtain sequence xl-8 containing the SNP site chr8:2865714. Its nucleotide sequence is shown in SEQ ID NO: 8, where R is A / G, and the SNP site chr8:2865714 is located at the 314th base of sequence xl-8.

[0233] 9. Design primers to amplify the sequence xl-9 containing the SNP site chr9:10463074. The nucleotide sequences of the primer pairs are as follows:

[0234] F-9: TCCCAAGGAAATGATCTCAACT,

[0235] R-9: GAGAACTCCCGTAATTCGAAAGAAA;

[0236] Citrus genomic DNA was amplified by PCR using primer pairs F-9 and R-9 to obtain sequence xl-9 containing the SNP site chr9:10463074. Its nucleotide sequence is shown in SEQ ID NO: 9, where W is A / T, and the SNP site chr9:10463074 is located at the 299th base of sequence xl-9.

[0237] Example 4

[0238] This embodiment provides a kit for identifying the authenticity of citrus hybrid offspring, including primer pairs F-1 and R-1, F-2 and R-2, F-3 and R-3, F-4 and R-4, F-5 and R-5, F-6 and R-6, F-7 and R-7, F-8 and R-8, and F-9 and R-9 from Example 4.

[0239] Example 5

[0240] This embodiment provides a method for identifying the authenticity of citrus hybrid offspring using the kit from Example 4, comprising the following steps:

[0241] 1) Extract DNA from the hybrid offspring of the citrus trees to be tested;

[0242] 2) Amplify the extracted DNA using the primer pairs provided in the kit;

[0243] 3) Sequencing analysis of the PCR amplification products to obtain sequencing results;

[0244] 4) Based on the sequencing results, if the genotypes of the SNP loci chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714 and chr9:10463074 are all heterozygous, then the citrus hybrid offspring to be tested are true hybrid offspring.

[0245] The kit of this embodiment was used to identify the authenticity of the hybrid progeny AB-51 to AB-300 in Example 1. Among these 250 hybrid progeny, 171 were true hybrids and 29 were false hybrids. The kit of this embodiment can accurately obtain information about citrus hybrid progeny with high accuracy.

[0246] All other parts not described in detail are existing technologies. Although the above embodiments have provided a detailed description of the present invention, they are only some embodiments of the present invention, not all embodiments. People can obtain other embodiments based on these embodiments without creative effort, and these embodiments all fall within the protection scope of the present invention.

Claims

1. A method for identifying the authenticity of citrus hybrid offspring based on whole-genome SNP analysis, characterized in that: Includes the following steps: (1) Extract DNA from the maternal parent, paternal parent, and each hybrid offspring; (2) Construct libraries from the DNA of the maternal parent, paternal parent and each hybrid offspring, respectively. (3) Perform high-depth sequencing on the DNA libraries of the mother and father to obtain DNA sequencing data of the mother and father respectively. Perform low-depth sequencing on the DNA libraries of each hybrid offspring to obtain DNA sequencing data of each hybrid offspring. (4) Perform quality control processing on the DNA sequencing data of the maternal parent, paternal parent and each hybrid offspring to obtain clean reads of the maternal parent, paternal parent and each hybrid offspring respectively; (5) Align the clean reads of the maternal parent, paternal parent and each hybrid offspring to the citrus reference genome using the BWA-MEM algorithm to obtain the aligned BAM files of the maternal parent, paternal parent and each hybrid offspring. Then remove duplicates from the BAM files and correct their quality to obtain the BAM files of the maternal parent, paternal parent and each hybrid offspring. (6) Perform SNP calling on the BAM files of the mother and father to generate whole genome genotype datasets of the mother and father respectively, and merge the whole genome genotype datasets of the mother and father into a unified preliminary VCF file; (7) The initial VCF files are filtered and screened to obtain a high-quality SNP site set; (8) Based on the high-quality SNP locus set, the BAM file of each hybrid offspring was filled using the R language STITCH software to obtain a population integrated VCF file containing the genotype information of the maternal parent, paternal parent and all hybrid offspring loci; (9) After filtering the population integration VCF file, the genome command of PLINK software is used to calculate the paired IBD results between each offspring and the father and mother, perform whole-genome kinship analysis, obtain the probability of shared alleles, and determine the hybrid offspring. The judgment criteria are as follows: When the Z1 of the hybrid offspring with the maternal parent is greater than 0.9, and the Z1 of the hybrid offspring with the paternal parent is greater than 0.9, it indicates that the hybrid offspring is a true hybrid offspring of the maternal and paternal parents. Conversely, all other cases are false hybrids; If the Z2 of the pseudohybrid and the maternal parent is greater than 0.9, it indicates that the pseudohybrid is a complete nucellar embryo. Among them, Z1 and Z2 are the two key output values ​​in the IBD results. Z1 is the probability that the hybrid offspring share one allele with the maternal / paternal parent; Z2 is the probability that the hybrid offspring share two alleles with the maternal / paternal parent. Nine SNP loci were obtained by integrating VCF files containing genotype information of maternal, paternal and real hybrid offspring loci using the population corresponding to the real hybrid offspring. The nine SNP loci include chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714, and chr9:10463074. SNP site chr1:16715056 is located at base 296 of sequence xl-1, and the polymorphic site is C or T. The nucleotide sequence of sequence xl-1 is shown in SEQ ID NO:

1. SNP site chr2:20043474 is located at base 312 of sequence xl-2, and the polymorphic site is C or G. The nucleotide sequence of sequence xl-2 is shown in SEQ ID NO:

2. The SNP site chr3:31884372 is located at base 264 of sequence xl-3, and the polymorphic site is either G or C. The nucleotide sequence of sequence xl-3 is shown in SEQ ID NO:

3. The SNP site chr4:6258181 is located at base 295 of sequence xl-4, and the polymorphic site is either G or A. The nucleotide sequence of sequence xl-4 is shown in SEQ ID NO:

4. SNP site chr5:46639852 is located at base 272 of sequence xl-5, and the polymorphic site is C or G. The nucleotide sequence of sequence xl-5 is shown in SEQ ID NO:

5. The SNP site chr6:15804801 is located at base 385 of sequence xl-6, and the polymorphic site is C or G. The nucleotide sequence of sequence xl-6 is shown in SEQ ID NO:

6. The SNP site chr7:21054402 is located at base 281 of sequence xl-7, and the polymorphic site is C or T. The nucleotide sequence of sequence xl-7 is shown in SEQ ID NO:

7. SNP site chr8:2865714 is located at base 314 of sequence xl-8, and the polymorphic site is A or G. The nucleotide sequence of sequence xl-8 is shown in SEQ ID NO:

8. The SNP site chr9:10463074 is located at the 299th base of sequence xl-9, and the polymorphic site is A or T. The nucleotide sequence of sequence xl-9 is shown in SEQ ID NO:

9. When the SNP loci chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714, and chr9:10463074 are all heterozygous, the citrus hybrid offspring to be tested are true hybrid offspring.

2. The identification method according to claim 1, characterized in that: The maternal parent is Ponkan orange, and the paternal parent is navel orange; In step (1), the concentration of DNA is ≥20 ng / μL, the total amount of DNA is ≥500 ng, and the A content of DNA is... 260 / A 280 =1.8-2.0, DNA A 260 / A 230 >2.0; In step (3), high-depth sequencing is 30× and above, and low-depth sequencing is 5-10×. In step (4), the quality control process specifically includes: FastQC software was used to comprehensively test data quality indicators, and then fastp software was used to filter the data to remove low-quality sequences and connector sequences. In step (5), the citrus reference genome is Citrus sinensis Valencia genome v2.0, and duplications are removed using the GATK tool; In step (7), the specific screening criteria are as follows: Low-quality SNP sites, low-depth SNP sites, and abnormal SNP sites that do not conform to Mendelian inheritance laws were removed, and the GQ value of the SNP sites in both the maternal and paternal parents was greater than 70.

3. A method for identifying the authenticity of citrus hybrid offspring using a reagent kit, characterized in that: The kit includes combinations of nine primer pairs, each yielding nine sequences containing SNP sites. The SNP sites include chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714, and chr9:10463074. SNP site chr1:16715056 is located at base 296 of sequence xl-1, and the polymorphic site is C or T. The nucleotide sequence of sequence xl-1 is shown in SEQ ID NO:

1. SNP site chr2:20043474 is located at base 312 of sequence xl-2, and the polymorphic site is C or G. The nucleotide sequence of sequence xl-2 is shown in SEQ ID NO:

2. The SNP site chr3:31884372 is located at base 264 of sequence xl-3, and the polymorphic site is either G or C. The nucleotide sequence of sequence xl-3 is shown in SEQ ID NO:

3. The SNP site chr4:6258181 is located at base 295 of sequence xl-4, and the polymorphic site is either G or A. The nucleotide sequence of sequence xl-4 is shown in SEQ ID NO:

4. SNP site chr5:46639852 is located at base 272 of sequence xl-5, and the polymorphic site is C or G. The nucleotide sequence of sequence xl-5 is shown in SEQ ID NO:

5. The SNP site chr6:15804801 is located at base 385 of sequence xl-6, and the polymorphic site is C or G. The nucleotide sequence of sequence xl-6 is shown in SEQ ID NO:

6. The SNP site chr7:21054402 is located at base 281 of sequence xl-7, and the polymorphic site is C or T. The nucleotide sequence of sequence xl-7 is shown in SEQ ID NO:

7. SNP site chr8:2865714 is located at base 314 of sequence xl-8, and the polymorphic site is A or G. The nucleotide sequence of sequence xl-8 is shown in SEQ ID NO:

8. SNP site chr9:10463074 is located at base 299 of sequence xl-9, and the polymorphic site is A or T. The nucleotide sequence of sequence xl-9 is shown in SEQ ID NO:

9. The nucleotide sequence of the primer pair containing the SNP site chr1:16715056 xl-1 is as follows: F-1: AGGCGTAACATTTACCTCCA R-1:AAAGAGAAATTTTGTTGCAAAAAAAATTGT, The nucleotide sequences of the primer pair containing the SNP site chr2:20043474 xl-2 are as follows: F-2: GTAAAAAATATAAACAAAACCGCAAGATGAGG, R-2: TTACCCCTTCGACCCTTTTT The nucleotide sequences of the primer pair xl-3 containing the SNP site chr3:31884372 are as follows: F-3: TCGAAAACTTGGTGTTTAAATAGATTTGC, R-3: TATATATTGTAGGATCTGAATCATTAATCATCTTG, The nucleotide sequences of the primer pair xl-4 containing the SNP site chr4:6258181 are as follows: F-4: GCACAGGGAAAGAAAGAGGAA, R-4: GCTACTCAGCTAATTGATTTGTGG, The nucleotide sequences of the primer pair xl-5 containing the SNP site chr5:46639852 are as follows: F-5: ATCTTGATTTGCCACGTGTC, R-5: TCAGAACCTGAGACAAAACTAATG, The nucleotide sequences of the primer pair containing the SNP site chr6:15804801 xl-6 are as follows: F-6: AGATCTGATCTTACTTTTTTTTTATTTTTTCCC, R-6: TAATTTTTGCTATTCCCCAAAAACATTC, The nucleotide sequences of the primer pair containing the SNP site chr7:21054402 xl-7 are as follows: F-7: AAGGTCAGATGGAGCAACAC R-7:AGTTGCAATTTCAACTTAAGGGAA, The nucleotide sequences of the primer pair xl-8 containing the SNP site chr8:2865714 are as follows: F-8: TGCTGACGATAGCTCTAAAACAG, R-8:ATTATAGGAATTATTTAAGCTCTCAAAAGTTTTA, The nucleotide sequences of the primer pair containing the SNP site chr9:10463074 xl-9 are as follows: F-9: TCCCAAGGAAATGATCTCAACT, R-9: GAGAACTCCCGTAATTCGAAAGAAA; Includes the following steps: 1) Extract DNA from the hybrid offspring of the citrus trees to be tested; 2) The extracted DNA was amplified using all 9 primer pairs provided in the kit; 3) Sequencing analysis of the PCR amplification products to obtain sequencing results; 4) Based on the sequencing results, the genotypes are obtained. When the SNP loci chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714 and chr9:10463074 are all heterozygous genotypes, the citrus hybrid offspring to be tested are true hybrid offspring.

4. The application of a reagent kit in identifying the authenticity of citrus hybrid offspring or in breeding citrus hybrid offspring, characterized in that: The kit includes combinations of nine primer pairs, each yielding nine sequences containing SNP sites; the SNP sites include chr1:16715056, chr2:20043474, chr3:31884372, chr4:6258181, chr5:46639852, chr6:15804801, chr7:21054402, chr8:2865714, and chr9:10463074. SNP site chr1:16715056 is located at base 296 of sequence xl-1, and the polymorphic site is C or T. The nucleotide sequence of sequence xl-1 is shown in SEQ ID NO:

1. SNP site chr2:20043474 is located at base 312 of sequence xl-2, and the polymorphic site is C or G. The nucleotide sequence of sequence xl-2 is shown in SEQ ID NO:

2. The SNP site chr3:31884372 is located at base 264 of sequence xl-3, and the polymorphic site is either G or C. The nucleotide sequence of sequence xl-3 is shown in SEQ ID NO:

3. The SNP site chr4:6258181 is located at base 295 of sequence xl-4, and the polymorphic site is either G or A. The nucleotide sequence of sequence xl-4 is shown in SEQ ID NO:

4. SNP site chr5:46639852 is located at base 272 of sequence xl-5, and the polymorphic site is C or G. The nucleotide sequence of sequence xl-5 is shown in SEQ ID NO:

5. The SNP site chr6:15804801 is located at base 385 of sequence xl-6, and the polymorphic site is C or G. The nucleotide sequence of sequence xl-6 is shown in SEQ ID NO:

6. The SNP site chr7:21054402 is located at base 281 of sequence xl-7, and the polymorphic site is C or T. The nucleotide sequence of sequence xl-7 is shown in SEQ ID NO:

7. SNP site chr8:2865714 is located at base 314 of sequence xl-8, and the polymorphic site is A or G. The nucleotide sequence of sequence xl-8 is shown in SEQ ID NO:

8. SNP site chr9:10463074 is located at base 299 of sequence xl-9, and the polymorphic site is A or T. The nucleotide sequence of sequence xl-9 is shown in SEQ ID NO:

9. The nucleotide sequence of the primer pair containing the SNP site chr1:16715056 xl-1 is as follows: F-1: AGGCGTAACATTTACCTCCA R-1:AAAGAGAAATTTTGTTGCAAAAAAAATTGT, The nucleotide sequences of the primer pair containing the SNP site chr2:20043474 xl-2 are as follows: F-2: GTAAAAAATATAAACAAAACCGCAAGATGAGG, R-2: TTACCCCTTCGACCCTTTTT The nucleotide sequences of the primer pair xl-3 containing the SNP site chr3:31884372 are as follows: F-3: TCGAAAACTTGGTGTTTAAATAGATTTGC, R-3: TATATATTGTAGGATCTGAATCATTAATCATCTTG, The nucleotide sequences of the primer pair xl-4 containing the SNP site chr4:6258181 are as follows: F-4: GCACAGGGAAAGAAAGAGGAA, R-4: GCTACTCAGCTAATTGATTTGTGG, The nucleotide sequences of the primer pair xl-5 containing the SNP site chr5:46639852 are as follows: F-5: ATCTTGATTTGCCACGTGTC, R-5: TCAGAACCTGAGACAAAACTAATG, The nucleotide sequences of the primer pair containing the SNP site chr6:15804801 xl-6 are as follows: F-6: AGATCTGATCTTACTTTTTTTTTATTTTTTCCC, R-6: TAATTTTTGCTATTCCCCAAAAACATTC, The nucleotide sequences of the primer pair containing the SNP site chr7:21054402 xl-7 are as follows: F-7: AAGGTCAGATGGAGCAACAC R-7:AGTTGCAATTTCAACTTAAGGGAA, The nucleotide sequences of the primer pair xl-8 containing the SNP site chr8:2865714 are as follows: F-8: TGCTGACGATAGCTCTAAAACAG, R-8:ATTATAGGAATTATTTAAGCTCTCAAAAGTTTTA, The nucleotide sequences of the primer pair containing the SNP site chr9:10463074 xl-9 are as follows: F-9: TCCCAAGGAAATGATCTCAACT, R-9: GAGAACTCCCGTAATTCGAAAGAAA.

Citation Information

Patent Citations

  • 20K liquid phase chip for citrus genotype identification and application of 20K liquid phase chip

    CN117305503A

  • Primer and method for identifying authenticity of distant hybridization offspring of clausena lansium and citrus

    CN120624696A