Method for identifying meiotic recombination in citrus f1 hybrid population based on next generation sequencing data

By using a method based on second-generation sequencing data, SNP sites in the F1 generation of citrus hybrid populations were screened and analyzed to accurately identify meiotic recombination, solving the difficulties in recombination analysis in citrus breeding and achieving efficient and low-cost recombination identification and accelerating the breeding process.

CN121931230BActive Publication Date: 2026-07-07HUAZHONG AGRI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAZHONG AGRI UNIV
Filing Date
2026-03-27
Publication Date
2026-07-07

Smart Images

  • Figure CN121931230B_ABST
    Figure CN121931230B_ABST
Patent Text Reader

Abstract

The application discloses a method for identifying meiotic recombination of citrus F1 hybrid population based on second-generation sequencing data. The parent and offspring are sequenced; the differential SNP sites which are homozygous in one parent and heterozygous in the other parent are screened; the paternal and maternal origin recombination differential SNP sites are distinguished; the control parent is combined to deduce the allele genotype of the offspring from the identification parent and to form the offspring chromosome haplotype; the offspring chromosome haplotype is clustered; the identification parent fitting chromosome haplotype is constructed according to the majority principle; and the identification parent fitting chromosome haplotype is compared with the offspring chromosome haplotype to determine the recombination event (when the father is analyzed, the father is the identification parent and the mother is the control parent; the mother is the same). The method breaks through the difficulty that the existing recombination analysis technology cannot identify the citrus hybrid population with less generation number, high parent heterozygosity and unknown haplotype, and makes it possible to identify the recombination of the perennial woody fruit tree population mainly with citrus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of genetic engineering technology, specifically to a method for identifying meiotic recombination in citrus F1 hybrid populations based on second-generation sequencing data. Background Technology

[0002] Meiotic recombination, a fundamental biological process in eukaryotes, begins with DNA double-strand breaks (DSBs) and is a crucial source of genetic diversity. Accurate identification of meiotic recombination events not only helps elucidate the molecular regulatory mechanisms of chromosome pairing and exchange but also enables the creation of high-resolution recombination maps, deepening our understanding of genomic dynamics and revealing how species adapt to environmental selection pressures by generating new genetic variations.

[0003] Before the large-scale application of modern biotechnology, citrus breeding in China was long characterized by "passive waiting, low efficiency, and weak predictability." Seedling selection and bud mutation selection methods relied heavily on chance, waiting for rare, superior variations to appear in nature. Although bud mutation selection successfully created a series of well-known varieties such as Newhall and Late Orange, and still has application value in specific historical stages and current scenarios, its inherent biological and technological limitations can no longer meet the core demands of modern industry for breeding efficiency and precision. With the establishment of modern horticultural science, conventional breeding centered on artificial hybridization has become mainstream, and design breeding is also rapidly advancing. In breeding practice, recombination is key to breaking the chain of positive and negative traits. By accurately identifying recombination events, it is possible to effectively predict and screen offspring individuals that have undergone beneficial recombination, thereby overcoming the chain of negative traits and accelerating the breeding process. Furthermore, in-depth analysis of recombination patterns helps to elucidate the genetic basis of heterosis, providing theoretical guidance for the selection of citrus parents and promoting the breeding of superior hybrid varieties.

[0004] In plant research, studies on meiotic recombination primarily focus on species such as Arabidopsis thaliana, rice, and maize. Compared to citrus, these species have advantages such as shorter generation cycles, larger population sizes, and clearer parental genetic backgrounds. The presence of numerous heterozygous sites in the citrus genome makes determining recombination breakpoints in offspring extremely complex, limiting the progress of related research.

[0005] Currently, methods for identifying recombination events can be mainly divided into two categories: one is sequencing analysis based on single pollen cells, and the other is resequencing analysis based on hybrid populations. Single pollen cell sequencing technology directly sequences sperm cells, directly obtaining haplotype information after meiosis, thus accurately determining the location of recombination events. This method has been successfully applied in maize and tea, but because it does not involve real progeny materials, it is more suitable for theoretical research on recombination patterns, and its application in practical breeding is relatively limited. On the other hand, resequencing analysis based on hybrid populations has been widely used in pure-line species, mainly including methods based on haplotype breakpoint analysis and methods based on linkage disequilibrium inference in population genetics. The linkage disequilibrium method does not rely on a pre-designed hybrid population, but rather indirectly infers historical recombination rates using existing genetic variation patterns in natural populations, which is obviously not suitable for recombination analysis in specific citrus hybrid populations. The haplotype breakpoint method is the most classic and intuitive research approach, based on the principle that recombination breaks the continuity of parental haplotypes. However, the high heterozygosity of citrus parents becomes the main challenge in the practical application of this method. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a method for identifying meiotic recombination in citrus F1 hybrid populations based on second-generation sequencing data. This method overcomes the difficulty of existing recombination analysis techniques in identifying citrus hybrid populations with few generations, high parental heterozygosity, and unknown haplotypes, making recombination identification possible for perennial woody fruit tree populations dominated by citrus.

[0007] To achieve the above objectives, the technical solution designed by the present invention is as follows:

[0008] This invention provides a method for identifying meiotic recombination in citrus F1 hybrid populations based on second-generation sequencing data, comprising the following steps:

[0009] S1. Extract genomic DNA from the maternal and paternal parents and perform genome sequencing to obtain their whole genome sequence information; perform low-depth genome sequencing on each offspring of the F1 generation hybrid population to obtain sequencing data for each offspring.

[0010] S2. Align the sequencing data of the maternal parent and the paternal parent to the reference genome, and perform SNP variation detection. Select SNP sites that are homozygous in one parent and heterozygous in the other parent as differential SNP sites. Multiple differential SNP sites form a differential SNP site set. S3. Select differential SNP sites that are heterozygous in the maternal parent and homozygous in the paternal parent as differential SNP sites for maternal recombination, and form a differential SNP site set for maternal recombination.

[0011] Meanwhile, the differential SNP loci that are heterozygous in the father and homozygous in the mother will be used as differential SNP loci for paternal recombination, forming a set of differential SNP loci for paternal recombination.

[0012] S4. Using chromosomes as units and homozygous paternal genotypes as background, based on the genotypes of each offspring at each maternal recombination-identifying SNP locus, infer the alleles obtained by each offspring from the mother at each maternal recombination-identifying SNP locus, and arrange them according to the physical location of chromosomes to form the offspring maternal chromosome haplotype.

[0013] In the same step, based on the homozygous maternal genotype, the paternal chromosome monotype of each offspring is formed according to the genotype of each offspring at each paternal recombination-identifying SNP locus.

[0014] S5. Using chromosomes as the unit, cluster the haplotypes of maternal chromosomes in all offspring into two maternal groups. One maternal group corresponds to one homologous chromosome of the maternal chromosome, and the other maternal group corresponds to the other homologous chromosome of the maternal chromosome. In the same step, cluster the haplotypes of paternal chromosomes in offspring into two paternal groups.

[0015] S6. Using chromosomes as units, based on the majority principle, the proportion of maternal alleles in each offspring of a maternal parent group in all offspring of that maternal parent group is used to infer the alleles at the maternal recombination-identifying SNP site of the maternal parent. The maternal parent fit chromosome haplotype is formed according to the order of the physical position of the chromosomes. The information of each offspring used to infer the maternal parent fit chromosome haplotype is regarded as the offspring chromosome haplotype of each offspring. Another maternal parent fit chromosome haplotype is obtained using another maternal parent group.

[0016] The two paternal chromosome haplotypes were obtained in the same steps.

[0017] S7. Using chromosomes as units, compare the two maternal chromosome haplotypes with their corresponding offspring chromosome haplotypes. If the alleles are consistent, mark it as True; if they are inconsistent, mark it as False. The transition between True and False indicates that the offspring have undergone meiotic recombination between the maternal recombination differential SNP sites.

[0018] In the same step, the two paternal chromosome haplotypes are compared with their corresponding offspring chromosome haplotypes. The True and False transitions are determined to indicate that the offspring have undergone meiotic recombination between the paternal recombination-identifying SNP sites.

[0019] Furthermore, in step S1, the sequencing depth of resequencing is at least 30 ×; the sequencing depth of low-depth sequencing is 2~10 ×.

[0020] The paternal parent is either a mandarin orange or a purple pomelo, and the maternal parent is either Ehime 30 or Guanxi honey pomelo.

[0021] Furthermore, in step S2, the reference genome is the Wenzhou mandarin orange genome or the Pingshan pomelo T2T genome;

[0022] The specific screening and identification SNP sites include: the SNP site is a bisected SNP site; the SNP site is not a deleted site in either the maternal or paternal parent; the sequencing depth of the SNP site is 10~60×, and the multiple bisected SNP sites selected constitute a bisected SNP site set.

[0023] Furthermore, step S2 also includes: using the selected biene SNP locus set to fill the genotypes of the sequencing data of all offspring, selecting SNP loci with a deletion rate of no more than 20% in the offspring, and obtaining multiple high-quality SNP loci to form a high-quality SNP locus set.

[0024] Furthermore, step S2 also includes verifying the Mendelian law of segregation for the differential SNP sites: using the chi-square goodness-of-fit test to evaluate whether the genotype distribution of the offspring of high-quality SNP sites meets the expected segregation ratio of 1:1, and retaining SNP sites with a p-value greater than 0.05.

[0025] Furthermore, in step S5, the clustering adopts the K-means clustering method, which forcibly divides the chromosomes into two groups, corresponding to the two homologous chromosomes of the maternal / paternal chromosomes respectively.

[0026] Furthermore, in step S6, most of the principles are as follows: when the proportion of maternal alleles in all offspring of the maternal parent group is greater than 50%, it is inferred that the maternal allele of the offspring is the allele at the maternal recombination-identifying SNP site of the maternal parent.

[0027] Furthermore, between steps S6 and S7, there is also a step of correcting the haplotype of the fitted chromosomes from the maternal / paternal parents:

[0028] Sa1. If the alleles at the corresponding differential SNP loci in the two maternal / paternal fitted chromosome haplotypes of the same chromosome are inconsistent, then the alleles of the two maternal / paternal fitted chromosome haplotypes are both correct.

[0029] Sa2. If one of the maternal / paternal matching chromosomes has a missing allele, it is supplemented by complementation based on the other maternal / paternal matching chromosome's allele.

[0030] Sa3. If the corresponding alleles in the two maternal / paternal fitted chromosomes are both missing, then delete the differential SNP site.

[0031] Sa4. For consecutive small fragments or isolated differential SNPs whose number of differential SNPs does not exceed 2% of the total number of differential SNPs on the corresponding chromosome, if the alleles of the two maternal / paternal fitted chromosomes are the same, then the differential SNPs shall be deleted.

[0032] Sa5. For a continuous long segment with the number of differential SNP sites exceeding 2% of the total number of differential SNP sites on the corresponding chromosome, if the alleles of the two maternal / paternal fitted chromosome haplotypes are the same, the alleles of the long segment in the maternal / paternal fitted chromosome haplotype should be complemented and corrected by the alleles corresponding to the haplotype of the other maternal / paternal fitted chromosome.

[0033] Furthermore, after step S7, the method further includes: clustering continuous SNP sites with the same judgment type into Bin blocks based on a set window, using the True and False transitions of the Bin marker as meiotic recombination events, and filtering out double crossover events in the offspring population that occur at a rate higher than a set threshold.

[0034] The window size is 200~400kb; the set threshold is 4%.

[0035] The present invention also provides an application of the aforementioned identification method in identifying meiotic recombination events in F1 generation hybrid populations of citrus or in breeding new citrus varieties.

[0036] The beneficial effects of this invention are:

[0037] The identification method of this invention is used to infer recombination events on each chromosome of the entire F1 generation of citrus hybrids. After identifying the discriminant SNP loci used to identify paternal and maternal recombination, the parental alleles of these loci in the offspring are inferred. The haplotypes of the paternal (maternal) chromosomes in the offspring are clustered, and their corresponding paternal (maternal) haplotypes are fitted according to the majority rule. After correction, the offspring haplotypes are compared with the paternal (maternal) haplotypes, and high double-crossover recombination rates are filtered out using Bin markers to reduce false-positive recombination events, thus completing the identification of recombination events on each chromosome of the entire offspring population. This method does not require high sequencing depth of the offspring data and does not require third-generation high-throughput sequencing to obtain parental haplotypes. It is low-cost, simple to operate, and can quickly identify recombination events on each chromosome of all offspring in the population. Attached Figure Description

[0038] Figure 1 The flowchart shows the identification method for meiotic recombination in citrus F1 hybrid populations based on second-generation sequencing data.

[0039] Figure 2Heatmap showing the effect of CHR9 test on the offspring haplotypes compared with the parents before correction for maternal parent AY30;

[0040] Figure 3 Heatmap showing the effect of CHR9 test and correction on the haplotypes of offspring compared with those of the parents for maternal parent AY30;

[0041] Figure 4 This is a landscape map of the ST group reorganization.

[0042] In the figure, the yellow-green recombination heatmap is the heatmap of chromosomes derived from the maternal parent AY30, and the pink-blue recombination heatmap is the heatmap of chromosomes derived from the paternal parent STJ. Different colors represent the four haplotypes of the parents, and the position of the yellow-green color change or the position of the pink-blue color change represents the occurrence of recombination events.

[0043] Figure 5 Landscape atlas of ZPGX population reorganization;

[0044] In the figure, the yellow-green recombination heatmap represents the maternal GXMY chromosome heatmap, and the pink-blue recombination heatmap represents the paternal ZPY chromosome heatmap. Different colors represent the four haplotypes of the parents.

[0045] Figure 6 This is a diagram showing the multiple sequence alignment results of the recombination site regions of the parents and offspring in Example 2;

[0046] Figure 7 This is a schematic diagram of the recombination mode of the recombinant single plant XL293 in Example 2. Detailed Implementation

[0047] The present invention will now be described in further detail with reference to specific embodiments, so that those skilled in the art can understand it.

[0048] Example 1

[0049] Detection of the quantity and distribution of meiotic recombination in the offspring of the AY30×STJ F1 population

[0050] In this embodiment, the maternal parent was Ehime 30 (AY30), and the paternal parent was Satsuma mandarin orange (STJ), resulting in 228 STJAY30 offspring. Combined with... Figure 1 As shown, the specific process is as follows:

[0051] 1. Genomic DNA was extracted from the maternal parent AY30, the paternal parent STJ, and 228 STJAY30 progeny. The genomic DNA of the maternal parent AY30 and the paternal parent STJ was resequencing, with a sequencing depth of at least 30 × (the average sequencing depth in this example was 33.1 ×). Sequencing data of the maternal parent AY30 and the paternal parent STJ were obtained. The genomic DNA of the 228 STJAY30 progeny were subjected to Illumina low-depth sequencing, with a sequencing depth of 2.07~8.42 × (the average sequencing depth in this example was 4.63 ×). Sequencing data of the 228 STJAY30 progeny were obtained.

[0052] 2. Analyze the sequencing data to obtain a set of discriminative SNP sites:

[0053] (1) The sequencing data of the maternal parent AY30 and the paternal parent STJ were aligned to the genome of Wenzhou mandarin orange (Citrus Pan-genome2breeding Database) using BWA software. Citrus reticulata 'Unshiu'), and performed SNP mutation detection (SNP calling) using GATK 4.0, using the default GATK filtering parameters (QD<2.0 || MQ<40.0 || FS>60.0 || SOR>3.0 || MQRankSum<-12.5 || ReadPosRankSum<-8.0).

[0054] (2) Use vcftools to filter according to the following criteria (--minQ 30 --remove-indels --max-alleles 2 --min-alleles 2 --max-missing 1 --min-meanDP 10 --max-meanDP 60):

[0055] ① Retain the biselequential SNP locus (a SNP (single nucleotide polymorphism) with only two allele forms in the population, i.e., (AA, Aa, aa)); ② The biselequential SNP locus is not a deletion locus in either the maternal or paternal parent; ③ The minimum sequencing depth of the biselequential SNP locus is not less than 10, and the maximum sequencing depth is not more than 60; ④ The biselequential SNP locus is homozygous in one parental genotype and heterozygous in the other parental genotype.

[0056] After screening according to the above criteria, the parent VCF format file and the parent VCF format file are obtained respectively, and then mixed to create a parent reference panel.

[0057] (3) The sequencing data of 228 STJAY30 progeny were aligned to the Wenzhou mandarin orange genome to obtain the bam files of 228 STJAY30 progeny. Using STITCH software, the genotypes were filled using the parental reference panel and the bam files of 228 STJAY30 progeny (default parameters, K=4, nGen=100) to obtain a population vcf file containing the genotypes of the maternal parent, paternal parent and 228 progeny loci. Using vcftools, multiple high-quality SNP loci were selected according to the following criteria to form a high-quality SNP locus set: the deletion rate of the biselary SNP loci after filtering in step (2) in the STJAY30 progeny was not higher than 20%.

[0058] (4) Ensure that the distribution of SNP loci genotypes in the population conforms to Mendel's law of segregation. For each SNP locus that is homozygous in one parent and heterozygous in another parent, the chi-square goodness-of-fit test is used to evaluate whether the genotype distribution of the offspring of the locus conforms to Mendel's expected segregation ratio (1:1). All SNP loci with a test p value > 0.05, i.e., loci that do not deviate significantly from the expected segregation ratio, are retained to ensure that subsequent analysis is based on conforming to Mendel's laws of inheritance.

[0059] 3. Identify differential SNPs used to determine chromosomal recombination events in offspring. Differential SNPs that are heterozygous in the mother and homozygous in the father are considered maternally recombinant differential SNPs; similarly, those that are heterozygous in the father and homozygous in the mother are considered paternally recombinant differential SNPs. For example, at a differential SNP locus, if the father's genotype is Aa and the mother's genotype is aa, then this SNP locus is a paternally recombinant differential SNP locus; or, if the father's genotype is aa and the mother's genotype is Aa, then this SNP locus is a maternally recombinant differential SNP locus. These two types of loci are used to determine whether the offspring's chromosomes are recombined from the mother and father, respectively. The identifying SNP loci are shown in Table 1. A total of 1,688,954 maternal recombination identifying SNP loci were screened on chromosomes 1 to 9 of citrus, forming the maternal recombination identifying SNP loci set on chromosome 1 to chromosome 9, respectively. A total of 525,552 paternal recombination identifying SNP loci were screened, forming the paternal recombination identifying SNP loci set on chromosome 1 to chromosome 9, respectively.

[0060] Table 1. Statistical table of identifiable SNP sites for AY30×STJ recombination.

[0061]

[0062] 4. Taking the set of maternal recombination-identifying SNP loci on chromosome 9 as an example, for a certain maternal recombination-identifying SNP locus, since the genotype of this SNP locus is heterozygous in the mother and homozygous in the father, and the genotypes of each offspring are known, the alleles obtained by each offspring from the father are known. Based on the genotypes of each offspring, the alleles obtained by each offspring from the mother can be inferred. For example, for a certain maternal recombination-identifying SNP locus, if the father's genotype is AA and the mother's genotype is Aa, if the offspring's genotype is AA, then the allele obtained by that offspring from the mother is A; if the offspring's genotype is Aa, then the allele obtained by that offspring from the mother is a. As another example, for a certain maternal recombination-identifying SNP locus, if the father's genotype is aa and the mother's genotype is Aa, if the offspring's genotype is aa, then the allele obtained by that offspring from the mother is a; if the offspring's genotype is Aa, then the allele obtained by that offspring from the mother is A.

[0063] Based on the physical location of the maternal recombination-identifying SNP loci on chromosome 9, the alleles obtained from the mother by each offspring are arranged in physical order to form the maternal haplotype of chromosome 9 for that offspring. This represents the maternal-derived alleles on each chromosome, composed of the A and a alleles (if the offspring genotype at a certain SNP locus is missing, the maternal allele at that SNP locus is recorded as missing). The same steps are used to obtain the maternal haplotypes of chromosome 1 through chromosome 8 for each offspring. In this example, a total of 228 offspring obtained maternal haplotypes of chromosome 1 through chromosome 9.

[0064] Using the paternal recombination identification SNP locus set of chromosome 1 to chromosome 9, the paternal chromosome haplotypes of 228 offspring were obtained in the same steps.

[0065] 5. Using chromosome 9 of the maternal parent as the physical unit, k-means clustering was performed on the maternal recombination identification SNP loci of chromosome 9 arranged sequentially on chromosome 9 of the 228 offspring chromosome 9 offspring maternal chromosome haplotypes, and they were forcibly divided into two maternal groups (the number of homologous chromosomes in diploid citrus is two, and regardless of whether recombination occurs or how it occurs, each offspring chromosome haplotype corresponds to one of the parental homologous chromosome haplotypes), that is, one maternal group corresponds to one homologous chromosome of the maternal parent chromosome 9, and the other maternal group corresponds to the other homologous chromosome of the maternal parent chromosome 9; in the same step, using chromosomes 1 to 8 of the maternal parent as the physical unit, k-means clustering was performed on the maternal chromosome haplotypes of chromosome 1 to chromosome 8 of the 228 offspring offspring to divide them into two maternal groups.

[0066] In the same step, using chromosomes 1 to 9 of the father as physical units, the paternal haplotypes of chromosome 1 to 9 of the 228 offspring were clustered to obtain two paternal groups corresponding to the two paternal chromatids respectively.

[0067] 6. Taking the maternal chromosome 9 as an example, according to the majority rule, the proportion of maternal alleles in each offspring of a maternal parent group among all offspring of that group is used to infer the allele at the corresponding maternal recombination-identifying SNP locus in that maternal parent. Alleles with a proportion > 50% are considered to be the alleles of their corresponding maternal parent; if the proportions are the same (= 50%), they are recorded as the allele missing at that SNP locus in the parent (for example, if one maternal parent group has 120 offspring, and at a certain maternal recombination-identifying SNP locus, 85 offspring have alleles missing). If the SNP genotype of the maternal allele is A, and the SNP genotype of the maternal allele in 35 offspring is a, then the SNP genotype of the allele at the corresponding SNP locus in the mother is A. For example, if one maternal parent group has 120 offspring, and at a certain maternal recombination-identifying SNP locus, the SNP genotype of the maternal allele in 60 offspring is A, and the SNP genotype of the maternal allele in the other 60 offspring is a, then the allele at the corresponding SNP locus in the mother is recorded as deleted. The alleles corresponding to each maternal recombination-identifying SNP locus in one homologous chromosome of the mother's chromosome 9 are deduced, and the maternal fitted chromosome haplotype is formed according to the physical position of the chromosome. Two maternal groups are used to infer the haplotypes of the two maternal fitted chromosomes. The allele information of each offspring used to infer the haplotype of the maternal fitted chromosome is considered as the haplotype of the offspring chromosome of each offspring.

[0068] In the same steps, the inference of two homologous chromosomes from chromosomes 1 to 8 of the maternal parent was completed, resulting in a total of 18 maternal parent-fitted chromosome haplotypes (citrus has 9 chromosomes, all of which are diploid, so there are 18 maternal parent-fitted chromosome haplotypes).

[0069] In the same steps, the inference of two chromatids in chromosomes 1 to 9 of the father was completed, resulting in a total of 18 paternal fitted chromosome haplotypes.

[0070] 7. Check and correct the haplotypes of the parental fitted chromosomes. Taking maternal chromosome 9 as an example, based on the heterozygous nature of the maternal recombination differential SNP loci in the maternal parent, the alleles at each differential SNP locus should be inconsistent in the two maternal fitted chromosome haplotypes of maternal chromosome 9. The corresponding alleles in the two maternal fitted chromosome haplotypes of maternal chromosome 9 should be compared one by one according to the following criteria:

[0071] (a) If the corresponding alleles in the two maternal fitted chromosome haplotypes are inconsistent, then at the maternal recombination identification SNP site, the alleles of the two maternal fitted chromosome haplotypes are correct.

[0072] (b) The allele of the haplotype missing in one of the maternal parent chromosomes is complemented by the allele of the haplotype missing in the other maternal parent chromosome. For example, at a certain maternal recombination identification SNP site, if the allele corresponding to the haplotype missing in one of the maternal parent chromosomes is missing, and the allele corresponding to the haplotype missing in the other maternal parent chromosome is a, then the missing allele is A.

[0073] (c) If the corresponding alleles in the two maternal fitted chromosome haplotypes are both missing, then delete the allele of the maternal recombination differential SNP site.

[0074] (d) For a continuous small fragment (the number of small fragments of differential SNP sites does not exceed 2% of the total number of differential SNP sites on the chromosome) and an isolated differential SNP site in a maternal fitted chromosome haplotype, if the alleles inferred from the two maternal fitted chromosome haplotypes are the same, then delete the allele of the maternal recombination differential SNP site.

[0075] (e) For a continuous long segment in a maternal fitted chromosome haplotype (the number of long segment SNP sites exceeds 2% of the total number of differential SNP sites on the chromosome), if the alleles inferred from the two maternal fitted chromosome haplotypes are the same, the alleles of the long segment in the maternal fitted chromosome haplotype should be complemented by the alleles corresponding to the haplotype of the other maternal fitted chromosome haplotype.

[0076] The same steps were used to check and correct the haplotypes of the 16 maternal chromosomes from chromosome 1 to chromosome 8.

[0077] The same steps were used to check and correct the haplotypes of the 18 paternal chromosomes (chromosomes 1-9).

[0078] 8. Taking the maternal chromosome 9 as an example, the two maternal fitted chromosome haplotypes are compared with the corresponding offspring chromosome haplotypes. Allele consistency is marked as "True" and inconsistency as "False". The ordered "True" and "False" at the chromosome level reflect the segment correspondence of the offspring chromosome haplotype from the maternal fitted chromosome haplotype. Theoretically, the transition between "True" and "False" is considered a recombination event. The physical distance between adjacent "True" and "False" sites at the transition is called the recombination interval, and the midpoint between two "True" and "False" sites is called the recombination breakpoint.

[0079] Visualizing the information in the form of a heatmap, using the maternal chromosome 9 as an example, the maternal allele information of the two progeny chromatids is plotted in heatmap form, such as... Figure 2 As shown, the vertical axis represents each offspring, and the horizontal axis represents each paternal recombination-identifying SNP locus. Red and blue represent the similarities and differences between the maternal fitted chromosome haplotype and the corresponding offspring chromosome haplotype (True (red) for similarities, False (blue) for differences). The transition between red and blue colors indicates a potential recombination event. The pink and gray heatmap at the bottom shows the comparison of the two offspring chromosome haplotypes on chromosome 9 for each maternal recombination-identifying SNP locus; pink indicates inconsistency in the predicted alleles, and gray indicates consistency. Figure 2 As can be seen above, the maternal recombination event is clearly abnormal. Correction is performed based on the complementation of the alleles corresponding to the other maternal group. Furthermore, since the two haplotypes of the maternal chromosomes are complementary, the reference object is set as one of the haplotypes of the chromosomes, and a new heatmap is drawn after comparison, as shown below. Figure 3 As shown.

[0080] The same steps were used to compare the maternal haplotypes of chromosomes 1 through 8.

[0081] The same steps were used to compare the paternal haplotypes of chromosomes 1 through 9.

[0082] 9. To reduce the impact of misclassification and deletion of individual SNP sites on recombination identification and to reduce the computational burden of redundant SNPs of the same type, this step clusters consecutive SNP sites with the same classification into Bin blocks based on a 250kb window. This step is performed using binmarker2 software (perl binmarkers-v2.3.pl file -m bin -w 250_000 |perl binmarkers-v2.3.pl -m fill -w 3 | binmarkers-v2.3.pl -m fill -w 5 |perlbinmarkers-v2.3.pl -m fill -w 7 |perl binmarkers-v2.3.pl -m fill2 -w 3 |perlbinmarkers-v2.3.pl -m correct -w 5 |perl binmarkers-v2.3.pl -m merge -ofile.bin.txt; the first -w parameter needs to be modified to define the window size, and the subsequent parameters are the default parameters). In the entire recombination inference of the STJ×AY30 F1 population, the Binmarker generated a total of 1325 bin markers. The average bin block length was 363.2kb, and each bin block had an average of 1579.9 SNPs. Each block was assigned a uniform "True" or "False" bin marker, and the transition between "True" and "False" bin markers was considered a recombination.

[0083] Recombination events with high double crossover rates are filtered out. Theoretically, the probability of double crossover occurring in the population is low. Therefore, double crossover events with an occurrence rate higher than 4% in the population are filtered out.

[0084] Recombination analysis was performed on each chromosome of both the maternal and paternal parents, and population recombination identification was completed. 3628 recombination events were identified in 228 F1 individuals, of which high-resolution recombination events (recombination interval less than 2kb) accounted for 58.1%, while fuzzy recombination events (recombination interval greater than 100kb) accounted for only 6.1%. Recombination maps were plotted, and the results are as follows: Figure 4 As shown.

[0085] Example 2

[0086] Detection of the quantity and distribution of meiotic recombination in the offspring of the F1 population of Guanxi honey pomelo × purple pomelo

[0087] I. The maternal parent used in this example was Guanxi Honey Pomelo GXMY, and the paternal parent was Purple Skin Pomelo ZPY, resulting in 78 offspring. The steps in Example 1 were used for testing. The genome compared was the Pingshan Pomelo T2T genome (Citrus Pan-genome2breeding Database genome). Citrus maxima 'Pingshan' (T2T genome)).

[0088] The identified differential SNP loci are shown in Table 2. A total of 848,273 maternal recombination differential SNP loci were screened on chromosomes 1 to 9 of citrus, forming a set of maternal recombination differential SNP loci. A total of 582,101 paternal recombination differential SNP loci were screened, forming a set of paternal recombination differential SNP loci.

[0089] Table 2. Statistical table of identifying SNP sites in GXMY×ZPY recombination.

[0090]

[0091] Recombination analysis was performed on each chromosome of both the maternal and paternal parents, and population recombination was identified. A total of 1051 recombination events were identified, of which high-resolution recombination events (recombination interval less than 2kb) accounted for 54.1%, and fuzzy recombination events (recombination interval greater than 100kb) accounted for only 11.5%. Recombination maps were plotted, and the results are as follows: Figure 5 As shown.

[0092] II. Sites were selected from the recombination identification results for experimental verification. SNP sites with recombination intervals less than 700 bp were selected for Sanger sequencing. Using DNA from the parental Guanxi pomelo and purple-skinned pomelo, the hybrid progeny XL293 and the control hybrid progeny XL121 and XL126 as templates, PCR amplification and alignment yielded the sequence information of the Pingshan pomelo T2T genome chr9:43480300-43480919. Single-clone strains were selected using an E. coli transformation system to obtain the two chromatid sequences at this location in the parental and hybrid progeny. Multiple sequence alignment results are shown below. Figure 6 and pattern results Figure 7 As shown, the results demonstrate that the progeny XL293 underwent one recombination event in the chromatids derived from the Guanxi pomelo parent within the range of chr9:43480453-43480687 (marker SNPs underwent an arrangement change on the chromosome, such as...). Figure 7 As shown in the figure, no recombination occurred in progeny XL121 and progeny XL126, and the fragments of progeny XL121 and progeny XL126 were derived from homologous fragments of the Guanxi pomelo parent. This demonstrates that the method of the present invention is reliable for studying recombination events in citrus F1 hybrid populations.

[0093] All other parts not described in detail are existing technologies. Although the above embodiments have provided a detailed description of the present invention, they are only some embodiments of the present invention, not all embodiments. People can obtain other embodiments based on these embodiments without creative effort, and these embodiments all fall within the protection scope of the present invention.

Claims

1. A method for identifying meiotic recombination in F1 generation hybrid populations of citrus based on second-generation sequencing data, characterized in that: Includes the following steps: S1. Extract genomic DNA from the maternal and paternal parents and perform genome sequencing on each parent to obtain their whole genome sequence information; Low-depth genome sequencing was performed on each offspring of the F1 generation hybrid population to obtain sequencing data for each offspring. S2. Align the sequencing data of the maternal and paternal parents to the reference genome and perform SNP variation detection. Select SNP sites that are homozygous in one parent and heterozygous in the other parent as differential SNP sites. Multiple differential SNP sites form a differential SNP site set. S3. The differential SNP loci that are heterozygous in the maternal parent and homozygous in the paternal parent are used as differential SNP loci for maternal recombination, forming a set of differential SNP loci for maternal recombination. Meanwhile, the differential SNP loci that are heterozygous in the father and homozygous in the mother will be used as differential SNP loci for paternal recombination, forming a set of differential SNP loci for paternal recombination. S4. Using chromosomes as units and homozygous paternal genotypes as background, based on the genotypes of each offspring at each maternal recombination-identifying SNP locus, infer the alleles obtained by each offspring from the mother at each maternal recombination-identifying SNP locus, and arrange them according to the physical location of chromosomes to form the offspring maternal chromosome haplotype. In the same step, based on the homozygous maternal genotype, the paternal chromosome monotype of each offspring is formed according to the genotype of each offspring at each paternal recombination-identifying SNP locus. S5. Using chromosomes as the unit, cluster the haplotypes of maternal chromosomes in all offspring, dividing them into two maternal groups. One maternal group corresponds to one homologous chromosome of the maternal chromosome, and the other maternal group corresponds to the other homologous chromosome of the maternal chromosome. In the same step, cluster the haplotypes of paternal chromosomes in offspring, dividing them into two paternal groups. S6. Using chromosomes as units, based on the majority principle, the proportion of maternal alleles in each offspring of a maternal parent group in all offspring of that maternal parent group is used to infer the alleles at the maternal recombination-identifying SNP site of the maternal parent. The maternal parent fit chromosome haplotype is formed according to the order of the physical position of the chromosomes. The information of each offspring used to infer the maternal parent fit chromosome haplotype is regarded as the offspring chromosome haplotype of each offspring. Another maternal parent fit chromosome haplotype is obtained using another maternal parent group. The two paternal chromosome haplotypes were obtained in the same steps. S7. Using chromosomes as units, compare the two maternal chromosome haplotypes with their corresponding offspring chromosome haplotypes. If the alleles are consistent, mark it as True; if they are inconsistent, mark it as False. The transition between True and False indicates that the offspring have undergone meiotic recombination between the maternal recombination differential SNP sites. In the same step, the two paternal chromosome haplotypes are compared with their corresponding offspring chromosome haplotypes. The True and False transitions are determined to indicate that the offspring have undergone meiotic recombination between the paternal recombination-identifying SNP sites. In step S1, the sequencing depth of resequencing is at least 30 ×; the sequencing depth of low-depth sequencing is 2~10 ×. The paternal parent is a sugar orange or a purple pomelo, and the maternal parent is Ehime 30 or Guanxi honey pomelo; In step S5, the clustering adopts the K-means clustering method, which forcibly divides the chromosomes into two groups, corresponding to the two homologous chromosomes of the mother / father chromosomes respectively. In step S6, most of the principles are as follows: when the proportion of maternal alleles in all offspring of the maternal parent group is greater than 50%, it is inferred that the maternal allele of the offspring is the allele at the maternal recombination-identifying SNP site of the maternal parent.

2. The identification method according to claim 1, characterized in that: In step S2, the reference genome is the Wenzhou mandarin orange genome or the Pingshan pomelo T2T genome; The specific screening and identification SNP sites include: the SNP site is a bisected SNP site; the SNP site is not a deleted site in either the maternal or paternal parent; the sequencing depth of the SNP site is 10~60×, and the multiple bisected SNP sites selected constitute a bisected SNP site set.

3. The identification method according to claim 2, characterized in that: Step S2 further includes: using the selected biene SNP locus set to fill the genotypes of the sequencing data of all offspring, screening SNP loci with a deletion rate of no more than 20% in the offspring, and obtaining multiple high-quality SNP loci to form a high-quality SNP locus set.

4. The identification method according to claim 3, characterized in that: Step S2 further includes verifying Mendelian segregation law for the differential SNP sites: using the chi-square goodness-of-fit test to evaluate whether the genotype distribution of the offspring of high-quality SNP sites meets the expected segregation ratio of 1:1, and retaining SNP sites with a test p value greater than 0.

05.

5. The identification method according to claim 1, characterized in that: Between steps S6 and S7, there is also a step of correcting the maternal / paternal fitted chromosome haplotype: Sa1. If the alleles at the corresponding differential SNP loci in the two maternal / paternal fitted chromosome haplotypes of the same chromosome are inconsistent, then the alleles of the two maternal / paternal fitted chromosome haplotypes are both correct. Sa2. If one of the maternal / paternal matching chromosomes has a missing allele, it is supplemented by complementation based on the other maternal / paternal matching chromosome's allele. Sa3. If the corresponding alleles in the two maternal / paternal fitted chromosomes are both missing, then delete the differential SNP site. Sa4. For consecutive small fragments or isolated differential SNPs whose number of differential SNPs does not exceed 2% of the total number of differential SNPs on the corresponding chromosome, if the alleles of the two maternal / paternal fitted chromosomes are the same, then the differential SNPs shall be deleted. Sa5. For a continuous long segment with the number of differential SNP sites exceeding 2% of the total number of differential SNP sites on the corresponding chromosome, if the alleles of the two maternal / paternal fitted chromosome haplotypes are the same, the alleles of the long segment in the maternal / paternal fitted chromosome haplotype should be complemented and corrected by the alleles corresponding to the haplotype of the other maternal / paternal fitted chromosome.

6. The identification method according to claim 1, characterized in that: The step S7 is followed by: clustering continuous SNP sites of the same type into Bin blocks based on a set window, taking the True and False transitions of the Bin marker as meiotic recombination events, and filtering out double crossover events in the offspring population that occur at a rate higher than a set threshold. The window size is 200~400kb; the set threshold is 4%.

7. The application of the identification method according to any one of claims 1 to 6 in identifying meiotic recombination events in F1 generation hybrid populations of citrus or in breeding new citrus varieties.