A method for detecting meiotic recombination events in a polyploid yeast genome

CN117153254BActive Publication Date: 2026-09-29NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310952380.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-31
Publication Date
2026-09-29
Estimated Expiration
2043-07-31

AI Technical Summary

Benefits of technology

[0032](1)基于全基因组测序提供的可靠的所有变异位点信息,有效地避免了多倍体中杂合突变reads波动性所导致的基因型鉴定难题,提供了高效、准确的重组事件检测方法,适用于多倍体酵母及其它以两种交配型亲本进行有性生殖的多倍体物种。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117153254B_ABST
    Figure CN117153254B_ABST
Patent Text Reader

Abstract

The application discloses a method for detecting meiotic recombination events in a polyploid yeast genome, comprising the following steps: (1) resequencing data evaluation; (2) determination of reliable screening markers between two parents; (3) further determination of heterozygous markers in the polyploid parent; (4) ploidy determination of the progeny sample; (5) genotype determination of the progeny sample; (6) construction of a progeny background panel; (7) confirmation of recombination events. The application can efficiently and accurately detect recombination events in a polyploid yeast.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for detecting meiotic recombination events in the genome of polyploid yeast, specifically involving the splicing and evaluation of second-generation sequencing data of polyploid yeast primitive genomic DNA, identification of screening markers between parents, identification of progeny ploidy and genotype, and confirmation of recombination events. Background Technology

[0002] Meiosis is the basis of sexual reproduction. Meiotic recombination is crucial for the precise segregation of chromosomes. It is usually initiated by programmed DNA double-strand breaks produced by SpO11 nucleases, and tends to produce recombination (CO) through the double-strand break repair (DSBR) pathway. The level of recombination is usually measured by the number and distribution of recombination events, which vary between species, between chromosomes within a species, and between different regions within a chromosome. The meiotic recombination rate is usually influenced by factors such as sex, gene or transposon density, GC nucleotide content, genomic sequence heterozygosity, chromosomal position of specific chromosomes and genes, genomic ploidy, chemical or temperature factors. The number of recombination events produced in each meiotic division is strictly regulated, including crossover interference, chromosome length, intrinsic characteristics of the recombination pathway, and population fitness. Recombination is not uniformly distributed on chromosomes, with obvious recombination hotspots and colds, and is influenced by DNA sequence characteristics, transcriptional activity, chromatin structure, nucleosome occupancy, transcription factors, and regulatory proteins.

[0003] Genomic polyploidy is widespread and a significant driving force behind species evolution and genome diversification. Based on whether the genomes originate from the same source, it can be divided into homologous polyploidy and allologous polyploidy. The sources of polyploid genomes may include chromosome doubling, fusion of unmeiotic gametes, and polyspermy. The generation of unmeiotic gametes is often related to the non-disjunction of homologous chromosomes during meiosis, heterozygosity between parental genomes, and external environmental factors. Currently, most research focuses on the mechanisms of polyploid formation, genetic breeding, genome structural evolution and its role in speciation and phenotypic trait formation, chromosome behavior and mechanisms during mitosis or meiosis, and the estimation of mutation rates and meiotic recombination rates.

[0004] Recombination in polyploids can occur through chromosomal rearrangement, genome invasion, nucleocytoplasmic interaction, or germplasm infiltration during meiosis. Regarding the meiotic recombination rate in polyploid species, some studies have constructed a pair of cis-linked fluorescent marker genes (GFP and RFP) on chromosome 3 of diploid Arabidopsis thaliana. Colchicine treatment yielded homotetraploids, and hybridization yielded allotetraploids. Genotypes and phenotypes of the seeds were statistically analyzed, and the recombination of the pair of genes was used to characterize the effect of polyploidization on the meiotic recombination rate. The results showed that compared to diploids, the recombination frequency was significantly increased after homo- or allotetraploidization, enhancing the rate of integration of beneficial genes and elimination of unfavorable genes. Furthermore, homologous chromosome pairing and the formation of multivalents were not the cause of the increased recombination rate. With the development of sequencing methods and genotyping techniques, genetic linkage mapping (four-point method) based on polymorphic molecular markers can also improve the estimation accuracy of recombination rates in polyploids (such as offspring from recombinant inbred lines of Arabidopsis thaliana) and provide information on the intensity of genetic interference between adjacent marker intervals. In addition, the HEI10 protein of the ZMM family, which contains a circular domain, is essential for the formation of class I recombination (sensitive to interference) and is present on the central elements of chromosome axes and synaptic complexes. The immunolocalization of HEI10 in pachytene can represent early selected recombination intermediates entering the ZMM pathway, and then correspond to the final class I COs.

[0005] In summary, considering that recombination is not uniformly distributed on chromosomes, assessing the whole-genome recombination rate of polyploid species using specific selection markers (number and chromosomal location) may be inaccurate. Furthermore, the widespread application of genetic linkage mapping may be limited by the creation of homozygous experimental materials, and the HEI10 immunolocalization method is complex and limited to certain species. Meanwhile, the detection of meiotic recombination events relies on reliable selection markers. Heterozygous markers in polyploid parents can also be used for screening, but heterozygous variant sites are often affected by sequencing variability and program parameter settings, increasing the difficulty of analysis. Therefore, this invention provides an efficient and accurate method for identifying recombination events based on polyploid yeast. This method is also applicable to other polyploid species that reproduce sexually using two mating-type parents. Summary of the Invention

[0006] In view of the lack of effective methods for identifying polyploid meiotic recombination events, the purpose of this invention is to provide an efficient and accurate method for detecting meiotic recombination events in polyploid yeast genomes.

[0007] To address the problems of existing technologies, the present invention provides the following technical solution: A method for detecting meiotic recombination events in the genome of polyploid yeast, comprising the following steps:

[0008] (1) Evaluation of resequencing data: Evaluation and assembly of the original genomic DNA sequencing files provided by the sequencing company;

[0009] (2) Determination of reliable screening markers between the two parents: Determine reliable screening markers between the two parents, especially heterozygous markers in polyploid parents;

[0010] (3) Further determination of heterozygous markers in polyploid parents: When determining the markers of parents, it should be clear that heterozygous markers can be introduced into polyploid parents;

[0011] (4) Determination of ploidy of offspring samples: The ploidy of offspring samples was determined using the Variant Allele Frequency (VAF) method.

[0012] (5) Genotype determination of offspring samples: Based on the ploidy and VAF of the offspring, the correspondence between the VAF interval of the polyploid given in this invention and the genotype is compared to determine whether the genotype of the offspring and the origin of its parents are clear;

[0013] (6) Constructing offspring background blocks: Only retain markers with clear parental origins, and use direct merging or filtering methods to merge adjacent marker sites of the same origin to construct offspring background blocks;

[0014] (7) Confirm the recombination event: and use IGV (Integrative Genomics Viewer), PCR and Sanger sequencing to check the authenticity of the recombination event.

[0015] Further, in step (1), check whether the data returned by the sequencing company contains clean data. If not, it needs to be filtered manually, and the resequencing data is quality checked and evaluated using FastQC software.

[0016] In step (2), the reference genome-based model refers to first aligning the reads in the clean data of the two parents to the reference genome, and then determining the differential sites between the two parents;

[0017] In step (3), further determination of heterozygous markers in polyploid parents: When determining the markers of parents, it should be clear that heterozygous markers can be introduced into polyploid parents; based on the ploidy of parents, a range of variable allele frequencies (VAF) is given, and corresponding genotypes are assigned, while filtering out variant sites outside the range; when at a certain site, one parent is a reliable heterozygous variant, and the other parent is a homozygous mutant or non-mutant, this site can be used as a screening marker;

[0018] In step (4), at each marker position, the reads in the clean data of the offspring sample are compared with the reference genome or the parent genome, and the ploidy of the offspring is determined based on the VAF distribution of the heterozygous markers inherited from the parents.

[0019] In step (5), based on step (4), the offspring genotype is determined according to the ploidy and frequency of variant alleles of the offspring samples, following the VAF interval and genotype correspondence rules described in step (3); and the parental origin of the variant sites at the corresponding positions in the offspring samples is discussed based on the genotypes of the parents at each marker position.

[0020] In step (6), only markers with a clear parental origin are retained, and adjacent marker sites of the same origin are merged using direct merging or filtering methods to construct the offspring background plate;

[0021] In step (7), if the parental origin of the background plate in the progeny changes once and there are at least two supporting markers, it is considered that a recombination has occurred. The recombination breakpoint and its surrounding area should have continuous reads to exclude ectopic interference, and the reads should be simple and clear in both parents. The insertion length of the reads near the breakpoint should be 450 bp. The authenticity of the markers before and after the breakpoint can be checked manually, and the recombination event can be verified by combining PCR and Sanger sequencing. If it is Saccharomyces cerevisiae, the type of recombination event can be refined according to whether the progeny sample comes from a set of tetraspores.

[0022] Furthermore, in step (1), if manual filtering is required to obtain clean data, read adapters, reads containing more than 10% N, and reads with low-quality bases accounting for more than 50% of the total number of reads must be removed; the resequencing data is quality checked and evaluated, and the resequencing data types include MD5 check value, data volume, GC content, and genome coverage.

[0023] Furthermore, in step (2), the parent-based model means that both hybrid parents have a fully assembled published genome, and the parent marker can be determined directly by whole-genome alignment; if the known genome is a closely related genome of the hybrid parents, the reads of the sequenced parents can be aligned to the closely related genomes, and the parent marker can be determined accordingly.

[0024] Further, in step (2), based on the reference genome model, clean data is aligned to the reference genome using the default parameters of BWA-mem; the SAM file is sorted using Picard's SortSam module to obtain the BAM file; abnormal duplicate reads are marked using Picard's MarkDuplicates tool, and then re-aligned using GATK's RealignerTargetCreator and IndelRealigner modules; SNVs and INDELs are identified using GATK's UnifiedGenotyper (UG) tool, and the "-rf MappingQuality-mmq 10" parameter is set to filter reads with an alignment quality less than 10 to obtain a VCF file containing all original variant sites. After evaluating the variant sites and using strict screening criteria, the parental markers are determined.

[0025] Furthermore, in step (3), the VAF range of variant allele frequencies in different ploidy parents is one of the core aspects of this invention. Specifically, if the parent is diploid, the VAF range and genotype correspondence are defined as follows: 0-0.1 (R / R), 0.35-0.65 (R / A), and 0.9-1 (A / A); if the parent is triploid, the ranges are 0-0.1 (R / R / R), 0.23-0.43 (R / R / A), 0.56-0.76 (R / A / A), and 0.9-1 (A / A / A); if the parent is tetraploid, the ranges are 0-0.1 (R / R / R / R), 0.15-0.35 (R / R / R / A), 0.35-0.65 (R / A / R / A), 0.65-0.85 (R / A / A / A), and 0.9-1 (A / A / A / A), where R represents consistency with the reference genome and A represents inconsistency.

[0026] Furthermore, in step (4), if the VAF peak value of the heterozygous marker inherited by the offspring is close to 0.5, then the offspring sample is diploid; if the VAF peak value is close to 0.33 and 0.66, it indicates that the heterozygous copy number is 1 or 2 copies, then the sample is triploid; and when the VAF peak value is close to 0.25, 0.5 and 0.75, the sample is tetraploid.

[0027] Furthermore, in step (5), determining the offspring genotype based on the correspondence between VAF intervals and genotypes, and identifying the parental origin of the variant sites at each marker position in the offspring sample, is another core aspect of this invention. Specifically, if only one parent is diploid, and the offspring is diploid, then regardless of the offspring's genotype, the parental origin can be clearly identified when the genotypes of the two parents are R / R and A, or A / A and R, respectively. For heterozygous markers of diploid parents, the offspring can be identified as originating from diploid parents only when the genotypes of the two parents are R / A and A (offspring are R / R), or R / A and R (offspring are A / A). If the offspring are triploid, there are 16 possible combinations. The parental origin can be identified in 12 cases, including 4 cases involving heterozygous markers of diploid parents. Specifically, when the genotypes of the two parents are R / A and A (offspring are A / A / A) or R / A and R (offspring are R / R / R), it can be determined that the offspring are derived from both parents.

[0028] Furthermore, in step (5), it can be determined that the offspring of the two parents with genotypes R / A and A are R / R, and the offspring of R / A and R are A / A. When the offspring of the two parents with genotypes R / A and A are A / A / A, and the offspring of R / A and R are R / R / R, it can be determined that the offspring are derived from both parents. When the offspring of the two parents with genotypes R / A and A are R / R / R, and the offspring of R / A and R are A / A / A, it can be determined that the offspring are derived from the diploid parents.

[0029] Furthermore, in step (6), when the number of available markers for the hybridization combination is small, the authenticity of the markers can be manually checked using IGV, and adjacent markers from the same source can be merged using the direct merging method. The filtering method requires a sufficient number of markers (≥2 / 10kb), and markers from the same source can be merged using the "--fill-gaps" function of VCFtools. The limiting parameters can be adjusted according to the actual situation, including the minimum fragment length, the number of fragment markers, and the marker ratio.

[0030] Beneficial effects: The method provided by this invention can efficiently and accurately detect recombination events in polyploid yeast, providing a powerful analytical tool for assessing the meiotic recombination rate of polyploid yeast. It can also be extended to other polyploid species that reproduce sexually with two mating-type parents, thus providing a certain molecular basis for studying the impact of recombination on the evolution of polyploid genomes.

[0031] Compared with the prior art, the present invention has the following advantages:

[0032] (1) Based on the reliable information on all variant sites provided by whole-genome sequencing, the genotype identification problem caused by the fluctuation of heterozygous mutant reads in polyploids is effectively avoided, and an efficient and accurate method for detecting recombination events is provided. It is applicable to polyploid yeast and other polyploid species that reproduce sexually with two mating types of parents.

[0033] (2) Compared with assessing recombination rates using a limited number of markers, which requires statistical analysis of a large amount of progeny genotype and phenotypic data and the use of mathematical models, genetic linkage mapping, and immunolocalization of HEI10 at pachytene stage, the method described in this invention is simple to operate and can provide more accurate whole-genome recombination rates and local recombination rates. It can also elucidate recombination distribution and recombination interference characteristics. In summary, the emergence of this invention makes up for the shortcomings of existing research methods and provides a very good approach for the rapid identification of meiotic recombination events in polyploid yeast and other species (sexual reproduction using two mating-type parents).

[0034] (3) The method provided by this invention can be used to efficiently identify and screen meiotic recombination events in polyploid yeast and other polyploid species that reproduce sexually with two mating-type parents. This not only provides a powerful analytical tool for assessing the polyploid meiotic recombination rate, but also allows for comparison with the identification methods for haploid meiotic recombination events, identifying similarities and differences, thereby optimizing and integrating the detection process for meiotic recombination events. Attached Figure Description

[0035] Figure 1 This is a schematic diagram of the process of the present invention. It is implemented in four steps, including the identification of parental markers (a), the determination of offspring ploidy (b), the determination of offspring genotypes (c), and the identification of recombination events (d). Models based on the reference genome (1), the parental genome (2), and (3) can all be used to determine markers between two parents, and the appropriate model can be selected based on the actual situation. Simultaneously, this step involves polyploid parents, and their heterozygous variation sites can also be used as screening markers between parents, such as... Figure 1 The bases in the boxes for Parent 1 and Parent 2 are shown.

[0036] Figure 2This is a distribution diagram of sexual spore recombination events in the single-chromosome yeast msh2 mutant hybrid combination (diploid × haploid) of the present invention (Example 1). The left side of the figure shows the sample name. The spores are divided into three categories according to ploidy: 2 haploids, 36 diploids, and 3 triploids. "centromere" indicates a centromere retained in the single-chromosome yeast. "gene" indicates the position of five markers, including chrV:463213, chrXIV:310185, chrXII:1031786, chrXV:147382-150276 (MSH2) and chrII:199071-201702 (MAT locus). In the diagram, red (A / A), blue (B / B), and gray (NA) represent the sources of Mat-α, Mat-a, and Mat-α / Mat-a (obtained after constructing the offspring background blocks), respectively. The conversion between the "A / A", "B / B", and "NA" blocks with clearly defined sources can be recorded as a single exchange and recombination.

[0037] Figure 3 This is a distribution diagram of sexual spore recombination events in the single-chromosome yeast msh2 mutant hybrid combination (diploid × diploid) of the present invention (Example 2). In this combination, there are 2, 9, and 1 diploid, triploid, and tetraploid spores, respectively, with recombination times of 37, 48, 14, 20, 14, 16, 18, 26, 14, 24, 10, and 12 times, respectively. Other illustrations are shown below. Figure 2 Consistent. Detailed Implementation

[0038] The present invention is further described below through two examples, but these are not intended to limit the scope of application of the invention.

[0039] To make the technical problems, technical solutions, and beneficial effects of this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0040] In this application, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0041] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, "at least one of a, b, or c", or "at least one of a, b, and c", can both mean: a, b, c, ab (i.e., a and b), ac, bc, or abc, where a, b, and c can be single or multiple.

[0042] It should be understood that in the various embodiments of this application, the order of the above processes does not imply the order of execution. Some or all steps may be executed in parallel or sequentially. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0043] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0044] The weights of the relevant components mentioned in the embodiments of this application can refer not only to the specific content of each component, but also to the proportional relationship between the weights of the components. Therefore, any scaling up or down of the content of the relevant components according to the embodiments of this application is within the scope disclosed in the embodiments of this application. Specifically, the mass in the embodiments of this application can be a well-known unit of mass in the chemical industry, such as μg, mg, g, or kg.

[0045] The terms "first" and "second" are used for descriptive purposes only, to distinguish objects, such as substances, from one another, and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. For example, without departing from the scope of the embodiments of this application, "first XX" may also be referred to as "second XX," and similarly, "second XX" may also be referred to as "first XX." Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of that feature.

[0046] The present invention provides a method for detecting meiotic recombination events in the genome of polyploid yeast, comprising the following steps:

[0047] (1) Evaluation of resequencing data: Check whether the data returned by the sequencing company contains clean data. If not, it needs to be manually filtered and the resequencing data should be quality checked and evaluated using FastQC software. If it is necessary to manually filter to obtain clean data, it is necessary to remove read adapters, reads with N ratio greater than 10%, and reads with low-quality bases accounting for more than 50% of the number of reads. The resequencing data should be quality checked and evaluated. The data types of resequencing include MD5 check value, data volume, GC content and genome coverage.

[0048] (2) Determination of reliable screening markers between the two parents: The reference genome-based model refers to first aligning the reads in the clean data of the two parents to the reference genome, and then determining the differential sites between the two parents; the parent-based model refers to the fact that both hybrid parents have fully assembled published genomes, and the markers of the parents can be determined directly by whole genome alignment; if the known genome is a closely related genome of the hybrid parents, the reads of the sequenced parents can be aligned to the closely related genomes, and then the markers of the parents can be determined accordingly.

[0049] In the reference genome-based model, clean data was aligned to the reference genome using the default parameters of BWA-mem. The SAM file was sorted using Picard's SortSam module to obtain the BAM file. Abnormal duplicate reads were marked using Picard's MarkDuplicates tool, and then re-aligned using GATK's RealignerTargetCreator and IndelRealigner modules. The UnifiedGenotyper (UG) tool of GATK was used to identify SNVs and INDELs, and the "-rf MappingQuality-mmq10" parameter was set to filter reads with an alignment quality of less than 10, resulting in a VCF file containing all original variant sites. After evaluating the variant sites and using strict screening criteria, the parental markers were determined.

[0050] (3) Further determination of heterozygous markers in polyploid parents: When determining the markers of parents, it should be clear that heterozygous markers can be introduced into polyploid parents; according to the ploidy of the parents, a range of variable allele frequencies (VAF) is given, and the corresponding genotype is assigned, while filtering out the variant sites outside the range; when at a certain site, one parent is a reliable heterozygous variant, and the other parent is a homozygous mutant or non-mutant, the site can be used as a screening marker; giving a range of variable allele frequencies (VAF) in parents with different ploidy is one of the core aspects of this invention. Specifically, if the parent is diploid, the VAF range and genotype correspondence are defined as follows: 0-0.1 (R / R), 0.35-0.65 (R / A), and 0.9-1 (A / A); if the parent is triploid, the ranges are 0-0.1 (R / R / R), 0.23-0.43 (R / R / A), 0.56-0.76 (R / A / A), and 0.9-1 (A / A / A); if the parent is tetraploid, the ranges are 0-0.1 (R / R / R / R), 0.15-0.35 (R / R / R / A), 0.35-0.65 (R / A / R / A), 0.65-0.85 (R / A / A / A), and 0.9-1 (A / A / A / A), where R represents consistency with the reference genome and A represents inconsistency.

[0051] (4) Determination of ploidy in offspring samples: At each marker position, the reads in the clean data of the offspring sample are compared with the reference genome or the parent genome. The ploidy of the offspring is determined based on the VAF distribution of the heterozygous markers inherited from the parents. If the VAF peak of the heterozygous markers inherited from the parents is close to 0.5, the offspring sample is diploid. If the VAF peak is close to 0.33 and 0.66, it indicates that the heterozygous copy number is 1 or 2 copies, and the sample is triploid. When the VAF peak is close to 0.25, 0.5 and 0.75, the sample is tetraploid.

[0052] (5) Genotype determination of offspring samples: Based on step (4), the offspring genotype is determined according to the ploidy and allele frequency of the offspring samples, following the VAF interval and genotype correspondence rules described in step (3); and the parental origin of the corresponding variant sites in the offspring samples is discussed based on the genotypes of the parents at each marker position; determining the offspring genotype and the parental origin of the variant sites at each marker position in the offspring samples based on the VAF interval and genotype correspondence is another core aspect of this invention. Specifically, if only one parent is diploid, and the offspring is diploid, the parental origin can be determined regardless of the offspring's genotype if the genotypes of the two parents are R / R and A, or A / A and R, respectively; for heterozygous markers of diploid parents, the offspring can be determined to originate from diploid parents only when the genotypes of the two parents are R / A and A (offspring are R / R), or R / A and R (offspring are A / A). If the offspring are triploid, there are 16 possible combinations. The parental origin can be clearly identified in 12 cases, including 4 cases involving heterozygous markers from diploid parents. Specifically, when the genotypes of the two parents are R / A and A (offspring are A / A / A), or R / A and R (offspring are R / R / R), it can be determined that the offspring originates from both parents. When the genotypes of the two parents are R / A and A (offspring are R / R / R), or R / A and R (offspring are A / A / A), it can be determined that the offspring originates from the diploid parent. Furthermore, this invention also provides methods for determining the genotypes and parental origins of diploid, triploid, and tetraploid offspring when both parents are diploid (Tables 1 to 3).

[0053] (6) Construction of progeny background blocks: Only markers with clear parental origin are retained. Adjacent marker loci from the same origin are merged using direct merging or filtering methods to construct progeny background blocks. When the number of available markers for a hybrid combination is small, the authenticity of the markers can be manually checked using IGV, and adjacent markers from the same origin can be merged using direct merging. Filtering requires a sufficient number of markers (≥2 / 10kb). The "--fill-gaps" function of VCFtools can be used to merge markers from the same origin. The limiting parameters can be adjusted according to the actual situation, including the minimum fragment length, the number of fragment markers, and the marker ratio.

[0054] (7) Confirmation of recombination event: A recombination is considered to have occurred when the parental origin of the background plate in the progeny changes once (with at least 2 supporting markers); the recombination breakpoint and its surrounding area should have continuous reads to exclude ectopic interference, and the reads should be simple and clear in both parents, with an insertion length of 450 bp near the breakpoint; the authenticity of the markers before and after the breakpoint can be checked manually, and the recombination event can be verified by combining PCR and Sanger sequencing; if it is Saccharomyces cerevisiae, the type of recombination event can be refined according to whether the progeny sample comes from a set of tetraspores.

[0055] Example 1

[0056] Identification of sexual spore recombination events in single-chromosome yeast msh2 mutant hybrid combinations (diploid × haploid).

[0057] 1. Resequencing of parents, F1 cells, and sexual spores of the msh2 mutant hybrid combination: Two mating types (Mat-α and Mat-a) of single-chromosome yeast were purchased from the Institute of Plant Physiology and Ecology (Synthetic Biology Elements and Database), Shanghai Institutes for Biological Sciences, Chinese Academy of Sciences. A Mat-α mating type msh2 mutant strain was constructed using the URA3 screening marker. After 83 plate subcultures, it was hybridized with the original Mat-a mating type strain to construct the single-chromosome yeast msh2 mutant hybrid combination mW1_2a. Subsequently, conjugated F1 monoclonal cells were induced to sporulate using potassium acetate liquid medium. Sporulation efficiency was observed under a microscope, and 41 sexual spores were identified from the products. Combining the advantages of Solarbio D1900-100T and traditional alcohol precipitation methods, genomic DNA was extracted from the parents, F1 cells, and sexual spores. The DNA was sent to Wuhan BGI Genomics Co., Ltd. for quality testing, and 44 samples meeting the library construction requirements were selected for whole-genome sequencing. Using the MGISEQ-2000RS high-throughput sequencing dataset (PE150), conventional library construction was employed, with fragment sizes ranging from 150 to 300 bp. Circular DNA molecular libraries were generated, and DNA nanospheres (DNBs) were synthesized. Fragment read lengths were 150 bp, and paired-end sequencing was used. The total data volume for each sample was 1 GB.

[0058] 2. Evaluation of Resequencing Data: If the returned data is raw data, it needs to be manually filtered. The method is to remove read adapters, reads containing more than 10% N, and reads with low-quality bases accounting for more than 50% of the read count. Subsequently, the returned sequencing data needs quality control. First, perform MD5 checksum verification to check for missing data. Then, calculate the GC content of the resequencing data. The GC content of *Saccharomyces cerevisiae* is approximately 40%, and a GC content of around 46% may indicate contamination by *Candida albicans* or other microorganisms. Finally, preliminarily calculate the sequencing data volume, estimate the sequencing depth, and assess whether the sample meets the requirements for subsequent analysis. Preliminary statistics show that the average sample data volume is 1.2 G, the average sequencing depth is 99.9-fold (specifically ranging from 90.5-fold to 170.8-fold), and the GC content is 40.8%. After aligning the sequencing reads to the reference genome S288C, the average coverage of regions with ≥5 reads is 94.76%, indicating that the data quality meets the requirements for subsequent analysis.

[0059] 3. Determination of reliable screening markers between the two parents: Based on the reference genome model, the Saccharomyces cerevisiae S288C genome and annotation files were downloaded from the National Center for Biotechnology Information (https: / / www.ncbi.nlm.nih.gov / genome / ?term=S288C) and used as the reference genome. Using the default parameters of the BWA-MEM algorithm, files with the .fq.gz extension from the clean data of both parents were aligned to the reference genome to obtain the original SAM file. Subsequently, the SAM file was sorted using the SortSam module of Picard version 1.114 to obtain the BAM file. Abnormal reads introduced by the sequencing process were marked using Picard's MarkDuplicates tool. Simultaneously, the RealignerTargetCreator and IndelRealigner modules of GenomeAnalysisTK version 3.7 were used to re-align regions near INDELs that might have assembly errors. Next, the coverage, depth, and base percentage of the BAM files were statistically analyzed. Then, the GATK UnifiedGenotyper (UG) tool was used to identify SNVs and INDELs. By setting "-rfMappingQuality-mmq10", a VCF file containing all original variant sites was obtained. After evaluating the variant sites and applying strict screening criteria, the parental markers were determined. The results showed that there were 925 markers between the two parents in this combination.

[0060] 4. Ploidy determination of progeny samples: At each marker position, reads in the clean data of the progeny samples were compared with the reference genome or parental genome. Ploidy was determined based on the VAF distribution of heterozygous markers inherited from both parents. Of the 41 sexual spores, mW1_2a-fsp18 and mW1_2a-fsp36 were haploid, and mW1_2a-fsp13, mW1_2a-fsp22, and mW1_2a-fsp34 were triploid; the remaining 36 were diploid.

[0061] 5. Determination of offspring genotypes: The method described in this invention provides the following correspondences between VAF ranges and genotypes in diploid samples: 0-0.1 (R / R), 0.35-0.65 (R / A), and 0.9-1 (A / A); if the sample is triploid, then 0-0.1 (R / R / R), 0.23-0.43 (R / R / A), 0.56-0.76 (R / A / A), and 0.9-1 (A / A / A). Here, R represents consistency with the reference genome, and A represents inconsistency. Taking triploid samples as an example, the number of markers clearly derived from the Mat-α parent, Mat-a parent, and both parents was counted. In mW1_2a-fsp13, there were 80, 94, and 429 markers, respectively; in mW1_2a-fsp22, there were 178, 7, and 256 markers, respectively; and in mW1_2a-fsp34, there were 66, 4, and 361 markers, respectively. The number of heterozygous markers involving the Mat-α parent in the three samples were 339, 157, and 177, respectively. The genotypes of the two parents were "R / A" and "R," respectively, and the offspring genotypes were "R / R / R" or "A / A / A."

[0062] 6. Construction of progeny background sections: Only markers with clear parental origin are retained; the number of markers with clear parental origin varies among different spores. Here, a filtering method is used to merge adjacent markers of the same origin from the mW1_2a combination progeny to determine the parental origin of each region in the progeny sample. The "--fill-gaps" function of VCFtools is used to merge markers of the same origin, with a minimum fragment length of 1kb and at least two markers per fragment.

[0063] 7. Confirmation of Recombination Events: The Mat-α mating type parent of the mW1_2a combination is diploid, and all 41 spores are random monospores. It cannot be determined whether any spores originated from the same set of tetraspores; therefore, recombination events other than simple exchange events (COs) cannot be obtained. Furthermore, considering the marker density and distribution of the mW1_2a combination, it is believed that the length of most NCO events is less than the interval between adjacent markers. Therefore, the impact of NCO events on the recombination rate assessment of this combination is negligible. Based on the background plates constructed in section 6, changes in background plates with clearly defined origins from the Mat-α parent, from the Mat-a parent, and from both parents can all be counted as one CO. The occurrence of recombination events is confirmed by manually checking the authenticity of markers before and after the breakpoint. According to the summary, except for mW1_2a-fsp38, the recombination times of mW1_2a-fsp1 to mW1_2a-fsp42 are as follows: 34 / 34 / 57 / 35 / 41 / 35 / 43 / 56 / 35 / 43 / 37 / 45 / 66 (triploid) / 56 / 40 / 34 / 50 / 15 / 52 / 51 / 49 / 45 (triploid) / 39 / 40 / 41 / 39 / 49 / 30 / 49 / 47 / 39 / 47 / 40 / 26 (triploid) / 31 / 0 / 45 / 30 / 42 / 40 / 54. Figure 2 The distribution of recombination breakpoints for each spore is given.

[0064] Example 2

[0065] Identification of sexual spore recombination events in single-chromosome yeast msh2 mutant hybrid combinations (diploid × diploid).

[0066] 1. Resequencing sample information: The source of the two parents of the hybrid combination m1_2e is as described in Example 1. Both parents are single-chromosome yeast msh2 mutants (diploid). After inducing sporulation in potassium acetate liquid medium, 12 random sporosomes were obtained. The genomic DNA extraction method was the same as in Example 1 and they were sent to BGI Genomics for next-generation sequencing.

[0067] 2. Evaluation of resequencing data: Analysis showed that the average sample size was 1.2G, the average sequencing depth was 94.8-fold (ranging from 91.2-fold to 101.8-fold), and the GC content was 40.6%. After aligning the sequencing reads to the reference genome S288C, the average coverage of regions with ≥5 reads was 94.75%, indicating that the data quality met the requirements for subsequent analysis.

[0068] 3. Determination of reliable screening markers between the two parents: Similarly, based on the pattern of the reference genome, the parental markers were determined, and a total of 1927 markers were finally identified between the two parents in this combination.

[0069] 4. Ploidy identification of offspring samples: It was determined that among the 12 sexual spores, there was 1 tetraploid, 9 triploids and 2 diploids.

[0070] 5. Determination of Offspring Genotypes: Considering that both parents in this combination are diploid, the corresponding genotypes can be determined according to the ploidy of the offspring, referring to Tables 1 to 3 respectively, and the origin of the parents can be discussed. Specifically, based on the VAF value of the tetraploid spore m1_2e-fsp1, genotypes were assigned according to 0-0.1 (R / R / R / R), 0.15-0.35 (R / R / R / A), 0.35-0.65 (R / A / R / A), 0.65-0.85 (R / A / A / A), and 0.9-1 (A / A / A / A). Referring to Table 3, the number of markers in the tetraploid spore m1_2e-fsp1 that clearly originated from the Mat-α parent, the Mat-a parent, and both parents were determined to be 82, 12, and 780, respectively. After determining the genotype of diploid spore m1_2e-fsp5 using VAF, as shown in Table 1, the number of markers from the three sources were 279, 22, and 150, respectively. Taking triploid spore m1_2e-fsp9 as an example (see Table 2), the number of markers from the three sources were 95, 28, and 190, respectively.

[0071] 6. Construction of progeny background blocks: Based on the filtering method, adjacent markers of the same origin in the m1_2e combination progeny are merged to determine the parental origin of each region in the progeny samples. The "--fill-gaps" function of VCFtools is used to merge markers of the same origin, with a minimum fragment length of 1kb and at least two fragment markers.

[0072] 7. Confirmation of Recombination Events: Both parents of the m1_2e combination are diploid, and all 12 spores are random monospores, resulting in only simple exchange recombination (COs) events. In summary, the tetraploid m1_2e-fsp1 underwent 12 recombinations, the diploids m1_2e-fsp5 and m1_2e-fsp7 underwent 37 and 48 recombinations respectively, and the triploids m1_2e-fsp2 / 3 / 4 / 6 / 8 / 9 / 10 / 11 / 12 underwent 14, 20, 14, 16, 18, 26, 14, 24, and 10 recombinations respectively. Figure 3 The distribution of recombination breakpoints for each spore is given.

[0073] Table 1 shows the genotype identification and parental origin determination of diploid offspring in hybrid combinations where both parents are diploid:

[0074] Table 1

[0075]

[0076]

[0077] Table 2 shows the genotyping of triploid offspring and determination of parental origin in hybrid combinations where both parents are diploid:

[0078] Table 2

[0079]

[0080]

[0081] The genotype identification and parental origin determination of tetraploid offspring in hybrid combinations where both parents are diploid are shown in Table 3.

[0082] Table 3

[0083]

[0084] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope. The scope of protection of the present invention is defined by the appended claims, specification, and their equivalents.

Claims

1. A method for detecting meiotic recombination events in the genome of polyploid yeast, characterized in that... Includes the following steps: (1) Evaluation of resequencing data: Evaluate and assemble the original genomic DNA sequencing files provided by the sequencing company; check whether the data returned by the sequencing company contains clean data. If not, manually filter the data and use FastQC software to perform quality control and evaluation on the resequencing data; if manual filtering is required to obtain clean data, remove read adapters, reads with N ratio greater than 10%, and reads with low-quality bases accounting for more than 50% of the reads; perform quality control and evaluation on the resequencing data. The resequencing data types include MD5 check value, data volume, GC content, and genome coverage. (2) Determination of reliable screening markers between the two parents: Determine reliable screening markers between the two parents; Based on the reference genome model, first align the reads in the clean data of the two parents to the reference genome, and then determine the differential sites between the two parents; In the reference genome model, use the default parameters of BWA-mem to align the clean data to the reference genome; Sort the SAM file using Picard's SortSam module to obtain the BAM file; Abnormal duplicate reads were marked using Picard's MarkDuplicates tool, and then re-aligned using GATK's RealignerTargetCreator and IndelRealigner modules. GATK's UnifiedGenotyper tool was used to identify SNVs and INDELs, and reads with alignment quality less than 10 were filtered by setting the "-rf MappingQuality -mmq 10" parameter to obtain a VCF file containing all original variant sites. After evaluating the variant sites and using strict screening criteria, the markers of the parents were determined. (3) Further determination of heterozygous markers in polyploid parents: When determining parental markers, it should be clear that heterozygous markers can be introduced into polyploid parents; based on parental ploidy, a VAF (variable allele frequency) range is given, and corresponding genotypes are assigned, while filtering out variant sites outside the range; when at a certain site, one parent has a reliable heterozygous variant, and the other parent has a homozygous mutant or non-mutant state, that site can be used as a selection marker; the correspondence between VAF range and genotype is defined as follows: if the parent is diploid, then 0-0.1 corresponds to R / R, 0.35-0.65 corresponds to R / A, and 0.9- 1 corresponds to A / A; if the parent is triploid, then 0-0.1 corresponds to R / R / R, 0.23-0.43 corresponds to R / R / A, 0.56-0.76 corresponds to R / A / A, and 0.9-1 corresponds to A / A / A; if the parent is tetraploid, then 0-0.1 corresponds to R / R / R / R, 0.15-0.35 corresponds to R / R / R / A, 0.35-0.65 corresponds to R / A / R / A, 0.65-0.85 corresponds to R / A / A / A, and 0.9-1 corresponds to A / A / A / A; where R represents consistency with the reference genome, and A represents inconsistency. (4) Determination of ploidy of offspring samples: The ploidy of offspring samples is determined by the variable allele frequency method; at each marker position, the reads in the clean data of the offspring sample are compared with the reference genome, and the ploidy of the offspring is determined according to the VAF distribution of the heterozygous markers inherited from the parents; if the VAF peak of the heterozygous markers inherited from the parents corresponds to 0.5, then the offspring sample is diploid; if the VAF peak corresponds to 0.33 and 0.66, it means that the heterozygous copy number is 1 or 2 copies, then the sample is triploid; and when the VAF peak corresponds to 0.25, 0.5 and 0.75, the sample is tetraploid; (5) Determination of genotype of offspring samples: Based on step (4), the offspring genotype is determined according to the ploidy and frequency of variant alleles of the offspring samples, and the correspondence between the VAF range and genotype described in step (3); and the parental origin of the variant sites at the corresponding positions in the offspring samples is discussed based on the genotypes of the parents at each marker position. (6) Constructing offspring background blocks: Only retain markers with clear parental origins, and use direct merging or filtering methods to merge adjacent marker sites of the same origin to construct offspring background blocks; (7) Confirm recombination event: Check the authenticity of recombination event with IGV, PCR and Sanger sequencing. If the parental origin of the background plate in the progeny changes once and there are at least 2 supporting markers, it is considered that a recombination has occurred. The recombination breakpoint and its surrounding area should have continuous reads to exclude ectopic interference, and the reads should be simple and clear in both parents. The insertion length of reads near the breakpoint should be 450 bp. Manually check the authenticity of the markers before and after the breakpoint, or combine PCR and Sanger sequencing to verify the recombination event.

2. The method for detecting meiotic recombination events in the genome of polyploid yeast according to claim 1, characterized in that: In step (5), if only one of the parents is diploid, and the offspring is diploid, the parental origin can be determined regardless of the genotype of the offspring, provided that the genotypes of the two parents are R / R and A, or A / A and R. For heterozygous markers of diploid parents, the offspring can be determined to originate from diploid parents only when the genotypes of the two parents are R / A and A and the offspring are R / R, or when the two parents are R / A and R and the offspring are A / A. If the offspring is triploid, there are 16 possible combinations, and the parental origin can be determined in 12 cases, of which 4 cases involve heterozygous markers of diploid parents.

3. The method for detecting meiotic recombination events in the genome of polyploid yeast according to claim 1, characterized in that: In step (5), when the offspring of parents with genotypes R / A and A are A / A / A ​​and the offspring of parents with genotypes R / A and R are R / R / R, it can be determined that the offspring are derived from both parents; when the offspring of parents with genotypes R / A and A are R / R / R and the offspring of parents with genotypes R / A and R are A / A / A, it can be determined that the offspring are derived from the diploid parents.

4. The method for detecting meiotic recombination events in the genome of polyploid yeast according to claim 1, characterized in that: In step (6), the filtering rule needs to have enough tags, enough to be ≥2 tags / 10kb. Use VCFtools' "--fill-gaps" to merge tags from the same source. The parameters are adjusted according to the actual situation, including minimum fragment length, number of fragment tags, and tag ratio.