A new algorithm for constructing a genetic map of the MAGIC population.

The new algorithm constructs a genetic map of the MAGIC population, which solves the problem of information loss caused by quality control screening, clarifies the parental origin of variant sites, and improves the accuracy of the genetic map and breeding efficiency. It is applicable to diploid and allopolyploid species.

CN122090930APending Publication Date: 2026-05-26ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2026-01-23
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies often require quality control screening when constructing MAGIC population genetic maps, which may remove some real genetic variations, leading to information loss. Furthermore, the lack of research on the parental origin of variation sites affects the breeding process.

Method used

A novel algorithm was used to analyze the whole-genome resequencing data of the MAGIC population, constructing a genetic linkage map with a total genetic distance closer to the theoretical value and more complete locus information. The parental origin of each variant locus was clarified, and the genetic map was constructed by calculating the parental genetic similarity index and recombination rate of chromosome segments.

Benefits of technology

It improves the accuracy and interpretability of genetic maps, enabling more accurate detection of minor genes controlling complex quantitative traits, clarifying the parental origin of variation sites, and improving the breeding process. It is applicable to diploid and allopolyploid species.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090930A_ABST
    Figure CN122090930A_ABST
Patent Text Reader

Abstract

This application discloses a novel algorithm for constructing a genetic map of a MAGIC population, comprising: collecting species data; segmenting chromosomes based on the collected data to form chromosome fragments; calculating the parental genetic similarity index of chromosome fragments to determine parental origin; obtaining a definite origin matrix and a fuzzy origin matrix; using definite fragments from a single parental origin in the definite origin matrix as anchor points; using the anchor points as starting points, examining adjacent fragments of fuzzy fragments in the fuzzy origin matrix to determine parental origin; merging fragments with the same parental origin and adjacent positions into a larger fragment to form a parental origin matrix; calculating recombination rate and genetic distance; constructing a genetic map; and using the novel algorithm to analyze the whole-genome resequencing data of a MAGIC population of a specified crop, thereby effectively detecting minor genes controlling complex quantitative traits, clarifying the parental origin of each variation site, and finding the required high-quality parents, thereby improving the breeding process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of genetic mapping technology, and in particular to a new algorithm for constructing genetic maps of MAGIC populations. Background Technology

[0002] Multiparent Advanced Generation Inter-Cross (MAGIC) populations are produced by crossbreeding multiple parents pairwise to obtain an individual containing information from all parents (Domínguez et al., 2025). This individual is then subjected to several generations of inbreeding or self-pollination to create a stable, homozygous population. Each offspring in this population has genetic information derived from all parents, and each locus is homozygous, meaning it originates from the same parent. Compared to traditional biparental populations, multiparental populations exhibit higher allelic and phenotypic diversity, enabling the detection of more frequent recombination events and achieving higher mapping accuracy. Compared to natural populations, the kinship among offspring produced by multiparental crosses is very clear, resulting in a well-defined population structure. Therefore, genetic studies of MAGIC populations have increased significantly in recent years. High-throughput sequencing (NGS), also known as next-generation sequencing technology, is a new technology for high-throughput parallel sequencing of large numbers of nucleic acid molecules (Liu et al., 2024). It boasts advantages such as low cost and the ability to generate large amounts of sequencing data in a short time. With the development of high-throughput sequencing technology, SNPs (Single Nucleotide Polymorphisms) and Indels (Insertion and Deletion) are being discovered more and more frequently. SNPs contain a wealth of genetic information, while Indels are a valuable complement to SNPs (Jia et al, 2025). Therefore, in genetic research, SNPs and Indels (hereinafter collectively referred to as variant sites) are often used as genetic markers for genetic linkage analysis, population genetics studies, and gene mapping. Constructing genetic maps is a crucial step in analyzing genetic linkage. A genetic map, or genetic linkage map, is an important component of genome research. It refers to a map showing the relative positions of genes and specific polymorphic markers within the genome, and can be used for genetic analysis and gene mapping of various diseases and traits.

[0003] Despite extensive research on MAGIC populations in the past, several challenges remain. Conventional MAGIC population studies often begin with quality control of the obtained variant sites, selecting a small number of representative sites from a large pool of polymorphic loci before proceeding with genetic map construction or correlation analysis (Oluwaseun J et al, 2025). However, due to computational limitations of software, it's impossible to input all loci from the entire genome to construct a genetic map. Therefore, current software for constructing MAGIC population genetic maps, such as mpMap (Huang et al, 2011) and GAPL (Zhang et al, 2018), first screens variant sites, removing strongly linked sites to select meaningful loci for map construction. Quality control itself can screen for loci with high effect values, while simultaneously reducing the computational burden of subsequent analyses and optimizing resource utilization. However, quality control may also present some challenges: it may remove some genuine genetic variations, especially those in low-frequency or rare variants, potentially leading to information loss; the post-quality-controlled data may require more complex interpretation, as it necessitates considering which variations were excluded and the reasons for their exclusion; and the post-quality-controlled data may not be suitable for direct sharing or reuse, as other researchers may require the original data for different analyses. Furthermore, when studying MAGIC populations, it is essential to understand not only the genotype information of variant sites but also their parental origin. Research on using MAGIC population sequence data to determine the parental origin of variant sites is limited, with most studies focusing on the association between genotype and phenotype, while studies on the association between parental origin and phenotype are less common. Summary of the Invention

[0004] This application provides a novel algorithm for constructing genetic maps of MAGIC populations. Analyzing whole-genome resequencing data of MAGIC populations of specified crops using this algorithm, a genetic linkage map is constructed that more closely approximates theoretical values ​​and contains more complete locus information without traditional quality control screening. This effectively detects minor genes controlling complex quantitative traits, clarifies the parental origin of each variant locus, and allows for direct identification of desired high-quality parents after effect calculation, thereby improving the breeding process. Combining information from all variant loci, the algorithm covers the entire chromosome and offers strong interpretability.

[0005] This application provides a novel algorithm for constructing genetic maps of MAGIC populations, including: S101, Collect diploid species data, and perform chromosome segmentation based on the collected data to form chromosome fragments. The collected diploid species data includes parental genotype data, MAGIC offspring genotype data, and reference genome data. Among them, the species include diploid species and allopolyploid species. S102, Calculate the parental genetic similarity index of chromosome segments based on chromosome segments, determine the parental origin based on the parental genetic similarity index, obtain a definite origin matrix and a fuzzy origin matrix, take the definite segments of a single parental origin in the definite origin matrix as anchor points, and use the anchor points as the starting point to check the adjacent segments of the fuzzy segments in the fuzzy origin matrix to determine the parental origin. In addition, for points that cannot be determined, randomly select a parent from the fuzzy origin matrix as the source. Finally, merge segments with the same parental origin and adjacent positions into a large segment to form a parental origin matrix. S103, Calculate recombination rate and genetic distance based on the merged large fragments, and construct a genetic map based on recombination rate and genetic distance, wherein the genetic map is a genetic map of a diploid species.

[0006] Preferably, the steps of chromosome segmentation are as follows: reading reference genome information, determining the length and start position of the chromosome, calculating the start and end coordinates of each window according to the preset window size, and assigning the mutation sites to the corresponding windows according to their physical locations to form chromosome segments.

[0007] Preferably, the method for determining parental origin is as follows: extracting genotypic information of all variation sites in each chromosome segment of the parents and offspring, calculating the homology between the offspring and each parent on the same segment, and calculating the parental genetic similarity index based on the homology. The formula for calculating the parental genetic similarity index is: Wherein, DST is the genetic similarity index between the parents, IBS0 is the number of loci with IBS=0, IBS1 is the number of loci with IBS=1, and IBS2 is the number of loci with IBS=2. The parent with the highest DST value is selected as the initial source of the offspring on this fragment, and the initial sources form a clear source matrix.

[0008] Preferably, the method for determining the parental origin further includes: when multiple parents have the same DST value and are all the highest values, recording the multiple parents as an initial source set to form a fuzzy source matrix, and determining the parental origin by extending the anchor point and randomly assigning parents.

[0009] Preferably, the formula for calculating the recombination rate is: in, For parental combination deductions, , For weighted parental combinations, ,in, The number of individuals in the offspring population from a specific parental combination. The recombination rate is defined as 0 ≤ r ≤ 0.5. This represents the total number of individuals in the offspring population, which is the sum of the number of individuals from all parental combinations. , representing the total number of individuals in all 64 combinations. = ,in, It refers to the combination of parental sources at adjacent loci for a single individual.

[0010] Preferably, the formula for calculating genetic distance is: in, Genetic distance, This represents the recombination rate.

[0011] Preferably, the method further includes the step of constructing an allopolyploid species map: S201. Collect allopolyploid parental data, construct a pan-genome map, screen for existing or missing variant blocks based on the pan-genome map, perform genotyping on the offspring of the MAGIC population, and obtain a three-dimensional genotype matrix. S202, the three-dimensional genotype matrix is ​​split into subgenome datasets, the parental origin of the subgenome datasets is determined, and collinear strands are constructed based on the parental origin of the subgenome datasets; S203, the formed collinear strands are corrected, and a hierarchical genetic map is constructed based on the corrected parental origin matrix. The hierarchical genetic map is the genetic map of allopolyploid species.

[0012] Preferably, the three dimensions of the three-dimensional genotype matrix are the number of offspring individuals, the total number of variant blocks present or absent, and the subgenome origin, respectively.

[0013] Preferably, the method for determining the parental origin of a subgenome dataset is as follows: for each present or missing variant block in each subgenome dataset, all offspring individuals are traversed; for each offspring individual, the genotype status of that present or missing variant block is counted; the genotype status of the offspring individual is compared with the genotype status of the parent in that present or missing variant block; the genotype similarity between the offspring individual and each parent is calculated; and the parent with the highest similarity is selected as the parental origin of the offspring individual in that present or missing variant block.

[0014] Preferably, the method for correcting the formed collinear chains is as follows: the correction includes point correction and line correction. The point correction involves traversing the collinear chains and units, checking the consistency of the presence or deletion variant blocks of the three subgenomes within the collinear unit, correcting fuzzy judgments, and marking irreconcilable conflicts as conflict points. The line correction involves finding anchor points for the remaining isolated fuzzy points or conflict points after point correction, reasoning and correcting the fuzzy points and conflict points according to topological relationships, and defining a large-scale continuous conflict region.

[0015] One or more technical solutions provided in this application have at least the following technical effects or advantages: the total genetic distance is closer to the real situation, the error in the simulation experiment can be less than 10%, it can more accurately reflect the actual distribution and inheritance law of genes on chromosomes, improve accuracy, use a new algorithm to analyze the whole genome resequencing data of the MAGIC population of a specified crop (taking tobacco as an example), and construct a genetic linkage map with a total genetic distance closer to the theoretical value and more complete locus information without traditional quality control screening, combine the neighboring loci to identify the parental source of the variant locus, provide more variation for the variant locus, thereby effectively detect the micro-effect genes controlling complex quantitative traits, clarify the parental source of each variant locus, and after calculating the effect, the required high-quality parents can be directly found, thereby improving the breeding process, and combining all variant locus information, covering the whole chromosome, with strong interpretability; This research achieves the construction of genetic maps at the allopolyploid level using structural variations (PAVs) as core markers, capturing functional genetic diversity at a larger scale. Through a collinear chain correction algorithm, the correction logic is upgraded from local voting to global topological reasoning, greatly improving the accuracy of parental origin determination. Point-level and line-level corrections can resolve most fuzzy judgments and conflicts, making the judgment results more reliable. A hierarchical, high-precision genetic map can be constructed that can clearly distinguish subgenome origins and uses structural variations as core markers. It can accurately locate important quantitative trait loci (QTLs) with subgenome specificity in structural variation regions, and study the evolutionary history, recombination patterns, and subgenome interaction mechanisms of allopolyploid genomes, providing a clear blueprint for precision design breeding of polyploid crops. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the new algorithm of the present invention for constructing a genetic map of a MAGIC population. Figure 2 This is a schematic diagram of the process for constructing a hierarchical genetic map according to the present invention. Detailed Implementation

[0017] To facilitate understanding of the present invention, a more complete description of this application will be given below with reference to the accompanying drawings, which illustrate preferred embodiments of the invention. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to enable a more thorough and complete understanding of the disclosure of the present invention.

[0018] It should be noted that the terms "vertical," "horizontal," "up," "down," "left," "right," and similar expressions used in this article are for illustrative purposes only and do not represent the only possible implementation.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention; the term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0020] Example 1: Figure 1 This is a flowchart illustrating a novel algorithm for constructing a MAGIC population genetic map according to an embodiment of the present invention, including: S101, Collect data, and perform chromosome segmentation based on the collected data to form chromosome fragments; the collected data includes parental genotype data, MAGIC offspring genotype data, and reference genome data; Specifically, we contacted a professional gene sequencing company to obtain whole-genome resequencing services for eight tobacco parents. The data obtained included tobacco variety, sample quantity, and sequencing depth. Parental samples were collected and sent to the sequencing company. After sequencing, the company returned the raw sequencing data. We used bioinformatics software BWA to align the sequencing data to the tobacco reference genome. After alignment, we used the variant detection tool GATK to detect variants, ultimately obtaining parental genotype data files in VCF format. Approximately 1500 offspring individuals were randomly selected from a high-generation cross (MAGIC) population of tobacco parents. Tissue samples, such as young leaves, were collected and sent to the sequencing company for whole-genome resequencing. The sequencing depth was determined by the parental genotype. To maintain consistency, after sequencing, the same processing procedure as the parental data was followed, namely, alignment to the reference genome and variant detection, to obtain VCF format genotype data files for the offspring individuals. Bioinformatics databases were accessed to search for the reference genome sequence and annotation files of the target tobacco crop. The searched sequences and files were downloaded using a download tool. The merge tool was used to merge the VCF files of 8 parents and about 1500 offspring into a unified VCF file containing all individuals and all variant sites. Plink was used to remove individuals and variant sites with excessively high genotype loss rates, and to remove sites located in clustered regions of the genome (such as telomeres and centromeres) or complex regions that are difficult to align, as the genotypes in these regions are usually unreliable. The reference genome information was read using the bioinformatics tool Bedtools to determine the length and start position of the chromosome. Based on a window size of 100kb, the start and end coordinates of each window were calculated. For example, for a chromosome of 1000kb in length, the first window is 1-100kb, the second window is 101-200kb, and so on. The variant sites were assigned to the corresponding windows according to their physical location to form chromosome segments. Each chromosome segment is uniquely defined by its chromosome number and physical start and end coordinates, and contains the genotypic data of all variant sites in all individuals (parents and offspring) within that segment.

[0021] S102, Calculate the parental genetic similarity index of chromosome segments based on chromosome segments, obtain a definite source matrix and a fuzzy source matrix based on the parental genetic similarity index, use definite segments from a single parent in the definite source matrix as anchor points, and examine adjacent segments of fuzzy segments in the fuzzy source matrix starting from the anchor points to determine the parental source. In addition, for points that cannot be determined, randomly select a parent from the fuzzy source matrix as the source. Finally, merge segments with the same parental source and adjacent positions into a large segment to form a parental source matrix.

[0022] Furthermore, based on chromosome segments, genotypic information for all variant sites within each segment of both parents and offspring was extracted from a file containing genotypic data for all parents and offspring. Using Plink 1.9, the IBS status (identical genetic similarity index) was calculated for each offspring individual, comparing its genotype on each segment with that of each parent on the same segment. IBS stands for homologous status, meaning that two individuals share the same DNA segment or alleles. This similarity may be inherited from a common ancestor or result from random matching. IBS status is categorized into three cases: IBS0, IBS1, and IBS2. IBS0 indicates that the offspring and parents do not share any alleles at a specific variant site in the segment; IBS1 indicates that the offspring and parents share one allele at a specific variant site in the segment; and IBS2 indicates that the offspring and parents share two alleles at a specific variant site in the segment. The parental genetic similarity index (DST) was calculated based on the IBS status using the following formula: ,in, It is the genetic similarity index between the parents. IBS1 represents the number of loci with IBS=0, and IBS2 represents the number of loci with IBS=1. For each locus with IBS=2, for each fragment of each progeny, check its DST value calculated with each parent, and select the parent with the highest DST value as the initial source for this fragment of the progeny. If multiple parents have the same and highest DST value for a certain fragment of a certain progeny, then record all these parents as the initial source set for that fragment, indicating that for this fragment, a unique parental source cannot be determined based on the DST value, and there are multiple possible parents. Based on the above steps, organize the initial sources and the initial source set into a definite source matrix and a fuzzy source matrix, respectively. The source matrix consists of two parts: a definite source matrix, which records fragments of parents with a single highest DST value (i.e., definite fragments), and a fuzzy source matrix, which records fragments of parents with multiple highest DST values ​​(i.e., fuzzy fragments). Fragments are selected from the definite source matrix by identifying adjacent fragments as fuzzy fragments in the fuzzy source matrix. The selected definite fragments are used as anchor points. Starting from the anchor point fragment, adjacent fragments are checked in both directions. For each fuzzy fragment, its fuzzy matrix is ​​checked to see if it contains the parent of the anchor point. If it does, the parent source of that fuzzy fragment is determined to be the parent of the anchor point fragment. This is because adjacent fragments... Genetics exhibits a certain continuity. If the possible sources of a fuzzy segment include the parent of the anchor segment, then logically, the fuzzy segment is more likely to originate from that parent. For example, segments 1, 2, and 3 are three consecutive segments on a chromosome. Segment 1 may originate from parents A and B, segment 2 from parent B, and segment 3 from parents C and D. In this case, since segment 2 is a definite segment in the definite source matrix, while segments 1 and 3 are fuzzy segments in the fuzzy source matrix, segment 2 will extend to help determine adjacent segments. Since segment 1 contains parent B, segment 1 will be identified as parent B. Fragment 3 originates from parent B, but does not contain parent B. Therefore, fragment 3 still retains its original parental origin, namely C and D. The process of extending the check to both sides and determining the parental origin of fuzzy fragments is repeated until all fragments that can be clearly identified through adjacency have been processed. After the above processing, there will still be fragments whose unique parental origin cannot be determined, namely the remaining fuzzy fragments. A parent is randomly selected from the fuzzy origin matrix as the source. The chromosomes of each offspring are traversed, and the parental origin information of each fragment is carefully checked. Consecutive fragments with the same parental origin and adjacent physical locations are merged into a large fragment (Bin Marker). By merging adjacent fragments, the complexity of the data can be reduced, and the parental origin situation of a larger region on the chromosome can be better reflected, forming a parental origin matrix. In the matrix, the rows represent offspring individuals, the columns represent Bin Markers, and each value in the matrix represents the parental origin of each offspring at each Bin Marker, represented by the parent's identifier (such as A, B, C...H, etc.).

[0023] S103: Calculate recombination rate and genetic distance based on the merged large fragments, and construct a genetic map based on recombination rate and genetic distance; Specifically, for any two adjacent large segments, the parental combination types appearing in the entire offspring population are counted, and the recombination rate between adjacent loci is calculated using the following formula: in, For parental combination deductions, , For weighted parental combinations, ,in, The number of individuals in the offspring population from a specific parental combination. and It reflects the distribution pattern of parental origin and is used to quantify the similarity or difference of genetic information between adjacent loci. Let r be the recombination rate, which is the probability of exchange occurring during meiosis, and its range is 0 ≤ r ≤ 0.5. This represents the total number of individuals in the offspring population, which is the sum of the number of individuals from all possible parental combinations. , representing the total number of individuals in all 64 combinations. = ,in, For a single individual, there are 64 possible combinations of adjacent Bin Markers, representing parental origin combinations at adjacent loci. Let b1-b8 represent individuals from A1A2 to A1H2. 57 -b 64 Let n1-n8 be the total number of individuals from H1A2 to H1H2 in the offspring population, and n be the number of individuals from A1A2 to A1H2 in the offspring population. 57 -n 64 Let n be the total number of individuals from H1A2 to H1H2 in the offspring population, then n1 = Σb1, n2 = Σb2, ..., n 64 =Σb 64 Then, r can be calculated, where r is the recombination rate between adjacent Bin Markers. Based on the calculated recombination rate, the genetic distance between Bin Markers is calculated using the Haldanemap function, as shown in the formula: in, Genetic distance, To calculate the recombination rate, the total genetic distance of the chromosome is obtained by summing the genetic distances between all adjacent Bin Markers. Genetic maps containing the physical location (start and end coordinates), genetic location, and parental origin information of each Bin Marker are generated according to the above method.

[0024] A specific example is as follows: Assuming the simulated population size is n and the number of markers is m, when simulating the generation of offspring individuals, the offspring population size is set to n = 1500 and the number of markers m is 1000, i.e., 500 SNP loci. The DST value is calculated using Plink 1.9, using plink --bfile. <filename>--genome --out <filename>This allows for the comparison of genotypic similarity between parental and offspring fragments, thereby calculating the DST value. Then, in a Linux system, code to extract this information is compiled to construct the parental origin matrix of the MAGIC offspring population. The resulting parental origin matrix is ​​imported into R for further precise locus determination. After identifying the unique parental origin for all loci, the results are output as a CSV file, imported into an Excel file, and formatted for GAPL. This is then imported into GAPL software for genetic map construction. Method 1: When manually setting recombination loci, the recombination rate between every 100 markers and the next marker is 10%, i.e., there is a 10% probability of recombination, r=0.1. No recombination rate is set for other loci. In Method 1, r=0.1 is substituted into the Haldane map function to calculate the theoretical genetic distance: X=- ln(1-2r) = Since recombination rates were set at 9 loci, the theoretical total genetic distance is... ×9=100.44 The genetic map of the MAGIC population constructed using the new algorithm has a total genetic distance of 101.31. The genetic distance between the recombination sites was 11.26 ± 0.62. The average error between the theoretical value and the actual value is Method 2: When setting recombination sites, the recombination rate between every 100 markers and the next marker is set to 10%, while the recombination rate between other SNPs is set to 0.01%. This allows for testing the accuracy of the algorithm when a small amount of recombination exists in the bin. For calculating the theoretical genetic distance, for every 100 marker intervals, with a recombination rate r = 0.1, the genetic distance calculated using the Haldane map function formula is still approximately 11.16 cM. Assuming there are also 9 such major recombination site intervals, the main part of the theoretical total genetic distance is still approximately 100.44 cM. If all sites are considered more precisely, since the recombination rate of other sites r² = 0.0001, the genetic distance is: X = - ln(1-2r) = The effect on the total genetic distance is minimal and negligible. The genetic map of the MAGIC population constructed by the new algorithm shows a total genetic distance of 107.98. The genetic distance between the recombination sites was 12.00 ± 0.67. , compared with the theoretical value (11.16) The average error is In contrast, the traditional algorithm, after quality control of Method 1, reduced the number of SNP markers from 500 to 103, resulting in a final genetic distance of 58.12 for the constructed genetic map. The error compared to the theoretical total genetic distance (approximately 100.44 cM according to Method 1) reached [a certain value]. For Method 2, after quality control, 184 markers remained after screening from 500 SNP markers, resulting in a total genetic distance of 72.62 cM for the final genetic map construction. This is significantly less than the theoretical total genetic distance (logically approximated to 100.44 cM based on the major recombination sites in Method 2). Therefore, it can be seen that the accuracy of the new algorithm is significantly higher than that of the traditional algorithm.

[0025] The technical solutions in the above embodiments of this application have at least the following technical effects or advantages: the total genetic distance is closer to the real situation, the error in the simulation experiment can be less than 10%, it can more accurately reflect the actual distribution and inheritance law of genes on chromosomes, improve accuracy, and use the new algorithm to analyze the whole genome resequencing data of the MAGIC population of a specified crop (taking tobacco as an example). Without traditional quality control screening, a genetic linkage map with a total genetic distance closer to the theoretical value and more complete locus information is constructed. Combined with neighboring loci, the parental source of the variant locus is identified, providing more variations for the variant locus, thereby effectively detecting the micro-effect genes that control complex quantitative traits, clarifying the parental source of each variant locus, and after calculating the effect, the required high-quality parents can be directly found, thereby improving the breeding process. Combining all variant locus information, covering the whole chromosome, it has strong interpretability.

[0026] Example 2: Example 1 mainly focused on diploid species, but for important economic crops such as wheat, cotton, and rapeseed, which are allopolyploids, a single SNP locus may simultaneously exhibit polymorphism in multiple subgenomes. This makes it impossible to directly determine which parent in which subgenome a variation originates from. Furthermore, many species have presence / deletion variants (PAVs), meaning a sequence is present in some individuals but missing in others. These large structural variations are important genetic resources, but they are easily lost in analyses based on a single reference genome. This example constructs a hierarchical genetic map to precisely locate important quantitative trait loci located in structural variation regions and possessing subgenome specificity, such as... Figure 2 As shown.

[0027] S201. Collect allopolyploid parental data, construct a pan-genome map, screen for existing or missing variant blocks based on the pan-genome map, perform genotyping on the offspring of the MAGIC population, and obtain a three-dimensional genotype matrix. Furthermore, eight representative allopolyploid crop parents were carefully selected from numerous allopolyploid crops. High-depth (>30X) whole-genome sequencing was performed on the selected parents. Bioinformatics tools were used to process the raw sequencing data. The DeBruijn graph algorithm, a pan-genome construction tool, was used to jointly assemble the whole-genome sequencing data of the eight parents. During the assembly process, the k-mer length was set, which determines the size of the basic unit for DeBruijn graph construction and has a significant impact on the quality of the assembly results. Through continuous parameter optimization and complex calculation and analysis, a pan-genome map was finally obtained. Each sequence block in the constructed pan-genome map was precisely aligned with known subgenome reference sequences. Through alignment analysis, the origin of each sequence block in the subgenome was determined, and a corresponding subgenome tag was assigned. After subgenome tagging, a detailed comparative analysis of parental sequence blocks was performed to screen for polymorphic presence or deletion variants (PAVs). Specifically, sequence blocks exhibiting differences in "presence" or "deletion" states between different parents were identified. For example, a sequence block present in parent A but missing in parent B was identified as a polymorphic PAV block. This screening process generated a list of polymorphic PAV blocks between parents. For each PAV block in the pan-genome, its "presence" or "deletion" status was determined in each offspring individual. Specifically, this was done by comparing the resequencing reads of each offspring individual with the alignment results on the pan-genome map (reads refer to MAGIC sequences). When resequencing progeny individuals in the population, their genomic DNA is randomly fragmented into numerous small segments, which are then sequenced. Each such segment is called a read. The coverage of each PAV block in the progeny individuals is calculated. If a PAV block has enough reads that align in the progeny, it is considered to "exist" in that progeny; otherwise, if almost no reads align, it is considered to be "missing". This process is repeated to genotype each progeny individual, resulting in a three-dimensional genotype matrix. The three dimensions of this matrix are: the number of progeny individuals (500), the total number of PAV blocks (determined based on screening results), and the subgenome origin (A, B, D). This three-dimensional genotype matrix comprehensively records the genotype information of each progeny individual in the MAGIC population on PAV blocks from different subgenome origins.

[0028] S202, the three-dimensional genotype matrix is ​​split into subgenome datasets, the parental origin of the subgenome datasets is determined, and collinear strands are constructed based on the parental origin of the subgenome datasets; Specifically, the three dimensions of the three-dimensional genotype matrix are the number of offspring individuals, the total number of PAV blocks, and the subgenome origin. Each element in the matrix represents the genotype information of a specific offspring individual on a specific PAV block from a specific subgenome origin. Data processing software is used to split the three-dimensional genotype matrix, creating three empty datasets (A, B, and D) to store the corresponding subgenome data. By iterating through the third dimension (subgenome dimension) of the three-dimensional genotype matrix, the data corresponding to each subgenome is extracted and stored in the corresponding dataset. Genotype data of eight parents on each PAV block are collected, also represented as "present" or "missing". For each subgenome dataset, step S102 is used to determine the parental origin of each subgenome dataset. Each PAV block is treated as an independent analysis segment. By comparing the state of the PAV blocks in the offspring individuals with the state between the parents, the most likely parental origin is determined. Specifically, for each PAV block in each subgenome dataset, all offspring individuals are iterated through, and for each offspring individual, the PAV block is counted. The genotypic status of the block ("present" or "missing") is compared with the genotypic status of the eight parents in the PAV block. The genotypic similarity between the offspring and each parent is calculated. For example, if the offspring and a parent have the same status in the PAV block (both are "present" or both are "missing"), the similarity is incremented by 1. The parent with the highest similarity is selected as the most likely parental source of the offspring in the PAV block. Using the genome-wide collinearity analysis tool MCScanX, collinearity analysis was performed on subgenomes A, B, and D. By comparing sequence homology and gene order among different subgenomes, the collinearity results were obtained. Based on the collinearity analysis results and the obtained parental origin information, a collinear chain was constructed. Specifically, in each subgenome, PAV blocks were traversed in physical position order. For each PAV block, based on the parental origin information, homologous PAV blocks with the same or similar parental origin in other subgenomes were searched. If PAV block units that are physically adjacent and homologous to each other in different subgenomes were found, they were connected to form a continuous collinear chain. Each chain has corresponding segments on three subgenomes, forming a validation chain. For example, assuming a collinearity dictionary `collinearity_dict` has been obtained, where the key is (subgenome 1, PAV block position 1) and the values ​​are [(subgenome 2, PAV block position 2), (subgenome 3, PAV block position 3)], representing homology information, an empty collinear chain list `collinear_chains` is initialized. All PAV blocks of a subgenome (e.g., subgenome A) are traversed: for the current PAV block, its homologous PAV blocks on other subgenomes are searched (according to `collinearity_dict`). If homologous PAV blocks are found, and the physical adjacency of these homologous PAV blocks on their respective subgenomes with the current PAV block on subgenome A satisfies the collinearity chain formation condition (e.g., consistent positional order), then these PAV blocks are connected to form a new collinear chain and added to `collinear_chains`.

[0029] S203, correct the formed collinear strands, and construct a hierarchical genetic map based on the corrected parental origin matrix; Furthermore, each constructed collinear chain is traversed. Each collinear chain consists of multiple collinear units, and each collinear unit contains three homologous PAV blocks: A, B, and D. For each collinear unit, the PAV block determination of the three subgenomes A, B, and D is checked. If the PAV block determination of a certain subgenome is ambiguous, for example, the possible parental origin of this block is {P1, P2}, while the homologous blocks of the other two subgenomes are clearly and consistently determined, for example, both being P1, then the ambiguous subgenome PAV block is corrected to the clearly determined P1. If an irreconcilable conflict occurs within the collinear unit, for example, the PAV block of subgenome A... Block P1 is identified, and homologous blocks of subgenome B are identified as P2. This collinear unit is then marked as a conflict point. After point correction, for the remaining isolated ambiguous points or conflict points, blocks with clear and high confidence levels on the collinear strand are identified as stable anchor points. For ambiguous points located between two anchor points with the same parental origin, since the parental origin of the anchor points is clear and consistent, the ambiguous point can be safely corrected to have the same parental origin as the anchor points based on topological relationships. For example, if the parental origin of both anchor points is P1, then the ambiguous point in the middle can be corrected to P1. For conflict points, analysis is performed based on their information with the anchor points on both sides, and a higher weight can be assigned to one side (such as the side closer to the parental origin of the anchor point). The judgment result of the conflict point is adjusted according to the weight. If there are large, continuous conflict regions on the collinear strand, these regions need to be precisely defined. These regions are likely "potential genomic recombination hotspots or structural variation boundaries," because a large conflict region may mean that genomic recombination or structural variation has occurred there, leading to PAV. The determination of the parentage of the blocks has become chaotic; Using a high-quality parental origin matrix corrected by the above collinearity chain, which records the parental origin information of PAV markers in each subgenome, the parental origin information of PAV markers in each subgenome (A, B, D) is organized separately. PAV markers in the same subgenome are arranged in a certain order (such as physical location order), and the parental origin corresponding to each marker is recorded to form an organized parental origin matrix. The recombination rate and genetic distance are calculated according to the formula in step S103. Based on the calculated recombination rate and genetic distance data, independent genetic maps of the three subgenomes A, B, and D are constructed respectively. During the construction process, PAV markers are used as nodes and recombination rate or genetic distance is used as edges to connect each marker according to their order and relationship on the genome, forming three independent genetic maps. The three independent genetic maps of the constructed subgenomes are integrated into a hierarchical genetic map, with physical location as the horizontal axis, and the genetic maps of the three genomes A, B, and D are displayed in parallel.

[0030] The technical solutions described in the embodiments of this application have at least the following technical effects or advantages: They enable the construction of genetic maps with structural variations (PAVs) as core markers at the allopolyploid level, capturing functional genetic diversity at a larger scale. Through a collinear chain correction algorithm, the correction logic is upgraded from local voting to global topological reasoning, greatly improving the accuracy of parental origin determination. Point-level and line-level corrections resolve most fuzzy judgment and conflict issues, making the judgment results more reliable. They construct a "hierarchical" high-precision genetic map that can clearly distinguish subgenome origins and uses structural variations as core markers, accurately locating important quantitative trait loci (QTLs) located in structural variation regions and possessing subgenome specificity. This allows for the study of the evolutionary history, recombination patterns, and subgenome interaction mechanisms of allopolyploid genomes, providing a clear blueprint for precision design breeding of polyploid crops.

[0031] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.< / filename> < / filename>

Claims

1. A novel algorithm for constructing genetic maps of MAGIC populations, characterized in that, include: S101, Collect diploid species data, and perform chromosome segmentation based on the collected data to form chromosome fragments. The collected diploid species data includes parental genotype data, MAGIC offspring genotype data, and reference genome data. Among them, the species include diploid species and allopolyploid species. S102, Calculate the parental genetic similarity index of chromosome segments based on chromosome segments, determine the parental origin based on the parental genetic similarity index, obtain a definite origin matrix and a fuzzy origin matrix, take the definite segments with a single parental origin in the definite origin matrix as anchor points, and examine the adjacent fuzzy segments with the anchor points as the starting point to determine the parental origin of these segments. In addition, for points that cannot be determined, randomly select a parent from the fuzzy origin matrix as the source. Finally, merge segments with the same parental origin and adjacent positions into a large segment to form a parental origin matrix. S103, Calculate recombination rate and genetic distance based on the merged large fragments, and construct a genetic map based on recombination rate and genetic distance, wherein the genetic map is a genetic map of a diploid species.

2. The novel algorithm for constructing a genetic map of a MAGIC population as described in claim 1, characterized in that, The steps for chromosome segmentation are as follows: read reference genome information, determine the length and start position of the chromosome, calculate the start and end coordinates of each window according to the preset window size, and assign the mutation sites to the corresponding windows according to their physical locations to form chromosome segments.

3. The novel algorithm for constructing a genetic map of a MAGIC population as described in claim 1, characterized in that, The method for determining parental origin is as follows: Genotypic information of all variant sites within each chromosome segment of both parents and offspring is extracted; homology between offspring and each parent on the same chromosome segment is calculated; and a parental genetic similarity index is calculated based on this homology. The formula for calculating the parental genetic similarity index is: DST = Wherein, DST is the genetic similarity index between the parents, IBS0 is the number of loci with IBS=0, IBS1 is the number of loci with IBS=1, and IBS2 is the number of loci with IBS=2. The parent with the highest DST value is selected as the initial source of the offspring on this fragment, and the initial sources form a clear source matrix.

4. The novel algorithm for constructing a genetic map of a MAGIC population as described in claim 3, characterized in that, The method for determining the parental origin further includes: when multiple parents have the same DST value and are all the highest values, the multiple parents are recorded as an initial source set to form a fuzzy source matrix, and the parental origin is determined by anchor point extension and random parent designation.

5. The novel algorithm for constructing a genetic map of a MAGIC population as described in claim 1, characterized in that, The formula for calculating the recombination rate is: in, For parental combination deductions, , For weighted parental combinations, ,in, The number of individuals in the offspring population from a specific parental combination. The recombination rate is defined as 0 ≤ r ≤ 0.

5. This represents the total number of individuals in the offspring population, which is the sum of the number of individuals from all parental combinations. , representing the total number of individuals in all 64 combinations. = ,in, It refers to the combination of parental sources at adjacent loci for a single individual.

6. The novel algorithm for constructing a genetic map of a MAGIC population as described in claim 1, characterized in that, The formula for calculating genetic distance is: in, Genetic distance, This represents the recombination rate.

7. The novel algorithm for constructing a genetic map of a MAGIC population as described in claim 1, characterized in that, It also includes the step of constructing an allopolyploid species map: S201. Collect allopolyploid parental data, construct a pan-genome map, screen for existing or missing variant blocks based on the pan-genome map, perform genotyping on the offspring of the MAGIC population, and obtain a three-dimensional genotype matrix. S202, the three-dimensional genotype matrix is ​​split into subgenome datasets, the parental origin of the subgenome datasets is determined, and collinear strands are constructed based on the parental origin of the subgenome datasets; S203, the formed collinear strands are corrected, and a hierarchical genetic map is constructed based on the corrected parental origin matrix. The hierarchical genetic map is the genetic map of allopolyploid species.

8. The novel algorithm for constructing a genetic map of a MAGIC population as described in claim 7, characterized in that, The three dimensions of the three-dimensional genotype matrix are the number of offspring individuals, the total number of variant blocks present or absent, and the subgenome origin.

9. The novel algorithm for constructing a genetic map of a MAGIC population as described in claim 7, characterized in that, The method for determining the parental origin of a subgenome dataset is as follows: For each present or missing variant block in each subgenome dataset, all offspring individuals are traversed. For each offspring individual, the genotype status of that present or missing variant block is counted. The genotype status of the offspring individual is compared with the genotype status of the parents in that present or missing variant block. The genotype similarity between the offspring individual and each parent is calculated. The parent with the highest similarity is selected as the parental origin of the offspring individual in that present or missing variant block.

10. The novel algorithm for constructing a genetic map of a MAGIC population as described in claim 7, characterized in that, The method for correcting the formed collinear chains is as follows: the correction includes point correction and line correction. The point correction involves traversing the collinear chains and units, checking the consistency of the presence or deletion variant blocks of the three subgenomes within the collinear unit, correcting fuzzy judgments, and marking irreconcilable conflicts as conflict points. The line correction involves finding anchor points for the remaining isolated fuzzy points or conflict points after the point correction, reasoning and correcting the fuzzy points and conflict points according to the topological relationship, and defining a large-scale continuous conflict region.