Multi-parent advanced generation inter-cross (mags) breeding for crop improvement
By using high-throughput whole-genome sequencing data and sliding window technology, genotype maps of multi-parental plants were constructed, solving the problems of accuracy and efficiency in genotype identification of multi-parental plants and achieving rapid and accurate genotype identification and genetic map construction.
Patent Information
- Application Number
- CN202110131330.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-30
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-01-30
AI Technical Summary
Existing genotyping methods suffer from low accuracy and time-consuming analysis in multi-parental plants, especially when the genotyping results are poor for plants with three or more parents.
Using a method based on high-throughput whole-genome sequencing data, we constructed SNP site information and recombination breakpoint analysis, combined with sliding window technology, to quantify the similarity between offspring and each parent, and constructed a genotype map of multi-parental plants, thus achieving rapid and accurate genotype identification.
It enables rapid and accurate identification of genotypes in multi-parental plants, improving the precision and efficiency of genotype identification, and is applicable to the construction of genetic maps and molecular design breeding for more crops and species.
Smart Images

Figure BDA0002925409000000101 
Figure BDA0002925409000000102 
Figure BDA0002925409000000103
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bio-information processing, in particular to multi-parental crop genotyping based on high-throughput whole genome sequencing. More particularly, the present application provides a method and device for identifying the genotype of multi-parental crops based on high-throughput whole genome sequencing data. BACKGROUND
[0002] At the end of last century, the use of DNA molecular markers greatly promoted the development of reverse genetics. With the progress of molecular biology technology, the types of markers and the methods of constructing genetic maps are gradually developing and improving. The advent of polymerase chain reaction (PCR) triggered an era of explosive application of molecular markers, because PCR can greatly simplify the experimental steps of marker design and result analysis. These DNA molecular markers are still widely used, but they also show more and more limitations in genome coverage, time consumption and cost. At present, the development of genomics and the gradual maturity of related technical methods provide the basis for replacing the marker-based mapping method with the high-throughput strategy based on genome.
[0003] Genomic sequences opened the door to high-throughput genotyping. Initially this was done using microarray chip technology, i.e. by hybridizing genomic DNA with oligonucleotides on gene chips to detect single nucleotide polymorphisms (SNPs). Since hundreds to thousands of markers can be detected in one hybridization, this method of genotyping greatly improves efficiency [1] . This method has been applied to some model biological systems such as humans, Arabidopsis and rice [2-4] . Although the goal of high-throughput has been achieved, the microarray-based method still has serious limitations, such as laborious, time-consuming, and high cost in the process of designing, producing and using microarrays.
[0004] The advent of second-generation sequencing technology has brought about a leap in methodology for genotyping and genetic mapping. The new sequencing technology not only increases the sequencing throughput by several orders of magnitude, but also allows the parallel sequencing of many samples [5-6] . The progress of these technologies has paved the way for the development of sequencing-based high-throughput genotyping methods. The new genotyping method combines the following advantages: fast and inexpensive, high-density marker coverage, high accuracy and high resolution, and is also suitable for more mapping populations and comparative genomics and genetic map construction between species.
[0005] Although there are some methods available for genotyping of 2 parent plants, there are obvious deficiencies in current methods for genotyping of polyparental plants involving 3 or more parents, such as low accuracy, long analysis time, etc.
[0006] Therefore, there is an urgent need in the art to provide a method and device for genotyping of polyparental plants involving 3 or more parents with fast analysis and accurate results. SUMMARY
[0007] The object of the present application is to provide a method and device for genotyping of polyparental plants with fast analysis and accurate results, so that the genotypes of polyparental constructed populations can be analyzed quickly, accurately and reliably.
[0008] In a first aspect of the present application, a method for genotyping of polyparental plants (such as crops) is provided, characterized in that the method comprises:
[0009] (a) providing sequencing data Df of a progeny plant to be genotyped and sequencing data Dp of the corresponding parent plants of the progeny plant for n parents and their progeny, wherein n is a positive integer ≥ 3;
[0010] (b) determining SNP site information of the parents and the progeny based on the sequencing data Df and the sequencing data Dp;
[0011] (c) judging the genotype of the progeny based on the SNP site information, thereby obtaining the evaluation results of each SNP of the progeny and the distribution information of recombination breakpoints on each chromosome of the whole genome of the progeny;
[0012] (d) constructing and / or mapping the genotype map of the progeny based on the SNP evaluation result information and the location information of the recombination breakpoints of the whole genome, thereby obtaining the genotyping results of the polyparental plants.
[0013] In another preferred embodiment, in step (c), the analysis of recombination breakpoint is based on SNP "string".
[0014] In another preferred embodiment, step (c) comprises analyzing the recombination breakpoints, thereby obtaining the analysis results of the recombination breakpoints,
[0015] and the analysis of the recombination breakpoints comprises:
[0016] (s1) constructing SNP "string", wherein the genotypes of all SNPs on each chromosome of the parents and the progeny are compressed in order into a string;
[0017] (s2) determining each sliding window corresponding to the SNP string according to a predetermined window size, and scoring each SNP site in each window, thereby obtaining a respective score value P for each parent within the window;
[0018] (s3) determining the genotype of each chromosomal region corresponding to the offspring based on the score value P obtained in step (s2).
[0019] In another preferred embodiment, in step (sl), all gaps between SNPs are removed regardless of the actual distance between the two adjacent SNPs.
[0020] In another preferred embodiment, in step (sl), the SNPs constituting the string are homozygous SNP sites of the parents.
[0021] In another preferred embodiment, in step (sl), the SNP sites included in the analysis are pre-screened, thereby excluding any SNP site in which either parent is heterozygous.
[0022] In another preferred embodiment, in step (s2), the scoring is performed according to the scoring rules of Table A.
[0023] In another preferred embodiment, in step (s3), the genotype of each chromosomal region corresponding to the offspring is determined based on the score value or score curve of each parent.
[0024] In another preferred embodiment, in step (s3), the genotype of each chromosomal region is determined based on the score value and standard deviation.
[0025] In another preferred embodiment, for a chromosomal region whose genotype is to be determined, if one parent A has a high score value (≥ 80% of the full score, preferably ≥ 80% of the full score) close to the full score, and the score value of this parent in this region is relatively stable without too much fluctuation, while the score values of the remaining parents are low (≤ 50% of the full score, preferably ≤ 30% of the full score) or have large fluctuations, then the genotype of the chromosomal region is determined as the genotype of the parent A.
[0026] In another preferred embodiment, in step (s3), by sliding the sliding window over the whole genome SNP sites, the score value of each parent on each chromosome is obtained, and the score curve of each parent is plotted with the position of each sliding window on the chromosome as the abscissa and the score value as the ordinate.
[0027] In another preferred embodiment, in step (s3), the sub-step of evaluating the heterozygous region is included.
[0028] (s3a) For the hybrid offspring of multiple parents, the parent origin of a certain chromosomal region is at most two parents, and whether this region is a heterozygous region is determined based on the score curves of the two parents in this region.
[0029] In another preferred embodiment, in step (s3), the degree of similarity of the offspring to each parent in this region is quantified, and the genotype of each region is determined according to the numerical characteristics (numerical value and standard deviation) of the score curve of each parent.
[0030] In another preferred embodiment, in step (s3), the genotype is determined in the following manner:
[0031] (Z1) If a certain parent has a higher score value in this region and its score value is stable (the score curve is close to the plateau), and the score curves of the remaining parents in this region have large fluctuations and large standard deviations (the score curve is a wave-shaped up and down), it is determined that this region is the homozygous genotype of this parent;
[0032] (Z2) When the number of parents is determined to be two, it can be inferred that this region is the heterozygous genotype of the two parents according to the fact that both of them have large numerical fluctuations and large standard deviations in this region;
[0033] In the case of multiple parents, it is possible to only find possible heterozygous regions, i.e., there is no parent with a high score value and stable in a certain region; because the scores of each parent in this region have large fluctuations, only the two most likely parents can be given according to the numerical characteristics.
[0034] (Z3) If two or more parents and the offspring are very similar in a certain region, and two or more parent curves with high scores and small standard deviations appear, it means that the analyzed several parents are very similar in this region and have little difference, so it can be temporarily determined (marked as "unknown region").
[0035] In another preferred embodiment, the method further comprises: if the genotypes on both sides of the unknown region are the same, determining this region as this genotype; and if the genotypes on both sides are different, regarding the middle position of this unknown region as a recombination breakpoint, and the unknown region on both sides is the genotype on both sides.
[0036] In another preferred embodiment, the offspring is a multiple-parent plant.
[0037] In another preferred embodiment, n is 3-6, more preferably 3, 4 or 5.
[0038] In another preferred embodiment, the sequencing data is selected from the group consisting of genomic sequencing data, RNA sequencing data, or a combination thereof.
[0039] In another preferred embodiment, the sequencing data is in fastq format.
[0040] In another preferred embodiment, the size of the sliding window is 170-500 consecutive SNP loci, preferably 200-400 consecutive SNP loci.
[0041] In another preferred embodiment, the sequencing depth of the sequencing data is 0.1x-10x, preferably 0.2x-5x.
[0042] In another preferred embodiment, the sequencing depth of the sequencing data is ≥1, preferably 1-5, more preferably 1.5-3.
[0043] In another preferred embodiment, for each chromosome, a respective parent score curve is obtained.
[0044] In another preferred embodiment, the SNP loci are used to determine the genotype of the individual.
[0045] In another preferred embodiment, in step (b), the sequencing data (e.g., fastq files) are aligned, processed by bwa and GATK software, and SNP information is obtained.
[0046] In another preferred embodiment, the SNP loci information includes position information and genotype information.
[0047] In another preferred embodiment, the SNP loci used to determine the genotype meet the following requirements:
[0048] I. The SNP loci cover the whole genome as much as possible and do not have missing regions in certain areas.
[0049] II. For any SNP loci, the SNP information (position information and genotype information) of the two parents and the offspring is known, and any one of the three cannot be known for the loci and should be deleted.
[0050] In another preferred embodiment, in step (c), the evaluation results of each SNP of the offspring are recorded in an rlt file, and the rlt file records the genotype determination of each SNP position.
[0051] The distribution information of recombination breakpoints on each chromosome of the whole genome of the offspring is recorded in a bin file, and the bin file records the distribution of recombination breakpoints on the 12 chromosomes of the whole genome.
[0052] In another preferred embodiment, in step (c), the SNPwindow script is used to read the genotype and determine the recombination breakpoint.
[0053] In another preferred embodiment, in step (d), the genotype map of m individuals of the offspring is simultaneously determined.
[0054] In another preferred embodiment, in step (d), the construction of recombination maps is performed by SNPwindow script, and the genetic map of each progeny individual is drawn by SNP2png script.
[0055] In another preferred embodiment, in step (d), the recombination map of each individual is also aligned to generate a recombination bin map by Bin2MCD script.
[0056] In another preferred embodiment, the resolution of the recombination bin map is 5-200 kb per bin, preferably 10-100 kb per bin.
[0057] In another preferred embodiment, the method further comprises processing the recombination bin map to obtain the genetic map of the progeny.
[0058] In another preferred embodiment, the method further comprises performing QTL analysis on the genetic map.
[0059] In another preferred embodiment, the method further comprises visualizing the genotypes of the entire population of parents and progeny, generating genotype data, and constructing a linkage map based on the genotype data.
[0060] In another preferred embodiment, the plant comprises a crop, preferably a crop of the family Poaceae.
[0061] In another preferred embodiment, the crop comprises rice, wheat, soybean, tobacco.
[0062] In a second aspect of the present application, a data analysis device for identifying the genotype of a multi-parent plant is provided, the device comprising:
[0063] a data input module for inputting the to-be-analyzed processed data, the processed data comprising sequencing data Df of a to-be-identified progeny plant and sequencing data Dp of a parent plant corresponding to the progeny plant;
[0064] a multi-parent plant genotype identification module configured to perform the method described in the first aspect of the present application to obtain the genotype identification result of the progeny;
[0065] and an output module for outputting the genotype identification result of the progeny.
[0066] In another preferred embodiment, the multi-parent plant genotype identification module comprises:
[0067] a SNP site information analysis submodule configured to determine SNP site information of the parents and the offspring based on the sequencing data Df and the sequencing data Dp;
[0068] a chromosome recombination breakpoint analysis submodule configured to determine the genotype of the offspring based on the SNP site information, thereby obtaining the evaluation results of each SNP of the offspring and the distribution information of the recombination breakpoints on each chromosome of the whole genome of the offspring;
[0069] a genotype map construction submodule configured to construct and / or draw a genotype map of the offspring based on the SNP evaluation result information and the position information of the whole genome recombination breakpoints of the offspring, thereby obtaining the genotype identification result of the multiple-parent plant.
[0070] In another preferred embodiment, the plant comprises a crop, preferably a crop of the family Poaceae.
[0071] In another preferred embodiment, the output module comprises a display, a printer, a pad, or the like.
[0072] It should be understood that, within the scope of the present application, each of the technical features described above and each of the technical features described in detail below (such as the embodiments) can be combined with each other to form new or preferred technical solutions. Due to the limited space, they will not be listed one by one here. BRIEF DESCRIPTION OF DRAWINGS
[0073] Figure 1 The whole genome recombination breakpoints of a two-parent simulated material are shown.
[0074] Figure 2 The whole genome recombination breakpoints of a four-parent simulated material are shown.
[0075] Figure 3 The genotype identification of a two-parent simulated offspring using the SNP-based sliding window method is shown.
[0076] Figure 4 The genotype identification of a two-parent simulated offspring using the SEG-Map software method is shown.
[0077] Figure 5 The genotype identification of a four-parent simulated offspring using the SNP-based sliding window method is shown.
[0078] Figure 6 The influence of different sliding window sizes on the accuracy of genotype identification results is shown.
[0079] Figure 7 The influence of different sequencing depths on the accuracy of genotype identification results is shown.
[0080] Figure 8 An analysis framework flow is shown for SNP-based sliding window genotype identification.
[0081] Figure 9 A genetic map is shown drawn using the SNP2png script.
[0082] Figure 10 A genotype identification set plot is shown for a rice population.
[0083] Figure 11 A genotype table is shown for a recombination segment map of a recombinant inbred line individual in one embodiment.
[0084] Figure 12 A SNP "string" of window size 15 is shown.
[0085] Figure 13 Four parent score curves are shown for simulated progeny of a rice chromosome 3.
[0086] Figure 14 Two parent score curves are shown for simulated progeny of a rice chromosome 11.
[0087] Figure 15 Parent score cases are shown for a three parent homozygous genotype call in one embodiment.
[0088] Figure 16 Parent score cases are shown for a two parent heterozygous genotype call in one embodiment.
[0089] Figure 17 Parent score cases are shown for a genotype call of unknown in one embodiment.
[0090] Figure 18 Subsequent genotype calls for unknown regions are shown in one embodiment.
[0091] Figure 19 A genotype identification plot is shown for a single individual in a DH population.
[0092] Figure 20 A genotype identification set plot is shown for a DH population.
[0093] Figure 21 Genotype identification is shown for a three parent material in one embodiment.
[0094] Figure 22 SEG-Map identification is shown for a three parent material in one embodiment.
[0095] Figure 23 Genotype identification is shown for a four parent simulated material in one embodiment.
[0096] Figure 24 The true genotypes of four parent analog materials in one embodiment are shown. DETAILED DESCRIPTION
[0097] Through extensive and in-depth research, the inventors of the present application have first developed a more rapid and accurate method for genotype identification, thereby achieving more effective genetic mapping and genome analysis. The method of the present application is particularly suitable for genotype analysis and identification of a multi-parent population with low-coverage sequencing. In the present application, the true SNP genotype information of a certain segment of a multi-parent and offspring is directly read, and then the similarity of the offspring to each parent in this segment is quantified, and according to the numerical characteristics (numerical value and standard deviation) of the score curve of each parent, a highly efficient, simplified and accurate method for genotype identification of a multi-parent plant (or multi-parent crop) is formed. On this basis, the present application is completed.
[0098] Specifically, the inventors of the present application have developed a high-throughput method for identifying the genotype of a recombinant population containing multiple parents based on whole-genome low-coverage sequencing data generated by second-generation sequencing technology. The inventors of the present application have designed a "sliding window" method to determine the genotype of this segment by comprehensive analysis of the genotypes of multiple single nucleotide polymorphisms (SNPs) in a local region of the genome, and further determine the specific position of the recombination breakpoint to construct a fine recombination map of a multi-parent population.
[0099] To verify this method, the inventors of the present application constructed simulated whole-genome sequencing data of a two-parent population and a multi-parent population, constructed a genetic linkage map using the method, and finally compared the identified genotype information with the true genotype of the simulated data. The genotype identification accuracy of the two-parent population can reach 89.61%, which is similar to the accuracy of the SEG-Map software method developed by the inventors of the present application for identifying the genotype of a two-parent population (the accuracy of the SEG-Map method is 89.32%). The genotype identification method newly developed by the inventors of the present application has an identification accuracy of 92.10% for a multi-parent population, which cannot be achieved by the SEG-Map software or method.
[0100] The method of the present application can effectively and rapidly analyze the genotype of each individual in the population, play a key guiding role in genome design breeding, and also provide rapid and accurate genotype data for QTL positioning of different crop multi-parent populations. At the same time, the inventors of the present application have tested the method using a real rice RIL genetic population, and used high-throughput sequencing-based genotype identification, and finally also obtained a high-precision recombination map.
[0101] Therefore, with the continuous development and improvement of sequencing technology, this genotype identification method based on low-coverage sequencing of genome can replace the traditional marker-based genotype identification method, and provide a powerful tool for large-scale genetic exploration research and solving more complex biological problems. The method of the present application is more suitable for genotype identification of multi-parent backcross populations that have undergone low-coverage sequencing, providing accurate genotype support for QTL mapping, and also helping molecular design breeding application of multi-parent populations.
[0102] Terms
[0103] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0104] As used herein, the term "comprising" or "including" can be open, semi-closed and closed. In other words, the term also includes "consisting essentially of" or "consisting of".
[0105] As used herein, the term "double parent" refers to 2 parents.
[0106] As used herein, the term "multi-parent" refers to 3 parents and more.
[0107] As used herein, the term "multi-parent plant" refers to a plant involving 3 parents and more, such as a progeny plant (e.g. a crop) involving 3, 4 or 5 parents.
[0108] Method for identifying genotype of multi-parent crop
[0109] The present application provides a method for identifying genotype of multi-parent crop. The method of the present application is a SNP site sliding window-based genotype identification method.
[0110] In the SNP site sliding window-based genotype identification method of the present application, the data processing is optimized. The optimized process can directly analyze and process the single-end or double-end short sequence sequencing results generated by the second-generation sequencing technology, and finally construct the genetic map of the recombinant population.
[0111] For a mapping population derived from two parents, the whole genome SNPs of the two parents need to be identified before the data analysis pipeline. The SNPs can be identified by high coverage whole genome sequencing, by the existing genome SNP information in the rice haplotype map, or by low coverage whole genome sequencing combined with missing genotype (SNP) imputation. Since the SNPs between the two parents can be identified by a fast and cost-effective approach, the sequencing-based genotype identification of a recombination population will mainly rely on the subsequent analysis, including read genotyping, recombination breakpoint determination, and genetic linkage map construction.
[0112] Functions, steps, and software (scripts) in data analysis can be found in Figure 8 .
[0113] The first step contains several tasks that can be processed simultaneously. A certain number of individuals in the recombination population and parental materials are subjected to second-generation high-throughput sequencing simultaneously. The obtained fastq files are aligned and processed by bwa and GATK software to obtain high-quality SNP information.
[0114] The SNP sites used for final genotype determination should meet the following requirements:
[0115] I. The SNP sites should cover the whole genome as much as possible and should not be missing in certain regions.
[0116] II. For any SNP site, the SNP information (position information and genotype information) of the two parents and the simulated offspring is known, and any of the three should be deleted if the site is unknown.
[0117] In addition, taking rice as an example, it is generally believed that rice parents are self-pollinated inbred lines, and there are basically no heterozygous sites in the genome, so if a heterozygous SNP site is found in the parents, it is generally considered that the site is not reliable, and therefore any SNP site with heterozygous parents can be deleted.
[0118] After screening the high-quality whole genome SNPs of the two parents and the offspring, a python script SNPwindow can be used to determine the genotype of the offspring. The script output will have two files, rlt file and bin file. The rlt file records the genotype determination of each SNP position, and the bin file records the distribution of recombination breakpoints on the 12 chromosomes of the whole genome. These two files are important basis for subsequent mapping and linkage analysis.
[0119] See Figure 9Generally, a perl script SNP2png can be used to draw a genotype map from rlt and bin files. The map is drawn according to the genotype information of the determined SNP sites and the position information of the whole genome recombination breakpoints, and different colors in the map represent different genotype types.
[0120] In actual work, a large number of offspring populations need to be processed, so the SNP information of each individual can be extracted, and then the SNPwindow script analysis is used to test, read the genotype, determine the recombination breakpoint, and construct the recombination map. The SNP2png is used to draw the genotype map of each individual, and the genotype of the individual is grasped as a whole, and then the script Bin2MCD is used to mate the recombination maps of all individuals to generate the recombination bin map.
[0121] Referring to Figure 10 A perl script can also be used to visually analyze the genotype of the whole population. The programs and scripts used in the analysis process are represented in italics, forming a series of analysis steps. The genotype data generated at the end of the analysis process can be directly used in other software (including MapMaker and JoinMap) to construct linkage maps.
[0122] The output data generated by the analysis software is the recombination bin, and the resolution is usually 100 kb per bin or even 10 kb per bin. The genotype results of the mapping population can be imported into MapMaker
[16] or JoinMap
[32] and other programs to construct a genetic map. With the genetic map, QTL analysis is performed.
[0123] Referring to Figure 11 The genetic map generated by the method of the present application is much more precise than the scale of the map generated by most traditional molecular markers.
[0124] The method of the present application involves determining recombination breakpoints, and the detailed process includes the following steps:
[0125] Step 1: Constructing SNP "string".
[0126] For example, when the window size is 15 (win=15). The genotypes of all SNPs on the 12 chromosomes of the two parents and the offspring are compressed into a string in order. Regardless of the actual distance between the two adjacent SNPs, all the gaps are removed.
[0127] For example, 12 chromosomes, the SNPs on the 12 chromosomes become 12 continuous strings (seeFigure 12 ). Blue color represents the genotype of parent 1, red color represents the genotype of parent 2, for a child of one parent, there are three possible genotypes for each SNP site, homozygous genotype of parent 1 (blue), homozygous genotype of parent 2 (red) and heterozygous genotype (yellow).
[0128] It is generally believed that the genome of artificially bred rice parent materials is highly homozygous, and for some multi-generation self-cross recombination rice populations, the genome is also relatively homozygous, only some heterozygous regions exist in part of the chromosome positions. Therefore, the SNP sites included in the analysis are first screened manually, and the SNP sites of which any parent is heterozygous are excluded, such sites cannot be accurately scored. In addition, if the sequencing depth of the child is not very high, the SNP site of which the child is heterozygous can be filtered out, because the heterozygous site judged according to the low depth is not highly reliable, and it is very likely to be a misjudgment caused by sequencing error.
[0129] Second step: scoring two parents in a window
[0130] According to Mendelian inheritance law, the scores of all SNP sites in a sliding window are calculated, and the total score of each parent is calculated as the score of the parent in the chromosome position of the sliding window. The scoring rule is shown in Table A or similar scoring rules. Preferably, the scoring rule is formulated according to the genetic law of the organism.
[0131] The genotype scoring rules in the present application are further combined with the theoretical model as follows.
[0132] For a sliding window composed of consecutive SNPs, the score value of the child to any one parent is composed of three parts: 1. the SNP site of the child identical to the parent 2. the site of the child different from the parent but in accordance with Mendelian inheritance law 3. the site of the child different from the parent and not in accordance with Mendelian inheritance law and the misjudgment site caused by various possible factors.
[0133] For a single parent A, the number of SNP sites of the child to be detected identical to the parent is m, the number of SNP sites of the child different from the parent but in accordance with Mendelian inheritance law is n, and the number of SNP sites of the child different from the parent and not in accordance with Mendelian inheritance law and the misjudgment site caused by various possible factors is e.
[0134] Then the score value S of the parent is A :
[0135] S A = s1*m + s2*n + s3*e
[0136] s1 is the score value of the SNP site which is the same as the parent.
[0137] s2 is the score value of the SNP site which is different from the parent but in accordance with Mendelian inheritance.
[0138] s3 is the score value of the SNP site which is different from the parent and not in accordance with Mendelian inheritance and the site caused by various possible factors.
[0139] In a continuous SNP frame with a size of N, there are i parents to be determined, for a given SNP site at the chromosome position k, the genotype of the offspring and the parent are g k and g' k The genotype of the gene of the rice pure line parent is generally 0 / 0, 0|0, 1 / 1, 1|1. The genotype of the offspring is generally 0 / 0, 0|0, 1 / 1, 1|1, 0 / 1, 0|1. The frequency of the allele of 2 is generally less, which is not considered for the time being.
[0140] For the i-th individual parent, the probability that the offspring is in accordance with its genotype is:
[0141]
[0142] The likelihood of the offspring belonging to the parent in this region is calculated using the Bayesian method:
[0143]
[0144] The genotype of the offspring in a certain region is determined to be the genotype of the parent with the highest probability of coincidence, that is, the maximum value of the coincidence probability in i parents is calculated:
[0145] P max =max{P1,P2,...,P i}
[0146] Preferably, in the present application, the following table is used for genotype scoring.
[0147] Table A Genotype Scoring Rule Table
[0148]
[0149] According to the scoring rules, s1=1, s2=1, and s3=0.
[0150] If the parent A is scored, the score is:
[0151]
[0152] According to this method, the score value of each parent is counted respectively, and then the sliding window calculation standard deviation std of the continuous parent score value is carried out.
[0153] Preferably, the following conditions are met when judging a certain region of a chromosome to be a homozygous genotype of parent A: the score S is the highest, and the standard deviation is the smallest.
[0154] Step 3: judging the genotype of a region of a chromosome according to the score value
[0155] By sliding the window on the whole genome SNP site, the score value of each parent on each chromosome can be obtained. The score value is taken as the vertical coordinate, and the position of each sliding window on the chromosome is taken as the horizontal coordinate to draw the score curve of each parent.
[0156] The genotype of each region of a chromosome is judged according to the characteristics of the score curves of different parents.
[0157] In an example, for example, see Figure 13 The sliding window scoring of the simulated offspring of four parents is shown, and the score curves of the four parents are drawn according to the score values of the four parents.
[0158] By observing the score curves of the four parents, in an example, it can be seen that the parent curve has two different distribution modes of a stable period and a fluctuation period in different regions of chromosome 11. Therefore, the genotype of the offspring in the region is judged by using the state of the curve of different parents in the same region.
[0159] As shown in the region of 1bp to 10780000bp in the figure, it can be observed that the score of the yellow curve (parent 4) in this region has a high score value close to full score, and the score value of this parent in this region is quite stable, without too much numerical fluctuation. The fluctuation of the score value is measured by the standard deviation in statistics. In this region, parent 4 has a very high score value and a small standard deviation in this region, and the score values of the other three parents in this region fluctuate in the range of 0-200, with a high standard deviation. Therefore, it can be judged that this region is the homozygous genotype of parent 4. By using a similar method, the genotype of the offspring in different regions of 12 chromosomes can be judged according to the score value.
[0160] In the figure, the rectangular bars corresponding to true correspond to the true genotype information of the simulated offspring of each chromosome segment, and the rectangular bars corresponding to judge represent the genotype of the offspring judged by the method of the present application. The information of the two is basically consistent.
[0161] Judgment of heterozygous regions
[0162] The judgment of the heterozygous region is illustrated by judging the genotype of the simulated offspring of the two parents. According to the genetic principle, even if the hybrid offspring is derived from multiple parents, the parent source of a certain chromosomal region is at most two parents. Therefore, whether this region is a heterozygous region can be judged according to the score curve of the two parents in this region.
[0163] In one example, as shown in Figure 14 , the genotype of the simulated offspring is identified, and the rectangular bar corresponding to true corresponds to the true genotype information of the simulated offspring of each chromosomal segment, and the rectangular bar corresponding to judge indicates the genotype of the offspring judged by the method of the application. Similarly, when the score of one parent is higher and the standard deviation is smaller, and the score of the other parent has considerable fluctuation and the standard deviation is larger, it is judged that this region is the homozygous genotype of the former (orange or blue region in the figure).
[0164] When the score values of the two parents in a certain segment (6000000bp to 15000000bp in the figure) are not much different from each other, and both have certain fluctuations and large standard deviations, this segment is judged to be the heterozygous genotype of the two parents (yellow region in the figure). This judgment result is also consistent with the true information.
[0165] Judgment based on quantified similarity
[0166] One of the core ideas of the method of the application is to directly read the true SNP genotype information of the parents and the offspring in a certain segment, then quantify the similarity of the offspring and each parent in this segment, and form a relatively simplified analysis model according to the numerical characteristics (numerical value and standard deviation) of the score curve of each parent, and then judge the genotype of each segment.
[0167] The judgment criteria of the application mainly include the following cases:
[0168] I. If the score value of a certain parent in the segment is higher and its score value is more stable (the score curve is close to the platform period), and the score curves of the remaining parents in the segment have greater fluctuations and larger standard deviations (the score curve is a wave peak shape with ups and downs), it is judged that this segment is the homozygous genotype of the parent Figure 15 ).
[0169] II. When the number of parents is two, according to the fact that both of the two parents have greater numerical fluctuations and larger standard deviations in a certain segment, it can be inferred that this region is the heterozygous genotype of the two parents Figure 16 ).
[0170] In the multiple parent determination, it is possible to only find the possible hybrid segment, i.e. no high score and stable parent is found in a certain segment. Because the score of each parent has a large fluctuation in the segment, only the two most possible parents can be given according to the numerical characteristics.
[0171] III. Because the determination method largely depends on the measurement of the genotype similarity between the offspring and the parents, if two or more parents are very similar to the offspring in a certain segment, two or more parent curves with high scores and small standard deviations appear, which indicates that the several parents analyzed in this segment are very similar and have no big difference, and the determination can be temporarily not made.
[0172] Referring to Figure 17 In some cases, the region to be determined is defined as "unknown" first, and the genotype of the region is determined by the genotypes of the regions on both sides.
[0173] Referring to Figure 18 If the genotypes on both sides are the same, the region is determined as this genotype, and if the genotypes on both sides are different, the middle position of the region is regarded as a recombination breakpoint, and the regions on both sides are the genotypes on both sides, respectively.
[0174] Second sliding window for genotype determination
[0175] In the present application, preferably, the genotype determination is performed by a second sliding window.
[0176] The genotype of the SNP is first determined by sliding window, and the parent score value in each window is counted.
[0177] The parent score value obtained is secondly determined by sliding window, and the height and standard deviation of the score value are detected.
[0178] The final genotype determination depends on the height and standard deviation of the score value obtained by the second sliding window, and the genotype determination is performed by the highest probability of a certain region of the offspring belonging to a certain parent.
[0179] In the present application, the genotype determination by the second sliding window can be faster and more accurate.
[0180] An example of the second sliding window is shown as follows:
[0181]
[0182] A device for identifying the genotype of a multiple parent crop
[0183] The present application also provides a device or an analysis device for identifying the genotype of a multiple parent crop for performing the method of the present application. Typically, the device comprises:
[0184] a data input module for inputting to-be-analyzed to-be-processed data, wherein the to-be-processed data comprises sequencing data Df of to-be-identified offspring plants and sequencing data Dp of parent plants corresponding to the offspring plants;
[0185] a multi-parent plant genotype identification module configured to perform the method according to the present application to obtain genotype identification results of the offspring plants;
[0186] and an output module for outputting the genotype identification results of the offspring plants.
[0187] The main advantages of the present application include:
[0188] (a) The present application provides, for the first time, a multi-parent plant genotype identification method based on high-throughput sequencing data. Prior to the present application, there is no systematic method for identifying the genotype of a multi-parent plant.
[0189] (b) The high-throughput genotype identification method of the present application can greatly simplify and accelerate the genetic positioning of quantitative traits in crops [37-39,20] .
[0190] (c) The theoretical method of the present application can be well combined with multi-parent populations for genotype identification, improving the accuracy and efficiency of QTL positioning, and making full use of the rich genetic variation in multi-parent populations. At the same time, it is also helpful for the improvement of crop genetic quality and molecular breeding design.
[0191] (d) In practical applications, the present application can be used for obtaining molecular markers closely linked to important agronomic traits, efficient screening of offspring in the breeding process, fine identification of genotype maps of improved varieties, etc., providing a fast and efficient means and platform for molecular marker-assisted selection breeding, and improving the efficiency and accuracy to a new level.
[0192] In summary, the high-throughput genotype identification method based on sequencing of the present application will provide convenience for solving complex biological problems and improving crop breeding.
[0193] The present application will be further described below in conjunction with specific examples. It should be understood that these examples are only used to illustrate the present application and not to limit the scope of the present application. The experimental methods in the following examples are generally carried out under conventional conditions or under the conditions recommended by the manufacturer, unless otherwise specified. Unless otherwise specified, percentages and parts are weight percentages and weight parts.
[0194] Example 1
[0195] 1.1 Whole genome simulation data making based on real sequencing data
[0196] 1.1.1 Rice materials and simulated data
[0197] The laboratory has a high-depth real rice material, indica 93-11 (Oryza sativa ssp. indica cv. 93-11), Shuohui 70, Wushan silk seedling and Huanghuazhan as parent materials, and the simulated offspring fastq data is obtained by aligning, screening and combining the real fastq data of the parent materials. In the simulated data of two parents, the inventors used 93-11 and Wushan silk seedling as simulated parents, simulated the homozygous region of each parent and the heterozygous region of the two parents on each chromosome of rice, to test the judgment of the recombination breakpoint and the heterozygous region of the method. In the multi-parent simulated data, 93-11, Shuohui 70, Wushan silk seedling, Huanghuazhan were used as simulated parents, and 100 times of data simulation was carried out, and different combinations of recombination breakpoints were simulated on each chromosome of rice. There are 4-6 recombination breakpoints on each chromosome of rice, and the total number of recombination breakpoints in the whole genome is 50-60, which aims to simulate the recombination within the genetic population of rice with multiple parent sources as much as possible and verify the accuracy of the method.
[0198] 1.1.2 Identification of simulated data SNPs
[0199] The sequencing data of 93-11, Shuohui 70, Wushan silk seedling, Huanghuazhan four parents were aligned with the complete sequence of japonica Nipponbare (japonica cv. Nipponbare) 12 chromosomes sequenced by International Rice Genome Sequencing Project (IRGSP) (IRGSP 1.0) http: / / rice.plantbiology.msu.edu / annotation_ pseudo_current.shtml )IRGSP 1.0, and the alignment software is bwa 0.7.17-r1188
[13] . Then use GATK software package
[14] HaplotypeCaller program in GATK software package (parameter is -ERC GVCF) to identify the candidate SNPs of the above-mentioned parents and simulated offspring. After obtaining the variant intermediate file g.vcf file of each parent and simulated offspring, all the variant intermediate files are merged using the GenomicsDBImport program in the GATK software package, and then the GenotypeGVCFs program in the GATK software package is used to export the merged variant file, the SelectVariants program is used to select the required SNP site information, and then the VariantFiltration program (parameters are --cluster-size 3 --cluster-window-size 10 --filter-expression "QD < 10.00" --filter-name lowQD --filter-expression "FS > 15.000" --filter-name highFS --genotype-filter-expression "DP > 50 || DP < 5" --genotype-filter-name InvalidDP) is used to filter all SNP sites to obtain high-quality SNP sites. Then, on this basis, genotype identification is performed on the simulated offspring.
[0200] 1.2 Program development of sequencing-based genotype identification process
[0201] In the sequencing-based genotype identification process, a large amount of data needs to be processed, a variety of different algorithms need to be applied, and some existing software such as sequence matching software and QTL analysis software are used. Therefore, the present inventors developed a plurality of perl and python scripts to realize the above-mentioned steps and make it a complete and easy-to-use process with wide universality.
[0202] After obtaining the SNP information of the parents and offspring through the GATK software, a python script is used to comprehensively analyze the SNP regionization identified for each individual by sliding a window along all SNP sites, read the genotype based on a fixed length sliding window, and then judge the recombination breakpoint and construct the recombination segment map. In addition, a perl script is used to generate a PNG format recombination segment map for each individual by using the intermediate file determined by the program, so that the user can intuitively browse the overall genotype. The GD module in Perl is required for mapping.
[0203] Next, another script Bin2MCD generates a recombination bin map from the recombination bin information of each individual, which is used to determine the recombination bin information of each individual.
[19] to facilitate subsequent QTL analysis. Once the phenotypes are evaluated and trait data are prepared, the output files can be directly used by some QTL analysis software packages to identify QTL, including Windows QTL Cartographer V2.5
[17] .
[0204] 1.3 Genotype identification based on rice DH population
[0205] 1.3.1 Rice DH population and triploid population
[0206] The rice DH population used in this study was constructed by the laboratory of National Center for Gene Research, Chinese Academy of Sciences. The two parents were Kasalath and japonica cv. Nipponbare. The DH population was generated from the F2 progeny after years of selfing and recombination. The present inventors selected several tens of lines for genotype identification. The rice triploid population used in this study was constructed by the laboratory of National Center for Gene Research, Chinese Academy of Sciences. The three parents were Wushan Siliao, 93-11 and Shukui 70. The plants in this population were generated from the hybridization of the three parents followed by selfing and recombination, and there were many recombination information in their genomes.
[0207] 1.3.2 Genotype identification of rice DH population and triploid population using sequencing-based method
[0208] The DH population of rice was genotyped using the method of the present application, and a high density map consisting of recombination bins was generated by Bin2MCD. To measure the accuracy of the method, the same genotype analysis and high density bin map were also performed using the method published in 2010. The high-depth (20-30x) sequencing data of the two parents Kasalath and Nipponbare were aligned to the Nipponbare reference genome IRGSP 1.0 using bwa software, and the high-quality SNP information of the two parents was found using GATK software, respectively. Then a perl script was used to replace the SNPs at the specified sites on the Nipponbare genome, thereby generating the pseudo reference of the two parents. The low-depth sequencing data of the DH population were aligned to the pseudo reference of the two parents, and then the genotype identification was performed.
[0209] Results
[0210] 2.1 Real parent-based rice multi-parent simulated data and genotype identification
[0211] 2.1.1 Simulated data of rice genome information
[0212] As shown in Figure 1 The expected genotype information of the simulated hybrid offspring from the two parents should be consistent with the standard graph. The simulated data was made in three cases: homozygous regions of Wuxiansimiao, homozygous regions of 93-11 and heterozygous regions of Wuxiansimiao and 93-11. The length of each region and the position of the recombination breakpoints are shown in the graph. The simulated data was made based on the real sequencing data of the two parents. First, the fastq data of the two parents were aligned to the rice Nipponbare genome, respectively. Then, the required fastq information was screened according to the alignment information (chromosome and position information) in the obtained sam file. Finally, the fastq data from the two parents was reformed into simulated hybrid offspring fastq data.
[0213] As shown in Figure 2 Similar methods were used to make simulated offspring data from four parents: Wuxiansimiao, Huanghuazhan, 93-11 and Suohui 70. The expected genomic genotype information and recombination breakpoints should be consistent with the information in the graph.
[0214] 2.1.2 Genotype identification of simulated data
[0215] The fastq data of the simulated offspring and the fastq data of the two parents Wuxiansimiao and 93-11 were aligned to the rice reference genome IRGSP 1.0, and then the whole genome variation information of the two parents and the simulated offspring was found by GATK software. After filtering and screening, high-quality SNP sites were obtained.
[0216] As shown in Figure 3 After obtaining the required SNPs, the method of "sliding window" was used to judge the SNPs in the whole genome. In a sliding window, the two parents were scored and compared. If one parent scored higher in this segment, this segment was determined to be the homozygous genotype of this parent (represented by red or blue in the graph). When the scores of the two parents were not significantly different, the segment was determined to be the heterozygous region of the two parents (represented by yellow in the graph). The inventors designed a quantitative method to measure the accuracy of the determination. The whole genome was divided into thousands of 100 kb small regions (or 20-200 kb small regions), and then the degree of consistency of the results obtained by the method of the present application with the standard graph could measure the accuracy of the method of the present application. According to this method, the genotype information of the identified simulated data was compared with the real simulated data genotype, and the accuracy of the identification of the two parents could reach 89.61%.
[0217] Meanwhile, the present inventors also used the published SEG-Map method to determine the genotype of the simulated offspring data. The fastq files of the simulated data were aligned to the pseudo reference of the two parents, and the parent-specific fastq sequences were screened using software. Then, the SNP site information was determined according to the sequence alignment position, and then the genotype information was determined using the sliding window method. This method has been theoretically verified and data simulated in the published article, and has high accuracy and feasibility. The accuracy of the SEG-Map software results was 89.32%, which was not much different from the method of the present application, indicating that the method of the present application indeed has high feasibility and accuracy.
[0218] The SEG-Map method has high reliability for genotype identification of two parents, and the present inventors have used this method for long-term genome analysis of rice materials. However, this method cannot be used for genotype identification of materials derived from multiple parents, and therefore the method of the present application is also used to solve the problem of genotype identification of multiple parents. Figure 5 As shown in the figure, the present inventors used the SNP-based sliding window method to identify the genotype of the simulated offspring derived from four parents. In a window, the four parents were scored, and if one parent scored the highest, the region was determined to be the homozygous region of the parent. In the figure, red, blue, green and yellow represent the homozygous regions of the four parents. The fastq data of the four parents was simulated 100 times, and the accuracy was quantified by dividing the genome into small regions. By comparing with the standard figure, the average simulation accuracy of the method of the present application for genotype identification of the simulated data of the four parents was 92.10%.
[0219] 2.1.3 Determination of recombination breakpoints
[0220] When the "window" slides along the chromosome, the genotype is read by the scores of two parent SNPs. A genotype will not change until a recombination breakpoint is encountered. The inventors found two types of breakpoints: one separates two homozygous genotypes, and the other separates a homozygous genotype from a heterozygous genotype; the former is the most common form in RILs, and the latter is mostly found in F2 populations. When a sliding window encounters a "homozygous / homozygous" breakpoint, the homozygous genotype temporarily changes to a heterozygous genotype and then changes back to a homozygous genotype. When a sliding window encounters a "homozygous / heterozygous" breakpoint, the homozygous genotype changes to a heterozygous genotype and then changes back to a homozygous genotype, at which point the boundary between the homozygous and heterozygous regions can be determined.
[0221] 2.1.4 Influence of different window sizes on the accuracy of multi-parent determination
[0222] When using this sequencing-based method for genotype identification studies, appropriate analysis parameters need to be set, first considering whether the size of the sliding window will affect the accuracy of genotype detection, such as the number of SNPs included in each window within a given physical length.
[0223] As shown in Figure 6 , the inventors used different window sizes to analyze the final SNP information obtained from the four-parent simulation data, and found that different sizes of the sliding window size indeed affected the final analysis accuracy. When the sliding window size was small (less than 199), the final accuracy was less than 90%, but when the sliding window size increased to 199, the genotype identification accuracy could reach 93.72%, but when the sliding window size continued to increase, the final accuracy did not change much, indicating that the accuracy of the determination result was not always increased with the increase of the size of the sliding window. For program running, larger sliding window size requires more computing resources and operation time, and when large-scale populations need to be processed, the time cost will be more prominent. Therefore, the inventors considered the time cost and accuracy comprehensively, and the sliding window size of 199 (or the sliding window size of 180-220) is a more reasonable choice.
[0224] 2.1.5 Influence of different sequencing depths on the accuracy of multi-parent determination
[0225] Then, considering that sequencing depth has a very important influence on genotype identification, and the SEG-Map software can perform relatively accurate genotype identification at a lower sequencing depth, the inventors conducted a depth test on the present method.
[0226] As shown in Figure 7 different depths, 0.2x, 1.5x and 3x, were tested. At each depth, 100 times of fastq data simulation were performed, then genotype identification was performed according to the variation information, and finally the accuracy of genotype identification was measured according to the compliance procedure with the standard map. The results show that with the increase of depth, the accuracy of genotype identification is slightly improved, but the improvement is not as expected, and the final genotype identification accuracy is not more than 95%. Therefore, the inventors have made further improvements.
[0227] 2.2 Genotype identification method based on SNP site sliding window
[0228] 2.2.1 Main steps of data analysis process
[0229] In order to enable this new sequencing-based genotype identification method to be widely used, the inventors have optimized the data processing process. This process can directly analyze and process the single-end or double-end short sequence sequencing results generated by the second-generation sequencing technology, and finally construct the genetic map of the recombinant population.
[0230] For a mapping population from two parents, before performing the data analysis process, the SNPs in the whole genome of the two parents need to be identified. This SNP identification work can be obtained from high-coverage whole genome sequencing, or from existing genomic SNP information in the rice haplotype map, or from low-coverage whole genome sequencing combined with missing genotype (SNP) filling. Since the SNP identification work between the two parent varieties can be obtained through a fast and cost-effective approach, the sequencing-based genotype identification of a recombinant population will mainly rely on the subsequent analysis work, including reading the genotype, determining the recombination breakpoint, and constructing the genetic linkage map.
[0231] Functions, steps, and software (scripts) in data analysis are as follows Figure 8The first step contains several tasks that can be processed simultaneously. A certain number of individuals in the recombined population and the parent materials are simultaneously subjected to second-generation high-throughput sequencing. The obtained fastq files are subjected to bwa and GATK software processing to obtain high-quality SNP information. The SNP sites used for final determination of genotypes should meet the following requirements: 1. The SNP sites should cover the whole genome as much as possible and should not be missing in some areas. 2. For any SNP site, the SNP information (position information and genotype information) of the two parents and the simulated offspring is known, and any one of the three should be deleted. 3. It is generally believed that the rice parent is a self-pollinated pure line, and there are basically no heterozygous sites in the genome, so if a heterozygous SNP site is found in the parent, it is generally believed that the site is not reliable, and therefore any SNP site with heterozygous parents is deleted.
[0232] After screening the high-quality whole genome SNPs of the two parents and the offspring, a python script SNPwindow is used to determine the genotype of the offspring. The script output will have two files, rlt file and bin file. The rlt file records the genotype determination of each SNP position, and the bin file records the distribution of recombination breakpoints on the 12 chromosomes of the whole genome. These two files are important basis for subsequent mapping and linkage analysis.
[0233] As shown in Figure 9 , generally, an rlt and bin file is first used to draw a genotype map by a perl script SNP2png, and the picture format is PNG format. The map is drawn according to the SNP site genotype information and whole genome recombination breakpoint position information determined by the program, and different colors in the map represent different genotype types.
[0234] In actual work, a large number of offspring populations need to be processed, so the SNP information of each individual can be extracted, and then analyzed and tested by the SNPwindow script to read the genotype, recombination breakpoint determination, and recombination map construction. First, the SNP2png is used to draw the gene map of each individual to have a overall grasp of the genotype of the individual, and then the script Bin2MCD is used to map all the recombination maps of the individuals to produce a recombination bin map.
[0235] As shown in Figure 10The genotype of the whole population can also be visualized by a perl script. The programs and scripts used in the analysis pipeline are in italics and form a series of analysis steps. The final output of the analysis pipeline is the genotype data which can be directly used in other software (including MapMaker and JoinMap) to construct the linkage map.
[0236] The output of the analysis software is the recombination bins, usually with a resolution of 100 kb per bin or even 10 kb per bin. The genotype results of the mapping population can be imported into MapMaker
[16] or JoinMap
[32] to construct the genetic map. With the genetic map, QTL analysis can be performed.
[0237] As Figure 11 shown, this genetic map is much more detailed than most of the maps generated by traditional molecular markers. The software package is compatible with multiple platforms (e.g. Unix, Linux and Windows). In addition to the perl environment itself, the GD module needs to be installed because of the mapping step in the pipeline.
[0238] 2.2.2 Detailed pipeline to determine the recombination breakpoints
[0239] Step 1: Construct SNP "strings".
[0240] Take the window size as 15 (win=15) as an example. The genotypes of all SNPs on the 12 chromosomes of the two parents and the progeny are compressed into a string in order. No matter how large the actual distance between two adjacent SNPs is, all the gaps are removed.
[0241] Thus, the SNPs on the 12 chromosomes become 12 continuous strings Figure 12 . In the figure, the blue color represents the genotype of parent 1, the red color represents the genotype of parent 2, and for a progeny, there are three possible genotypes at each SNP site, the homozygous genotype of parent 1 (blue), the homozygous genotype of parent 2 (red), and the heterozygous genotype (yellow).
[0242] It is generally believed that the genome of the artificially cultivated rice parent material is highly homozygous, and for some multi-generation self-cross recombination rice population, the genome is also relatively homozygous, only some heterozygous regions exist in part of the chromosome position. Therefore, the SNP sites included in the analysis are first artificially screened, and the SNP sites in which any parent is heterozygous are excluded. Such sites cannot be accurately scored. In addition, if the sequencing depth of the offspring is not very high, the SNP sites of the offspring can be filtered, because the heterozygous sites determined according to the low depth are not highly reliable, and it is very likely that the misjudgment is caused by sequencing errors.
[0243] Second step: scoring two parents in a window
[0244] According to Mendelian inheritance law, the scores of all SNP sites in a sliding window are calculated, and the total score of each parent is calculated as the score of the parent in the chromosome position of the sliding window. According to the score of each parent, the degree of conformity of the offspring to the parent is measured. The preferred scoring rule is shown in Table A. The scoring rule of the present application is formulated according to the genetic law of organisms.
[0245] Table A Genotype Scoring Rule Table
[0246]
[0247] Third step: judging the genotype of the chromosome region according to the score value
[0248] By sliding the sliding window on the whole genome SNP site, the score value of each parent on each chromosome can be obtained. The score curve of each parent is drawn by taking the score value as the vertical coordinate and the position of each sliding window on the chromosome as the horizontal coordinate. The genotype of each chromosome is determined according to the characteristics of the score curves of different parents. Figure 13 As shown, the sliding window scoring of the offspring of the simulated four-parent source was performed, and the score curves of the four parents were drawn according to the score values of the four parents.
[0249] By observing the score curves of the four parents, it can be seen that the parent curve has two different distribution modes of platform stable period and fluctuation period in different regions of chromosome 11. Therefore, the genotype of the offspring in the region is determined by using the state of the curve of different parents in the same region.
[0250] As shown in the 1bp to 10780000bp region in the figure, it can be observed that the score of the yellow curve (parent 4) in this region has a high score value close to full score, and the score value of this parent in this segment is quite stable, without too much numerical fluctuation, which is measured by the standard deviation in statistics. In this region, parent 4 has a very high score value and a small standard deviation in this region, while the score values of the other three parents in this region fluctuate up and down in the range of 0-200, with a high standard deviation, so it is judged that this segment is the homozygous genotype of parent 4. Using a similar method, the genotype of the offspring in different regions of the 12 chromosomes can be determined according to the score value. In the figure, the rectangular bar corresponding to true corresponds to the true genotype information of the simulated offspring of each chromosome segment, and the rectangular bar corresponding to judge indicates the genotype of the offspring determined by the method of the application, and the information of the two is basically consistent.
[0251] Fourth step: judgment of heterozygous region
[0252] The judgment of the heterozygous region is explained by the genotype judgment of the simulated offspring of two parents. According to the principles of genetics, even if the hybrid offspring is derived from multiple parents, the parent source of a certain segment of the chromosome is at most two parents. Therefore, whether this segment is a heterozygous region can be determined according to the score curve of the two parents in this region. As shown in the figure, the genotype of the simulated offspring is identified, and the rectangular bar corresponding to true corresponds to the true genotype information of the simulated offspring of each chromosome segment, and the rectangular bar corresponding to judge indicates the genotype of the offspring determined by the method of the application. Similarly, when the score of one parent is higher and the standard deviation is smaller, and the score of the other parent has considerable fluctuation and a large standard deviation, it is judged that this segment is the homozygous genotype of the former (orange or blue region in the figure). Figure 14
[0253] When the score values of the two parents in a certain segment (6000000bp to 15000000bp in the figure) are not much different from each other, and both have certain fluctuations and large standard deviations, this segment is judged to be the heterozygous genotype of the two parents (yellow region in the figure). This judgment result is also consistent with the true information.
[0254] Fifth step: explanation of judgment standard
[0255] The core idea of the method of the application is to directly read the true SNP genotype information of the parents and the offspring in a certain segment, then quantify the similarity of the offspring and each parent in this segment, form a relatively simplified analysis model according to the numerical characteristics (numerical value and standard deviation) of the score curve of each parent, and then determine the genotype of each segment. The judgment standard mainly includes the following cases:
[0256] I. If one parent has a higher score in this segment and its score is stable (the score curve is close to the platform), while the rest of the parents have a greater fluctuation in the score curve of this segment, the standard deviation is large (the score curve is a wave peak shape up and down), it is judged that this segment is the homozygous genotype of the parent. Figure 15 ).
[0257] II. When the number of parents is judged to be two, according to the fact that both parents have a greater fluctuation in a certain segment, the standard deviation is large, it can be inferred that this region is the heterozygous genotype of the two parents. Figure 16 ).
[0258] In the case of multiple parents, it is likely that only the possible heterozygous segment can be found, that is, a high score and stable parent cannot be found in a certain segment. Because the scores of each parent have a large fluctuation in this segment, only the two most likely parents can be given according to the numerical characteristics.
[0259] As shown in Figure 17 , the region to be judged is defined as "unknown" first, and the genotype of this region is determined by the genotype of its two adjacent regions.
[0260] As shown in Figure 18 , if the genotypes of the two adjacent regions are the same, the region is determined to be this genotype, and if the genotypes of the two adjacent regions are different, the middle position of the region is regarded as a recombination breakpoint, and the two adjacent regions are the genotypes of the two adjacent regions.
[0261] 2.3 Genotype identification of rice DH population based on sequencing
[0262] 2.3.1 Rice DH population
[0263] The two parents of the rice DH population used are Kasalath and Nipponbare. It is a population formed by inducing haploid and doubling of F1 generation by double cross. Its plants are homozygous, and the offspring after self-crossing is pure line, which can be repeated for many years and many points, and is an ideal material for studying genotype and environment interaction.
[0264] 2.3.2 SNP identification between two parents
[0265] The high-depth sequencing data (20x-30x) of two parent Kasalath and Nipponbare materials were aligned to the rice reference genome IRGSP 1.0 by bwa software, and then the required high-quality SNP information was found by GATK software. The SNP information of the parents and the SNP information of all offspring were combined into the same vcf file, which facilitated the extraction of the required variation site information from it.
[0266] 2.3.3 Genotype identification of DH population
[0267] The average sequencing depth of each offspring in the DH population was about 0.02x, which belonged to low-depth sequencing data. The SNP information of each offspring was extracted separately from the parent information, and then SNPwindow script was used for judgment to obtain the rlt file and bin file judged by each offspring.
[0268] As Figure 19 shown, using SNP2png script, the genotype identification results were visualized using the result file obtained in the previous step. In the figure, the homozygous genotype of the two parents (Kasalath in red and Nipponbare in blue) can be observed. For a multi-year self-recombination inbred population, the reliability of the heterozygous region is low, which may be caused by sequencing error or low polymorphism of the two parents in this region.
[0269] Then Bin2MCD script was used to calculate the map file of the overall genotype distribution using the bin file of the entire population as input. The map file divides the whole genome into many small bins, and the genotype of each bin is determined according to the genotype identification results of the individual.
[0270] As Figure 20 shown, after the map file, a perl script was used to visualize the genotype information of the entire population, and the proportion of different genotypes at each bin position was also calculated, which is an important parameter for population genetics research. Figure 20 The proportion of the three genotypes in different bins is represented by the proportion of red and blue in part of the figure. The visualization of this step is convenient for a quick and direct understanding of the population genotype. Combined with the phenotype output map file, winQTL and other analysis software can be directly used for QTL mapping analysis.
[0271] 2.3.4 Genotype identification of rice three-parent materials
[0272] As Figure 16As shown, the three-parental progeny materials cultivated in the laboratory were subjected to multi-parental genotype identification. The sequencing data amount of the progeny was about 0.2x, and the three parent materials were Wushan Siliao, 93-11 and Suohui 70, and the sequencing depth of the three parents was about 20x-30x. The progeny and the SNP information of the three parents were integrated into the same vcf, and then the final high-quality SNPs were further screened. Then the SNPwindow script was used to determine the genotype of the progeny. In a window, if the score of a parent is the highest, the region is determined as the homozygous genotype of the parent.
[0273] As shown in Figure 21 , the bin file of the recombination breakpoint was determined by the determination degree of the multi-parental, and a perl script was used to visualize the determination results. We can directly see the genotype information on the 12 chromosomes through the picture. The red region corresponds to parent one Wushan Siliao, the blue region corresponds to parent two 93-11, the green region corresponds to parent three Suohui 70, and the yellow region is the heterozygous region.
[0274] As shown in Figure 22 , the SEG-Map method was used to determine the genotype of the material, so the genotype determination of the plants in the laboratory before was mainly based on the genotype identification of the two parents Wushan Siliao and 9311. Three genotypes can be determined, Wushan Siliao homozygous genotype, 93-11 homozygous genotype and the heterozygous genotype of the two. According to the determination results of the three parents, the inventors found that the heterozygous segment determined for the two-parental material is likely to correspond to the homozygous genotype of the third parent. Therefore, the method of the present application can make up for the shortcomings of the previous SEG-Map software under the condition of ensuring accuracy, and solve the problem of multi-parental genotype determination.
[0275] 2.3.5 Genotype identification of four-parental simulated rice materials
[0276] As shown in Figure 23 , the inventors used the real sequencing fastq data of four real rice materials 93-11, Suohui 70, Wushan Siliao and Huanghuazhan as four parents, then according to the alignment results, the reads of the corresponding region were segmented, and a simulated progeny data was manufactured by manual combination and screening. The real genotype information and recombination breakpoint of the progeny are clear, so the simulated data can be used to evaluate the feasibility and accuracy of the present application.
[0277] According to the determination method, the inventors identified dozens of recombination breakpoints in the whole genome, and the different chromosomal regions determined were roughly consistent with the real genotype results.
[0278] As shown in Figure 24 , the figure shows the real genotype information of the simulated progeny of the present application.
[0279] Comparing the detailed regions of the two, the inventors found some differences in certain local areas. They examined the intermediate output RLT file of the judgment process to investigate the reasons for these differences. Possible reasons are as follows: 1. Because the sequencing data depth of the offspring is not very high, only a portion of the variation information of the whole genome can be captured, which may miss some important parental differentiating sites, resulting in the inability to distinguish the true parent in certain regions. 2. The two or more identified parents are very similar in certain regions, and there is no parental polymorphism. This situation is not due to sequencing errors or sequencing depth. For such highly similar regions, a judgment may be temporarily omitted, and the genotype determination depends on the genotype information of the two adjacent genotypes. Therefore, it may be possible that some regions cannot be accurately judged, and the genotype is assigned to the most likely parental genotype on both sides.
[0280] The analysis results clearly show that different regions of the whole genome can be identified as having the most likely true parent, which is difficult to achieve with previously published SEG-Map software and traditional analysis methods. This also provides a multi-parent analysis method for the breeding of different crops.
[0281] III. Discussion
[0282] Multiparental populations hold great promise for genetic analysis. By selecting multiple parents, population genetic diversity can be increased. Hybridization and self-pollination (or inbreeding) can be used to merge multiple parents into a single population, increasing the number of recombinations. Multiparental populations not only increase the frequency of recombinations and uncover the genetic basis behind complex traits, but also possess significant potential for breeding applications due to the richness of the selected parents' genetic makeup. Compared to biparental populations, multiparental populations have a larger number of parents, increasing the richness of population variation, including allelic and phenotypic diversity, improving mapping accuracy and precision, and enhancing QTL detection efficiency. The large number of accumulated recombination events improves QTL localization resolution. Because the parental selection in multiparental populations is more refined (i.e., the criteria are more stringent), and the increased genetic diversity from multiple parents allows for the application of QTL results in breeding research. Furthermore, compared to natural populations, multiparental populations are constructed through a homogeneous mixture of multiple parents. Because pedigree relationships are known in natural populations, detailed information about population construction is available. From an experimental design perspective, this avoids population stratification and controls false positives in localization results.
[0283] Recombination populations are the basis of Mendelian genetics experiments and have been the key to the study of genes, genomes and genetic variation. However, genotyping a mapping population has always been a laborious, time-consuming process, involving costly and tedious marker development and genotyping hundreds of individuals with hundreds of markers. Moreover, the resolution of the map obtained using such methods is relatively low [34-36] By applying next-generation sequencing technology, the present inventors have developed a fast, efficient, low-cost, high-information-content and reliable method for genotyping. With this new method, the ultra-high resolution genotyping of a typical mapping population containing several hundred individuals can be completed by a genome sequencing service center within several weeks, rather than months or even years as required by traditional marker-based methods.
[0284] The present inventors have developed a new method for high-throughput genotyping by detecting SNPs through whole-genome low-coverage resequencing. This type of SNP data differs from traditional genetic markers in two main aspects. First, in a recombination population, not all strains can obtain information on a SNP site through random sequencing. Second, a single SNP site is not a reliable marker or site for genotyping because of potential sequence errors.
[0285] To handle these SNP data with unique properties generated by next-generation sequencing, the present inventors have further developed a new analysis framework, i.e. a "sliding window" method, to determine the genotype of a segment based on the genotypes of multiple SNPs in a local position.
[0286] Based on this theory, the present inventors have also developed a set of program processing and analysis pipeline, named SEG-Map (Sequencing Enabled Genotyping for Mapping recombination populations), meaning a sequencing-based recombination population mapping pipeline. Using SEG-Map, starting from the analysis of single-end or paired-end short sequence data generated by Illumina Genome Analyzer II (GAII), through multiple steps of analysis, a genetic map of a recombination population can be constructed. This method is suitable for a two-parent constructed recombination population.
[0287] The present inventors have developed a novel procedure for analyzing genotype data and corresponding methods and devices. The procedure of the present application can optimize the steps of the previous SEG-Map procedure, is compatible with current bioinformatics analysis software and different types of high-throughput sequencing data, and most importantly, can quickly, accurately and reliably analyze the genotypes of a multi-parental constructed population.
[0288] The method of the present application can help better application of multi-parental populations in crop breeding, can accurately identify more QTL sites in multi-parental populations, and can help genome prediction of multi-parental populations, thereby providing a basis for direct application of the multi-parental populations as germplasm resources in varieties.
[0289] In the current analysis procedure, the program for reading genotypes and determining recombination breakpoints is designed to adapt to various types of mapping populations and completely connect with the previous identification of SNPs and the subsequent construction of recombination bin maps. After combining these functions, the analysis software takes short sequences generated by the second-generation sequencing technology as input, performs a series of operations, and outputs recombination bins, which can be analyzed by existing software for constructing genetic linkage maps and QTL (quantitative trait loci) analysis.
[0290] The present inventors used a high-throughput sequencing-based method to identify the genotypes of a rice recombinant inbred line and showed the advantages of the new genotype identification method over the commonly used PCR-based method. Before developing the sequencing-based method to identify the F 11 Before developing the sequencing-based method to identify the F
[0291] The coverage of resequencing can be easily adjusted, which also allows the inventors to choose the appropriate marker density level and resolution of recombination breakpoints at the shortest time and least resource input. When new scientific questions arise that require higher marker density or more accurate determination of recombination breakpoints, the inventors can increase the coverage of resequencing for the whole or a part of the mapping population. It is particularly noted that the recombination breakpoints can be determined very accurately using this method, and theoretically within 1 kb if the coverage of resequencing is high enough. Such a fine resolution can detect "double crossing over" events that cannot be identified by other types of genetic markers. Ultimately, this method can improve the accuracy of QTL detection and increase the efficiency and success rate of gene cloning. The accurately identified recombination breakpoints also enable the study of genomic regions with special genetic properties, such as recombination hotspots.
[0292] In summary, the high-throughput genotyping method combined with the second-generation sequencing technology greatly simplifies and accelerates the genetic mapping of quantitative traits in crops [37-39,20] The theoretical method proposed by the inventors can be well combined with multi-parent populations for genotyping, improving the accuracy and efficiency of QTL mapping, and fully utilizing the rich genetic variation present in multi-parent populations. At the same time, it is also helpful for crop genetic quality improvement and molecular breeding design. In practical applications, this method can be used for obtaining molecular markers tightly linked to important agronomic trait genes, efficient screening of offspring in the breeding process, fine identification of genotype maps of improved varieties, etc., providing a fast and efficient means and platform for molecular marker-assisted selection breeding, and improving efficiency and accuracy to a new level. In summary, this high-throughput genotyping method based on sequencing will provide convenience for solving complex biological problems and crop breeding improvement.
[0293] All documents mentioned in the present application are incorporated herein by reference as if each document were individually incorporated. In addition, it is to be understood that various alterations and modifications can be made to the present application upon reading and understanding the above lecture of the present application, and these equivalent forms also fall within the scope defined by the appended claims.
[0294] REFERENCES
[0295] [1] Winzeler, E. A. et al. (1998) Direct allelic variation scanning of the yeast genome. Science, 281 : 1194-1197.
[0296] [2] Meaburn, E., Butcher, L.M., Schalkwyk, L.C., & Plomin, R. (2006) Genotyping pooled DNA using 100K SNP microarrays: a step towards genome wide association scans. Nucleic Acids Res., 34: e27.
[0297] [3] Singer, T. et al. (2006) A high-resolution map of Arabidopsis recombinant inbred lines by whole-genome exon array hybridization. PLoS Genet., 2: e144.
[0298] [4] Jeremy, E. et al. (2008) Development and evaluation of a high-throughput, low-cost genotyping platform based on oligonucleotide microarrays in rice. Plant Methods, 4: 13.
[0299] [5] Craig, D.W. et al. (2008) Identification of genetic variants using bar-coded multiplexed sequencing. Nat. Methods, 5: 887-893.
[0300] [6] Cronn, R. et al (2008). Multiplex sequencing of plant chloroplast genomes using Solexa sequencing-by-synthesis technology. Nucleic Acids Res., 36: e122.
[0301] [7] Doi K, Iwata N, Yosiiimura A (1997). The construction of chromosome substitution lines of African rice (Oryza glaberrima Steud.) in the background of japonica rice (O. sativa L.). Rice Genet. Newsl., 14:39-41.
[0302] [8] Wan XY, Wan JM, Su CC, Wang CM, Shen WB, Li JM et al. (2004). QTL detection for eating quality of cooked rice in a population of chromosome segment substitution lines. Theor. Appl. Genet., 110:71-79.
[0303] [9] Ebitani T, Takeuchi Y, Nonoue Y, Yamamoto T, Takeuchi K, Yano M (2005). Construction and evaluation of chromosome segment substitution lines carrying overlapping chromosome segments of indica rice cultivar 'Kasalath' in a genetic background of japonica elite cultivar 'Koshihikari'. Breed Sci., 55:65-73.
[0304]
[10] Hao W, Jin J, Sun SY, Zhu MZ, Lin HX (2006). Construction of chromosome segment substitution lines carrying overlapping chromosome segments of the whole wild rice genome and identification of quantitative trait loci for rice quality. J. Plant Physiol. Mol. Biol., 32:354-362.
[0305]
[11] Huang X, Wei X, Sang T, Zhao Q, Feng Q, Zhao Y, Li C, Zhu C, Lu T, Zhang Z, Li M, Fan D, Guo Y, Wang A, Wang L, Deng L, Li W, Lu Y, Weng Q, Liu K, Huang T, Zhou T, Jing Y, Li W, Lin Z, Buckler ES, Qian Q, Zhang Q, Li J, Han B. (2010) Genome-wide association studies of 14 agronomic traits in rice landraces. Nature Genet., 42:961-967
[0306]
[12] Li, H., Ruan, J., & Durbin, R. (2008) Mapping short DNA sequencing reads and calling variants using mapping quality scores. Genome Res., 18: 1851-1858.
[0307]
[13] Kurtz, S. et al. (2004) Versatile and open software for comparing large genomes. Genome Biol., 5:R12.
[0308]
[14] Rice, P., Longden, I., & Bleasby, A. (2000) EMBOSS: The European molecular biology open software suite. Trends in Genetics, 16: 276-277.
[0309]
[15] Ning, Z., Cox, A. J., & Mullikin, J. C. (2001) SSAHA: a fast search method for large DNA databases. Genome Res., 11 : 1725-1729.
[0310]
[16] Lincoln, S. E. & Lander, S. L. (1993) Mapmaker / exp 3.0 and mapmaker / qtl 1.1. Technical report. Whitehead Institute of Medical Research, Cambridge, MA.
[0311]
[17] Wang, S., Basten, C. J. & Zeng, Z. B (2007). Windows QTL Cartographer 2.5. Department of Statistics, North Carolina State University, Raleigh, NC.
[0312]
[18] Li, R. et al. (2009) SOAP2: an improved ultrafast tool for short read alignment. Bioinformatics, 25: 1966-1967.
[0313]
[19] Van Os, H. et al. (2006) Construction of a 10,000-marker ultradense genetic recombination map of potato: Providing a framework for accelerated gene isolation and a genome wide physical map. Genetics, 173: 1075-1087.
[0314]
[20] Xu J, Zhao Q, Du P, Xu C, Wang B, Feng Q, Liu Q, Tang S, Gu M, Han B, Liang G. (2010) Developing high throughput genotyped chromosome segment substitution lines based on population whole-genome re-sequencing in rice (Oryza stative L.). BMC Genomics, 11:656.
[0315]
[21] Paterson AH, Bowers JE, Bruggmann R, Dubchak I, Grimwood J, Gundlach H, Haberer G, Hellsten U, Mitros T, Poliakov A, Schmutz J, Spannagl M, Tang H, Wang X, Wicker T, Bharti AK, Chapman J, Feltus FA, Gowik U, Grigoriev IV, Lyons E, Maher CA, Martis M, Narechania A, OtiUar RP, Penning BW, Salamov AA, Wang Y, Zhang L, Carpita NC, Freeling M, Gingle AR, Hash CT, Keller B, Klein P, Kresovich S, McCann MC, Ming R, Peterson DG, Mehboob-ur-Rahman, Ware D, Westhoff P, Mayer KF, Messing J, Rokhsar DS. (2009) The Sorghum bicolor genome and the diversification of grasses. Nature, 457(7229):551-6.
[0316]
[22] Stein LD, Mungall C, Shu S, Caudy M, Mangone M, Day A, Nickerson E, Stajich JE, Harris TW, Arva A, Lewis S. (2002) The generic genome browser: a building block for a model organism system database. Genome Res., 12(10): 1599-610.
[0317]
[23] Rice Annotation Project. (2007) Curated genome annotation of Oryza sativa ssp. japonica and comparative genome analysis with Arabidopsis thaliana. Genome Res., 17: 175-83.
[0318]
[24] Rice Annotation Project. (2008) The rice annotation project database (RAP-DB): 2008 update. Nucleic Acids Res., 36: D1028-D1033.
[0319]
[25] The Rice Full-Length cDNA Consortium. (2003) Collection, mapping, and annotation of over 28,000 cDNA clones from japonica rice, Science, 301: 376-379.
[0320]
[26] Liu, X., Lu, T., Yu, S., et al. (2007) A collection of 10,096 indica rice full-length cDNAs reveals highly expressed sequence divergence between Oryza sativa indica and japonica subspecies, Plant Mol. Biol., 65: 403-415.
[0321]
[27] International Rice Genome Sequencing Project (2005). The map-based sequence of the rice genome. Nature, 436: 793-800.
[0322]
[28] Yu, J. et al. (2005) The Genomes of Oryza sativa: A history of duplications. PLoS Biol., 3: 266-281.
[0323]
[29] Dohm, J. C, Lottaz, C, Borodina, T., & Himmelbauer, H. (2008) Substantial biases in ultra-short read data sets from high-throughput DNA sequencing. Nucleic Acids Res., 36: e105.
[0324]
[30] Mouse Genome Sequencing Consortium. (2002) Initial sequencing and comparative analysis of the mouse genome. Nature, 420: 520-562.
[0325]
[31] Frazer, K. A., Eskin, E., Kang, H. M., Bogue, M. A., Hinds, D. A., Beilharz, E. J., Gupta, R. V., Montgomery, J., Morenzoni, M. M., Nilsen, G. B., et al. (2007) A sequence-based variation map of 8.27 million SNPs in inbred mouse strains. Nature, 448: 1050-1053
[0326]
[32] Stam P. (1993) Construction of integrated genetic linkage maps by means of a new computer package: JoinMap. Plant J., 3: 739-44.
[0327]
[33] Sasaki, A. et al. (2002) A mutant gibberellin-synthesis gene in rice. Nature, 416:701-702.
[0328]
[34] Eshed, Y. and Zamir, D (1995). An introgression line population of Lycopersicon pennellii in the cultivated tomato enables the identification and fine mapping of yield-associated QTL. Genetics, 141 : 1147-1162.
[0329]
[35] Loudet, O., Chaillou, S., Camilleri, C, Bouchez, D. and Daniel-Vedele, F. (2002) Bay-0 x Shahdara recombinant inbred line population: a powerful tool for the genetic dissection of complex traits in Arabidopsis. Theor. Appl. Genet., 104: 1173-1184.
[0330]
[36] Simon, M., Loudet, O., Durand, S., Berard, A., Brunel, D., Sennesal, F.-X., Durand-Tardif, M., Pelletier, G. and Camilleri, C. (2008) Quantitative trait loci mapping in five new large recombinant inbred line populations of Arabidopsis thaliana genotyped with consensus single-nucleotide polymorphism markers. Genetics, 178: 2253-2264.
[0331]
[37] Huang, X. et al (2009). High-throughput genotyping by whole-genome resequencing. Genome Res., 19: 1068-1076.
[0332]
[38] Xie W, Feng Q, Yu H, Huang X, Zhao Q, Xing Y, Yu S, Han B, Zhang Q. (2010) Parent-independent genotyping for constructing an ultrahigh-density linkage map based on population sequencing. Proc. Natl. Acad. Sci. U S A., 107(23): 10578-83. Epub 2010 May 24.
[0333]
[39] Zhao Q, Huang X, Lin Z, Han B. (2010) SEG-Map: A novel software for genotype calling and genetic map construction from next-generation sequencing. Rice, 3: 98-102.
Claims
1. A method of identifying the genotype of a multiple parent plant, comprising, The method comprises: (a) providing sequencing data Df of a progeny plant to be identified and sequencing data Dp of a parent plant corresponding to the progeny plant for n parents and their offspring, wherein n is a positive integer greater than or equal to 3; (b) determining SNP site information of the parents and the offspring based on the sequencing data Df and the sequencing data Dp; (c) determining the genotype of the offspring based on the SNP site information, thereby obtaining the evaluation results of each SNP of the offspring and the distribution information of recombination breakpoints on each chromosome of the whole genome of the offspring; (d) constructing and / or drawing a genotype map of the offspring based on the SNP evaluation result information and the position information of the recombination breakpoints of the whole genome, thereby obtaining the genotype identification result of the multiple parent plants; wherein step (c) comprises analyzing the recombination breakpoint, thereby obtaining the analysis result of the recombination breakpoint, and the recombination breakpoint analysis comprises: (s1) constructing an SNP "string", wherein the genotypes of all SNPs on each chromosome of the parents and the offspring are compressed into a string in order; (s2) determining each sliding window corresponding to the SNP string according to a predetermined window size, and scoring the SNP sites in each window, thereby obtaining the score value P of each parent in the window; (s3) determining the genotype of each chromosome region of the offspring based on the score value P obtained in step (s2); wherein in step (s3), by sliding the sliding window on the whole genome SNP site, the score value of each parent on each chromosome can be obtained, and the score curve of each parent can be drawn with the score value as the vertical coordinate and the position of each sliding window on the chromosome as the horizontal coordinate; In step (s3), the genotype is evaluated in the following manner: (Z1) if a certain parent has a higher score value in the segment and its score value is relatively stable, the score curve approaches the plateau period, and the score curves of the remaining parents in the segment have greater fluctuations with larger standard deviations, and the score curve is a wave-shaped up and down, it is judged that this segment is the homozygous genotype of the parent; (Z2) when the number of parents is judged to be two, according to the fact that both parents have greater numerical fluctuations and larger standard deviations in a certain segment, it can be inferred that the region is the heterozygous genotype of the two parents; (Z3) if two or more parents are similar to the offspring in a certain segment, two or more parent curves with high scores and small standard deviations appear, indicating that the analyzed several parents are very similar in this segment and have little difference, which can be temporarily judged as "unknown region"; wherein in step (b), the SNP sites used to determine the genotype meet the following requirements: I. The SNP sites cover the whole genome as much as possible and do not have missing in some regions; II. For any SNP site, the SNP information of the corresponding two parents and the offspring is known, i.e., the position information and the genotype information are known, and any one of the three is unknown, the site should be deleted.
2. The method of claim 1, wherein, In step (c), the analysis of recombination breakpoints is based on SNP "string".
3. The method of claim 1, wherein, In step (s1), the gap between SNPs is removed regardless of the actual distance between the two adjacent SNPs.
4. The method of claim 1, wherein, In step (s1), the SNPs constituting the string are homozygous SNP sites of the parents.
5. The method of claim 1, wherein, In step (s1), the SNP sites included in the analysis are pre-screened, thereby excluding any SNP site with a heterozygous parent.
6. The method of claim 1, wherein, In step (s3), for each chromosome region of the offspring, the genotype corresponding to each chromosome region of the offspring is determined based on the score value or score curve of each parent.
7. The method of claim 1, wherein, For a chromosome region to be judged, if one parent A has a high score value close to full score, i.e. ≥80% of full score, and the score value of this parent in this region is relatively stable without too much fluctuation, while the score value of the remaining parent is low, i.e. ≤50% of full score or there is a large numerical fluctuation, the genotype of this chromosome region is determined as the genotype of the parent A.
8. The method of claim 1, wherein, In step (s3), the evaluation of heterozygous regions includes the following sub-step: (s3a) For a hybrid offspring of multiple parents, the parent source of a certain chromosome region is still set to be at most two parents, and whether this region is a heterozygous region is determined based on the score curve of the two parents in this region.
9. The method of claim 1, wherein, In step (s3), the similarity of the offspring to each parent in this segment is quantified, and the genotype of each segment is determined according to the numerical characteristics of the score curve of each parent, including the numerical value and the standard deviation.
10. The method of claim 1, wherein, The method further comprises: if the genotypes on both sides of the unknown region are the same, the region is determined as this genotype; and if the genotypes on both sides are different, the middle position of the unknown region is regarded as a recombination breakpoint, and the genotypes on both sides of the unknown region are respectively the genotypes on both sides.
11. The method of claim 1, wherein, The offspring is a multiple-parent plant.
12. The method of claim 1, wherein, n is 3-6.
13. The method of claim 1, wherein, The sequencing data is selected from the group consisting of genomic sequencing data, RNA sequencing data, or a combination thereof.
14. The method of claim 1, wherein, The size of the sliding window is 170-500 consecutive SNP sites. The sequencing depth of the sequencing data is 0.1x-10x.
15. The method of claim 14, wherein, The size of the sliding window is 200-400 consecutive SNP sites.
16. The method of claim 14, wherein, The sequencing depth of the sequencing data is 0.2x-5x.
17. The method of claim 1, wherein, In step (b), the sequencing data is aligned and processed by bwa and GATK software to obtain SNP information.
18. The method of claim 17, wherein, The SNP site information includes position information and genotype information.
19. The method of claim 1, wherein, In step (c), the genotype is read and the recombination breakpoint is determined by the SNPwindow script.
20. The method of claim 1, wherein, In step (c), the evaluation results of each SNP of the offspring are recorded in the rlt file, which records the genotype determination of each SNP position.
21. The method of claim 1, wherein, In step (d), the construction of the recombination map is performed by the SNPwindow script, and the genetic map of each offspring individual is drawn by the SNP2png script.
22. The method of claim 1, wherein, In step (d), the recombination bin map of each individual is aligned by a Bin2MCD script to generate a recombination bin map.
23. The method of claim 1, wherein, The resolution of the recombination bin map is 5-200 kb per bin.
24. The method of claim 23, wherein, The resolution of the recombination bin map is 10-100 kb per bin.
25. The method of claim 1, wherein, The method further comprises processing the recombination bin map to obtain the genetic map of the offspring.
26. The method of claim 1, wherein, The plant comprises a crop.
27. The method of claim 26, wherein, The plant comprises a crop of the family Poaceae.
28. The method of claim 26, wherein, The crop comprises rice, wheat, soybean, and tobacco.
29. A data analysis device for identifying the genotype of a multi-parental plant, the device comprising: a data input module for inputting to-be-analyzed to-be-processed data, the to-be-processed data comprising sequencing data Df of a to-be-identified offspring plant and sequencing data Dp of a parent plant corresponding to the offspring plant; a multi-parental plant genotype identification module configured to execute the method of claim 1 to obtain a genotype identification result of the offspring; and an output module for outputting the genotype identification result of the offspring.
30. The apparatus of claim 29, wherein, The multi-parental plant genotype identification module comprises: a SNP site information analysis submodule configured to determine SNP site information of parents and offspring based on the sequencing data Df and the sequencing data Dp; a chromosome recombination breakpoint analysis submodule configured to determine the genotype of the offspring based on the SNP site information to obtain an evaluation result of each SNP of the offspring and distribution information of recombination breakpoints on each chromosome of the whole genome of the offspring; a genotype map construction submodule configured to construct and / or draw a genotype map of the offspring based on the SNP evaluation result information and the location information of the whole genome recombination breakpoints to obtain the genotype identification result of the multi-parental plant.