A high-throughput genotype intelligent analysis method
By introducing a high-throughput intelligent genotype analysis method with multiple quality assessments and batch correction, the problems of inconsistent genotype data and batch-to-batch differences have been solved, achieving efficient and accurate breeding guidance.
Patent Information
- Application Number
- CN202510990937.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-07-18
AI Technical Summary
In existing technologies, high-throughput genotype data suffers from inconsistent formats and varying quality, lacking effective quality control standards and evaluation systems. This leads to significant systematic differences between batches, and traditional methods cannot effectively handle linkage disequilibrium structures and population kinship estimation biases, affecting breeding efficiency and accuracy.
We introduce multiple quality assessment criteria, batch effect correction, linkage disequilibrium network analysis, and hybrid vigor relationship index modeling. By standardizing data format, removing batch effects, and identifying SNPs as molecular markers, we construct a population kinship matrix and a hybrid vigor prediction model.
It significantly improved the quality of genotype data, reduced the number of molecular markers, lowered computational complexity, and improved breeding accuracy and efficiency, with a hybrid vigor prediction accuracy of over 85%.
Smart Images

Figure CN120656541B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of breeding analysis, and particularly relates to a high-throughput genotype intelligent analysis method. BACKGROUND
[0002] With the rapid development of genomics and information technology, various high-throughput sequencing platforms such as Illumina, PacBio, Oxford Nanopore, etc. are emerging, and the data output per sequencing is rapidly growing from GB to TB, providing unprecedented massive genomic information for biological breeding. These large-scale genomic data provide strong support for breeding techniques such as Genomic Selection, Marker-Assisted Selection and genetic diversity evaluation.
[0003] However, the genotype data formats from different sequencing platforms are not unified, the data quality is uneven, there is a lack of effective quality control standards and evaluation system, and with the increase of sequencing batches, the batch effect is increasingly significant, which brings serious challenges to data integration and analysis; traditional SNP (single nucleotide polymorphism) screening methods such as p-value-based screening or fixed interval sampling cannot fully consider the linkage disequilibrium structure between SNP sites, resulting in information redundancy or loss of key variations; at the same time, the calculation complexity is too high when analyzing millions of SNP sites in a large population, which seriously affects the analysis efficiency; in terms of population kinship estimation, traditional methods such as VanRaden method or GCTA-GRM method often ignore the influence of population structure, resulting in biased kinship estimation, especially in the presence of obvious population stratification, this bias is particularly significant, which is not conducive to breeders to make strain selection and hybrid design.
[0004] Therefore, an intelligent high-throughput genotype analysis method is urgently needed to overcome the above technical bottlenecks and improve breeding efficiency and accuracy. SUMMARY
[0005] The application provides a high-throughput genotype intelligent analysis method, which introduces multiple quality evaluation standards, batch effect correction, linkage disequilibrium network analysis and hybrid advantage relationship index modeling to improve the quality of genotype data, optimize molecular marker screening, accurately estimate population kinship and accurately predict hybrid advantage.
[0006] To achieve the above application purposes, the specific technical solutions are as follows:
[0007] A high-throughput intelligent genotyping method, the method comprising the following steps:
[0008] Step S1: Collect high-throughput genotype data, standardize the sample data format, and perform quality assessment and batch effect removal on the sample data to obtain preprocessed genotype data.
[0009] The method for quality assessment of sample data includes: calculating the sample genotype coverage, calculating the conversion / transversion ratio in the genotype data, and calculating the genotype error rate.
[0010] Samples were screened based on the quality assessment results, and genotype coverage was excluded. The sample.
[0011] Elimination conversion / transformation ratio Abnormal deviation from the population The sample, and These represent the expected and variance of the population transition / transition ratio, respectively.
[0012] Genotype error rate elimination Sequencing batch data.
[0013] Step S2: Using the preprocessed genotype data, calculate the genetic similarity matrix between samples and perform linkage disequilibrium analysis to identify tag SNPs as molecular markers.
[0014] Step S3: Estimate population kinship and predict hybrid vigor based on the obtained molecular markers.
[0015] Furthermore, the standardized sample data format is specifically as follows:
[0016] Genotype data obtained from different sequencing platforms were uniformly converted into a standard format. Genotype coding adopted... Numerical representation, where 0 represents the homozygous major allele AA, 1 represents the heterozygous genotype Aa, and 2 represents the homozygous minor allele aa.
[0017] Missing data are filled using a multiple imputation method based on the locally linked unbalanced LD structure. The calculation formula is as follows: ,in Indicates sample At the site The genotype values to be inserted on the above, Indicates the location At LD threshold The set of all non-missing sites within the range, Indicates site with site The weights between them , For site with site The LD coefficient between them.
[0018] Furthermore, the sample genotype coverage The formula for calculation is: ,in, For the sample Number of effective genotype loci in the middle This represents the total number of sites.
[0019] Transformation / transversion ratio in genotype data The formula for calculation is: ,in, This represents the number of conversion mutations. The conversion mutations include: A G or C T, The number of transversion mutations, including: A C, A T, G C or G T. The left and right sides represent the genes before and after the mutation, respectively.
[0020] Genotype error rate The formula for calculation is: ,in, This represents the number of samples that were repeated for sequencing. For duplicate samples Number of inconsistent sites For duplicate samples The total number of sites in the community.
[0021] Furthermore, the method for removing batch effects includes:
[0022] Constructing the position effect index ,in Indicates batch midpoint The average genotype, and They are the sites Mean and standard deviation across all batches For batch The sequencing time is in days. and These represent the mean and standard deviation of sequencing time for all batches.
[0023] right Batch correction is performed on the sites, and the correction formula is: ,in For site The batch effect coefficient is obtained by minimizing the corrected inter-batch variance: , Indicates site Genotype value vectors across all batches.
[0024] Furthermore, calculating the genetic similarity matrix between samples includes the following steps:
[0025] Based on the preprocessed genotype data, for any two samples and Calculate samples and samples At the site correlation on :
[0026] When two samples have identical genotypes When sharing an allele Completely different times .
[0027] Calculate the sum of associations between two samples across all common non-missing sites: ,in This represents the number of common non-deleted SNP sites.
[0028] Calculate genetic similarity based on allele frequency weighting: ;
[0029] in, , for site The weight, For site Secondary allele frequencies.
[0030] Furthermore, the linkage disequilibrium analysis specifically involves: using the preprocessed genotype data, calculating the linkage disequilibrium coefficient between any two SNP loci. : ,in, , For the frequency of haplotype AB, and The frequencies of the major alleles at loci A and loci B are respectively. and The frequencies of the minor alleles at loci A and loci B are respectively.
[0031] Constructing a chain-disequilibrium decay model: ,in, The effective population size is given by d, which represents the genetic distance between two loci in moles.
[0032] Through chain imbalance half-life Assess the extent of linkage disequilibrium in the population, i.e. At that time, the chain imbalance half-life is: .
[0033] Furthermore, the identification tag SNP as a molecular marker specifically involves constructing an SNP site network based on linkage disequilibrium analysis. Where the vertex set V represents all SNP sites, and the edge set E represents linkage disequilibrium, when two sites are in a state of linkage disequilibrium... Greater than the threshold At that time, there is an edge between these two sites.
[0034] Calculate the degree centrality of each SNP site: ,in This is an indicator function.
[0035] For each connected component, the SNP site with the highest degree centrality is selected as the label SNP of the connected component.
[0036] If multiple loci have the same degree centrality, the locus with the minor allele frequency closest to 0.5 is selected as the tag SNP.
[0037] Calculate the capture rate of the tag SNP: Where T is the set of label SNPs, This represents the total number of SNP sites.
[0038] Furthermore, the estimation of population kinship based on the obtained molecular markers includes: using the identified label SNPs as molecular markers to calculate the common ancestor coefficient matrix between samples. :
[0039] ,in For the number of SNPs in the label, Indicates sample In the Genotype values on each tag SNP For the first Secondary allele frequencies of tag SNPs.
[0040] Calculate the kinship matrix for population structure correction ,in The population structure matrix was obtained using the population structure analysis software STRUCTURE. This is a diagonal matrix of differentiation coefficients between different subgroups.
[0041] A rootless tree is constructed based on a phylogenetic tree derived from a kinship matrix. The branch lengths of the tree reflect the genetic distance between samples. .
[0042] Furthermore, the method for predicting hybrid vigor specifically involves calculating the expected heterozygosity of the F1 hybrids based on the parental genotypes. : ,in, and Representing the parents Japanese marriage At the site The genotype value.
[0043] Calculate the relationship index between genetic distance between parents and hybrid vigor: ,in For parent and The genetic distance between them.
[0044] Constructing a model for predicting hybrid vigor: ,in The predicted hybrid vigor value, , , and The regression coefficients are estimated using the least squares method based on known hybridization phenotypic data and corresponding molecular marker data.
[0045] Compared with the prior art, the beneficial effects of this invention are:
[0046] This invention significantly improves the reliability of genotype data by establishing a triple quality assessment standard of coverage, transformation / transversion ratio, and error rate. It introduces a position effect index to assess batch effects and effectively eliminates systematic bias between different sequencing batches through an adaptive correction algorithm, reducing inter-batch variance by 85% after correction. Based on linkage disequilibrium network analysis, it identifies tag SNPs, reducing the number of molecular markers by 70% while maintaining a capture rate of over 90%, significantly reducing the computational complexity of subsequent analyses. This invention constructs a hybrid vigor relationship index based on heterozygosity and genetic distance, achieving a hybrid vigor prediction accuracy of over 85%, providing precise guidance for efficient breeding. Attached Figure Description
[0047] Figure 1 This is a flowchart of a high-throughput intelligent genotype analysis method according to the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention are described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of this invention, not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0049] like Figure 1 As shown, this invention provides a high-throughput intelligent genotyping method, which includes the following steps:
[0050] Step S1: Collect high-throughput genotype data, standardize the sample data format, and perform quality assessment and batch effect removal on the sample data to obtain preprocessed genotype data.
[0051] High-throughput genotyping data can be derived from various sequencing platforms, including but not limited to Illumina HiSeq / NovaSeq series, PacBio Sequel system, Oxford Nanopore PromethION, etc. Each platform generates raw data in different formats. For example, Illumina typically outputs FASTQ format, but after genotyping analysis, it may be in VCF or BCF format; PacBio may output BAM format; and Oxford Nanopore may output FAST5 format. This method first converts these different formats into a unified tabular form containing sample ID, chromosome location, reference allele, variant alleles, and genotype values. The standardized sample data format is specifically as follows:
[0052] Genotype data obtained from different sequencing platforms were uniformly converted into a standard format. Genotype coding adopted... Numerical representation, where 0 represents the homozygous major allele AA, 1 represents the heterozygous genotype Aa, and 2 represents the homozygous minor allele aa.
[0053] For example, for VCF files from the Illumina platform, tools such as VCFtools or bcftools can be used to extract SNP information and convert it to a standard format. For long-read data from PacBio or Oxford Nanopore, SNP detection must first be performed using tools such as GATK HaplotypeCaller, and then the format must be standardized. Standardized data typically includes columns for chromosome number (Chr), physical location (Pos), reference allele (Ref), variant allele (Alt), and genotype values (0 / 1 / 2) for each sample, facilitating data processing and analysis in environments such as Python and R.
[0054] Missing data are filled using a multiple imputation method based on the locally linked unbalanced LD structure. The calculation formula is as follows: ,in Indicates sample At the site The genotype values to be inserted on the above, Indicates the location At LD threshold The set of all non-missing sites within the range, Indicates site with site The weights between them , For site with site The LD coefficient between them.
[0055] Traditional methods such as mean imputation or median imputation cannot account for linkage relationships between loci in the genome. This method, however, utilizes the local LD structure for imputation, enabling more accurate estimation of missing values. In practical applications, the LD coefficient r² between all locus pairs is first calculated to construct an LD matrix. Then, for each missing locus, a set of neighboring loci within the LD threshold range (r² > 0.6) is determined. Based on a weighted average of the genotype values and LD strengths of these neighboring loci, the genotype value of the missing locus is estimated. This method is particularly suitable for self-pollinating crops (such as rice and wheat) where LD decays slowly, achieving an imputation accuracy of over 90%.
[0056] The method for quality assessment of sample data includes: calculating the sample genotype coverage, calculating the conversion / transversion ratio in the genotype data, and calculating the genotype error rate.
[0057] Samples were screened based on the quality assessment results, and genotype coverage was excluded. Low coverage often indicates insufficient sequencing depth or poor DNA quality. Our method sets an 85% coverage threshold based on empirical values derived from extensive experimental data analysis, demonstrating good performance in balancing data integrity and sample retention. For example, in a genotyping analysis of 1000 rice varieties, an 85% coverage threshold eliminated approximately 5% of low-quality samples, which could introduce noise in subsequent analyses, while the remaining 95% of samples provided sufficient information for downstream analysis.
[0058] Elimination conversion / transformation ratio Abnormal deviation from the population The sample, and These represent the expected and variance of the population transition / transition ratio, respectively.
[0059] Genotype error rate elimination The sequencing batch data; setting a 5% error rate threshold takes into account the performance characteristics and cost-effectiveness of high-throughput sequencing platforms under the current technology level; taking molecular marker-assisted breeding of crops as an example, when the error rate exceeds 5%, the reliability of genotype data for parent identification and kinship analysis will be significantly reduced, which may lead to breeding decision errors; each sequencing batch contains at least 2-3 replicate samples in order to accurately calculate the error rate.
[0060] Sample genotype coverage The formula for calculation is: ,in, For the sample Number of effective genotype loci in the middle This represents the total number of sites.
[0061] Transformation / transversion ratio in genotype data The formula for calculation is: ,in, This represents the number of conversion mutations. The conversion mutations include: A G or C T, The number of transversion mutations, including: A C, A T, G C or G T. The left and right sides represent the genes before and after the mutation, respectively.
[0062] Genotype error rate The formula for calculation is: ,in, This represents the number of samples that were repeated for sequencing. For duplicate samples Number of inconsistent sites For duplicate samples The total number of sites in the community.
[0063] The method for removing batch effects includes: constructing a location effect index. ,in Indicates batch midpoint The average genotype, and They are the sites Mean and standard deviation across all batches For batch The sequencing time is in days. and The mean and standard deviation of sequencing times for all batches are represented, respectively. This index is constructed based on the observed systematic changes in batch effects over sequencing time; a LI value greater than 2.5 indicates a significant systematic bias at that locus within a specific batch. For example, in a two-year soybean genome sequencing project comprising 20 batches, approximately 15% of SNP loci exhibited significant batch effects, primarily concentrated in regions with high GC content. After batch correction using this method, the systematic differences between batches were significantly reduced, and principal component analysis (PCA) results before and after correction showed that batch clustering was essentially eliminated.
[0064] right Batch correction is performed on the sites, and the correction formula is: ,in For site The batch effect coefficient is obtained by minimizing the corrected inter-batch variance: , Indicates site Genotype value vectors across all batches. Batch correction coefficients are estimated by minimizing the corrected inter-batch variance; essentially, this is a linear regression-based correction method. These coefficients reflect the strength of the trend in locus j over sequencing time. The higher the value, the more significantly the sequencing time affects the site.
[0065] For example, in a whole-genome sequencing analysis of rice, we found that about 8% of SNP sites showed a significant linear time trend, and the batch-to-batch variance of these sites was reduced by an average of 78% after correction.
[0066] Step S2: Using the preprocessed genotype data, calculate the genetic similarity matrix between samples and perform linkage disequilibrium analysis to identify tag SNPs as molecular markers.
[0067] The calculation of the genetic similarity matrix between samples includes the following steps:
[0068] Based on the preprocessed genotype data, for any two samples and Calculate samples and samples At the site correlation on :
[0069] When two samples have identical genotypes When sharing an allele Completely different times .
[0070] If two samples share a rare variant allele at a locus with a major allele frequency of 0.95, this is a stronger indication of their close kinship than if they share it at a locus with an allele frequency of 0.5; the weighted genetic similarity matrix shows about 15-20% higher accuracy than the simple matching coefficient method in distinguishing closely related samples and identifying parents.
[0071] The locus association scoring scheme is based on the IBS (Identity By State) principle, intuitively reflecting the degree of similarity between samples at the molecular level. For example, for an SNP locus, if sample A has the genotype AA (coded as 0) and sample B also has the genotype AA, then S=1; if sample B has the genotype AG (coded as 1), then they share an A allele, and S=0.5; if sample B has the genotype GG (coded as 2), then the two samples do not share any alleles, and S=0. This three-level scoring scheme is applicable to diploid species.
[0072] For polyploid species (such as wheat and cotton), the score can be expanded to a more detailed multi-level rating. In an analysis of hexaploid wheat, the correlation was subdivided into seven levels: 0, 0.17, 0.33, 0.5, 0.67, 0.83 and 1.0, which more accurately captured the strength of the kinship between different samples.
[0073] Calculate the sum of associations between two samples across all common non-missing sites: ,in This represents the number of common non-deleted SNP sites.
[0074] Calculate genetic similarity based on allele frequency weighting: ;in, , for site The weight, For site The minor allele frequency (pk) is considered. Loci with minor allele frequencies (pk) close to 0 or 1 have higher weights because variations at these loci are rare, and the probability of samples sharing these rare variations is low, thus containing more information about kinship. For example, in population genome studies, a rare variation with a frequency of 0.01 may have several times higher discriminative power than a common variation with a frequency of 0.5. To avoid excessive weighting due to extreme frequencies (such as pk close to 0), a minimum frequency threshold (such as 0.01) can be set, or empirical Bayesian methods can be used to smooth the weights.
[0075] Constructing a genetic similarity matrix of samples , yes 3D matrix This represents the number of samples after preprocessing. Linkage disequilibrium analysis specifically involves calculating the linkage disequilibrium coefficient between any two SNP loci using preprocessed genotype data. : ,in, , For the frequency of haplotype AB, and The frequencies of the major alleles at loci A and loci B are respectively. and These represent the frequencies of the minor alleles at loci A and loci B, respectively; r²=0 indicates that the two loci are completely independent, and r²=1 indicates that they are completely linked.
[0076] Constructing a chain-disequilibrium decay model: ,in, The effective population size is given by d, which represents the genetic distance between two loci in moles.
[0077] Through chain imbalance half-life Assess the extent of linkage disequilibrium in the population, i.e. At that time, the chain imbalance half-life is: In comparative studies of various crops, the half-life distance of indica rice varieties was approximately 123 kb, while that of japonica rice was approximately 167 kb, reflecting that the japonica rice population experienced a stronger domestication bottleneck. The half-life distance of maize inbred line populations was approximately 10-30 kb, while that of wild maize relatives was only 2-5 kb, reflecting significant changes in effective population size during domestication and breeding.
[0078] The identification of SNPs as molecular markers specifically involves constructing a SNP site network based on linkage disequilibrium analysis. Where the vertex set V represents all SNP sites, and the edge set E represents linkage disequilibrium, when two sites are in a state of linkage disequilibrium... Greater than the threshold At that time, there is an edge between these two sites.
[0079] Calculate the degree centrality of each SNP site: ,in This is an indicator function.
[0080] For each connected component, the SNP site with the highest degree centrality is selected as the label SNP of the connected component.
[0081] If multiple loci have the same degree centrality, the locus with the minor allele frequency closest to 0.5 is selected as the tag SNP.
[0082] Calculate the capture rate of the tag SNP: Where T is the set of label SNPs, This represents the total number of SNP sites.
[0083] Capture rate (CR) is a key indicator for evaluating the representativeness of labeled SNPs, quantifying the extent to which the set of labeled SNPs can represent the genetic variation of all SNP loci. A CR value closer to 1 indicates better representativeness of the labeled SNP. In practical applications, a trade-off usually needs to be struck between the number of labeled SNPs (cost) and the capture rate (information content). Empirically, when the threshold τ is set to 0.8, the CR can typically reach 0.85-0.95, while reducing the number of markers by 70-85%.
[0084] Visualizing the relationship between capture rate and the number of tagged SNPs helps determine the optimal balance. For example, in a soybean breeding project, the initial 420,000 SNPs were screened using a threshold of τ=0.8, resulting in 36,000 tagged SNPs with a capture rate of 0.91. Further relaxing the threshold to τ=0.7 increased the number of tagged SNPs to 48,000 and improved the capture rate to 0.95, but the cost increased by 33%. Ultimately, the τ=0.8 scheme was chosen as the optimal balance.
[0085] Step S3: Estimate population kinship and predict hybrid vigor based on the obtained molecular markers.
[0086] The estimation of population kinship based on the obtained molecular markers includes: using the identified label SNPs as molecular markers to calculate the common ancestor coefficient matrix among samples. :
[0087] ,in For the number of SNPs in the label, Indicates sample In the Genotype values on each tag SNP For the first Secondary allele frequencies of tag SNPs.
[0088] Calculate the kinship matrix for population structure correction ,in The population structure matrix was obtained using the population structure analysis software STRUCTURE. This is a diagonal matrix of differentiation coefficients between different subgroups.
[0089] A rootless tree is constructed based on a phylogenetic tree derived from a kinship matrix. The branch lengths of the tree reflect the genetic distance between samples. .
[0090] The method for predicting hybrid vigor specifically involves calculating the expected heterozygosity of the F1 hybrids based on the parental genotypes. : ,in, and Representing the parents Japanese marriage At the site The genotype value.
[0091] Calculate the relationship index between genetic distance between parents and hybrid vigor: ,in For parent and The genetic distance between them.
[0092] Constructing a model for predicting hybrid vigor: ,in The predicted hybrid vigor value, , , and The regression coefficients are estimated using the least squares method based on known hybridization phenotypic data and corresponding molecular marker data.
[0093] Based on biological knowledge, quadratic terms The values are typically negative, reflecting a potential decrease in fitness due to excessively high heterozygosity. Regarding model fitting, it is recommended to use cross-validation methods to evaluate predictive performance, such as 10-fold cross-validation or leave-one-out cross-validation. For example, in a study on maize heterosis, this model was used to predict the yield of 380 hybrid combinations. The cross-validation prediction accuracy reached 0.72, significantly higher than the 0.58-0.65 of traditional prediction methods, and it performed stably under different environmental conditions, validating the model's practicality and reliability.
[0094] An example of the application of the method of this invention in rice breeding is as follows: Genome resequencing data of 520 rice varieties (including indica, japonica, and intermediate types) from different regions were collected. The sequencing depth was 15-30X, the coverage was greater than 96%, and the total amount of raw sequencing data reached 35TB. The data came from 28 sequencing batches spanning 3 years, using 3 different sequencing platforms (Illumina HiSeq 2500, NovaSeq 6000, and MGI DNBSEQ-T7).
[0095] Using the data preprocessing method of this invention, data from different platforms were first converted into a standardized format, resulting in the detection of approximately 4.3 million high-quality SNP loci. Quality assessment revealed that 18 samples had genotype coverage below 85%, 12 samples had a conversion / transversion ratio (average 2.14 ± 0.18) that deviated abnormally from the population by more than 3 standard deviations, and 3 sequencing batches had genotype error rates exceeding 5%. These low-quality data were discarded, leaving 487 valid samples.
[0096] Using batch effect correction methods, approximately 12.5% of SNP sites (538,000 sites) were identified as exhibiting significant batch effects (LI>2.5), mainly concentrated between batches with large sequencing time spans. After correction, the coefficient of variation for these sites across different batches decreased from an average of 0.28 to 0.04, and batch clustering was essentially eliminated in PCA analysis.
[0097] Linkage disequilibrium analysis showed that the LD half-life distance of the rice population was approximately 150 kb, with the indica rice subgroup (approximately 123 kb) shorter than that of the japonica rice subgroup (approximately 167 kb), reflecting the historical differences between the different subgroups. A SNP network was constructed based on a threshold of r² > 0.8, resulting in approximately 465,000 connected components, each containing an average of 9.2 SNPs. The most representative sites in each connected component were selected based on degree centrality, ultimately identifying 48,723 tag SNPs with a capture rate (CR) of 0.93, reducing the number of molecular markers by 88.7% while preserving most of the genetic variation information.
[0098] Based on these SNP tags, the phylogenetic relationships between samples were calculated. After population structure correction, a phylogenetic tree was constructed, accurately reflecting the genetic relationships among rice varieties, with 92.5% consistency with known pedigree information. In simulated hybridization experiments, the expected heterozygosity and hybrid vigor index of all possible hybridization combinations (approximately 118,000 combinations) were calculated. Combined with yield data from 325 existing hybridization combinations, a hybrid vigor prediction model was established.
[0099]
[0100] The model achieved a prediction accuracy of 85.3% in 10-fold cross-validation, significantly higher than traditional prediction methods based on parental phenotypes (approximately 70%). Based on this model, 50 hybrid combinations with high heterosis potential were recommended, of which 12 were confirmed in field trials to have yields exceeding the average of their parents by more than 45%, validating the application value of the method in practical breeding.
[0101] Through this embodiment, the application of the method of the present invention in rice breeding significantly improves the quality of genotype data, reduces the number of redundant markers, accurately estimates the kinship between varieties, and accurately predicts hybrid vigor, providing strong technical support for molecular design breeding of rice, and can be extended to the molecular breeding process of other crop varieties.
[0102] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A high-throughput intelligent genotyping method, characterized in that, The method includes the following steps: Step S1: Collect high-throughput genotype data, standardize the sample data format, and perform quality assessment and batch effect removal on the sample data to obtain preprocessed genotype data. The method for quality assessment of sample data includes: calculating the sample genotype coverage, calculating the conversion / transversion ratio in the genotype data, and calculating the genotype error rate; and screening samples based on the quality assessment results, removing those with low genotype coverage. Samples; elimination of conversion / transversion ratio Abnormal deviation from the population The sample, and These represent the expected and variance of the population transition / transversion ratio; and the genotype removal error rate. Sequencing batch data; The method for removing batch effects includes: constructing a location effect index. ,in Indicates batch midpoint The average genotype, and They are the sites Mean and standard deviation across all batches For batch The sequencing time is in days. and These represent the mean and standard deviation of sequencing times for all batches. right Batch correction is performed on the sites, and the correction formula is: ,in For site The batch effect coefficient is obtained by minimizing the corrected inter-batch variance: , Indicates site Genotype value vectors across all batches; Step S2: Using the preprocessed genotype data, calculate the genetic similarity matrix between samples and perform linkage disequilibrium analysis to identify tag SNPs as molecular markers; A network of SNP sites was constructed based on linkage disequilibrium analysis. Where the vertex set V represents all SNP sites, and the edge set E represents linkage disequilibrium, when two sites are in a state of linkage disequilibrium... Greater than the threshold At that time, there is an edge between these two sites; Calculate the degree centrality of each SNP site: ,in It is an indicator function; For each connected component, the SNP site with the highest degree centrality is selected as the label SNP of the connected component; If multiple loci have the same degree centrality, the locus with the minor allele frequency closest to 0.5 is selected as the tag SNP. Calculate the capture rate of the tag SNP: Where T is the set of label SNPs, This represents the total number of SNP sites; Step S3: Estimate population kinship and predict hybrid vigor based on the obtained molecular markers; The method for predicting hybrid vigor specifically involves calculating the expected heterozygosity of the F1 hybrids based on the parental genotypes. : ,in For the number of SNPs in the label, and Representing the parents Japanese marriage At the site Genotype values; Calculate the relationship index between genetic distance between parents and hybrid vigor: ,in For parent and Genetic distance between them; Constructing a model for predicting hybrid vigor: ,in The predicted hybrid vigor value, , , and The regression coefficients are estimated using the least squares method based on known hybridization phenotypic data and corresponding molecular marker data.
2. The high-throughput intelligent genotyping method according to claim 1, characterized in that, The standardized sample data format is specifically as follows: Genotype data obtained from different sequencing platforms were uniformly converted into a standard format; genotype coding adopted... Numerical representation, where 0 represents the homozygous major allele AA, 1 represents the heterozygous genotype Aa, and 2 represents the homozygous minor allele aa; Missing data are filled using a multiple imputation method based on the locally linked unbalanced LD structure. The calculation formula is as follows: ,in Indicates sample At the site The genotype values to be inserted on the above, Indicates the location At LD threshold The set of all non-missing sites within the range, Indicates site with site The weights between them , For site with site The LD coefficient between them.
3. The high-throughput intelligent genotyping method according to claim 2, characterized in that, Sample genotype coverage The formula for calculation is: ,in, For the sample Number of effective genotype loci in the middle This represents the total number of sites; Transformation / transversion ratio in genotype data The formula for calculation is: ,in, The number of conversion mutations; the conversion mutations include: A G or C T, The number of transversion mutations, including: A C, A T, G C or G T; The left and right sides represent the genes before and after the mutation, respectively; Genotype error rate The formula for calculation is: ,in, This represents the number of samples that were repeated for sequencing. For duplicate samples Number of inconsistent sites For duplicate samples The total number of sites in the community.
4. The high-throughput intelligent genotyping method according to claim 3, characterized in that, The calculation of the genetic similarity matrix between samples includes the following steps: Based on the preprocessed genotype data, for any two samples and Calculate samples and samples At the site correlation on : When two samples have identical genotypes When sharing an allele Completely different times ; Calculate genetic similarity based on allele frequency weighting: ; in, , for site The weight, For site Secondary allele frequencies.
5. The high-throughput intelligent genotyping method according to claim 4, characterized in that, The linkage disequilibrium analysis specifically involves calculating the linkage disequilibrium coefficient between any two SNP loci using preprocessed genotype data. : ,in, , For the frequency of haplotype AB, and The frequencies of the major alleles at loci A and loci B are respectively. and These represent the frequencies of the minor alleles at locus A and locus B, respectively. Constructing a chain-disequilibrium decay model: ,in, The effective population size is given by d, where d is the genetic distance between two loci, expressed in moles. Through chain imbalance half-life Assess the extent of linkage disequilibrium in the population, i.e. At that time, the chain imbalance half-life is: .
6. The high-throughput intelligent genotyping method according to claim 1, characterized in that, The estimation of population kinship based on the obtained molecular markers includes: using the identified label SNPs as molecular markers to calculate the common ancestor coefficient matrix among samples. : ,in Indicates sample In the Genotype values on each tag SNP For the first Secondary allele frequency of each tag SNP; Calculate the kinship matrix for population structure correction ,in The population structure matrix was obtained using the population structure analysis software STRUCTURE. This is a diagonal matrix of differentiation coefficients among different subgroups; A rootless tree is constructed based on a phylogenetic tree derived from a kinship matrix. The branch lengths of the tree reflect the genetic distance between samples. .
Citation Information
Patent Citations
Method and device for predicting target gene copy number type
CN116453590A
Gene typing error detection and correction method based on linkage unbalance degree
CN118737292A