Genotype filling optimization method and reference panel for picea norvegensis
By constructing a haplotype reference panel and optimizing parameters, and using IMPUTE5 software and genetic maps, the accuracy and efficiency issues of genotype filling in Norway spruce were solved, achieving high-precision genotype filling and filling a technological gap in forest species.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING FORESTRY UNIV
- Filing Date
- 2025-12-24
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, genotype filling in Norway spruce is inaccurate and inefficient, and lacks a systematic evaluation and optimization process, which prevents researchers from effectively utilizing the genomic data of this tree species.
By constructing a haplotype reference panel, using IMPUTE5 software and integrating the Norway spruce genetic map, setting a strict quality control threshold of Rsqsoft ≥ 0.8, optimizing the genotype filling process, and establishing a standardized method.
It achieved high-precision, whole-genome resolution genotype filling, solving the problems of the huge genome and numerous repetitive sequences in Norway spruce, and providing reliable data support for genetic map construction and breeding research.
Smart Images

Figure CN121963845A_ABST
Abstract
Description
A genotypic filling optimization method for Norway spruce and a reference panel Technical Field
[0001] This invention belongs to the field of forest tree genetics technology, specifically relating to a genotype filling optimization method and reference panel for Norway spruce. Background Technology
[0002] Genotype imputation is a bioinformatics technique that uses statistical models to infer missing genotypes in target samples based on haplotype structure and linkage disequilibrium information from reference panels. Its core value lies in its ability to bridge the data density gap between different genotyping technology platforms (such as GBS or SNP microarrays) and whole-genome resequencing at extremely low cost. Supported by high-quality reference panels such as the 1000 Genomes Study and the 1000-Bull Genomes Study, researchers can imput low-density data to near-genome-wide levels, significantly enhancing the utilization value of the data. This technique has demonstrated enormous potential in several fields: in genome-wide association studies (GWAS), it increases marker density, thereby improving the accuracy of quantitative trait locus localization; in genomic selection (GS), predictive models based on imputed high-density genotypes have been proven to consistently improve the accuracy of breeding value estimation. Therefore, genotype imputation has become a standard and crucial analytical step in human and plant / animal genome research.
[0003] However, the successful application of genotype filling heavily relies on large-scale, high-quality haplotype reference panels that match the genetic background of the target population. This is a major bottleneck for forest species, especially conifers like Norway spruce. Norway spruce possesses a large and highly repetitive genome, leading to low accuracy in sequence alignment and variant detection, posing a fundamental challenge to constructing high-quality reference panels. Furthermore, there is currently a complete lack of publicly available haplotype reference resources covering its entire distribution area with a large sample size.
[0004] More importantly, establishing a reliable and efficient genotyping process cannot be achieved solely through a reference panel. The accuracy of genotyping is profoundly influenced by a series of key parameters and technological choices. For example, different statistical models exhibit varying performance under specific genomic structures; the impact of integrating genetic map information on genotyping integrity is significant; and the quality filtering threshold for genotyping results must be set to ensure the reliability of downstream data. Currently, for Norway spruce, there is a complete lack of systematic evaluation and clear guidance on how these key parameters affect genotyping accuracy. This technological gap means that researchers attempting genotyping in Norway spruce lack optimized, reproducible standard procedures to follow, forcing them to rely on experience from other species or engage in blind parameter experimentation. This severely limits the scientific value and application potential of genotyping technology in this important tree species.
[0005] In summary, the industry urgently needs to develop a systematically evaluated and optimized genotype filling process specifically for Norway spruce, which is no less urgent than the construction of the reference panel itself. Summary of the Invention
[0006] In view of this, the purpose of this application is to provide a genotype filling optimization method and reference panel for Norway spruce, so as to solve the problems of low accuracy and low efficiency of genotype filling in the prior art.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] Firstly, this invention provides a method for genotypic optimization of Norway spruce. The core of this method lies in determining an optimal combination of parameters through systematic evaluation and establishing a standardized process based on this. Specifically, it includes the following steps:
[0009] First, a haplotype reference panel was constructed, its key feature being its sufficient size to encompass the major haplotype diversity of the Norway spruce population. This reference panel was built based on whole-genome resequencing data from unrelated Norway spruce individuals. Second, low-density genotype data of the Norway spruce samples to be filled were obtained. Subsequently, using the IMPUTE5 software and integrating the Norway spruce genetic map, genotype filling was performed on the low-density genotype data based on the constructed reference panel. Finally, the filling results underwent rigorous quality control, retaining SNPs with Rsqsoft scores no lower than 0.8, thereby obtaining high-quality whole-genome filled genotype data.
[0010] In some embodiments, to ensure that the reference panel is sufficiently saturated, the haplotype reference panel may be constructed based on data from no less than 700 individuals.
[0011] In some embodiments, the low-density genotype data may be cost-effective genotyping data such as SNP chip data or exon capture sequencing data.
[0012] In some embodiments, the haplotype reference panel can be efficiently and accurately inferred and constructed using SHAPEIT v5.1.1 software.
[0013] Secondly, the present invention provides a Norway spruce genotype filling reference panel for implementing any of the above methods. This reference panel is constructed based on whole-genome resequencing data of unrelated Norway spruce individuals, and its size is sufficient to cover the major haplotype diversity of the Norway spruce population, thereby providing a stable and complete information basis for genotype filling.
[0014] In some embodiments, the reference panel may be specifically constructed based on data from no fewer than 700 individuals.
[0015] Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the methods described in the first aspect above.
[0016] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described in the first aspect above.
[0017] Compared with the prior art, the beneficial effects of this application are as follows:
[0018] This invention provides a method for optimizing genotype filling in Norway spruce and a reference panel. Through systematic experiments, it reveals for the first time the key factors affecting the accuracy of genotype filling in Norway spruce and establishes the first validated standardized filling process for this species. The core of this invention lies in the discovery and integration of three key optimization elements: the use of IMPUTE5 software, the integration of Norway spruce-specific genetic maps, and the setting of a strict quality control threshold of Rsqsoft ≥ 0.8. The synergistic effect of these three elements significantly improves the accuracy and completeness of the filling results. More importantly, this invention clarifies the existence of a "scale saturation effect" in Norway spruce reference panels. The dedicated reference panel constructed based on this effect can efficiently cover the main genetic diversity of the population with an optimal sample size (approximately 1000 individuals), fundamentally solving the technical bottleneck of establishing high-quality reference resources due to the large genome and numerous repetitive sequences of this tree species. This invention enables the efficient and reliable upgrading of economical low-density data (such as chip or targeted sequencing data) to whole-genome resolution, completely filling the gap in the field of high-precision genotype filling technology for Norway spruce, and providing an indispensable basic tool and data support for its large-scale genetic map construction, gene mining and molecular design breeding research. Attached Figure Description
[0019] Figure 1 shows a comparison of the performance of three filler software programs (IMPUTE5, Beagle, and Minimac4) under two scenarios: "using genetic mapping (+Map)" and "not using genetic mapping (-Map)," as the Rsqsoft threshold increases. (a) Number of SNPs retained after filtering; (b) Genotype concordance; (c) Correlation. );
[0020] Figure 2 shows the filling performance of 50 individuals resequencing on chromosome 11 at different marker densities, based on 997 reference panels.
[0021] Figure 3 shows the trends in genotype filling consistency and correlation under different minor allele frequency ranges. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention is further described below with reference to specific embodiments. Unless otherwise described in detail, the technical means used in the following embodiments are all conventional means well known to those skilled in the art. Alternatively, they may be carried out according to the kit and product instructions. Unless otherwise specified, the materials and reagents used in the following embodiments are commercially available.
[0023] Example 1
[0024] 1. Reference panel construction and genotype filling optimization methods
[0025] 1.1 Accuracy assessment of genotype filling
[0026] To systematically evaluate the reliability of genotype filling in Norway spruce and identify its key regulatory factors, this application, based on the construction of a reference panel of 1,047 unrelated high-depth resequencing haplotypes, used 50K chips and exon capture data as filling targets, and designed five gradient experiments to examine the core variables that may affect filling accuracy. The specific scheme is as follows:
[0027] Software integration and genetic mapping: SHAPEIT v5.1.1 was used to construct a haplotype reference panel from 997 unrelated individuals with high-depth resequencing. Subsequently, 50 overlapping samples with 50K microarray data were selected as fillers. IMPUTE5 v1.2.1, Beagle v5.4, and Minimac4 v4.1.6 were run on chromosome 11 to systematically compare the effects of integrating the genetic map on the number of SNPs obtained, genotypic concordance, and correlations. The influence of ) was investigated to determine the optimal software configuration for filling the Norway spruce genotype.
[0028] Marker density: To examine the impact of marker density on imputation accuracy, 1K, 2K, 3K, 4K, 5K, 10K, 25K, 50K, and 100K SNPs were randomly selected from the resequencing data of chromosome 11 of 50 overlapping individuals, with each density being independently replicated three times. Using a haplotype reference panel constructed from 997 samples as a baseline, IMPUTE5 v1.2.1 was used to imput the genotypes of the target individuals, and the consistency and correlation between the obtained genotypes and the actual data were calculated.
[0029] Reference population size: 50, 100, 400, and 700 individuals were randomly selected from 1,047 unrelated resequencing samples. Haplotype reference panels of different sizes were constructed using SHAPEIT v5.1.1. Genotyping of chromosome 11 was then performed on 1,297 exon capture individuals using IMPUTE5 v1.2.1. Each experiment was independently repeated three times: 500 SNPs from the exon capture data were initially left blank, and the remaining 13,612 loci were used as the filling backbone. Finally, the consistency and correlation between the filled results and the true genotypes were calculated.
[0030] Filler population size: A haplotype reference panel was constructed using 1,047 unrelated samples. From 1,297 exon-captured individuals, 200, 400, 600, 800, or 1,000 individuals were randomly selected as filler targets for chromosome 11. Each round of experiments was repeated three times, masking 500 SNPs and retaining 13,612 loci for filler. The concordance and correlation between the obtained genotypes and the actual data were assessed. ).
[0031] Minor allele frequencies (MAFs): To assess the impact of minor allele frequencies (MAFs) on imputation accuracy, 700 haplotype reference panels were randomly selected from 1,047 unrelated individuals, and the remaining 347 individuals were used as the target sample. Approximately 4,000 SNP loci were randomly selected for imputation. After imputation, the loci were divided into 18 left-closed and right-open intervals according to MAFs: the intervals were divided with a step size of 0.01 in the range of 0-0.1 and with a step size of 0.05 in the range of 0.1-0.5. The consistency and correlation between the imputation results and the true genotypes were calculated.
[0032] 1.2 Calculation of Filling Accuracy
[0033] To comprehensively evaluate the accuracy of genotype filling, this application uses two complementary indicators: "consistency" and "correlation".
[0034] Genotypic consistency. Based on the resequencing genotype, the maximum a posteriori probability genotype output by IMPUTE5 is compared with it. The genotype coding rule is: 0 / 0 → 0, 0 / 1 or 1 / 0 → 1, 1 / 1 → 2. Locus-level consistency is defined as:
[0035]
[0036] Further averaging across all sites yields overall consistency:
[0037]
[0038] Genotype correlation. The resequencing genotypes (GT field) of all samples are converted into allele count vectors. (Values 0, 1, 2); The maximum posterior probability genotype (GT field) output by IMPUTE5 is also converted into an allele count vector. Then calculate and The square of the Pearson correlation coefficient :
[0039]
[0040] 1.3 Genotyping and Quality Control
[0041] Based on 1,047 unrelated high-depth resequencing samples, this application first constructed a haplotype reference panel using SHAPEIT v5.1.1, and after comparing different filling schemes and filtering thresholds, determined to use IMPUTE5 v1.2.1 for exon capture data genotyping. Before filling, the following quality control steps were performed on the data: (1) removing insertion / deletion markers, (2) retaining only biallelic loci, (3) marking loci with genotype quality value (GQ) < 10 as deleted, (4) removing loci with a detection rate < 70%, (5) removing loci with MAF < 0.05, and (6) for the entire superior tree species population, removing SNPs with excessive heterozygosity and deviations from the Hardy-Weinberg equilibrium test (HWE test) (P value < 1×10⁻⁶). -5 (7) Remove individuals with a detection rate < 50%. After the above filtering, a total of 221,168 loci with MAF > 0.05 were obtained for filling. After completing the genotyping, the obtained data were further filtered, and only those with Rsqsoft ≥ 0.8, MAF > 0.05 and HWE P value > 1×10 were retained. -5 The study identified 11,958,254 high-quality SNPs for subsequent GWAS. For the SNP microarray data, sites not anchored to autosomes were first removed, yielding 43,267 SNPs. Subsequently, genotyping was performed on the filtered microarray data using IMPUTE5v1.2.1, and sites with Rsqsoft < 0.8 after imputation were further removed, resulting in 3,165,176 high-quality SNPs for subsequent GS.
[0042] 2. Results
[0043] 2.1 The Influence of Filling Software and Genetic Mapping on Filling Accuracy
[0044] To evaluate the impact of the filtering threshold on imputation performance, the performance of three genotype imputation software programs (IMPUTE5, Beagle, and Minimac4) was compared under both integrated and unintegrated genetic maps conditions. Specifically, the effect of different Rsqsoft values on SNP count, genotype consistency, and relevance (R²) was analyzed, and the results are shown in Figure 1. With increasing Rsqsoft, the number of SNPs retained by all three methods decreased. However, after introducing the genetic map, the number of SNPs retained by IMPUTE5 and Beagle significantly increased, indicating that genetic information can effectively improve imputation integrity at low thresholds. In contrast, Minimac4 showed a stable upward trend in consistency and relevance under different thresholds and with and without the use of a genetic map, demonstrating strong robustness. For Beagle and IMPUTE5, although the overall imputation accuracy improved with increasing threshold, when Rsqsoft was between 0.1 and 0.6, the imputation accuracy without the genetic map was slightly higher than that with the map, which may be related to the difference in the number of SNPs within this threshold range. When Rsqsoft is set to ≥ 0.8, considering the three indicators of SNP number, consistency and correlation (Table 1), IMPUTE5 shows better overall performance in genotyping of Norwegian spruce, and is therefore recommended as the preferred filling tool in this application.
[0045] Table 1. Number of SNPs, Filling Consistency, and Correlation Obtained by Different Filling Software when Rsqsoft ≥ 0.8
[0046]
[0047] 2.2 The effect of marker density on filling accuracy
[0048] With a reference panel consisting of 997 unrelated individuals, this application systematically evaluated the impact of marker density on chromosome 11 of the filling population on filling accuracy. As shown in Figure 2, both genotypic concordance and correlation (R²) monotonically increased with the number of markers; when the number of markers increased to approximately 25,000 loci, they climbed to 0.963 and 0.901, respectively, after which the increase tended to plateau. This result indicates that, under the current conditions of the Norway spruce population, chromosomal region, and reference panel, a density of approximately 25,000 markers is sufficient to effectively capture the linkage disequilibrium (LD) structural information necessary to support high-precision genotyping, bringing the filling accuracy close to the potential upper limit of this technology. While further increasing the marker density may further improve the theoretical limit, its marginal benefit to filling accuracy is already very limited. From a cost-effectiveness perspective, this density can serve as a reasonable threshold for balancing filling performance and genotyping costs under similar conditions.
[0049] 2.3 The impact of the size of the reference population on the accuracy of the filling
[0050] With a fixed population size of 1,297, this application compared the impact of different reference panel sample sizes (50, 100, 400, 700, and 1,047) on the accuracy of filling exon capture data on chromosome 11 (Table 2). The results showed that genotype concordance and correlation (R²) fluctuated only slightly with increasing reference sample size, with no statistically significant differences. These results suggest that when the reference panel is constructed based on high-depth whole-genome sequencing, its haplotype library is already sufficiently saturated, and further increasing the reference population size is unlikely to significantly improve the filling accuracy of Norway spruce.
[0051] Table 2. Changes in chromosome 11 filling accuracy with reference panel sample size (fill population fixed at 1,297 cases)
[0052]
[0053] 2.4 The impact of the size of the infill population on infill accuracy
[0054] After determining the reference panel size to be 1,047 cases, the impact of the filler population sample size (200, 400, 600, 800, 1,000, and 1,297) on the accuracy of chromosome 11 filling was further evaluated (Table 3). The results showed that both concordance and correlation (R²) remained constant with increasing filler population size, without significant fluctuations, confirming that IMPUTE5 is also insensitive to filler cohort size even with a high-depth reference panel. Combined with the aforementioned result that "filler accuracy remained high and showed no significant improvement when the reference panel sample size increased from 100 to 1,047," this indicates that once the reference haplotype library reaches sufficient saturation based on high-depth whole-genome sequencing data, neither further expanding the reference sample nor increasing the filler cohort can significantly improve the genotype filling accuracy of Norway spruce.
[0055] Table 3. Effect of the filler population sample size on the accuracy of chromosome 11 filling when the reference panel is fixed at 1,047 cases.
[0056]
[0057] 2.5 Impact of minor allele frequencies on filling accuracy
[0058] To understand the impact of MAF on filling accuracy, this application randomly selected 700 cases from 1,047 high-depth resequencing samples to construct a reference panel, and the remaining 347 cases as the filling target population. 4,000 loci were randomly selected for filling experiments. The MAF was divided into 18 left-closed, right-open intervals: [0, 0.01), [0.01, 0.02), … [0.09, 0.10), [0.10, 0.15), …, [0.45, 0.50). The filling performance of different frequency intervals was systematically evaluated. As shown in Figure 3, genotypic concordance and correlation (R²) generally increased with increasing MAF. The filling accuracy of low-frequency variants (MAF < 0.05) was extremely sensitive to frequency changes, with R² showing a steep increase. However, when MAF > 0.05, the curve tended to flatten, indicating that the filling accuracy of common variants was close to saturation, and further increasing the frequency had limited contribution to accuracy.
[0059] 3. Technical effects achieved by the present invention
[0060] Genotyping technology can efficiently deduce whole-genome information from economic genotyping data. However, the performance and effectiveness of this technology in forest species (such as Norway spruce) have not been systematically analyzed. This application is the first to comprehensively evaluate the key parameters affecting genotyping performance for Norway spruce and establish an optimized and reliable filling process. The main technical effects achieved by this application are as follows: (1) It clarifies that genetic map, filling software, quality control threshold (Rsqsoft statistic), marker density, reference and filling population size, and MAF are the key factors determining filling accuracy, and determines their optimal configuration; (2) It constructs a high-quality, large-scale haplotype reference panel specifically for Norway spruce.
[0061] 3.1 The impact of imputation method, genetic map, and different thresholds on imputation accuracy
[0062] This application systematically compared the performance of three imputation methods: IMPUTE5, Beagle, and Minimac4. Since true genotypes are often unavailable, these software programs generate an internal quality index, Rsqsoft, after imputation. The results of this application show that, when pursuing high-precision imputation (e.g., retaining sites with Rsqsoft ≥ 0.8), IMPUTE5 outperforms the other two software programs in terms of the number of SNPs obtained, genotype consistency, and correlation, thus identifying it as the preferred tool for genotype imputation of Norway spruce.
[0063] Integrating genetic maps significantly improves the filling integrity of IMPUTE5 and Beagle, allowing them to retain more SNPs at the same Rsqsoft threshold. Minimac4, however, offers limited improvement because its algorithm does not rely on genetic maps. Without genetic maps, IMPUTE5 and Beagle linearly convert physical distance to genetic distance by default, which can easily lead to inaccurate filling window settings and systematically underestimate the Rsqsoft values of low-frequency sites, resulting in a significant reduction in the number of filtered SNPs.
[0064] For common variants under Hardy-Weinberg equilibrium, the Rsqsoft values of different software are comparable; however, for rare variants (MAF ≤ 0.05), Minimac4's evaluation is often more conservative, with generally lower Rsqsoft values. The results of this application are consistent with this: when only sites with Rsqsoft ≥ 0.4 are retained, as the filtering criteria become more stringent, Minimac4 retains significantly fewer SNPs than other methods.
[0065] 3.2 Effects of marker density, reference population and impregnated population size, and minor allele frequency on impregnation accuracy
[0066] After determining that IMPUTE5 was a suitable tool, this application further investigated the effects of marker density, population size, and MAF. The results showed that filling accuracy improved with increasing marker density, reaching a plateau at a density of 25K. This indicates that this density is sufficient to effectively capture most of the key linkage disequilibrium information on chromosome 11 of Norway spruce. This finding provides crucial design guidance for subsequent large-scale genetic studies: in practice, microarray design or targeted sequencing protocols for Norway spruce do not need to blindly pursue excessively high marker densities, but should focus on selecting appropriate densities and ensuring uniform marker distribution to achieve optimal filling results.
[0067] The size of the filler population (increasing from 200 to 1,297 individuals) had little effect on filler consistency and correlation. Similarly, when the reference population size was increased from 50 to 1,047, the gain in genotype filler consistency was only 1.2%, and the gain in correlation was 4.0%, indicating that filler accuracy was less dependent on the size of the reference population.
[0068] Furthermore, the accuracy of imputation increases with increasing MAF. Genotype imputation is more difficult for inferring rare variants (MAF ≤ 0.05), while the accuracy for imputation of common variants remains relatively stable.
[0069] The above description is illustrative only and not restrictive of the present invention. Those skilled in the art will understand that many modifications, variations or equivalents can be made without departing from the spirit and scope defined by the appended claims, and all such modifications, variations or equivalents will fall within the protection scope of the present invention.
Claims
1. A method for genotypic optimization of Norway spruce, characterized in that, Includes the following steps: A haplotype reference panel was constructed, which covers the major haplotype diversity of the Norway spruce population. The reference panel was constructed based on whole-genome resequencing data of unrelated Norway spruce individuals. Low-density genotype data of the Norway spruce samples to be filled were obtained. Using the IMPUTE5 software and integrating the Norway spruce genetic map, genotype filling was performed on the low-density genotype data based on the haplotype reference panel. The filling results were quality controlled, retaining SNPs with Rsqsoft scores not lower than 0.8 to obtain high-quality filled genotype data.
2. The method according to claim 1, characterized in that, The haplotype reference panel was constructed based on whole-genome resequencing data from no fewer than 700 unrelated individuals of Norway spruce.
3. The method according to any one of claims 1 to 3, characterized in that, The low-density genotype data includes SNP chip data or exon capture sequencing data.
4. The method according to any one of claims 1 to 3, characterized in that, The haplotype reference panel was constructed using SHAPEIT v5.1.1 software.
5. A Norway spruce genotype-filled reference panel for implementing the method of any one of claims 1 to 4, characterized in that, The reference panel was constructed based on whole-genome resequencing data of unrelated Norway spruce individuals and was large enough to cover the major haplotype diversity of the Norway spruce population.
6. The reference panel according to claim 5, characterized in that, It was constructed based on whole-genome resequencing data from no fewer than 700 unrelated individuals of Norway spruce.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1 to 4.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4.