Method and system for integrating chip and filling data to optimize genome selection

By adopting the integrated strategy of 'chip anchoring and filling expansion', the problem of the prediction accuracy of high-density genotype filling data decreasing instead of increasing in genomic selection was solved, and the accuracy of genomic prediction was improved. In particular, a standardized methodology was provided for the budding period and six-year-old tree height traits of Norway spruce.

CN122090929APending Publication Date: 2026-05-26NANJING FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING FORESTRY UNIV
Filing Date
2025-12-31
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies have shown that when using high-density genotype-filled data for genome selection, the prediction accuracy decreases rather than increases, mainly due to noise and marker redundancy, and the effect is particularly unsatisfactory in species with complex genomes.

Method used

An integrated strategy of 'chip anchoring and filling expansion' was adopted. Significant sites validated by GWAS were screened from chip data as the core SNP set, and combined with independently associated SNPs in whole genome filling data to construct an optimized SNP set for genome selection.

Benefits of technology

It significantly improves the accuracy of genome prediction, for example, the accuracy of predicting the budding period and six-year-old tree height of Norway spruce is improved by 21.7% and 6.2%, respectively, and provides a standardized methodology for efficient and reliable genome selection of species with complex genomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090929A_ABST
    Figure CN122090929A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for integrating a chip and filling data to optimize genome selection, and belongs to the technical field of genome breeding. The method comprises the following steps: firstly, acquiring chip data and whole genome filling data; then executing a chip anchoring and filling expansion strategy: screening a core SNP set from chip data based on whole genome association analysis, and screening an additional significant SNP set from filling data; and finally, combining the two sets to form an optimized SNP set and carrying out genome selection analysis. The internal bottleneck that high-quality filling data is directly used for genome selection to cause reduction of accuracy is disclosed for the first time, through an innovative integration strategy, the prediction robustness is ensured by utilizing chip data, new genetic signals are explored by utilizing the filling data, advantage complementation is realized, the breeding value prediction accuracy is remarkably improved, and the breeding value prediction method is suitable for large-scale popularization and application. An effective scheme is provided for solving the application problem of high-density genotype data in breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of genome breeding technology, specifically relating to a method and system for optimizing genome selection by integrating chip and filler data. Background Technology

[0002] Genotype imputation is a bioinformatics technique that uses statistical models to infer missing genotypes in target samples based on haplotype structure and linkage disequilibrium information in a reference panel. This technique effectively bridges the data density gap between low-density genotyping platforms (such as SNP chips) and whole-genome resequencing, thereby significantly enhancing the value of data in genetic research. With the maturation of related technologies, how to efficiently utilize the high-density data generated by imputation has become a new research focus in the field of genomic selection breeding.

[0003] In genomic selection, predictive models based on high-density genotypes after genotype stuffing are highly anticipated. Theoretically, the increased marker density should capture more genetic variations controlling traits, thereby steadily improving the accuracy of breeding value estimation. However, in practical applications, especially in complex genomic species such as forest trees, the results of directly using all stuffed data for prediction are not ideal; their accuracy may even be lower than using the original microarray data. This phenomenon constitutes a pressing technical contradiction in the field: while genotype stuffing expands the data scale, it may also introduce a large number of inference errors and redundant information between markers. This "noise" dilutes the true genetic signal and leads to overfitting of the predictive model, ultimately limiting the practical effectiveness of this technology in genomic selection.

[0004] Therefore, the current technological bottleneck has shifted from "how to obtain filler data" to "how to efficiently utilize filler data." Simply possessing high-density data is insufficient to guarantee improved prediction accuracy. There is an urgent need for an innovative strategy that can effectively integrate the advantages of genotype data from different sources (such as reliable microarray data and exploratory filler data) to solve the aforementioned data noise and redundancy problems, thereby truly unleashing the application potential of genotype filler technology in genome selection and achieving a substantial breakthrough in the accuracy of breeding value prediction. Summary of the Invention

[0005] To address the aforementioned technical problems in the existing technology, the present application aims to provide a method and system for optimizing genome selection by integrating chip and filler data.

[0006] To solve the above-mentioned technical problems, the technical solution of this application is as follows:

[0007] In a first aspect, the present invention provides a method for optimizing genome selection by integrating chip and filler data, comprising the following steps:

[0008] Obtain microarray genotype data and whole-genome-filled genotype data from Norway spruce samples;

[0009] Based on the locus significance statistics obtained from genome-wide association analysis (GWAS), SNPs associated with the target trait are screened from the microarray data to form a core SNP set;

[0010] In the filled data, the loci in the core SNP set are included as covariates (fixed effects) in the GWAS of the filled data. Additional associated SNPs independent of the core set are selected based on the locus significance statistics obtained from the GWAS to form the filled SNP set.

[0011] The core SNP set and the filled SNP set are merged to form an optimized SNP set;

[0012] Genomic selection analysis is performed based on the optimized SNP set to predict breeding values.

[0013] Furthermore, when screening the core SNP set, the chip data is further subjected to chain imbalance (LD) pruning, wherein a threshold is used to determine chain imbalance. No greater than 0.2.

[0014] Furthermore, the screening strategy for the core SNP set is determined based on the genetic architecture of the target trait:

[0015] If the target trait is an oligogenic genetic structure, then based on the GWAS significance statistic, a preset number of SNPs with the highest significance are selected from the chip data;

[0016] If the target trait has a polygenic genetic structure, then based on GWAS significance statistics and linkage disequilibrium pruning, a set of SNPs that meet the preset association criteria is selected from the chip data.

[0017] Further, the step of screening SNPs associated with the target trait from the chip data and / or the filling data includes: selecting SNPs from the corresponding dataset in descending order of GWAS significance statistics.

[0018] Furthermore, the preset quantity or the preset association standard is determined by cross-validation optimization on the training set.

[0019] Furthermore, the whole-genome-filled genotype data was obtained through the following optimized process: genotype filling was performed using haplotype-aware filling software and species-specific genetic maps based on a large-scale, unrelated Norway spruce haplotype reference panel, followed by quality control of the filling results.

[0020] Secondly, the present invention provides a Norway spruce genome selection system, characterized in that it comprises:

[0021] The data input module is used to receive microarray genotype data and whole-genome-filled genotype data;

[0022] The SNP integration module is configured as follows:

[0023] Based on the locus significance statistics obtained from GWAS, SNPs associated with the target trait are screened from chip data to form a core SNP set.

[0024] In the imputed data, the core SNP set is used as a covariate for correction, and based on the analysis results of this condition, SNPs that are independently associated with the target trait are selected to form the imputed SNP set.

[0025] The core SNP set and the filled SNP set are merged into an optimized SNP set;

[0026] The predictive analysis module is used to run a genomic selection model based on the optimized SNP set and output the breeding value prediction results.

[0027] Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described in the first aspect above.

[0028] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect above.

[0029] Compared with the prior art, the beneficial effects of this application are as follows:

[0030] This invention provides an optimized genomic selection method and system. Addressing the technical bottleneck of existing technologies that directly use whole-genome filling data for genomic selection, where inference noise and marker redundancy lead to decreased prediction accuracy, this invention achieves a significant improvement in breeding value prediction accuracy through an integrated strategy of "chip anchoring and filling expansion" combined with precise screening based on trait genetic architecture. Specific improvements are as follows:

[0031] 1) Overcoming the bottleneck in the application of high-density filler data in genome selection: This invention, for the first time, combines the advantages of high-reliability microarray data with high-coverage filler data through a "microarray anchoring, filler expansion" strategy. This method constructs a robust prediction foundation based on GWAS-validated significant sites in the microarray data, and then utilizes filler data to discover new, minor association signals across the entire genome. After adopting this integration strategy, the genome prediction accuracy for budburst and six-year-old tree height (H6) of Norway spruce was improved to 0.764 and 0.429, respectively; compared with directly using whole-genome filler data for genome selection, the prediction accuracy was significantly improved by 21.7% (from 0.628 to 0.764) and 6.2% (from 0.404 to 0.429), respectively, thus effectively alleviating the problem of limited predictive performance when high-quality filler data is directly used for genome selection.

[0032] 2) A precise and reproducible screening process centered on the "ranking method" was established: This invention clarifies and verifies that SNP screening according to the order of GWAS significance statistics from high to low (i.e., the "ranking method") is a key operation for achieving efficient genomic selection. Based on this, optimizations are made for traits with different genetic structures: for oligogenic traits (such as Budburst), a smaller number of the most significant top SNPs are screened; for polygenic traits (such as H6), a larger number of SNPs that have undergone linkage disequilibrium pruning are screened. Significant SNPs with a variance ≤ 0.2 were identified. This process uses cross-validation to determine the optimal parameters, ensuring the objectivity and adaptability of the strategy.

[0033] 3) This invention provides a complete and standardized technical solution: It not only proposes innovative strategies but also specifies the complete steps from data preparation (obtaining high-quality filled data based on a large-scale reference panel and optimized workflow), core screening (ranking method), parameter optimization (cross-validation), to final model construction. This systematically validated standardized workflow provides an efficient, reliable, and reproducible tool for genomic selection breeding of Norway spruce and other complex genomic species, greatly improving the accuracy and efficiency of breeding prediction. Attached Figure Description

[0034] Figure 1 Manhattan plots for genome-wide association analysis of the budburst trait in Norway spruce. A) Manhattan plot based on exon capture data before genome-wide autofill; B) Manhattan plot based on exon capture data after genome-wide autofill.

[0035] Figure 2Manhattan plots for genome-wide association analysis of diameter at breast height (DBH) in Norway spruce. A) Manhattan plot based on exon capture data before genome-wide filling; B) Manhattan plot based on exon capture data after genome-wide filling.

[0036] Figure 3 Manhattan plots for genome-wide association analysis of frost damage (FD) in Norway spruce. A) Manhattan plot based on exon capture data before genome-wide filling; B) Manhattan plot based on exon capture data after genome-wide filling;

[0037] Figure 4 A graph showing the changes in the genomic prediction accuracy of the Budburst and H6 traits in Norway spruce under different SNP screening strategies (cross-family validation comparison).

[0038] Figure 5 A comparison of the ABLUP, RRBLUP, and DeepGS prediction accuracy of the Budburst and H6 traits of Norway spruce under chip data and chip + QTN joint site. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the invention is further described below with reference to specific embodiments. Unless otherwise described in detail, the technical means used in the following embodiments are all conventional means well known to those skilled in the art, or are performed according to the kit and product instructions. Unless otherwise specified, the materials and reagents used in the following embodiments are commercially available.

[0040] Example 1

[0041] 1. Method

[0042] 1.1 Genotype Data Sources and Basic Analysis

[0043] To evaluate the utility of genotype-filled data in GWAS and GS, this application first generated genome-wide genotype-filled data based on a high-quality reference panel (1047 unrelated individuals) and an optimized workflow (using IMPUTE5 software and integrating genetic maps), and then performed GWAS and GS on the filled data and the original data.

[0044] The process for optimizing and obtaining whole-genome-filled genotype data in this application is as follows:

[0045] 1. Haplotype reference panel construction: Based on whole-genome resequencing data from 1,047 unrelated Norwegian spruce individuals, haplotype reference panels were constructed using SHAPEIT v5.1.1.

[0046] 2. Genotyping execution: Using the IMPUTE5 genotyping software and integrating the genetic map information of Norway spruce, and based on the haplotype reference panel mentioned above, genome-wide genotyping was performed on low-density genotype data from 2,201 target samples, of which 1,297 were exon capture sequencing data and 904 were 50K chip sequencing data.

[0047] 3. Quality Control of Filler Results: Strict quality control was performed on the filler results, retaining SNPs with Rsqsoft scores no lower than 0.8, and further filtering out SNPs with a minor allele frequency (MAF) below 0.05 and those deviating from Hardy-Weinberg equilibrium (HWEP value < 1 × 10⁻⁶). -5 The loci were identified, ultimately yielding high-quality whole-genome genotype data.

[0048] In this embodiment, the whole genome filling genotype data used in the "filling expansion" step is obtained by filling the microarray sequencing data through the above optimization process.

[0049] 1.2 GWAS

[0050] Genome-wide association analyses before and after genotype filling were performed using the `-mlma-loco` module in GCTA (v1.94.1) and a mixed linear model. The same phenotypic data and software parameters were used in both analyses to ensure comparability of results. The p-value was set at ≤ 1×10⁻⁶. -7 The significance threshold for the whole genome was used, and based on this, the Manhattan plot and QQ plot were plotted using the R package CMplot (v4.5.1).

[0051] 1.3 Evaluation of SNP screening strategy based on filled genotypes

[0052] This application evaluated eight different SNP screening strategies to examine their impact on the accuracy of genome selection. All analyses were based on an initial dataset containing 3,165,176 imputed SNPs.

[0053] Strategy 1: Random SNP screening (chromosome balance)

[0054] Without considering the p-value, an equal number of SNPs are randomly selected from each chromosome. This strategy was gradient-tested from 12 SNPs to the entire dataset.

[0055] Strategy 2: p-value-guided SNP screening (chromosome balance)

[0056] An equal number of SNPs were selected from each chromosome in ascending order of association p-values ​​(with preference given to the lowest p-value). The test range covered 12 SNPs up to the entire dataset.

[0057] Strategy 3a: Linkage disequilibrium pruning - chromosome balanced SNP screening

[0058] Linkage disequilibrium pruning was performed on genome-wide SNPs (threshold r² ≤ 0.2, retaining at least 300,000 SNPs), followed by the selection of an equal number of SNPs within each chromosome based on the minimum p-value. This strategy was evaluated using standard population-wide cross-validation.

[0059] Strategy 3b: Linkage disequilibrium pruning - genome-wide p-value SNP screening

[0060] After linkage disequilibrium pruning (r² ≤ 0.2, minimum retention of 300,000 SNPs), SNPs were strictly screened according to the minimum p-value across the entire genome, without any requirement for uniform chromosome distribution. This strategy was tested using cross-validation within a standard population.

[0061] Strategy 4a: Random SNP screening at a gene center

[0062] From chain-linked unbalanced pruning ( In a dataset containing ≤ 0.2 and ≥ 300,000 SNPs, the range of SNPs was limited to the genome and flanking regions (±5 kb), and then an equal number of SNPs were randomly drawn from the genome center dataset according to chromosome.

[0063] Strategy 4b: SNP screening with p-value at the gene center

[0064] In the linkage-disequilibrium-pruned gene center dataset (gene body ± 5 kb), SNPs with the smallest p-value were screened while maintaining a balanced selection of chromosomes.

[0065] Strategy 5a: Cross-family prediction - random SNP

[0066] The same chromosome balanced random selection method as Strategy 1 was adopted, but it was validated through a cross-family prediction scheme: 90% of families were used as the training set and the remaining 10% as the validation set.

[0067] Strategy 5b: Cross-family prediction - linkage disequilibrium pruning p-value-guided SNP

[0068] The same SNP screening method as strategy 3a (linkage disequilibrium pruning - chromosome balance - minimum p-value) was used, but validation was performed through cross-family prediction: 90% of families were used for training and 10% for validation.

[0069] 1.4 The impact of SNP array filling on GS

[0070] This application employs three models—ABLUP, RRBLUP, and DeepGS—to conduct phenotypic prediction analysis based on genotype data before and after SNP chip filling, aiming to evaluate the impact of genotype filling on the accuracy of genome prediction. All models utilize 10-fold cross-validation. The Pearson correlation coefficient is calculated between the estimated breeding values ​​obtained from training and the actual phenotypic values, and the average of the ten cross-validation results is used as the prediction accuracy of the breeding values. In the ABLUP model, a kinship matrix between individuals is constructed based on recorded pedigree relationships, and the pedigree breeding value prediction accuracy is calculated using ASReml-R v4.1.0.176. The RRBLUP model estimates the SNP marker effect using the R package RRBLUP v4.6.3, and evaluates the prediction accuracy based on the Pearson correlation coefficient between the estimated genome breeding value (GEBV) and the actual phenotypic values ​​in the validation population. For the deep learning model, the R package DeepGS v1.2 is used as the prediction framework. GEBV is calculated using machine learning methods based on genotype markers, and its accuracy is evaluated in the same manner. In addition, a test set was introduced during the DeepGS modeling process to evaluate the model's fit. Ten-fold cross-validation was still used during the training phase, and a total of 100 model training iterations were completed.

[0071] To obtain an optimization strategy based on GWAS results, the `-mlma-loco` module of GCTA v1.94.1 was used to perform association analysis on each training set in the 10-fold cross-validation of the SNP microarray data. Based on the p-values ​​obtained from GWAS, different numbers of significant loci were selected for RRBLUP and DeepGS modeling. Taking the H6 and Budburst traits as examples, the analysis revealed an optimal number of SNPs that maximized genome prediction accuracy; for example, in the Budburst trait, both RRBLUP and DeepGS models achieved the best predictive performance when containing approximately 200 SNPs. Based on the core SNP set obtained above, it was converted into a dose matrix using PLINK v1.90b4.9. In the whole-genome genotype data excluding the core SNP set, GWAS was performed using the `-mlma` module of GCTA, and the core SNP dose matrix obtained in the previous step was included as a fixed-effect covariate in the model. Subsequently, based on the GWAS results, the most significant quantitative trait nucleotide loci were selected and merged into the optimal SNP set of the microarray, and RRBLUP and DeepGS prediction analyses were performed again.

[0072] 1.5 Integrated Optimization Strategy of "Chip Anchoring and Filling Expansion"

[0073] To address the aforementioned issue of decreased prediction accuracy caused by directly using all filler data, this invention proposes an integrated optimization strategy called "chip anchoring, filler expansion." The core of this strategy lies in selecting and combining high-quality loci from both chip data and filler data, specifically including the following steps:

[0074] A. Core SNP Set Screening (Chip Anchoring): Based on GWAS results, SNPs significantly associated with the target trait are screened from the raw chip data to form a core SNP set. This set is then subjected to linkage disequilibrium pruning. (≤ 0.2) to remove redundant information.

[0075] B. Filling SNP set screening (filling expansion): Exclude (delete) loci from the core SNP set from the whole genome filling genotype data to obtain the processed filling dataset; include the loci from the core SNP set as covariates (fixed effects) in the GWAS of the processed filling dataset; screen out additional associated SNPs that are independent of the core set based on the locus significance statistics obtained from this GWAS to form the filling SNP set.

[0076] C. Optimization of SNP set construction and genome selection: The core SNP set and the filler SNP set are merged to form the final optimized SNP set, and genome selection analysis is performed based on this set to obtain breeding value prediction results.

[0077] 2. Results

[0078] 2.1 Filler data reveal novel genetic signals in GWAS

[0079] Compared to the original exome capture dataset, whole-genome filling significantly increased the number of testable variant sites. For all traits, this method revealed novel and significant signal clusters that were not detected in the exome capture dataset. Figure 1 Taking Budburst as an example, these new signals translated into the discovery of 17 new quantitative trait loci (QTLs). After filling, the main peak on chromosome 12 showed better continuity, and the intensity of associated signals was significantly improved. Most notably, a distinct and significant peak appeared for the first time on chromosome 2. Figure 1 B), this peak corresponds to three newly discovered loci. Similar enhancements were observed in other traits. Genome-wide association studies following filler analysis of DBH and FD identified 3 and 7 new QTLs, respectively, that were missing in exome capture analysis. Figure 2 , 3 ).

[0080] 2.2 Performance Comparison of Different SNP Screening Strategies and Confirmation of the Optimal Strategy

[0081] The impact of different SNP screening strategies on the accuracy of genome prediction varies significantly. Figure 4 Strategy 1 (random sampling) showed a continuous increase in precision within the 12–9,996 loci range, subsequently coinciding with the "full SNP" baseline, indicating that random sampling needs to reach a certain density to capture most of the genetic signal. Strategy 2 (selection of sites by GWAS significance ranking) initially showed higher precision than Strategy 1, but the optimal values ​​for the two traits diverged: Budburst reached a peak of 0.661 at 7,992 loci, significantly better than the full SNP's 0.628; H6, however, consistently failed to surpass the full SNP precision (0.404). Strategy 3a applied LD filtering (r² ≤ 0.2) to Strategy 2, further improving overall performance: Budburst and H6 achieved 0.671 and 0.425 respectively, both higher than the full SNP baseline, and the curves were more stable. Comparing Strategy 3a and 3b (with or without mandatory uniform chromosome sampling) revealed that Budburst was superior to Strategy 3b in the first 1,992 loci, peaking at 0.749, after which it overlapped with Strategy 3a. H6, on the other hand, was slightly superior to Strategy 3a across all regions. Ultimately, both achieved the highest precision (0.425) on the entire set filtered by LD. Focusing on gene regions and their upstream and downstream 5 kb for Strategy 4a / 4b showed that defining functional regions did not provide additional benefits: the Strategy 4a curve for Budburst almost overlapped with Strategy 1, and 4b peaked at 0.654 at 9,996 loci, still lower than Strategy 3b. The precision of Strategy 4a for H6 (0.372) was significantly lower than that of random sampling (0.405), and Strategy 4b decreased after reaching 0.381 at 3,996 loci. The trends of Strategy 5a / 5b validated across families were consistent with those of Strategy 1 / 3a, but Strategy 5b (LD filtering + significant loci + cross-family) achieved the highest or near-highest predictive accuracy in both traits, confirming that the strategy still has a robust advantage in independent family scenarios.

[0082] 2.3 The "chip anchoring, fill expansion" strategy overcomes the GS bottleneck and improves prediction accuracy.

[0083] This application is based on a prior systematic evaluation of various alternative methods. Figure 4 Genomic selection based on an optimized mixed SNP strategy was designed and implemented for the Budburst and H6 traits. This evaluation confirmed the effectiveness of linkage disequilibrium screening (…). SNP selection strategies with a p-value of ≤0.2 and based on p-value information can achieve the best balance between prediction accuracy and computational efficiency when using filler data for genome selection. Based on this principle, we have developed a trait-specific "chip anchoring, filler expansion" analysis workflow to address the technical bottleneck of limited predictive performance when high-quality filler data is directly used for genome selection.

[0084] To validate this strategy, we extracted SNP subsets of different sizes from both microarray and whole-genome filling data, based on phylogenetic kinship uptake (ABLUP). We then combined RRBLUP with the DeepGS deep learning model to perform genome prediction and accuracy comparisons for the Norway spruce Budburst and H6 traits. The results first confirmed the existence of the aforementioned technical bottleneck: if RRBLUP analysis is performed directly using all filling data, the prediction accuracy is slightly lower than that of the model based on the original microarray data.

[0085] Therefore, the first step was to implement a "chip anchoring" step, which involved screening the chip data according to the aforementioned optimization strategy. The results showed that for the Budburst trait, the 200 most significant SNPs (based on genome-wide association analysis) were selected. The model performed optimally when the filter was < 0.2 (DeepGS: 0.786; RRBLUP: 0.761); for the H6 trait, all samples need to be included. Chip SNP values ​​with a filter of <0.2 are required to achieve optimal performance (DeepGS: 0.404; RRBLUP: 0.427).

[0086] The "filling expansion" step then proceeds, incorporating the selected chip SNPs as fixed covariates into the genome-wide association analysis model. Additional significant SNPs are further identified from the genome-wide filler data, constructing the final integrated prediction model. For the Budburst trait, the model combines the 200 optimal chip SNPs with 500 additional significant SNPs from the filler data; for the H6 trait, it combines all filtered chip SNPs with 200 additional significant SNPs from the filler data.

[0087] The verification results show that the integration strategy has achieved significant results. Figure 5For the Budburst trait, the prediction accuracy of rrBLUP-array-best+QTN reached 0.764, which is 21.0% higher than ABLUP, 12.0% higher than rrBLUP-array, 21.6% higher than rrBLUP-imputed, and 0.4% higher than rrBLUP-array-best. For the H6 trait, the prediction accuracy of rrBLUP-array-best+QTN reached 0.429, which is 11.7% higher than ABLUP, 1.7% higher than rrBLUP-array, 6.1% higher than rrBLUP-imputed, and 0.5% higher than rrBLUP-array-best. Furthermore, the DeepGS model showed significant differences in performance on the two traits: its prediction accuracy (0.786) was better than the best linear method for the budding trait, and it could be further improved to 0.797 after adding a significant SNP from the filler source; however, for the six-year-old tree height trait, its accuracy (0.404) was only better than ABLUP, and it did not improve after adding an additional SNP.

[0088] 3. Technical effects achieved by this application

[0089] This application first obtained high-quality whole genome filling data of Norwegian spruce based on an optimized genotype filling process, providing a reliable data foundation for subsequent genetic analysis of this species. Under this premise, this application explored in depth the application bottlenecks and solutions of filling data in genome selection, and the core technical effects achieved are as follows: (1) It revealed the inherent mechanism that direct use of filling data for genome selection under the premise of high-quality filling data will still lead to a decrease in prediction accuracy, and clarified that data noise and marker redundancy are inherent technical bottlenecks; (2) It proposed an integrated strategy of "chip anchoring and filling expansion", and through complementary advantages, successfully broke through the above bottlenecks and achieved a stable improvement in prediction accuracy; (3) It constructed a precise optimization scheme for different genetic architecture traits, providing complete methodological support for achieving efficient and accurate genome selection breeding.

[0090] 3.1 The Value of High-Quality Filled Data and the Discovery of Bottlenecks in GS Applications

[0091] Analysis based on high-quality filled data shows that it exhibits different value dimensions in GWAS and GS.

[0092] This application demonstrates that a genotype imputation strategy starting with exon capture data significantly improves the detection power of the Budburst trait in GWAS. This is manifested in two aspects: First, the imputation data fills the information gaps between the original markers, significantly enhancing the continuity and statistical significance of the known major signal peaks on chromosome 12. More significantly, on chromosome 2, the imputation data reveals a novel, independently associated signal cluster. This discovery proves that the imputation strategy provided in this application can unlock previously overlooked key genomic regions regulating phenological traits.

[0093] Based on existing pedigree information, this application systematically evaluated the efficacy of different SNP screening strategies and statistical models for genomic prediction of the Budburst and H6 traits in Norway spruce. The results show that data quality and SNP screening are key factors affecting prediction accuracy. Notably, directly using all filled SNPs for RRBLUP prediction resulted in lower accuracy than using high-quality microarray data. This phenomenon stems primarily from two reasons: First, while genotype filling increases marker density, it introduces unavoidable errors and uncertainties. This "noise" dilutes the true genetic effect signal, thus reducing the efficiency of the prediction model. Second, widespread linkage disequilibrium (LD) among high-density SNPs can lead to model overfitting; for models like RRBLUP that assume marker effects follow a normal distribution, redundant marker information not only fails to provide additional genetic information but also interferes with the accurate estimation of key causal site effects. The results of this application provide evidence for the above mechanisms: Figure 2 As shown, for strategy 3a (which is based on strategy 2 with LD filtering), the prediction accuracy of breeding values ​​for the H6 and Budburst traits improved from 0.404 and 0.628 to 0.425 and 0.635, respectively, compared to using all loci.

[0094] In stark contrast, applying GWAS significance screening and LD filtering to the raw microarray data correspondingly improved prediction accuracy. The success of this strategy lies in the fact that GWAS screening preferentially retains sites most strongly associated with the target trait, while LD filtering… A value < 0.2 ensures the independence of the selected sites and minimizes information redundancy. Therefore, compared to blindly pursuing a large number of markers, ensuring the quality and independence of markers through rigorous statistical screening is crucial for building robust and efficient genome prediction models.

[0095] This application found that, even after rigorous screening, the predictive performance of loci based on filler data still falls short of the best loci in high-quality microarray data. This clarifies that the core value of filler data lies not in its sheer quantity, but in its ability to "discover" true genetic signals missed by microarray data through broader genome coverage.

[0096] 3.2 Impact of different SNP-filling screening strategies on the accuracy of genome selection

[0097] This application systematically evaluated the impact of eight SNP screening strategies on the accuracy of GS based on imputed data. The results show that the selection of the optimal strategy is highly dependent on the genetic architecture of the target trait, marker density, and prediction scenario (such as cross-family).

[0098] The results of this application clearly demonstrate that there is no universally optimal SNP screening strategy. The starkly different performances of the Budburst (an oligogenic quantitative trait) and H6 (a minor polygenic quantitative trait) constitute one of the core findings of this application. For oligogenic traits like Budburst, significance-based screening (Strategy 2) performs exceptionally well at low to medium densities because GWAS can effectively capture a limited number of loci with significant effects. Conversely, for polygenic traits like H6, whose genetic variance is diffusely distributed across the genome and contributed by a large number of minor loci, random sampling (Strategy 1) or uniform sampling with LD trimming (Strategy 3a) is more advantageous at high densities because they can more fairly capture minor variations across the genome.

[0099] A comparison of Strategy 3a and 3b reveals subtle trade-offs in strategy design. LD pruning itself, by removing redundant markers, increases information density and is a key step in improving the performance of all strategies. However, the effect of the "chromosome balance" constraint varies depending on the trait. For H6, chromosome balance (Strategy 3a) provides a stable performance guarantee. For Budburst, removing this constraint (Strategy 3b) unlocks the model's potential at low density, allowing it to focus on selecting the most significant sites, thus achieving peak predictive accuracy.

[0100] The results of Strategy 4a and 4b indicate that limiting SNP screening to the genome body and its upstream and downstream 5 kb regions did not significantly improve the accuracy of genome prediction. Therefore, given the current level of annotation and understanding, a genome-wide balanced screening strategy (such as Strategy 3a / 3b) demonstrates higher reliability and stability in practice compared to relying on potentially incomplete prior functional information.

[0101] The absolute accuracy (up to approximately 0.2) observed in cross-family prediction (Strategy 5a and 5b) was significantly lower than that in intra-population prediction. However, this application was still able to distinguish the relative effectiveness of different strategies: Strategy 5b (LD filtering + significant loci) maintained the highest or near-highest prediction accuracy for both traits. This finding suggests that although absolute accuracy is limited by population genetic differentiation, robust strategies employing LD pruning and significant loci enrichment can still maximize the extraction and utilization of transferable genetic signals.

[0102] 3.3 Implementation Effect of the "Chip Anchoring, Filling and Expanding" Integration Strategy

[0103] This application proposes an integrated strategy of "chip anchoring and filler expansion." This strategy successfully addresses the accuracy bottleneck of directly using filler data for genome sequencing (GS) and achieves unexpected technical results. The logical basis of this strategy is as follows: First, validated high-quality significant loci in the chip data are used as "anchors" to fix their effects, ensuring that the model has a robust predictive foundation. Then, leveraging the broad coverage of the filler data, new, low-impact association signals are systematically "expanded" across the entire genome to search for those in regions not covered by the chip or that were not effectively captured by the chip due to linkage disequilibrium (LD) structural differences. Finally, the marker set constructed through this complementary approach outperforms any single data source in predictive performance, demonstrating the effectiveness of positioning filler data as a "supplement and discovery" tool.

[0104] The differentiated responses of different traits to genome prediction strategies are an external manifestation of the essential differences in their genetic architecture. This application observes a stark contrast in the Budburst and H6 traits. For the oligogenic Budburst, the "few but precise" SNP strategy and the complex DeepGS model are the optimal solutions; while for the polygenic H6, the "saturation" strategy relying on a large number of markers and the robust RRBLUP model show their superiority.

[0105] The above description is illustrative only and not restrictive of the present invention. Those skilled in the art will understand that many modifications, variations or equivalents can be made without departing from the spirit and scope defined by the appended claims, and all such modifications, variations or equivalents will fall within the protection scope of the present invention.

Claims

1. A method for optimizing genome selection by integrating microarray and filler data, characterized in that, Includes the following steps: Obtain microarray genotype data and whole-genome-filled genotype data from Norway spruce samples; Based on the locus significance statistics obtained from genome-wide association analysis, SNPs associated with the target trait are screened from the microarray data to form a core SNP set; In the filled data, the loci in the core SNP set are included as covariates of fixed effects in the GWAS of the filled data. Additional associated SNPs independent of the core set are screened out based on the locus significance statistics obtained from the GWAS to form the filled SNP set. The core SNP set and the filled SNP set are merged to form an optimized SNP set; Genomic selection analysis is performed based on the optimized SNP set to predict breeding values.

2. The method according to claim 1, characterized in that, When screening the core SNP set, the chip data is further subjected to chain imbalance pruning, wherein the threshold used to determine chain imbalance is... No greater than 0.

2.

3. The method according to claim 2, characterized in that, The selection strategy for the core SNP set is determined based on the genetic architecture of the target trait: If the target trait is an oligogenic genetic structure, then based on the GWAS significance statistic, a preset number of SNPs with the highest significance are selected from the chip data; If the target trait has a polygenic genetic structure, then based on GWAS significance statistics and linkage disequilibrium pruning, a set of SNPs that meet the preset association criteria is selected from the chip data.

4. The method according to any one of claims 1 to 3, characterized in that, The step of screening SNPs associated with the target trait from the chip data and / or the fill data includes: selecting SNPs from the corresponding dataset in descending order of GWAS significance statistic.

5. The method according to claim 3, characterized in that, The preset quantity or the preset association standard is determined by cross-validation optimization on the training set.

6. The method according to claim 1, characterized in that, The whole-genome-filled genotype data were obtained through the following optimized process: genotype filling was performed using haplotype-aware filling software and species-specific genetic maps based on a large-scale, unrelated Norway spruce haplotype reference panel, followed by quality control of the filling results.

7. A Norway spruce genome selection system, characterized in that, include: The data input module is used to receive microarray genotype data and whole-genome-filled genotype data; The SNP integration module is configured to: screen SNPs associated with the target trait from microarray data based on the locus significance statistics obtained by GWAS to form a core SNP set; screen SNPs not included in the core SNP set but associated with the target trait from the filler data to form a filler SNP set; merge the core SNP set and the filler SNP set into an optimized SNP set; and a prediction analysis module is used to run a genomic selection model based on the optimized SNP set and output breeding value prediction results.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.