A method of genomic selection analysis considering significance of local genetic correlation of traits
By using a genomic selection analysis method that considers the significance of local genetic correlations between traits, the problem of neglecting local genetic correlations between traits in existing technologies has been solved, achieving higher accuracy in genomic prediction and breeding selection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG AGRICULTURAL UNIVERSITY
- Filing Date
- 2026-03-12
- Publication Date
- 2026-06-05
AI Technical Summary
Existing genomic selection methods ignore local genetic correlations between traits, resulting in limited prediction accuracy. The assumption of global genetic correlation in multi-trait models is not met, which limits the efficiency of breeding selection.
A genomic selection analysis method that considers the significance of local genetic correlations of traits was adopted. By processing phenotypic data, processing genotype data, conducting genome-wide association analysis, and performing genomic selection, a genomic relationship matrix was constructed to estimate local genetic correlations. This matrix was then incorporated into a dual-trait prediction model to estimate genomic breeding values.
It improves the accuracy of genome prediction models and the efficiency of breeding selection, enabling a more detailed characterization of genetic relationships between traits and enhancing prediction precision.
Smart Images

Figure CN122157762A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of animal genetics and breeding technology, specifically relating to a genomic selection analysis method that considers the significance of local genetic correlation of traits. Background Technology
[0002] Genome selection ( Genomic Selection , GS As an efficient breeding technology, traditional genome prediction methods are mostly based on single-trait models, ignoring the genetic correlation between traits, which limits prediction accuracy. To improve prediction accuracy, multi-trait genome prediction models (MTM) are being developed. Multi - trait GBLUP , MTGBLUP ) was proposed, which integrates global genetic correlations among traits ( Global genetic correlation This is used to jointly estimate breeding values. However, global genetic correlation assumes that all genomic regions contribute uniformly to the covariance among traits, which is inconsistent with actual biological mechanisms.
[0003] Due to the existence of gene pleiotropy and linkage disequilibrium, the contributions of different genomic regions to the genetic correlation between traits exhibit significant heterogeneity. Some regions may show significant correlations, while others may not. This local genetic correlation ( Local genetic correlation , LGC The differences in traits are ignored in traditional multi-trait models, which limits the accuracy of prediction models and the efficiency of breeding selection.
[0004] Existing genome selection platforms lack a complete system for using LGC saliency in genome selection models. This invention fills this gap. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a genomic selection analysis method that considers the significance of local genetic correlation of traits, thereby resolving the issues in the prior art. The technical solution adopted by this invention is as follows: A genomic selection analysis method that considers the significance of local genetic correlation of traits, including Step 1, Phenotypic Data Processing, includes missing value detection and evaluation of the original phenotypic data, and outlier identification and processing; Step 2, genotype data processing, including quality control of sequencing data, sequence alignment, genetic variation detection based on alignment results, genotype filling based on sequencing genotypes, and quality control of detection rate, minor allele frequency, and Hardy-Weinberg balance. Step 3: Generate summary statistics, including genome-wide association analysis and summary data extraction; Step 4, genome selection, includes constructing a genome relation matrix and estimating the local genetic relevance of all blocks.P The value is obtained by incorporating local genetic correlations into the prediction of dual traits and estimating the genomic breeding value. Furthermore, in step 1, the phenotypic data processing performs quality control on the original phenotypic data through descriptive statistical analysis and distribution visualization.
[0006] Furthermore, step 2 includes quality assessment and filtering of the raw sequencing data; aligning the filtered sequencing sequences to a reference genome to obtain the alignment results of each sample on the reference genome; performing genetic variation detection based on the alignment results to identify single nucleotide variants and small fragment insertion / deletion variants; and performing multi-layer quality control on the detected genetic variation sites, including filtering rules based on variant detection rate, minor allele frequency, and Hardy-Weinberg balance test to obtain genotype data.
[0007] Furthermore, step 3 includes: performing a genome-wide association analysis based on the quality-controlled genotype and phenotypic data; and extracting genetic effects, standard errors, and significance levels from the genome-wide association analysis results.
[0008] Furthermore, step 4 includes: Step 4.1: Based on genotype data, perform linkage disequilibrium structure analysis on the whole genome to divide the genome into mutually independent local LD blocks; Step 4.2: Based on the statistical data summarized by genome-wide association analysis, the LAVA method is used to estimate the local genetic correlations and their significance among traits within each LD block; Step 4.3: Classify linkage disequilibrium segments based on the significance level of local genetic correlations; Step 4.4: Introduce local genetic information of different categories of segments into the bitrait genomic selection model to estimate the genomic breeding value.
[0009] Furthermore, the bitrait genomic selection model considering the significant local genetic correlation of traits in step 4.4 is expressed as follows: The total heritability of two traits is defined as: .
[0010] in , It is a vector of phenotypic values for trait 1 and trait 2. , This is the fixed-effects design matrix for trait 1 and trait 2. , It is the fixed effects vector of trait 1 and trait 2. , It is significant LGC Genetic value design matrix for the region, , It is significant LGC The region's genomic genetic value vector, , Significance LGC Genetic value design matrix for the region, , Significance LGC The region's genomic genetic value vector, , It is the vector of total genetic values for trait 1 and trait 2. Based on salient regions SNPs The constructed genome relationship matrix, It is significant LGC The genetic variance-covariance matrix of the region It is the genetic variance of trait 1 in the significant region; It is the genetic variance of trait 2 in the significant region; , It is the genetic covariance of traits 1 and 2 in the significant region; and The corresponding terms are for non-significant regions. , Let be the vector of random residual effects for traits 1 and 2. It is the residual variance-covariance matrix of the two traits.
[0011] The present invention has the following beneficial effects: This invention successfully integrates local genetic correlations into a multi-trait genomic prediction model for dairy cows and systematically verifies its significant effect on improving prediction accuracy. By revealing and utilizing the genetic sharing structure of local regions of the genome, it is possible to characterize the genetic relationships between traits more precisely than traditional methods. Attached Figure Description
[0012] Figure 1 For flowcharts; Figure 2 This is a schematic diagram showing the distribution of locally genetically related segments on chromosomes; Detailed Implementation
[0013] The following will be described in conjunction with embodiments of the present invention. Figure 1 - 2The technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.
[0014] This invention proposes a genomic selection analysis system that considers the significance of local genetic correlation of traits, comprising: The phenotypic data processing module is used to detect and evaluate missing values in the original phenotypic data, and to identify and process outliers. The genotype data processing module is used for quality control of sequencing data, sequence alignment, genetic variation detection based on alignment results, genotype filling based on sequencing genotypes, and quality control of detection rate, minor allele frequency, and Hardy-Weinberg balance. The summary statistics generation module is used for genome-wide association analysis and summary data extraction; The genome selection module is used to construct a genome relation matrix and estimate the local genetic correlations of all blocks. P The value is used to estimate the genomic breeding value by incorporating local genetic correlations into the prediction of dual traits.
[0015] like Figure 1 The present invention also proposes a genomic selection analysis method that considers the significance of local genetic correlation of traits, including: Step 1, phenotypic data processing, includes: outlier identification and processing of the raw phenotypic data to reduce the impact of outlier observations on the results of subsequent genomic selection analysis.
[0016] The phenotypic data processing module first receives raw phenotypic data from external data sources or user input, and performs descriptive statistical analysis on the phenotypic data to obtain the basic distribution characteristics of the phenotypic data.
[0017] Based on this, and considering the distribution characteristics of the phenotypic data, observations that significantly deviate from the overall distribution are identified, and the identified outliers are further verified. Outliers that lack biological rationale or are clearly due to measurement errors or data entry errors are removed or corrected, thereby obtaining high-quality phenotypic data for subsequent analysis.
[0018] Optionally, outliers are selected based on interquartile range (IMR). inter quartile range , IQR The rules are used for identification, among which When the observed values satisfy: If an outlier is detected, it will be marked as an outlier. All outlier records will be manually checked before further analysis and removed once it is confirmed that they do not have biological rationale or are measurement / entry errors.
[0019] Step 2, genotype data processing, includes: quality assessment and filtering of the raw sequencing data to remove low-quality sequencing fragments and adapter sequences; alignment of the filtered sequencing sequences to a reference genome to obtain the alignment results for each sample on the reference genome; genetic variation detection based on the alignment results to identify single nucleotide variants and small fragment insertion / deletion variants; and, based on this, multi-layer quality control is performed on the detected genetic variation sites, including filtering rules based on variant detection rate, minor allele frequency, and Hardy-Weinberg equilibrium test, thereby obtaining high-quality genotype data that meets the requirements of subsequent genetic analysis. The genotype data processing module first performs quality assessment and filtering on the acquired raw high-throughput sequencing data to remove low-quality sequencing fragments and adapter contamination sequences. The raw sequencing data is... FASTQ The format is stored, and the quality assessment includes detecting the distribution of sequencing base quality values, sequencing length distribution, and adapter contamination. A quality filtering process is used to remove sequencing fragments with an average sequencing quality below a preset threshold, and to cut or discard sequencing fragments containing adapter sequences.
[0020] Optionally, adopt FastQC The software performs quality assessment on the raw sequencing data, using... fastp The software filters and cuts low-quality sequencing fragments and adapter sequences to obtain high-quality sequencing sequence data.
[0021] The sequence alignment process involves aligning the quality-controlled sequencing sequences to determine their position within a reference genome. The quality-filtered sequencing sequences are then aligned to the reference genome of the corresponding species to obtain the alignment results for each sample. The alignment process comprehensively considers match accuracy, mismatch count, and insertion / deletion cases to ensure the accuracy of the alignment results.
[0022] Optionally, adopt BWA - MEM The alignment algorithm aligns the sequencing sequences and generates... SAM The comparison file is in a specific format; then it is used... SAMtools The software compares and performs format conversion, sorting, and indexing on the results to generate... BAM The comparison result file is in a specific format.
[0023] The aforementioned genetic variation detection based on the sequence alignment results aims to identify single nucleotide variants and small fragment insertion / deletion variants. BAMThe alignment results file is formatted as follows. A probabilistic model-based variant detection method is used to calculate the sequencing depth and allele ratio supporting each candidate variant site, thereby identifying single nucleotide variants (SNPs). SNP ) and insertion / deletion variants ( InDel ).
[0024] Optionally, adopt GATKHaplotypeCaller The software performs mutation detection and generates a dataset containing information about the mutation sites. VCF Formatted file.
[0025] The aforementioned multi-layer quality control is performed on the detected genetic variant sites to further improve the reliability of genotype data. The quality control includes the following steps: (1) Variant detection rate filtering: the proportion of genotypes successfully detected at each site in all samples is statistically analyzed, and variant sites with detection rates below a preset threshold are filtered; (2) Secondary allele frequency filtering: the secondary allele frequency of each site is calculated, and variant sites with secondary allele frequencies below a preset threshold are filtered to reduce the impact of extremely low-frequency variants on the analysis results; (3) Hardy-Weinberg equilibrium test: the Hardy-Weinberg equilibrium test is performed on each variant site, and variant sites that significantly deviate from Hardy-Weinberg equilibrium are filtered.
[0026] Optionally, adopt PLINK The software performs statistical and filtering processing on variant detection rate, minor allele frequency, and Hardy-Weinberg equilibrium to obtain a high-quality genotype dataset.
[0027] After the above quality control steps, high-quality genotype data that meet the requirements of subsequent genome-wide association analysis and genome selection calculations are obtained and used as input data for subsequent modules.
[0028] Step 3, generating summary statistics, including genome-wide association analysis and summary data extraction; specifically, this includes: performing genome-wide association analysis based on quality-controlled genotype and phenotypic data; and extracting summary statistics such as genetic effects, standard errors, and significance levels from the genome-wide association analysis results for subsequent local genetic correlation analysis and genome selection modeling.
[0029] The summary statistics generation module performs genome-wide association analysis based on high-quality genotype and phenotypic data, and extracts summary statistics from the association analysis results to provide input data for subsequent local genetic correlation estimation and genomic selection model construction.
[0030] The genome-wide association analysis is based on quality-controlled genotype and phenotypic data, and performs association analysis on each genetic variation site across the entire genome to assess the association between each genetic variation site and the target trait.
[0031] Optionally, adopt GMAT The software executes a linear mixture model to test genetic variation sites. The model may include a fixed-effects term to control for environmental factors, and introduce a random-effects term to control for genetic correlations and population structure effects among individuals, thus completing the association analysis.
[0032] After completing the genome-wide association analysis, the summary statistics generation module extracts summary statistics from the association analysis result file for subsequent analysis.
[0033] The summarized statistical data includes, but is not limited to: the effect estimate for each genetic variation locus; the standard error corresponding to the effect estimate; and the significance level of the statistical test. P Values; genomic location information and allele information of the variant sites.
[0034] Optionally, use Python The data processing script, written in a language, extracts and formats the summarized statistical data.
[0035] Step 4, the genome selection module, includes constructing a genome relationship matrix and estimating the local genetic correlations of all blocks. P The value is obtained by incorporating local genetic correlations into a bitrait genomic selection model to estimate the genomic breeding value.
[0036] Furthermore, step 4 includes: Step 4.1, perform linkage disequilibrium analysis on the whole genome based on genotype data ( LD Structural analysis divides the genome into mutually independent local segments. LD Block; Step 4.2, based on genome-wide association analysis ( GWAS ) Summarize statistical data, using LAVA Methods estimate each LD Local genetic correlations among traits within a block and their significance; Step 4.3: Classify linkage disequilibrium segments based on the significance level of local genetic correlations; Step 4.4: Introduce local genetic information of different categories of segments into the bitrait genomic selection model to estimate the genomic breeding value.
[0037] Furthermore, the bitrait genomic selection model considering the significant local genetic correlation of traits described in step one is expressed as follows: The total heritability of two traits is defined as: .
[0038] in , It is a vector of phenotypic values for trait 1 and trait 2. , This is the fixed-effects design matrix for trait 1 and trait 2. , It is the fixed effects vector of trait 1 and trait 2. , It is significant LGC Genetic value design matrix of the region ( SIG (representing "significant") , It is significant LGC The region's genomic genetic value vector, , Significance LGC Genetic value design matrix of the region ( NON (representing "non-significant") , Significance LGC The region's genomic genetic value vector, , It is the vector of total genetic values for trait 1 and trait 2. Based on salient regions SNPs The constructed genome relationship matrix, It is significant LGC The genetic variance-covariance matrix of a region includes: It represents the genetic variance of trait 1 in the significant regions. It represents the genetic variance of trait 2 in the significant regions. , It is the genetic covariance of traits 1 and 2 in the significant region. and Non-significant region ( NON The corresponding item for ). , Let be the vector of random residual effects for traits 1 and 2. It is the residual variance-covariance matrix of the two traits.
[0039] The following uses dairy cows as an example to illustrate the specific implementation of this invention: Step 1, the specific steps for tabular data processing are as follows: The reference population consisted of 6,649 Chinese Holstein dairy cows. Phenotypic data were derived from official breeding estimates from the China Dairy Cattle Association. EBV ), targeting three important milk production traits: milk yield ( MY ), milk fat percentage ( FP ), milk protein percentage ( PP Check for missing records for each trait.
[0040] Calculate the mean, standard deviation, first quartile, and third quartile for each trait; records that do not conform to the interquartile range rule are considered outliers and removed.
[0041] Step 2, the specific content of the genotype data processing module is as follows: Genotyping data was generated using genome sequencing data. FastQC The software performs quality assessment and filtering on the raw sequencing data, removing low-quality sequencing fragments and adapter sequences; the filtered sequencing sequences are then used... BWA - MEM Alignment to a reference genome was performed to obtain the alignment results for each sample on the reference genome; GATK Genetic variation detection is performed to identify single nucleotide variants and small insertion / deletion variants; multi-layered quality control is implemented for the detected genetic variation sites. Elimination of minor allele frequencies ( MAF ≤0.05 SNP Remove Hardy-Weinberg equilibrium ( HWE )test P Value ≤ 1 e -6 of SNP Remove those with a call rate of less than 90%. SNP After quality control, 11,106,834 high-quality samples were ultimately retained. SNP For subsequent analysis.
[0042] Step 3: Perform genome-wide association analysis based on the genotype and phenotypic data after quality control; extract summary statistics such as genetic effects, standard errors, and significance levels from the genome-wide association analysis results for subsequent local genetic correlation analysis and genome selection modeling.
[0043] Step 4, the specific details of genome selection are as follows: Step 4.1: Using a segmentation method based on linkage disequilibrium structure, the entire genome was divided into 2,902 independent or semi-independent linkage disequilibrium segments, with each segment having an average length of approximately 0.8–1.0 mm. Mb .
[0044] Step 4.2: Based on the summary statistics from the genome-wide association analysis, estimate the local genetic correlations between target traits, and perform a statistical significance test on these local genetic correlations to obtain the significance level of the local genetic correlations corresponding to each linkage disequilibrium segment. The linkage disequilibrium segments are divided into segments with significant local genetic correlations and segments without significant local genetic correlations. Examples of the number of significant segments for different trait pairs are shown in Table 1.
[0045] Table 1. Number of significantly correlated local genetic regions for different trait pairs.
[0046] Step 4.3: Introduce segments with significant local genetic correlation and segments without significant local genetic correlation into the dual-trait genomic selection model to predict genomic breeding values.
[0047] The total heritability of two traits is defined as: .
[0048] in , It is a vector of phenotypic values for trait 1 and trait 2. , This is the fixed-effects design matrix for trait 1 and trait 2. , It is the fixed effects vector of trait 1 and trait 2. , It is significant LGC Genetic value design matrix of the region ( SIG (representing "significant") , It is significant LGC The region's genomic genetic value vector, , Significance LGC Genetic value design matrix of the region ( NON (representing "non-significant") , Significance LGC The region's genomic genetic value vector, , It is the vector of total genetic values for trait 1 and trait 2. Based on salient regions SNPsThe constructed genome relationship matrix, It is significant LGC The genetic variance-covariance matrix of a region includes: It represents the genetic variance of trait 1 in the significant regions. It represents the genetic variance of trait 2 in the significant regions. , It is the genetic covariance of traits 1 and 2 in the significant region. and Non-significant region ( NON The corresponding item for ). , Let be the vector of random residual effects for traits 1 and 2. It is the residual variance-covariance matrix of the two traits.
[0049] The above implementation demonstrates that the improved multi-trait genome prediction method based on local genetic correlation provided by this invention is practical and efficient, and can accelerate the genetic breeding process.
[0050] Specific embodiments of the present invention are as follows: I. Subjects and Data Sources: A dairy herd was selected as the analysis subject. The dairy herd consisted of 6,649 Holstein dairy cows, each with corresponding phenotypic and genotypic data.
[0051] The phenotypic data includes multiple traits related to production performance. In this embodiment, milk yield, milk fat percentage, and milk protein percentage are selected as exemplary analytical traits. The genotypic data are derived from high-throughput genotyping or sequencing platforms, covering a large number of single nucleotide polymorphism markers across the entire genome. After quality control, 11,106,834 high-quality genotypes were retained. SNP Loci are used for subsequent analysis. Basic information on phenotypic and genotypic data is shown in Table 1.
[0052] Table 1. Basic information on phenotypic and genotypic data of dairy cow populations.
[0053] II. Phenotypic data processing, i.e., step 1 of this invention; First, the raw phenotypic data of the dairy cow herd is processed. Specifically, the raw phenotypic observations for each individual are obtained, and outliers are identified based on the statistical distribution characteristics of the phenotypic data.
[0054] For phenotypic observations that significantly deviate from the overall distribution, further verification is conducted in conjunction with biological rationale. Outlier observations confirmed to be due to measurement errors or data entry mistakes are removed, thereby obtaining high-quality phenotypic data for subsequent analysis.
[0055] III. Genotype data processing, i.e., step 2 of this invention; After completing the phenotypic data processing, the genotypic data of the dairy cow population is processed.
[0056] Specifically, this includes: first, performing quality assessment and filtering on the raw sequencing or genotyping data to remove low-quality sequencing fragments and adapter contamination sequences; then, aligning the filtered sequencing sequences to a reference genome to obtain the alignment results for each individual on the reference genome; and based on these alignment results, performing genetic variation detection to identify single nucleotide variants and small fragment insertion / deletion variants.
[0057] Furthermore, multi-layered quality control is performed on the detected genetic variation sites, including filtering rules based on variation detection rate, minor allele frequency, and Hardy-Weinberg equilibrium test, to obtain high-quality genotype data that meet the requirements of subsequent genetic analysis.
[0058] IV. Generate summary statistical data, i.e., step 3 of this invention; After obtaining high-quality phenotypic and genotypic data, genome-wide association analysis was performed.
[0059] Specifically, based on the genotype and phenotypic data, association tests are performed on each genetic variation site across the entire genome to obtain the genetic effect estimate, standard error, and significance level corresponding to each genetic variation site.
[0060] Subsequently, the above statistics were extracted from the genome-wide association analysis results to construct a summary statistical data set for subsequent local genetic association analysis.
[0061] V. Identification of local genetic significance, i.e., steps 4.1–4.2 of this invention; In this embodiment, the significance of local genetic correlations across the entire genome is identified based on the summarized statistical data.
[0062] First, a segmentation method based on linkage disequilibrium structure was used to divide the entire genome into 2,902 independent or semi-independent linkage disequilibrium segments, with each segment having an average length of approximately 0.8–1.0 mm. Mb .
[0063] Subsequently, within each linkage disequilibrium segment, based on the summary statistics of single-trait genome-wide association analysis, the local genetic correlations between target traits are estimated, and the statistical significance of the local genetic correlations is tested to obtain the significance level of the local genetic correlations corresponding to each linkage disequilibrium segment.
[0064] Based on the significance level, linkage disequilibrium segments are divided into segments with significant local genetic correlation and segments without significant local genetic correlation. Examples of the number of significant segments for different trait pairs are shown in Table 2.
[0065] Table 2. Examples of the number of significantly correlated local genetic regions for different trait pairs.
[0066] VI. Segment-weighted bitrait genomic selection analysis, i.e., step 4.3 of this invention; After classifying linkage disequilibrium segments, segments with significant local genetic correlation and segments without significant local genetic correlation were introduced into the bitrait genomic selection model, respectively.
[0067] Specifically, within the framework of a two-trait linear mixture model, independent genetic effect terms are constructed for different category segments, and different genetic variance-covariance structures are configured for these genetic effect terms. By estimating the model parameters, the genetic effects of the two traits in significant and insignificant segments are obtained, and the genetic effects of each segment are summarized to obtain the genomic breeding values corresponding to the two traits.
[0068] By employing the above-described embodiments, the present invention can characterize the differential contribution of different genomic segments to the genetic correlation between traits during genomic selection analysis, and introduce local genetic correlation significance information into the dual-trait genomic selection model, thereby improving the accuracy and reliability of multi-trait genomic breeding value prediction while maintaining model stability.
[0069] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Any modifications, alterations, alterations, or substitutions made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A genomic selection analysis method considering the significance of local genetic correlation of traits, characterized in that, include Step 1, Phenotypic Data Processing, includes missing value detection and evaluation of the original phenotypic data, and outlier identification and processing; Step 2, genotype data processing, including quality control of sequencing data, sequence alignment, genetic variation detection based on alignment results, genotype filling based on sequencing genotypes, and quality control of detection rate, minor allele frequency, and Hardy-Weinberg balance. Step 3: Generate summary statistics, including genome-wide association analysis and summary data extraction; Step 4, genome selection, includes constructing a genome relation matrix and estimating the local genetic relevance of all blocks. P The value is used to estimate the genomic breeding value by incorporating local genetic correlations into the prediction of dual traits.
2. The genomic selection analysis method considering the significance of local genetic correlation of traits according to claim 1, characterized in that, In step 1, phenotypic data processing uses descriptive statistical analysis and distribution visualization to perform quality control on the original phenotypic data.
3. The genomic selection analysis method considering the significance of local genetic correlation of traits according to claim 1, characterized in that, Step 2 includes quality assessment and filtering of the raw sequencing data; aligning the filtered sequencing sequences to a reference genome to obtain the alignment results of each sample on the reference genome; performing genetic variation detection based on the alignment results to identify single nucleotide variants and small fragment insertion / deletion variants; and performing multi-layer quality control on the detected genetic variation sites, including filtering rules based on variant detection rate, minor allele frequency, and Hardy-Weinberg balance test to obtain genotype data.
4. The genomic selection analysis method considering the significance of local genetic correlation of traits according to claim 1, characterized in that, Step 3 includes: performing genome-wide association analysis based on the quality-controlled genotype and phenotypic data; and extracting genetic effects, standard errors, and significance levels from the genome-wide association analysis results.
5. The genomic selection analysis method considering the significance of local genetic correlation of traits according to claim 1, characterized in that, Step 4 includes: Step 4.1: Based on genotype data, perform linkage disequilibrium structure analysis on the whole genome to divide the genome into mutually independent local LD blocks; Step 4.2: Based on the statistical data summarized by genome-wide association analysis, the LAVA method is used to estimate the local genetic correlations and their significance among traits within each LD block; Step 4.3: Classify linkage disequilibrium segments based on the significance level of local genetic correlations; Step 4.4: Introduce local genetic information of different categories of segments into the bitrait genomic selection model to estimate the genomic breeding value.
6. The genomic selection analysis method considering the significance of local genetic correlation of traits according to claim 5, characterized in that, The bitrait genomic selection model considering the significance of local genetic correlation in step 4.4 is expressed as follows: The total heritability of two traits is defined as: 。 in , It is a vector of phenotypic values for trait 1 and trait 2. , This is the fixed-effects design matrix for trait 1 and trait 2. , It is the fixed effects vector of trait 1 and trait 2. , It is significant LGC Genetic value design matrix for the region, , It is significant LGC The region's genomic genetic value vector, , Significance LGC Genetic value design matrix for the region, , Significance LGC The region's genomic genetic value vector, , It is the vector of total genetic values for trait 1 and trait 2. Based on salient regions SNPs The constructed genome relationship matrix, It is significant LGC The genetic variance-covariance matrix of the region It is the genetic variance of trait 1 in the significant region; It is the genetic variance of trait 2 in the significant region; , It is the genetic covariance of traits 1 and 2 in the significant region; and The corresponding terms are for non-significant regions. , Let be the vector of random residual effects for traits 1 and 2. It is the residual variance-covariance matrix of the two traits.