Animal hybridization effect analysis and prediction method based on genome and application thereof
By constructing hybridization experimental populations, conducting genotype detection and differentiation analysis, quantifying hybridization effects, and identifying hybridization advantage regions, the problem of explaining the relationship between genome differentiation and hybridization effects in animal breeding has been solved, achieving efficient and accurate breeding guidance and reducing costs and time.
Patent Information
- Application Number
- CN202510917871.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies in animal breeding lack a systematic integration of genomic information and phenotypic data, making it impossible to accurately predict hybridization effects. In particular, they cannot explain the relationship between the degree of genomic differentiation and hybridization effects, resulting in significant risks in breeding decisions. Furthermore, existing methods are costly, time-consuming, and inefficient.
By constructing hybridization experimental populations, collecting and cleaning phenotypic data, performing genotype detection, inferring haplotype information, conducting genome differentiation analysis, quantifying hybridization effects using Fisher geometric models, identifying hybridization advantage regions, establishing FST-λ relationship curves, identifying sex-specific hybridization advantage regions, and providing sex-specific breeding strategies.
It has realized the quantitative relationship between the degree of genome differentiation and hybridization effect, identified hybridization-sensitive regions, guided breeding practices, reduced resource waste, improved breeding efficiency, shortened the breeding cycle, and provided a basis for scientific breeding decisions.
Smart Images

Figure CN120877873A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of animal breeding and genetics technology, specifically to a genome-based method for analyzing and predicting animal hybridization effects and its applications. Background Technology
[0002] Hybrid breeding is a widely used technique in animal breeding. By crossing different breeds or strains of animals, offspring with heterosis (also known as hybrid vigor) can be produced. Heterosis refers to the phenomenon that hybrid offspring exhibit better traits than the average level of their parents, or even better than the best parent, in certain traits, such as increased growth rate, enhanced disease resistance, and improved fertility. However, not all hybrid combinations produce ideal heterosis; some hybridizations can even lead to hybrid inferiority.
[0003] In plant breeding, there has been considerable research on predicting hybridization benefits. For example, Chinese patent CN109727640B discloses a genome-wide prediction method based on automated machine learning technology. This method examines the genotype and phenotypic data of hybrids in the population used for modeling. It utilizes the H2O tool within the automated machine learning framework to construct and automatically select the optimal model / parameters, evaluating the impact of each marker on crop phenotypes in different ecoregions. Then, based on the parental genotypes, the hybrid genotypes are deduced. By integrating the genotypic effects of various molecular markers in the hybrids, phenotypic data is predicted, and promising hybrid combinations are recommended for breeders to choose from in breeding practices, thereby reducing uncertainty in breeding pair selection and improving breeding success rates. However, the hybridization effects of plants and animals differ greatly, and plant breeding techniques cannot be directly applied to animal breeding.
[0004] Currently, the prediction of hybridization effects in animal breeding mainly relies on the following methods:
[0005] Direct experimentation: This method involves conducting numerous hybridization experiments, observing the offspring's performance, and selecting the optimal hybridization combination. While highly accurate, this method is extremely costly, time-consuming (especially for large animals like pigs and cattle), resource-intensive, and inefficient.
[0006] Statistical analysis based on phenotypic data: This method uses historical hybridization data to build statistical models and predict the effects of new hybrid combinations. However, it requires a large amount of historical data and has low accuracy in predicting new combinations without available historical data.
[0007] Marker-assisted selection: This method uses molecular markers associated with the target trait for selection, but it often only targets a single or a few traits and is difficult to comprehensively assess the hybridization effect.
[0008] Chinese patent CN118480613A discloses the application of a low-density 5K SNP chip for the entire pig genome. By combining genotype and phenotypic data obtained from the chip, a breeding value prediction model is established, thereby improving the accuracy of genetic assessment. However, gene chips predict individual breeding values based on specific SNP markers, addressing the question of "which individual to select?" This falls under the category of genomic selection technology. Moreover, gene chips themselves are expensive and cannot address the question of "which hybrid combination to select?", nor can they effectively explain the relationship between the degree of genomic differentiation and hybridization effects.
[0009] In summary, current technologies in animal breeding lack a method that can systematically integrate genomic information and phenotypic data to comprehensively predict hybridization effects. In particular, existing technologies have not been able to effectively explain the relationship between genomic differentiation and hybridization effects, nor can they accurately predict the critical point of transition from heterosis to heterosis, leading to significant risks in breeding decisions. Therefore, there is an urgent need to develop a hybridization effect analysis and prediction method based on genomic data to improve breeding efficiency, reduce costs, and shorten the breeding cycle. Summary of the Invention
[0010] To address the above problems, this invention proposes a genome-based method for analyzing and predicting animal hybridization effects and its applications.
[0011] The genome-based method for analyzing and predicting animal hybridization effects provided by this invention includes the following steps:
[0012] S1. Construct a hybridization experimental population;
[0013] S2. Collect phenotypic data and perform cleaning processing;
[0014] S3. Genotyping of all individuals in the experimental population;
[0015] S4. Use specialized software to infer the haplotype information of all individuals in the experimental population;
[0016] S5. Genome differentiation analysis;
[0017] It also includes S6, hybridization effect analysis, which specifically includes the following steps:
[0018] S6-1. Differentiate genotype combinations: Based on the haplotype information inferred in step S4, classify the genotypes of F2 generation individuals into the following four genotype combinations within each fixed genomic window: A / A, B / B, A / B, B / A.
[0019] S6-2. Determine the optimal phenotypic value: For each trait, calculate the average phenotypic value of the male and female F2 populations respectively, which is taken as the optimal phenotypic value for that sex.
[0020] S6-3. For each trait, the deviation between the phenotypic expression and the phenotypic optimum of each genotype combination is quantified by phenotypic distance, which is divided into absolute phenotypic distance (d) and relative phenotypic distance (D).
[0021] The absolute phenotypic distance (d) is: d = |phenotype (genotype) - phenotype (total sex)|, where phenotype (genotype) is the phenotypic average of each genotype combination, and phenotype (total sex) is the overall phenotypic average of the F2 population of that sex.
[0022] The relative phenotypic distance (D): D = d / phenotype (gender population);
[0023] S6-4. For each trait, calculate the hybridization effect size for each genotype.
[0024] The hybridization effect size = |D(genotype) - D(control genotype)|, where D(genotype) is the relative phenotypic distance of each of the four genotypes, and D(control genotype) is the relative phenotypic distance of the genotype that differs from D(genotype) by only one allele substitution; the hybridization effect size covers four cases, namely: |D(A / A)-D(A / B)|, |D(A / A)-D(B / A)|, |D(A / B)-D(B / B)|, and |D(B / A)-D(B / B)|.
[0025] S6-5. Obtain the λ parameter: Fit the data with a hybridization effect size ≤ 0.3 to obtain the λ parameter; the larger the λ parameter, the more dominant the small effect, indicating heterosis; the smaller the λ parameter, the more dominant the large effect, indicating heterosis.
[0026] Furthermore, the fitting process for the λ parameter is as follows: Select data with a hybridization effect size ≤ 0.3 to eliminate the interference of extreme values on the distribution fitting, use the fitdistr function of the MASS package in R language to fit the exponential distribution, estimate the λ parameter and its confidence interval, evaluate the goodness of fit through the Kolmogorov-Smirnov test to ensure the rationality of the exponential distribution hypothesis, and finally obtain the λ parameter.
[0027] Furthermore, the genome differentiation analysis includes performing differentiation analysis on the genome data of the F0 ancestor, using specialized software to calculate the FST value across the entire genome, quantifying the degree of genome differentiation between parents, dividing the entire genome into fixed-size windows, calculating the FST value at the window level, and dividing the entire genome window into multiple gradient levels based on the size of the FST value.
[0028] Furthermore, the entire genome was divided into fixed windows of 100kb each.
[0029] Furthermore, it also includes using genomic differentiation data to identify regions of hybrid vigor, specifically including the following steps:
[0030] 1) Establish the relationship between FST and the λ parameter: Based on the FST value and the λ parameter, analyze the association between the two within a fixed window across the whole genome, including:
[0031] 1-1) FST grading: The FST value of the whole genome is divided into 20 equal quantile intervals. The part with FST less than 0 is recorded as 0. From low differentiation (FST≈0) to high differentiation (FST≈1), it represents the gradient of genetic differences between parents.
[0032] 1-2) Summary of λ parameters: For each FST interval, collect the λ parameters of the corresponding window and calculate the mean and standard deviation.
[0033] 1-3) Relationship modeling: Use a nonlinear regression model to fit the relationship between the FST value and the λ parameter, plot the FST-λ relationship curve, and identify the critical point of the hybridization effect of the trait from heterosis to heterosis by the inflection point or slope change of the curve.
[0034] 2) Identifying regions of hybrid vigor: Based on the FST-λ relationship curve, windows with λ parameters significantly higher than the average level of the whole genome are marked as regions of hybrid vigor. The specific steps are as follows:
[0035] 2-1) Threshold calculation: Calculate the mean (μ) and standard deviation (σ) of the λ parameter;
[0036] 2-2) Determine the core heterosis threshold: Tcore = μ + kσ (k = 1.0);
[0037] 2-3) Determine the extended heterosis threshold: T_extension = max(T_core, μ−cσ) (c=0.5);
[0038] 2-4) Adaptive adjustment mechanism: When max(λ) < (μ + 1.5σ), k is automatically set to 1.0; when max(λ) ≥ (μ + 1.5σ), k is set to 1.5.
[0039] 3) Trait category analysis: Plot FST-λ relationship curves for different traits to identify the hybridization advantage regions of each trait.
[0040] Furthermore, it also includes identifying sex-specific hybrid vigor regions: analyzing the FST-λ relationship curves of male and female F2 populations separately, plotting FST-λ relationship curves for different sexes for different traits, and identifying sex-specific hybrid vigor regions.
[0041] Furthermore, it also includes visualization and result output: using R scripts to generate FST-λ relationship curves and sex comparison plots, and outputting a detailed list of hybrid vigor regions through calculation and analysis, including chromosome location, λ value, FST value and annotated genes.
[0042] The genome-based method for analyzing and predicting animal hybridization effects disclosed in this invention is applied in animal breeding.
[0043] The beneficial effects of this invention are as follows:
[0044] 1. The method of this invention establishes for the first time a quantitative relationship between the degree of genome differentiation and hybridization effect, reveals the genetic basis of heterosis and heterosis, provides a new perspective for understanding the hybridization mechanism, and solves the problem of "which hybridization combination to choose?"
[0045] 2. The method of the present invention can effectively explain the relationship between the degree of genome differentiation and the hybridization effect. By analyzing the differences in hybridization effects in different genome regions, it can identify genome regions that are sensitive to hybridization and directly guide molecular marker-assisted selection in breeding practices.
[0046] 3. The method of this invention effectively integrates genomic information and phenotypic data, providing a systematic method for evaluating hybridization effects. It can identify regions of hybridization advantage through a single whole-genome analysis, making up for the shortcomings of existing technologies that rely solely on phenotypic data or single molecular markers. It can accurately predict hybridization effects, provide long-term guidance for subsequent multi-generational breeding, provide a scientific basis for animal breeding decisions, avoid the waste of resources from a large number of inefficient hybridization combination experiments, significantly reduce the number of actual hybridization experiments, save breeding costs and time, and accelerate the breeding process.
[0047] 4. The method of this invention provides a sex-specific analysis method, revealing the sex differences in hybridization effects. This can be used to formulate sex-specific breeding strategies and improve the selection efficiency of target traits. In practical breeding, the optimal strategy is to first use the method of this invention to determine the best hybridization combinations and key genomic regions, and then develop precise SNP markers targeting these regions for individual selection, achieving full-chain breeding optimization from macro-hybridization strategies to micro-individual selection.
[0048] 5. The method of this invention is not only applicable to the hybridization breeding of traditional economic animals such as pigs, cattle, sheep, and chickens, but can also be extended to genetic improvement in aquaculture, breeding of special economic animals, and wildlife protection, and has broad industrial application value. Attached Figure Description
[0049] Figure 1 This is a flowchart illustrating the overall process of the method of the present invention.
[0050] Figure 2This is a co-occurrence verification diagram for Example 2;
[0051] Figure 3 Partial results of the effect size of allele substitution on traits calculated in Example 3;
[0052] Figure 4 This is a schematic diagram showing the relationship between the hybridization effect and the degree of differentiation (FST value) of all traits in the genome of the third-generation pig hybrid population in Example 4;
[0053] Figure 5 This is a schematic diagram showing the relationship between the hybridization effect of all traits of male and female individuals in the genome of a third-generation pig hybrid population and the FST value in Example 4;
[0054] Figure 6 This is a schematic diagram showing the relationship between the hybridization effect of the hip circumference trait in male and female individuals of the third-generation pig hybrid population in Example 5 and the FST value;
[0055] Figure 7 This is a schematic diagram showing the relationship between the hybridization effect and FST value of the skeletal-related traits (mid-trunk back bone, forequarter bone, and hindquarter bone) of the male and female individuals in the third-generation hybrid pig population of Example 5. Detailed Implementation
[0056] The present invention will be further described below with reference to the embodiments.
[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] Example 1: A genome-based method for analyzing and predicting animal hybridization effects:
[0059] See attached Figure 1 The present invention provides a genome-based method for analyzing and predicting animal hybridization effects, comprising the following steps:
[0060] Step S1: Constructing the hybridization experimental population:
[0061] Select two or more varieties / lines with sufficient genetic differentiation as parents to design a hybridization experiment. Based on this, construct an experimental population of three or more generations, including F0, F1, and F2 generations. The F0 generation must be a purebred ancestor. Other varieties can be introduced into the F1 and subsequent hybridization lines, but the introduced varieties must be purebred. F1 is the first generation of hybridization, F2 is the second generation of hybridization, and so on, to propagate multiple generations and record complete family information.
[0062] Step S2: Collect phenotypic data:
[0063] Phenotypic data were collected from the animals in the experimental group, including but not limited to: growth and development related indicators (such as body weight, height, and length at different ages), muscle characteristics (such as muscle fiber type and area), fat deposition indicators, skeletal characteristics, and organ weight. For time-series traits, measurements were taken at multiple key time points to capture dynamic changes in the traits.
[0064] After collecting the data, phenotypic data cleaning is performed, specifically including:
[0065] S2-1. Outlier Handling: Perform box plot analysis on all collected phenotypic data to identify potential outliers, and mark data points that exceed the mean ± 3 standard deviations. Based on biological knowledge and expert judgment, decide whether to remove, correct, or retain these outliers.
[0066] S2-2, Missing value handling: For missing data in some measurements, imputation is performed reasonably based on the average value of individuals of the same parity and sex, or statistical models are used to predict missing values;
[0067] S2-3. Data standardization: Perform batch effect correction on data from different batches to eliminate systematic errors caused by factors such as measurement time and environmental conditions;
[0068] S2-4. Consistency test: Through correlation analysis and regression analysis, test the biological rationality between related traits and identify potential measurement or recording errors.
[0069] Step S3, whole-genome genotyping:
[0070] Genotyping was performed on all individuals in the experimental population using high-density SNP chips or whole-genome sequencing technology. Simultaneously, whole-genome resequencing was performed on the F0 ancestor to obtain more precise information on genomic variations.
[0071] Step S4, Haplotype Inference and Recombination Identification:
[0072] Phase analysis of genomic data was performed using specialized software (such as SHAPEIT2) combined with pedigree information to infer haplotype information for all individuals in the experimental population. By comparing and analyzing haplotype differences between F1 and F2 generations, recombination events and their loci were identified.
[0073] Step S5, Genome Differentiation Analysis:
[0074] Differentiation analysis was performed on the genomic data of the F0 ancestor, and FST values across the entire genome were calculated using estimators such as Weir and Cockerham to quantify the degree of genomic differentiation between parents. The genome was divided into fixed-size windows (e.g., 100kb), and FST values were calculated at the window level. The genomic windows were then divided into multiple gradient levels based on the size of the FST values.
[0075] Step S6, Hybridization effect analysis:
[0076] Hybridization effect analysis is the core step of this method, aiming to quantify the phenotypic differences in hybrid offspring under different genotype combinations using the Fisher geometric model (FGM), revealing the genetic basis of heterosis or heterosis. This step utilizes genotype and phenotypic data from F2 individuals, combined with haplotype information, to systematically evaluate the magnitude of hybridization effects in various regions of the genome, providing a quantitative basis for subsequent identification of regions of hybrid vigor. Specifically, this includes:
[0077] S6-1. Differentiating Genotype Combinations:
[0078] Based on the haplotype information inferred in step S4, the genotypes of F2 generation individuals are classified into the following four combinations within each genomic window (100kb):
[0079] A / A: Both haplotypes originate from parent A (e.g., MIN / MIN, two haplotypes from the Min pig).
[0080] B / B: Both haplotypes originate from parent B (e.g., LW / LW, haplotypes from two Large White pigs).
[0081] A / B: The maternal haplotype comes from parent A, and the paternal haplotype comes from parent B (e.g., MIN / LW).
[0082] B / A: The maternal haplotype comes from parent B, and the paternal haplotype comes from parent A (e.g., LW / MIN).
[0083] By distinguishing between two heterozygous genotypes, A / B and B / A, and considering the origin of parental haplotypes, possible maternal or paternal effects were analyzed. To ensure the accuracy of genotype classification, R language scripts were used in conjunction with pedigree information to verify the correctness of haplotype assignments window by window, and potential haplotype inference errors were marked and corrected.
[0084] S6-2. Determine the optimal value of the phenotype:
[0085] For each trait, the phenotypic mean of the male and female F2 populations is calculated separately and taken as the sex-specific phenotypic optimum (i.e., fitness optimum). The determination of the phenotypic optimum is based on the assumption of the Fisher geometric model, that is, the population phenotypic mean is close to the optimal fitness value under natural selection.
[0086] S6-3. For each trait, quantify the deviation between the phenotypic expression and the optimal phenotypic value for each genotype combination using phenotypic distance, divided into absolute phenotypic distance (d) and relative phenotypic distance (D):
[0087] The absolute phenotypic distance (d) is: d = |phenotype (genotype) - phenotype (total sex)|, where phenotype (genotype) is the phenotypic average of each genotype combination, and phenotype (total sex) is the overall phenotypic average of the F2 population of that sex.
[0088] Relative phenotypic distance (D): To eliminate dimensional differences, it is standardized to relative phenotypic distance.
[0089] The relative phenotypic distance (D): D = d / phenotype (gender population);
[0090] S6-4. For each trait, calculate the hybridization effect size for each genotype:
[0091] The hybridization effect size is defined as |D(genotype) - D(control genotype)|, where D(genotype) is the relative phenotypic distance between the four genotypes, and D(control genotype) is the relative phenotypic distance between the genotype and D(genotype) that differs from D(genotype) by only one allele substitution. This covers four cases (A / A vs A / B, A / A vs B / A, A / B vs B / B, B / A vs B / B). The hybridization effect size is taken as the absolute value, thus covering four cases: |D(A / A) - D(A / B)|, |D(A / A) - D(B / A)|, |D(A / B) - D(B / B)|, and |D(B / A) - D(B / B)|. Standardization eliminates dimensional differences, making the phenotypic effect sizes comparable.
[0092] S6-5. Obtain the λ parameter: Fit the data with a hybridization effect size ≤ 0.3 to obtain the λ parameter; the larger the λ parameter, the more dominant the small effect, indicating heterosis; the smaller the λ parameter, the more dominant the large effect, indicating heterosis.
[0093] The λ parameter reflects the distribution characteristics of the hybridization effect size, assuming that the effect size follows an exponential distribution. The fitting process is as follows:
[0094] Data filtering: Select data with effect sizes ≤ 0.3 to exclude the interference of extreme values on the distribution fit.
[0095] Exponential distribution fitting: Using the fitdistr function of the MASS package in R, fit the exponential distribution and estimate the λ parameter and its confidence interval.
[0096] Model validation: The goodness of fit was evaluated using the Kolmogorov-Smirnov test to ensure the reasonableness of the exponential distribution hypothesis.
[0097] Once the λ parameter is obtained, genomic differentiation data can be used in animal breeding to find regions of hybrid vigor and to develop hybridization effect analysis systems for application.
[0098] Example 2: Verifying that the population phenotypic mean is close to the optimal fitness value under natural selection:
[0099] 2.1 Theoretical Basis and Verification Objectives:
[0100] In Fisher's geometric model (FGM), the optimal fitness value is conceptual. This invention selects the population phenotypic mean as the phenotypic optimal value for each trait, and the rationality of this selection needs to be verified. The verification method is to determine whether the population mean is close to the optimal fitness value by quantifying the phenotypic distance between the population size variation of genotypes and their relative population mean.
[0101] The experimental population used in this embodiment was constructed at the Southwest Domestic Pig Molecular Breeding Base, using European Large White (LW) and East Asian Min pigs (MIN) as parent breeds to create an LW-MIN crossbred population. Specifically, 5 male LW pigs and 15 female MIN pigs were selected as F0 grandparents and crossbred using artificial insemination to obtain the F1 generation. From the F1 generation, 9 males and 45 females were selected for mating, producing 578 F2 individuals, including 294 males and 284 females. A complete family pedigree was established to record the kinship of all individuals.
[0102] 2.2 Verification of the range of traits:
[0103] All 135 phenotypic traits were used in the validation, including:
[0104] 131 continuous traits: body weight (30, 60, 120, 180, 240 days of age), body length, organ weight, skin and backfat thickness, bone size, eye muscle shape, and physiological indicators, etc.
[0105] Four discrete traits were identified: ear shape (erect ear = 3, semi-erect ear = 2, droopy ear = 1), skin color (black = 2.5, white = 0.5, mottled = 1.5), number of ribs, and number of vertebrae.
[0106] To verify the rationality of using the population phenotypic mean as the optimal value for fitness, the relationship between population size change (ΔN) and phenotypic distance (d) relative to a specific genotype and a control genotype was analyzed. The specific genotype refers to a heterozygous genotype (A / B or B / A), and the control genotype refers to a homozygous genotype (A / A or B / B). The four comparison cases are as follows:
[0107] A / A serves as the control genotype, while A / B serves as the specific genotype.
[0108] A / A serves as the control genotype, while B / A serves as the specific genotype.
[0109] B / B serves as the control genotype, while A / B serves as the specific genotype.
[0110] B / B serves as the control genotype, while B / A serves as the specific genotype.
[0111] In this LW-MIN hybrid population, the genotype correspondence with the control genotype is as follows:
[0112] A / A = MIN / MIN, A / B = MIN / LW, B / A = LW / MIN, B / B = LW / LW.
[0113] 2.3 The verification process includes the following steps:
[0114] 2.3.1 Population Size Change Analysis: Calculate the change in the number of individuals with a specific genotype relative to the control genotype (ΔN), defined as: ΔN = (N... 特定基因型 -N 对照基因型 ) / N 总体;
[0115] Where, N 特定基因型 and N 对照基因型 N represents the number of individuals with the specific genotype and the control genotype, respectively. 总体 This represents the total number of individuals within the window. A positive value ΔN indicates an increase in the frequency of a specific genotype, while a negative value indicates a decrease in the frequency.
[0116] 2.3.2 Phenotypic Distance Analysis: The phenotypic distance (d) for all genotypes is calculated, defined as the absolute difference between the phenotypic mean and the mean of the sex-specific population.
[0117] d = |phenotype (genotype) - phenotype (total sex)|;
[0118] Wherein, phenotype (genotype) is the mean phenotype corresponding to the genotype subgroup, and phenotype (total sex) is the mean overall phenotype of the F2 population of that sex. Therefore, we can obtain:
[0119] d(A / A) = |phenotype (A / A genotype) - phenotype (total sex)|;
[0120] d(A / B) = |phenotype (A / B genotype) - phenotype (total sex)|;
[0121] d(B / B) = |phenotype (B / B genotype) - phenotype (total sex)|;
[0122] d(B / A) = |phenotype (B / A genotype) - phenotype (total sex)|;
[0123] If the expected d value for a specific genotype is less than that for the control genotype, it indicates that its phenotype is closer to the population mean.
[0124] Phenotypic distance difference calculation:
[0125] Calculate the phenotypic distance difference between a specific genotype and a control genotype based on four genotype comparison combinations:
[0126] ① Compare combination 1, MIN / MIN and MIN / LW:
[0127] Phenotypic distance difference = d(MIN / MIN) - d(MIN / LW);
[0128] ② Compare combination 2, MIN / MIN and LW / MIN:
[0129] Phenotypic distance difference = d(MIN / MIN) - d(LW / MIN);
[0130] ③ Compare combination 3, LW / LW with MIN / LW:
[0131] Phenotypic distance difference = d(LW / LW) - d(MIN / LW);
[0132] ④ Compare combination 4, LW / LW and LW / MIN:
[0133] Phenotypic distance difference = d(LW / LW) - d(LW / MIN).
[0134] Hypothesis verification: If the population phenotypic mean is close to the fitness optimum, the expected phenotypic distance *d* between heterozygous genotypes (MIN / LW, LW / MIN) and homozygous genotypes (MIN / MIN, LW / LW) is smaller than that between homozygous genotypes (MIN / MIN, LW / LW), meaning the heterozygous genotype phenotype is closer to the population mean. Therefore, the phenotypic distance difference in the above four comparisons should be positive, indicating that the heterozygous genotype has a phenotypic expression closer to the population mean than the homozygous genotype.
[0135] Data standardization: To facilitate statistical analysis, phenotypic distance data were standardized: the phenotypic distance of the control genotype was multiplied by 1000 and rounded up to obtain the position coordinates; the phenotypic distance of the specific genotype was multiplied by 1000 and rounded up to obtain the difference coordinates. This method preserves the relative relationships of the data while facilitating subsequent grouping statistics and visualization analysis. Through this systematic phenotypic distance analysis, the degree of deviation of different genotypes from the population mean can be quantitatively assessed, providing data support for verifying the rationality of the population mean as the optimal value for fitness.
[0136] 2.3.3 Co-occurrence Verification: Analyze whether a specific genotype simultaneously exhibits a positive ΔN (increased frequency) and a smaller d (closer to the mean). A two-dimensional analysis plot is constructed, with the distance to the control genotype on the x-axis and the distance to the specific genotype on the y-axis, and population size changes color-coded. A proportionally scaled reference line is added; when a data point is below the reference line, it indicates that the specific genotype is closer to the population mean.
[0137] Verification results: such as Figure 2 As shown in Figure -A, analysis of the male F2 population revealed that when a specific genotype (heterozygote) was closer to the population mean than the control genotype (homozygote), the population size change was mostly positive, indicating an increase in the proportion of heterozygous genotypes in the population. Conversely, when a specific genotype was farther from the population mean, the population size change was mostly negative, indicating a decrease in its proportion within the population. Figure 2 As shown in -B, the validation results of female individuals showed a consistent pattern with those of males, also indicating that heterozygous genotypes closer to the population mean had higher relative fitness.
[0138] Statistical validation: Spearman rank correlation analysis was used to verify the correlation between ΔN and the phenotypic distance difference (d control - d specific). The analysis results showed a significant negative correlation (r = -0.43, p = 1.2 × 10⁻⁻⁻⁴). 8 This indicates that the closer a genotype is to the population mean, the greater the increase in its population frequency. This result strongly supports the conclusion that the population phenotypic mean is close to the optimal value for fitness.
[0139] Example 3: Application of the method of the present invention in pig breeding:
[0140] Includes the following steps:
[0141] 3-1. Construction of the pig hybridization experimental population: The experimental population used in this example is the same as that in Example 2.
[0142] 3-2. Phenotypic Data Collection:
[0143] Comprehensive phenotypic data were collected from all individuals in the experimental group, covering a total of 135 traits, including:
[0144] ① Growth and development related indicators: birth weight, weight at 30 days of age, weight at 60 days of age, weight at 120 days of age, weight at 180 days of age, weight at 240 days of age, height, length, etc.
[0145] ②Muscle characteristics: muscle fiber type, muscle fiber area, muscle content, etc.;
[0146] ③Indicators of fat deposition: back fat thickness, abdominal fat thickness, intramuscular fat content, etc.
[0147] ④ Skeletal characteristics: bone size, weight, density, etc.;
[0148] ⑤ Organ weight: The weight of major organs such as the heart, liver, spleen, lungs, and kidneys.
[0149] For time-series traits such as body weight, measurements were taken at six time points: at birth (0 days), 30 days, 60 days, 120 days, 180 days, and 240 days. For other traits, measurements were taken at 240±3 days of age. All measurements were performed according to standardized operating procedures to ensure the accuracy and consistency of the data.
[0150] 3-3. Genome-wide genotyping:
[0151] Genomic DNA was extracted from all laboratory animals and genotyped using the Illumina PorcineSNP60 microarray, which contains approximately 60,000 SNP markers covering the entire pig genome. Simultaneously, whole-genome resequencing was performed on the F0 grandparents at a sequencing depth of 30X to obtain higher-density genomic variation information.
[0152] Genome sequencing was performed using the HiSeq2000 platform, and sequencing libraries with 500 bp inserts were constructed. Sequencing data were aligned to a Duroc reference genome using BWA software, and SAMtools was used for sorting, merging, and PCR duplication removal of the BAM files. To reduce erroneous alignments of short reads in repetitive regions, genomic reads from secondary alignments were removed. SNP detection was performed using the GATK toolkit.
[0153] 3-4. Haplotype Inference and Recombination Identification:
[0154] Phase analysis was performed on the SNP genotype data of all individuals using SHAPEIT2 software to infer haplotypes. During haplotype inference, known family information was incorporated for error correction. The haplotypes of F2 individuals were analyzed to determine paternal and maternal origin based on sequence similarity to the four haplotypes of the F1 parents.
[0155] Develop an R-based analysis script to identify recombination events and their sites during F1 gamete formation by comparing the differences between F1 haplotypes and the haplotypes passed on to F2. Record the chromosomal location, size, and involved gene regions for each recombination event.
[0156] 3-5. Genomic differentiation analysis:
[0157] The genome differentiation analysis includes performing differentiation analysis on the genome data of the F0 ancestor, using specialized software to calculate the FST value across the entire genome, quantifying the degree of genome differentiation between parents, dividing the entire genome into fixed-size windows, calculating the FST value at the window level, and dividing the entire genome window into multiple gradient levels based on the size of the FST value.
[0158] In this specific embodiment, the degree of genomic divergence between the LW and MIN varieties is calculated based on the resequencing data of the F0 ancestor. Specifically, the FST estimator of Weir and Cockerham is used to calculate the FST value across the entire genome within a 100kb sliding window.
[0159] Based on the FST value, the 100kb window of the whole genome was divided into 20 quantile intervals, representing the gradient from low to high differentiation. Each differentiation level window was labeled and numbered in preparation for subsequent analysis.
[0160] 3-6. Hybridization effect analysis:
[0161] The hybridization effect was analyzed based on the Fisher geometric model (FGM). First, the genotypes of F2 individuals were divided into four types within each 100kb window according to haplotype information: MIN / MIN (two haplotypes from MIN sources), MIN / LW (maternal MIN haplotype and paternal LW haplotype), LW / MIN (maternal LW haplotype and paternal MIN haplotype), and LW / LW (two haplotypes from LW sources).
[0162] For each 100kb window and for each trait, calculate the phenotypic differences for the following four sets of comparisons:
[0163] ① Differences between MIN / MIN and MIN / LW; ② Differences between MIN / MIN and LW / MIN; ③ Differences between LW / LW and MIN / LW; ④ Differences between LW / LW and LW / MIN.
[0164] Phenotypic differences were standardized relative to the phenotypic mean for each sex to eliminate the influence of sex effects. Based on the Fisher geometric model (FGM), we used the following method to quantify the hybridization effect:
[0165] 1) Calculating the phenotypic optimum: For each trait, the phenotypic mean of the male and female F2 populations was calculated separately as the sex-specific phenotypic optimum. To verify the rationality of using the population mean as the fitness optimum, we compared the relationship between the population size change of a specific genotype relative to the control genotype and the phenotypic distance (relative to the population mean) to confirm the closeness between the phenotypic mean and the fitness optimum.
[0166] 2) Calculate the absolute phenotypic distance (d): For a specific genotype in each genomic window, calculate the difference between its phenotypic mean and the overall phenotypic mean for the corresponding sex, using the following formula:
[0167] The absolute phenotypic distance (d) is: d = |phenotype (genotype) - phenotype (total sex)|, where phenotype (genotype) is the phenotypic average of each genotype combination, and phenotype (total sex) is the overall phenotypic average of the F2 population of that sex.
[0168] 3) Calculate the relative phenotypic distance (D): The relative phenotypic distance (D) is: D = d / phenotype (sex population). This standardization process eliminates the influence of dimensional differences between different traits, making the effect sizes of different traits comparable. This provides a basis for subsequent calculations that combine all traits.
[0169] 4) Calculate the size of the allele substitution effect: Hybridization effect size = |D(genotype) - D(control genotype)|; where the genotype and control genotype differ by only one allele substitution.
[0170] Based on the principle of single allele substitution, the following comparison combinations are formed:
[0171] Comparison of MIN / MIN and MIN / LW (maternal allele substitution);
[0172] Comparison of MIN / MIN and LW / MIN (paternal allele substitution);
[0173] Comparison of MIN / LW and LW / LW (paternal allele substitution).
[0174] Comparison of LW / MIN and LW / LW (maternal allele substitution).
[0175] In each comparison, the two genotypes differ by only one allele from the paternal or maternal lineage, ensuring that the analysis focuses on a pure, single-step allele substitution effect. These paired comparisons allow for a systematic assessment of the impact of allele substitutions in different directions and backgrounds on hybridization effects.
[0176] The specific calculation of the hybridization effect between males and females is as follows:
[0177] Calculation of female hybridization effect size: Hybridization effect data of female individuals were extracted from the genomic database. The specific steps were as follows: A genomic window with chromosomes 1 to 18 (i.e., autosomes) was selected, ensuring that the window had corresponding records in both the mutation effect statistics table and the FST differentiation table, and that the weighted FST value of the window fell within the specified differentiation interval. For windows meeting the criteria, the absolute difference between the phenotypic deviation of a specific female genotype and that of the female control genotype was calculated as the effect size, and data points with an effect size not exceeding 0.3 were selected for subsequent analysis.
[0178] Male effect size calculation: Using the same screening conditions and calculation logic as for females, hybridization effect data for male individuals were extracted from the genomic database. Specifically, under the same chromosome range and FST interval constraints, the absolute difference between the phenotypic deviation of male-specific genotypes and that of male control genotypes was calculated to obtain male-specific effect size data. Data points with effect sizes ≤0.3 were also screened to ensure the same quality standards and analytical scope as the female data.
[0179] The effect size of allele substitution on a trait is defined as |D|. (基因型)- D (对照基因型) The genotypes differed from the control genotypes by only one allele substitution. By combining the allele substitution effect sizes across all 135 traits genome-wide, an effect size distribution was formed, with some data shown in [reference needed]. Figure 3 .
[0180] 5) λ parameter estimation: Using the 'fitdistr' function from the MASS package in R, an exponential distribution was fitted to the allele substitution data with effect sizes ≤ 0.3 to obtain the distribution parameter λ. The λ parameter reflects the distribution characteristics of the hybridization effect size. A larger λ value indicates greater benefit; a larger λ parameter indicates that small effects dominate, resulting in heterosis; a smaller λ parameter indicates that large effects dominate, resulting in heterosis. The genome-wide distribution of λ values was observed, with peak positions representing regions of hybridization advantage.
[0181] The fitting process for the λ parameter is as follows: Select data with a hybridization effect size ≤ 0.3 to eliminate the interference of extreme values on the distribution fitting, use the fitdistr function of the MASS package in R language to fit the exponential distribution, estimate the λ parameter and its confidence interval, evaluate the goodness of fit through the Kolmogorov-Smirnov test to ensure the rationality of the exponential distribution hypothesis, and finally obtain the λ parameter.
[0182] Example 4: Identifying regions of hybrid vigor using genomic differentiation data:
[0183] 4.1 Establishing the relationship between FST and the λ parameter: Based on the FST value and the λ parameter, analyze the association between the two within a fixed window across the whole genome, including:
[0184] 4.1.1 FST grading: The FST value of the whole genome is divided into 20 equal quantile intervals. The part with FST less than 0 is recorded as 0. From low differentiation (FST≈0) to high differentiation (FST≈1), it represents the gradient of genetic differences between parents.
[0185] 4.1.2 Summary of λ parameters: For each FST interval, collect the λ parameters of the corresponding window and calculate the mean and standard deviation.
[0186] 4.1.3 Relationship Modeling: Use a nonlinear regression model to fit the relationship between the FST value and the λ parameter, plot the FST-λ relationship curve, and identify the critical point of the hybridization effect of this trait from heterosis to heterosis by the inflection point or slope change of the curve.
[0187] 4.2 Identifying Hybridization Dominance Regions: Based on the FST-λ relationship curve, windows with λ parameters significantly higher than the genome-wide average were marked as hybridization dominance regions. The specific steps are as follows:
[0188] 4.2.1 Threshold Calculation: Calculate the mean (μ) and standard deviation (σ) of the λ parameter;
[0189] 4.2.2 Determine the core heterosis threshold: Tcore = μ + kσ (k = 1.0);
[0190] 4.2.3 Determine the extended heterosis threshold: T_extension = max(T_core, μ−cσ) (c=0.5);
[0191] 4.2.4 Adaptive adjustment mechanism: When max(λ) < (μ + 1.5σ), k is automatically set to 1.0; when max(λ) ≥ (μ + 1.5σ), k is set to 1.5.
[0192] 4.3 Trait Category Analysis: Plot FST-λ relationship curves for different traits to identify the hybridization advantage regions of each trait.
[0193] In this specific implementation, the hybridization effect parameter λ was calculated for each of the 20 FST quantile intervals, and the results are shown in Table 1. A curve showing the relationship between FST values and the λ parameter was plotted to identify the FST regions of hybridization dominance. Figure 4 .Depend on Figure 4 It can be seen that as FST (genetic differentiation degree) increases, a distinct inverted U-shaped distribution of λ is observed:
[0194] In the region FST < 0.095: λ value is low; in the region 0.095 < FST < 0.286: λ reaches its peak; in the region FST > 0.286: λ value decreases.
[0195] Table 1: Correspondence between FST-λ for overall traits
[0196]
[0197] The whole genome analysis showed that the mean μ = 33.46, the standard deviation σ = 0.436, the maximum value max(λ) = 34.04, and the minimum value min(λ) = 32.30.
[0198] Adaptive threshold calculation: Validation condition: max(λ) = 34.04 < (33.46 + 1.5 × 0.436) = 34.114, enable adaptive adjustment: k = 1.0.
[0199] Calculate the hybrid vigor region and check the coverage. Core heterosis threshold: T 核心 = 33.46 + 0.436 = 33.896; Extended dominance threshold: T 扩展 = max(33.5, 33.46-0.218) = 33.5. The calculated heterosis region table is shown in Table 2.
[0200] Table 2: Hybrid Advantage Regions
[0201]
[0202] After calculating the threshold range, the threshold standard for expanding the hybrid vigor region was found to be λ > 33.5. Referring back to Table 1: FST-λ correspondences, the range of 0.09454525 to 0.2604634 in intervals 6 and 16 is found to be FST-λ. ST The turning point, the region of hybrid vigor is F ST The range is from 0.09454525 to 0.2604634. This indicates that in pig crossbreeding, the FST range of 0.09454525 to 0.2604634 represents the optimal degree of genetic differentiation between parents. In breeding practice, breeds with genomic differentiation falling within this range should be selected for crossbreeding. Avoid selecting parental combinations with FST < 0.095 (insufficient differentiation) or FST > 0.260 (over-differentiation). Breed pairing strategy: Direct crossbreeding is not recommended for closely related breeds with excessively low differentiation, as the heterosis effect is limited; for distantly related breeds with excessively high differentiation, careful evaluation is necessary, as hybrid disadvantage may occur; focus should be placed on breed combinations with moderate differentiation, which are expected to produce significant heterosis.
[0203] The method of the present invention further includes finding sex-specific hybrid vigor regions: analyzing the FST-λ relationship curves of male and female F2 populations respectively, plotting FST-λ relationship curves for different sexes for different traits, and identifying sex-specific hybrid vigor regions.
[0204] Specifically, in this embodiment, the above analysis was performed on both male and female individuals, comparing the differences in hybridization effects between the sexes and their relationship with genome differentiation, resulting in a curve illustrating the relationship between male and female differences. Figure 5 This allows for a more precise localization of the hybrid vigor regions in different sexes. Figure 5It can be observed that, for all 135 traits, the λ value of males is generally higher than that of females. This indicates that when improving these 135 traits, males are more effective than females in using hybrid vigor to improve them.
[0205] On the other hand, it also helps in the precise selection of genomic regions. When identifying dominance markers, priority should be given to genomic windows within the FST heterosis region, as the SNP markers contained therein have higher breeding value. Molecular markers from these regions can be developed into a core marker set for heterosis prediction, establishing a parental evaluation system based on markers from heterosis regions. In marker-assisted selection, priority should be given to detecting marker genotypes within heterosis regions, and the dominance performance of hybrid offspring can be predicted based on the genotypic differences between parents in these regions. Developing rapid, low-cost marker detection schemes can improve breeding efficiency.
[0206] Example 5: Calculate the possible different optimal FST intervals for different economic traits:
[0207] The method of this invention can also analyze different categories of traits (such as hip circumference, bone structure, growth traits, muscle traits, fat deposition, etc.) separately, exploring the hybridization effect patterns of different trait categories by using specific traits. Using the method of this invention, different optimal FST intervals can be calculated for different economic traits (growth, reproduction, disease resistance, etc.).
[0208] After obtaining the λ parameter using the method in Example 1, the hybridization effect patterns of different single traits or different categories of traits were analyzed. It was found that various traits have different FST-λ relationship characteristics and heterosis expression patterns, which provides a scientific basis for formulating trait-oriented precision breeding strategies.
[0209] Specifically, in this embodiment, a single-target analysis of the hip circumference trait of the experimental group in Example 3 can be performed to obtain... Figure 6 According to the method in Example 4, the area above the dotted line in the figure is the region of hybrid vigor for this trait. Female: T 核心 =103.5984, Male: T 核心 =116.052. The specific corresponding FST-λ is shown in Table 3:
[0210] Table 3: Correspondence between FST-λ for hip circumference traits
[0211]
[0212] As shown in Table 3, in pig crossbreeding, for the rump girth trait, the region with the greatest hybrid vigor in female parents is the FST range of 0.09454525 to 0.2604634, and breeds with genomic differentiation falling within this range should be selected for crossbreeding. Similarly, for the rump girth trait, the regions with the greatest hybrid vigor in male parents are the FST ranges of 0.09472785 to 0.108589, 0.108589 to 0.12206225, 0.1347032 to 0.14782405, and 0.17536825 to 0.1900842, and breeds with genomic differentiation falling within these ranges should be selected for crossbreeding. Based on these conclusions, a specific crossbreeding program can be developed for pig breeding of the rump girth trait.
[0213] A joint analysis of the skeletal-related traits (mid-trunk back bones, anterior bones, and hind bones) of the experimental group in Example 3 yielded the following results: Figure 7 According to the method in Example 4, the area above the dotted line in the figure is the region of hybrid vigor for this trait. Female: T 核心 =25.8296, Male: T 核心 =27.11041. The specific corresponding FST-λ is shown in Table 4:
[0214] Table 4: Correspondence between FST-λ for skeletal traits
[0215]
[0216] As shown in Table 4, in pig crossbreeding, for skeletal-related traits (mid-body backbone, forequarters, and hindquarters), the FST range of 0.14782405 to 0.1612275 and 0.1900842 to 0.20604825 for female parents represents the region with the greatest hybrid vigor, and breeds with genomic differentiation falling within this range should be selected for crossbreeding. Similarly, for skeletal-related traits (mid-body backbone, forequarters, and hindquarters), the FST range of 0.2241914 to 0.243097 for male parents also represents the region with the greatest hybrid vigor, and breeds with genomic differentiation falling within this range should be selected for crossbreeding. Based on these conclusions, a specific crossbreeding program can be developed for pig breeding of skeletal-related traits (mid-body backbone, forequarters, and hindquarters).
[0217] In actual breeding, different parent selection and hybridization strategies can be adopted according to different breeding objectives: For breeding projects with growth performance as the main improvement objective, the focus should be on the genotype combination in the hybridization advantage region of growth traits, and parents with appropriate FST values in this region should be selected for hybridization; for muscle quality improvement, it is necessary to identify the hybridization advantage region specific to muscle traits, which may differ from the optimal FST interval of growth traits; for the improvement of fat deposition traits, it is also necessary to develop a special hybridization scheme based on the specific FST-λ relationship of this trait category to find a more accurate FST interval.
[0218] The results of this invention can be visualized and output through the development of computer programs: FST-λ relationship curves and sex comparison diagrams can be generated using R scripts, and a detailed list of hybrid vigor regions can be output through calculation and analysis, including chromosome location, λ value, FST value, and annotated genes. However, even without developing a dedicated analysis system, hybrid vigor regions can be identified using statistical software and bioinformatics tools.
[0219] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of equivalents of the claims be included within the present invention.
[0220] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A genome-based method for analyzing and predicting animal hybridization effects, comprising the following steps: S1. Construct a hybridization experimental population; S2. Collect phenotypic data and perform cleaning processing; S3. Genotyping of all individuals in the experimental population; S4. Use specialized software to infer the haplotype information of all individuals in the experimental population; S5. Genomic differentiation analysis; Its characteristic is that it further includes S6, hybridization effect analysis, which specifically includes the following steps: S6-1. Differentiate genotype combinations: Based on the haplotype information inferred in step S4, classify the genotypes of F2 generation individuals into the following four genotype combinations within each fixed genomic window: A / A, B / B, A / B, B / A. S6-2. Determine the optimal phenotypic value: For each trait, calculate the average phenotypic value of the male and female F2 populations respectively, which is taken as the optimal phenotypic value for that sex. S6-3. For each trait, quantify the deviation between the phenotypic expression and the optimal phenotypic value for each genotype combination using phenotypic distance, divided into absolute phenotypic distance (d) and relative phenotypic distance (D): The absolute phenotypic distance (d) is: d = |phenotype (genotype) - phenotype (total sex)|, where phenotype (genotype) is the phenotypic average of each genotype combination, and phenotype (total sex) is the overall phenotypic average of the F2 population of that sex. The relative phenotypic distance (D): D = d / phenotype (gender population); S6-4. For each trait, calculate the hybridization effect size for each genotype: The hybridization effect size = |D(genotype) - D(control genotype)|, where D(genotype) is the relative phenotypic distance of each of the four genotypes, and D(control genotype) is the relative phenotypic distance of the genotype that differs from D(genotype) by only one allele substitution; the hybridization effect size covers four cases, namely: |D(A / A)-D(A / B)|, |D(A / A)-D(B / A)|, |D(A / B)-D(B / B)|, and |D(B / A)-D(B / B)|. S6-5. Obtain the λ parameter: Fit the data with a hybridization effect size ≤ 0.3 to obtain the λ parameter; the larger the λ parameter, the more dominant the small effect, indicating heterosis; the smaller the λ parameter, the more dominant the large effect, indicating heterosis.
2. The method for analyzing and predicting animal hybridization effects based on genomes according to claim 1, characterized in that, The fitting process for the λ parameter is as follows: Select data with a hybridization effect size ≤ 0.3 to eliminate the interference of extreme values on the distribution fitting, use the fitdistr function of the MASS package in R language to fit the exponential distribution, estimate the λ parameter and its confidence interval, evaluate the goodness of fit through the Kolmogorov-Smirnov test to ensure the rationality of the exponential distribution hypothesis, and finally obtain the λ parameter.
3. The method for analyzing and predicting animal hybridization effects based on genomes according to claim 1, characterized in that, The genome differentiation analysis includes performing differentiation analysis on the genome data of the F0 ancestor, using specialized software to calculate the FST value across the entire genome, quantifying the degree of genome differentiation between parents, dividing the entire genome into fixed-size windows, calculating the FST value at the window level, and dividing the entire genome window into multiple gradient levels based on the size of the FST value.
4. The method for analyzing and predicting animal hybridization effects based on genomes according to claim 3, characterized in that, The entire genome was divided into fixed windows of 100kb each.
5. The method for analyzing and predicting animal hybridization effects based on genomes according to claim 3, characterized in that, It also includes using genomic differentiation data to identify regions of hybrid vigor, specifically including the following steps: 1) Establish the relationship between FST and the λ parameter: Based on the FST value and the λ parameter, analyze the association between the two within a fixed window across the whole genome, including: 1-1) FST grading: The FST value of the whole genome is divided into 20 equal quantile intervals. The part with FST less than 0 is recorded as 0. From low differentiation (FST≈0) to high differentiation (FST≈1), it represents the gradient of genetic differences between parents. 1-2) Summary of λ parameters: For each FST interval, collect the λ parameters of the corresponding window and calculate the mean and standard deviation. 1-3) Relationship modeling: Use a nonlinear regression model to fit the relationship between the FST value and the λ parameter, plot the FST-λ relationship curve, and identify the critical point of the hybridization effect of the trait from heterosis to heterosis by the inflection point or slope change of the curve. 2) Identifying regions of hybrid vigor: Based on the FST-λ relationship curve, windows with λ parameters significantly higher than the average level of the whole genome are marked as regions of hybrid vigor. The specific steps are as follows: 2-1) Threshold calculation: Calculate the mean (μ) and standard deviation (σ) of the λ parameter; 2-2) Determine the core heterosis threshold: Tcore = μ + kσ (k = 1.0); 2-3) Determine the extended heterosis threshold: T_extension = max(T_core, μ−cσ) (c=0.5); 2-4) Adaptive adjustment mechanism: When max(λ) < (μ + 1.5σ), k is automatically set to 1.0; when max(λ) ≥ (μ + 1.5σ), k is set to 1.
5. 3) Trait category analysis: Plot FST-λ relationship curves for different traits to identify the hybridization advantage regions of each trait.
6. The method for analyzing and predicting animal hybridization effects based on genomes according to claim 5, characterized in that, It also includes finding sex-specific hybrid vigor regions: analyzing the FST-λ relationship curves of male and female F2 populations separately, plotting FST-λ relationship curves for different sexes for different traits, and identifying sex-specific hybrid vigor regions.
7. The method for analyzing and predicting animal hybridization effects based on genomes according to claim 5, characterized in that, It also includes visualization and results output: using R scripts to generate FST-λ relationship curves and sex comparison plots, and outputting a detailed list of hybrid vigor regions through calculation and analysis, including chromosome location, λ value, FST value and annotated genes.
8. The application of the genome-based animal hybridization effect analysis and prediction method according to any one of claims 1-7 in animal breeding.
Citation Information
Patent Citations
Genome-wide prediction method and device based on automated machine learning technology
CN109727640B
Pig whole genome low-density 5K SNP chip and application
CN118480613A