Genomic selection methods for machine learning of additive and non-additive effects

By constructing a genome relationship matrix using UKin unbiased kinship estimation and machine learning methods, the problems of matrix singularity and bias in genome selection are solved, improving the accuracy and unbiasedness of genome prediction, especially showing good results in the prediction of complex traits.

CN119108012BActive Publication Date: 2026-04-21INSTITUTE OF ANIMAL SCIENCES OF CHINESE ACADEMY OF AGRICULTURAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSTITUTE OF ANIMAL SCIENCES OF CHINESE ACADEMY OF AGRICULTURAL SCIENCES
Filing Date
2024-08-27
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing genome selection methods, the singularity, bias, and redundancy of the genome relationship matrix affect accuracy. Traditional methods are unable to effectively consider the interaction effects between genes, especially non-additive effects, leading to biases in genetic analysis.

Method used

The UKin unbiased kinship estimation method was used to construct a genome relationship matrix. Multinomial kernel ridge regression and support vector machine regression in machine learning were combined to increase non-additive effects. The genome relationship matrix between individuals was constructed through kernel functions to capture the additive and non-additive effects of traits.

Benefits of technology

It improves the accuracy and unbiasedness of genome prediction, better explains genetic variation, and enhances the accuracy and efficiency of genome selection, especially showing good results in the prediction of complex traits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119108012B_ABST
    Figure CN119108012B_ABST
Patent Text Reader

Abstract

The application discloses a genomic selection method for machine learning of additive and non-additive effects, and by adding non-additive effects to a GBLUP model, the prediction accuracy of the GBLUP model is greatly improved. By applying a kernel function in the GBLUP, the model can better simulate the complex genetic relationship between genes and the nonlinear combination of gene effects, thereby improving the explanation ability of genetic variation. In addition, a machine learning strategy K P RR is constructed based on a polynomial kernel function and a kernel ridge regression. By combining the advantages of the polynomial kernel function and the kernel ridge regression method, the K P RR strategy allows remote data points to contribute to the value of the model, which enables the interaction in the genomic data to be more effectively contributed, and the genomic prediction effect can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of genome selection technology, and more specifically to a machine learning-based genome selection method for additive and non-additive effects. Background Technology

[0002] Genomic selection (GS) utilizes genome-wide single nucleotide polymorphism (SNP) markers to predict an individual's genomic estimated breeding value (GEBV), thereby achieving efficient selection for important economic traits. The genomic relation matrix measures the closeness of kinship between individuals, playing a crucial role in genomic selection. However, genomic relation matrices constructed using conventional methods currently suffer from several problems:

[0003] 1. Singularity of the matrix. When the number of loci is finite, or two individuals have the same genotype on all markers, or the number of markers is less than the number of individuals, the genomic relationship matrix may be singular, making it impossible to solve genomic selection models;

[0004] 2. Differences between genomic relation matrices constructed using different methods can directly impact genomic selection decisions and genetic progress. For example, the presence of different populations or subpopulations within a sample can lead to biased or misleading results in the genomic relation matrix. Therefore, a suitable genomic relation matrix is ​​needed to ensure the speed and direction of genetic improvement. However, genomic relation matrices constructed using existing methods are prone to affecting the accuracy of genomic estimations of breeding values.

[0005] In addition, traditional genome selection methods, such as those based on Bayesian theory, have the following problems:

[0006] 1) Genomic selection relies on high-density genotypes to predict the genetic value of an individual. A larger number of markers means that more accurate genomic estimates of breeding values ​​can be obtained, but at the same time, more redundant information is introduced, which will affect the accuracy of traditional genomic selection methods.

[0007] 2) The statistical models and algorithms used in genome selection have certain limitations. For example, the efficiency of Bayesian methods decreases significantly when the number of markers increases, which limits the further application of Bayesian methods.

[0008] 3) Traditional genomic selection methods are based on linear models and cannot take into account the interaction effects between genes. Therefore, the complex genetic background of quantitative traits cannot be fully considered.

[0009] Furthermore, in the early stages of genomic selection, many methods only consider additive effects in genomic data (the sum of genotype values ​​of multiple minor genes affecting quantitative traits). In reality, non-additive effects (effects arising from interactions between alleles or non-allelic genes) play a central role in hybrid vigor, polymorphism, and genomic selection. Ignoring non-additive genetic effects can bias multiple genetic analysis processes. Therefore, genomic prediction models integrating non-additive effects can achieve better predictive performance. However, how to integrate non-additive effects into the genomic selection process to improve its accuracy remains a pressing technical problem in this field. Summary of the Invention

[0010] This invention aims to improve the accuracy of genome selection by providing a machine learning-based genome selection method that addresses both additive and non-additive effects.

[0011] To achieve this objective, the present invention adopts the following technical solution:

[0012] A machine learning-based genome selection method for additive and non-additive effects is provided, comprising the following steps:

[0013] S1, construct the genome relationship matrix using the UKin unbiased kinship estimation method;

[0014] S2, add non-additive effects to the linear model and construct a non-additive genomic relationship matrix;

[0015] S3 uses machine learning to capture the additive and non-additive effects of traits, and uses the learned machine model to predict individual traits and output the prediction results.

[0016] Preferably, in step S2, the statistical model that adds non-additive effects to the linear model is expressed by the following formula (2-10):

[0017] y = 1μ + Z a g a +D d g d +e (2-10)

[0018] In formula (2-10), y represents the observed value of the target trait;

[0019] μ represents the population mean;

[0020] g a Z represents the value of a sports species. a This indicates a sports species value g. a The correlation matrix;

[0021] g dIndicates the dominant effect. Defined as a normal distribution followed by dominant inheritance; D is the variance of dominant inheritance; D is the marker-based dominant genome relationship matrix.

[0022] d d It is g d The correlation matrix;

[0023] e represents the remaining residual.

[0024] Preferably, in step S2, the statistical model for adding the non-additive effects to the genome relationship matrix is ​​expressed by the following formula (2-11):

[0025] y = 1μ + Z a g a +D ep g ep +e (2-11)

[0026] In formula (2-11), y is the observed value of the target trait;

[0027] μ is the population mean;

[0028] g a Z represents the value of a sports species. a This indicates a sports species value g. a The correlation matrix;

[0029] d ep It is g ep The correlation matrix;

[0030] g ep This is a superordinate effect. Defined as a normal distribution that epistatic genetic effects follow; E is the epistatic variance; E is the epistatic genomic relationship matrix based on SNP markers.

[0031] Preferably, in step S1, the genomic relationship matrix constructed using the UKin unbiased kinship estimation method is expressed by the following expression:

[0032]

[0033] It can be seen from formulas (1-4)-(1-6) that The expected value of ρ is independent of the chosen SNP, therefore, ρ ii' The estimated value can be expressed as:

[0034]

[0035] Where ρ ii′It is the true affinity coefficient between individuals i and i′; It is ρ ii′ The estimated value; Represents ρ ii′ The true calculated value; k and l represent individuals other than individuals i and i′, X ij and X i′j This refers to the number of minor alleles; It is the average minor allele frequency of this SNP. This represents the genotypic variance; j represents the j-th SNP; E represents the expected value; n represents the number of n individuals in the population; m is the number of SNPs; X kj X represents the j-th SNP locus of the k-th individual; lj This represents the j-th SNP locus of the l-th individual; express The expected value.

[0036] 5. The genome selection method for machine learning targeting additive and non-additive effects according to claim 4, characterized in that, in step S3, genome prediction is performed using multinomial kernel ridge regression, specifically including the following steps:

[0037] A1, construct a multinomial kernel ridge regression prediction model as expressed in the following formula (2-1):

[0038]

[0039] In formula (2-4), y(x) represents the genomic prediction value for the input individual x;

[0040] α i This represents the coefficient corresponding to the i-th training sample;

[0041] n represents the number of training samples;

[0042] K represents the kernel matrix composed of the kernel function values ​​of each training sample;

[0043] K(x,x i ) represents the i-th training sample x i The kernel function value between x;

[0044] A2, input individual x into the polynomial kernel ridge regression prediction model, and the model predicts the genome prediction results for x.

[0045] Preferably, the loss function of the polynomial kernel ridge regression prediction model is expressed by the following formula (2-2):

[0046] L(w)=‖φ(X)wy‖ 2 +λ‖w‖ 2 (2-2)

[0047] In formula (2-2), L(w) represents the loss function;

[0048] φ(X) represents the result of mapping the original data X to a high-dimensional space through a kernel function;

[0049] w is the weight vector;

[0050] y is the target variable;

[0051] λ is the regularization parameter;

[0052] The goal of kernel ridge regression is to minimize the loss function expressed by formula (2-2).

[0053] As a preferred option, the optimal solution of the loss function expressed by formula (2-2) is obtained by solving the linear equation expressed by the following formula (2-3):

[0054] (K+λI)α=y (2-3)

[0055] In formula (2-3), I is the identity matrix;

[0056] α is the coefficient vector.

[0057] Preferably, genome prediction is performed using multinomial kernel support vector machine regression, the steps of which include:

[0058] B1, construct a support vector machine regression prediction model as expressed by the following formula (2-8):

[0059]

[0060] In formula (2-8), f(x) represents the genomic prediction value for the input individual x;

[0061] K(x,x i ) represents the i-th training sample x i The kernel function value between the input x and the input x;

[0062] This represents the positive weights of the predicted values ​​for each individual.

[0063] α i This represents the positive weights in the training set; n represents the number of training samples.

[0064] b indicates paranoia;

[0065] B2, input individual x into the support vector machine regression prediction model, and the model predicts the genome prediction results for x.

[0066] Preferably, the loss function of the support vector machine regression prediction model is expressed by the following formula (2-5):

[0067]

[0068] In formula (2-5), Loss ∈ Represents the loss function;

[0069] y represents the true value of the genome corresponding to individual x;

[0070] ∈ represents the threshold.

[0071] As a preferred option, the solution to the optimal solution of the loss function expressed by formula (2-5) is constrained by the following formulas (2-6) and (2-7):

[0072]

[0073] The following conditions must be met:

[0074]

[0075] In formulas (2-6) and (2-7), ξ,ξ * These are slack variables;

[0076] ξ i , Let be the slack variable for the i-th training sample;

[0077] C is the regularization parameter;

[0078] w is the weight;

[0079] b represents paranoia;

[0080] φ(x i ) to use training samples x i Functions mapped to higher-dimensional space;

[0081] y i Let i be the ground truth value of the genome corresponding to the i-th training sample;

[0082] w T It is the transpose of the weight w.

[0083] The present invention has the following beneficial effects:

[0084] 1. Traditional genomic relationship matrices often deviate from the true kinship of individuals, affecting subsequent genetic analyses. This embodiment proposes an unbiased estimation method, the UKin method, which corrects this bias by integrating genetic information from the entire population. The method's ability to improve heritability estimation accuracy is demonstrated. Specifically, the UKin method's predictive performance in the GBLUP model was evaluated on simulated and multiple real-world datasets. The results show varying degrees of variation in both prediction accuracy and unbiasedness. Overall, the UKin method exhibits similar predictive accuracy to traditional methods on both simulated and real-world datasets. These results demonstrate that the UKin method addresses the biased estimation problem of traditional methods while maintaining the accuracy and unbiasedness of genomic predictions, showing valuable application potential in predicting complex traits. This superior performance may stem from two reasons. First, it more accurately reflects the genetic relationships between individuals. Unbiased estimation ensures that the estimated kinship values ​​do not deviate from the true values ​​due to systematic bias, which is particularly important for predicting breeding values ​​in genomic prediction. Secondly, the UKin method effectively handles negative values ​​in the genome relation matrix, which can be a cause of bias in traditional methods. By optimizing the handling of these negative values, the UKin method can more accurately capture genetic variations, thereby ensuring the accuracy and unbiasedness of predictions.

[0085] 2. This embodiment introduces the concept of kernel functions from machine learning to study methods for constructing genomic relationship matrices between individuals. It utilizes machine learning kernel functions to construct genomic relationship matrices between individuals, replacing the traditional relationship matrix in GBLUP, and for the first time studies this method on different types of datasets. The information required for constructing genomic relationship matrices using kernel functions is the same as traditional methods; only the genotype data of the population is needed. By applying kernel functions in GBLUP, the model can better simulate complex genetic relationships between genes and the nonlinear combinations of gene effects, thereby improving the explanatory power of genetic variation. Furthermore, the kernel function method provides a variety of kernel function options, allowing the model to flexibly select the most suitable kernel function based on the characteristics of the dataset and the research objectives.

[0086] 3. By adding non-additive effects to the GBLUP model using formulas (2-10) and (2-11), the prediction accuracy of the GBLUP model was significantly improved.

[0087] 4. This application constructs a machine learning strategy K based on multinomial kernel functions and kernel ridge regression. PKRR was first applied to genome prediction research and tested on simulated and real datasets. The multinomial kernel function plays a crucial role in the prediction process, mapping the input genotype data to a kernel space, which differs from linear models. This application combines the advantages of multinomial kernel functions and kernel ridge regression methods. P The RR strategy allows distant data points to contribute to the model's values, which enables interactions in genomic data to contribute more effectively and can significantly improve genomic prediction performance. Attached Figure Description

[0088] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0089] Figure 1 This is a diagram illustrating the implementation steps of the genome selection method for machine learning targeting additive and non-additive effects provided in this embodiment of the invention.

[0090] Figure 2 This is a comparison chart of the performance of the Ukin method and other genome relation matrix construction methods in GBLUP prediction;

[0091] Figure 3 This is a comparison chart of the unbiasedness of GBLUP predictions using different genome relationship matrices in a simulated dataset;

[0092] Figure 4 This is a comparison chart of the accuracy of GBLUP predictions using different genome relationship matrices in the wheat dataset;

[0093] Figure 5 This is a comparison chart of the accuracy of GBLUP predictions using different genome relationship matrices in the dairy cow dataset and the large white pig dataset;

[0094] Figure 6 This is a comparison chart of the unbiasedness of GBLUP predictions using different genome relationship matrices in real datasets;

[0095] Figure 7 This is a comparison chart showing the accuracy of GBLUP predictions when using different kernel functions to construct genome relationship matrices in a simulated dataset.

[0096] Figure 8 This is a comparison chart showing the unbiasedness of GBLUP predictions when constructing genome relation matrices using different kernel functions in a simulated dataset.

[0097] Figure 9This is a comparison chart showing the accuracy of GBLUP predictions when constructing genome relation matrices using different kernel functions in the wheat dataset;

[0098] Figure 10 This is a comparison chart showing the accuracy of GBLUP predictions when constructing genome relationship matrices using different kernel functions in Holstein cattle and Large White pig datasets;

[0099] Figure 11 This is a comparison chart showing the unbiasedness of GBLUP predictions when constructing genome relation matrices using different kernel functions on real datasets.

[0100] Figure 12 This is a comparison chart of the prediction accuracy of different prediction methods used on the wheat dataset;

[0101] Figure 13 This is a comparison chart of the predictive accuracy of epistatic genetic background traits;

[0102] Figure 14 This is a comparison chart of the predictive accuracy of dominant effect genetic background traits;

[0103] Figure 15 This is a comparison chart of the computation time of all methods on all traits. Detailed Implementation

[0104] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0105] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual images. They should not be construed as limiting the scope of this patent. To better illustrate the embodiments of the present invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0106] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "inner," and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present patent. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0107] In the description of this invention, unless otherwise explicitly specified and limited, the term "connection" or similar designation indicating a connection between components should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0108] The genome selection method based on additive and non-additive effects based on machine learning provided in this invention embodiment, such as... Figure 1 As shown, the steps include:

[0109] S1, construct the genome relationship matrix using the UKin unbiased kinship estimation method;

[0110] S2, add non-additive effects to the linear model and construct a non-additive genomic relationship matrix;

[0111] S3 uses machine learning to capture the additive and non-additive effects of traits, and uses the learned machine model to predict individual traits and output the prediction results.

[0112] I. Constructing a genome relation matrix using the UKin method

[0113] One of the key steps in genomic selection is constructing a genomic relationship matrix among individuals. While traditional methods for constructing genomic relationship matrices have been widely applied and validated in various biological contexts, these methods have proven to produce biased estimates of true kinship. When the genetic structure of a population is complex, the bias in the genomic relationship matrix constructed using traditional methods can affect the accuracy and reliability of genomic selection models. To address this issue, this embodiment employs an unbiased kinship (UKin) method, an improvement on traditional methods, to correct for biases in the traditional construction process by comprehensively utilizing the genetic information of the entire population. This application also tests the technical advantages of the genomic relationship matrix constructed using the UKin unbiased kinship estimation method in the GBLUP (Genomic best linear unbiased prediction) model and compares it with traditional construction methods.

[0114] The following details the process of constructing a genome relationship matrix using the UKin unbiased kinship estimation method in this embodiment, as well as the process of testing its technical advantages using the GBLUP model and comparing it with traditional construction methods. Specifically, it includes the following technical steps:

[0115] 1.1 Materials and Methods

[0116] 1.1.1 Test Methods

[0117] 1.1.1.1 Best linear unbiased estimation of the genome

[0118] The GBLUP method assumes that all markers follow the same normal distribution and contribute equally to the genetic variance. It replaces the pedigree-based relation matrix in the mixed linear model with a marker-based genetic relation matrix. The GBLUP model can be written as:

[0119] y = 1μ + Zg a +e (1-1)

[0120] Where y is the phenotypic value vector; μ is the population mean; Z is g a The correlation matrix; It is a vector of random additive genetic effects; It is the residual effect vector. and These represent the additive genetic variance and residual variance, respectively; I is the identity matrix; and G is the label-based additive genomic relation matrix. This model was implemented using the R package "sommer". Subsequent relation matrices were constructed using custom-written code in R.

[0121] 1.1.1.2 Traditional methods for constructing genome relation matrices

[0122] Currently, the traditional genomic relation matrix construction method frequently used in the GBLUP model is expressed by the following equations (1-2) and (1-3):

[0123]

[0124] In equations (1-2) and (1-3), for the j-th (1≤j≤m) SNP, It is an estimate of the kinship coefficient between individuals i and i' (1≤i, i'≤n); p j is the gene frequency of the j-th SNP. m is the number of SNPs; n is the number of individuals; x ij x is the genotype value of the j SNP loci of individual i; i′j Let represent the genotype values ​​of the j SNP loci of individual i′; Z represents the corrected genotype matrix, Z′ represents the transpose of the Z matrix; and G represents the kinship matrix. In this embodiment, the expression expressed by formula (1-2) is defined as the VanRaden method polynomial, and the expression expressed by formula (1-3) is defined as the Yang method polynomial.

[0125] 1.1.1.3 Unbiased Affinity Estimation Method

[0126] In this embodiment, the genome relationship matrix constructed by the unbiased kinship estimation method is expressed by the following expressions (1-4), (1-5), (1-6), and (1-7):

[0127]

[0128] Where ρ ii' X is the true kinship coefficient between individuals i and i'; k and l represent other individuals besides individuals i and i', X ij and X i'j This refers to the number of minor alleles; It is the average minor allele frequency of this SNP. is the genotype variance; j represents the j-th SNP locus; E represents the expected value. Simplifying the above formula, we get:

[0129]

[0130] As can be seen from the above formula, The expected value is independent of the chosen SNP. Therefore, ρ ii' The estimated value can be expressed as:

[0131]

[0132] It can be seen from equations (1-6) and (1-7) that It is valid for any SNP. The expected value is still ρ ii' This means It is ρ ii' An unbiased estimate. Combining the above equations, we can find... This can actually be expressed as a Yang method polynomial. This means that calculating the unbiased estimated UKin matrix only requires the same genomic information as the Yang method. The mathematical relationship between the UKin method and the Yang method can be derived as follows:

[0133]

[0134] Therefore, the UKin method is a traditional method expression. It is a linear combination of constant terms, meaning its computational complexity does not increase significantly. Note that the introduction and derivation of the UKin method here are based on the Yang method. Similarly, the UKin method can also be based on the VanRaden method. Therefore, in the following study, this embodiment uses "Ukin+Yang" to represent the UKin unbiased method based on the Yang method; and "Ukin+VanRaden" to represent the UKin unbiased method based on the VanRaden method.

[0135] 1.1.1.4Weir method

[0136] The Weir method uses the method of moments to estimate the kinship coefficient θ between individuals i and i'. ii' This method constructs a relational matrix based on genealogy for comparison. The method is expressed as follows:

[0137]

[0138] Among them, M ii' M is the proportion of identical alleles carried by individuals i and i' (1≤i, i'≤n). ii' It can be calculated as follows:

[0139]

[0140] Where X ij and X i'j This represents the reference allele doses for individuals i and i' at SNP j. For a population of size n, it is the average pairing of all distinct individual pairs:

[0141]

[0142] In fact, the ratio of the expected value of the numerator to the denominator... That is, the objective parameter (θ) ii' -θ S ) / (1-θ S This method has been shown to achieve lower mean squared error and performs better when the number of SNPs is large.

[0143] 1.1.1.5 Polynomial Kernel Function Method

[0144] A polynomial kernel transforms data points in the original feature space using polynomials, enabling the model to capture nonlinear relationships within the data.

[0145] The genome relation matrix constructed using the polynomial kernel function method is expressed by the following expression (1-12):

[0146]

[0147] In formula (1-12), γ, a, and d are adjustable parameters;

[0148] 'a' is a positive real scalar that can affect the middle and higher order terms and the lower order terms of a polynomial. The middle and higher order terms are those with a middle to high exponent (exponent greater than or equal to the exponent threshold) in the polynomial, and the lower order terms are those with a low exponent (exponent less than the exponent threshold) in the polynomial.

[0149] d is a positive integer;

[0150] x i This represents the genotype vector of the i-th individual;

[0151] x j Represents the genotype vector of the j-th individual;

[0152] K(x i ,x j ) represents x i and x j The kernel function value is used to adjust the scaling and nonlinearity by adjusting the parameter d. The parameter a is used to adjust the generated polynomial, representing the relative influence of higher-order and lower-order terms.

[0153] 1.1.1.6 Cosine Kernel Function Method

[0154] The cosine kernel assesses the similarity between two vectors by measuring the cosine of the angle between them. In calculating the genomic relation matrix, the genotype vectors of two individuals are considered as input vectors. The value of the kernel is defined as the quotient of the dot product of these two vectors and their modulo.

[0155] The genome relation matrix constructed using the cosine kernel function method is expressed by the following expression (1-13):

[0156]

[0157] In formula (1-13), x i ·x j The genotype vector x of the i-th individual i The genotype vector x of the j-th individual j The dot product;

[0158] ||x i ‖ and ||x j || are vectors x i and x j The Euclidean norm;

[0159] K(x i ,x j ) represents x i and x j The kernel function value between.

[0160] Cosine similarity does not involve parameter selection, making it relatively simple and efficient, especially when dealing with large-scale data.

[0161] 1.1.1.7 Gaussian Kernel Function Method

[0162] The Gaussian kernel, also known as the radial basis function kernel, determines its value based on the Euclidean distance (i.e., the Euclidean norm of the difference) between the input vectors, rather than the dot product. This kernel is popular in machine learning models, but it is sensitive to parameter selection and prone to overfitting during training.

[0163] The genome relation matrix constructed using the Gaussian kernel function method is expressed by the following expression (1-14):

[0164]

[0165] In formula (1-14), the positive real scalar γ is the width parameter of the kernel function, which controls the decay rate of the function;

[0166] ||x i -x j || 2 There are two vectors x i and x j The square of the Euclidean distance between them;

[0167] K(x i ,x j ) represents x i and x j The kernel function value represents the genotypic similarity between individuals.

[0168] 1.1.1.8 Hyperbolic Tangent Kernel Function Method

[0169] The hyperbolic tangent kernel (sigmoid kernel) is a commonly used kernel function in support vector machines (SVMs) and is also frequently used as an activation function in deep learning. With proper parameter tuning, this function can exhibit very complex nonlinear relationships. However, it can introduce model instability in certain situations, so careful evaluation and selection of appropriate parameters are necessary before application. Furthermore, in some parameter settings, the hyperbolic tangent kernel is similar to the Gaussian kernel. Therefore, calculating the genome-wide similarity matrix may require significant computational resources.

[0170] The genome relation matrix constructed using the hyperbolic tangent kernel function method is expressed by the following expression (1-15):

[0171]

[0172] In formula (1-15), x i This represents the genotype vector of the i-th individual;

[0173] x j Represents the genotype vector of the j-th individual;

[0174] tanh is the hyperbolic tangent function;

[0175] α is the slope parameter of the kernel function; b is the intercept parameter of the kernel function.

[0176] K(x i ,x j ) represents x i and x j The kernel function value represents the genotypic similarity between individuals, i.e., the element in the matrix.

[0177] 1.1.2 Test Materials

[0178] 1.1.2.1 Simulated Dataset

[0179] The simulation process was based on the genetic background and statistical history of pigs. The existing QMSim software was used for the simulation, and the parameter settings and breeding structure were based on existing research. The entire simulation process was repeated 5 times, mainly including the following three parts: (1) Simulating historical ancestors. Linkage disequilibrium was established from the historical ancestral population, and mutation drift equilibrium was established. There were 200 generations, with 420 animals raised in each generation, including 20 males and 400 females. The number of offspring from each mating was 10, with an equal sex ratio. (2) Simulating the recent population. The last generation of the historical ancestral population was selected as the founding animal. The number of litters was 10, and the number of generations was 10. The female replacement rate was 0.8, and the male replacement rate was 1.0. The selection of animals was based on the breeding values ​​generated by the BLUP (Best linear unbiased prediction) method. A total of 40,421 offspring populations with pedigree, phenotypic and genomic information were obtained. Genomic prediction was carried out using 4,000 animals from the last generation. (3) Genomic parameters. The simulated genome consists of 18 pairs of chromosomes, each pair being 100 cM in length. Two genomic scenarios were simulated: one where the trait is controlled solely by minor polygenes, and the other where the trait is controlled by both major and minor polygenes, with effect values ​​following a normal distribution. Each chromosome contains 25 QTLs (Quantitative Trait Loci) and 1,000 SNPs. All markers are biallelic and randomly distributed across all chromosomes. Three heritability values ​​were set for both scenarios: 0.1, 0.3, and 0.5. For ease of subsequent analysis, these traits are represented by T1, T2, T3 and T4, T5, T6, respectively. Genomic parameters are shown in Table 2-1.

[0180] Table 1-1 Genomic information parameters of the simulated dataset

[0181]

[0182] 1.1.2.2 Large White Pig Dataset

[0183] A total of 3,290 Shanghai Xinong Large White pigs were selected, including 1,836 Canadian and 1,454 French. All animals underwent performance testing at 100 kg body weight. Live weight and backfat thickness were obtained and corrected to age at 100 kg body weight (AGE100) and corrected backfat thickness at 100 kg body weight (BF100). The phenotypic values ​​used were the corrected phenotypes. The Canadian and French strains were bred on two different farms with essentially the same nutritional management levels. The two strains were unrelated. The pedigree of these animals comprised 4,719 Large White pigs, recorded from 2012 to 2020. The mean generation for genome and phenotype was 6.12 generations.

[0184] Genotyping was performed using the GeneSeek GGP Porcine HD chip. This chip was based on Susscrofa 10.2, so it was updated to Susscrofa 11.1 before subsequent analysis. Quality control was performed using PLINK v1.90 software: (1) individual genotype detection rate >90%; (2) marker genotype detection rate >90%; (3) minimum allele frequency >0.05; (4) only autosomal SNP markers were retained. After quality control, 35,172 SNPs and 3,290 individuals were retained. The traits used in the analysis were expressed as AGE100 and BF100.

[0185] 1.1.2.3 Wheat Dataset

[0186] This dataset comprises 599 wheat lines from the Global Wheat Program of the International Centre for Maize and Wheat Improvement, tested in various wheat production environments divided into four basic target environment sets. Genotyping was performed using a diversity array technique, which uses two values ​​to represent their presence or absence. Markers with minor allele frequencies below 0.05 were removed, and missing genotypes were imputed using samples from marginally distributed marker genotypes. After these quality controls, 1,279 markers remained. The dataset is available from CROSSA et al. (2010). Phenotypic values ​​follow a normal distribution, and the four traits in this dataset are represented by W1, W2, W3, and W4, respectively.

[0187] 1.1.2.4 German Holstein cattle

[0188] The German Holstein cattle genome prediction population comprised 5,024 bulls, genotyped using the Inllumina bovine SNP50BeadChip. After quality control, 42,551 SNPs remained. Based on previous research and their representative genetic structures, estimated breeding values ​​for three traits—milk yield (MY), milk fat percentage (MFP), and somatic cell score (SCS)—were used for further analysis. The MFP trait is highly heritable, controlled by one major gene and numerous genes with smaller effects. The MY trait is moderately heritable, also controlled by one major gene and several genes with moderate and small effects. The SCS trait is a low-heritable trait, controlled by numerous genes with smaller effects. Phenotypic values ​​were corrected to follow a normal distribution. Tables 1-2 provide a statistical summary of all datasets in this example.

[0189] Table 1-2 Statistical summary of all datasets

[0190]

[0191]

[0192] 1.1.3 Evaluation Indicators

[0193] In this embodiment, a 5-fold cross-validation procedure with 5 replicates is used for analysis. Through cross-validation, the dataset for each analysis is randomly divided into 5 parts: 80% of the individuals are used as the training set, and 20% are used as the validation set. Predictive accuracy and unbiasedness are used as evaluation metrics for the model. Accuracy is the Pearson correlation coefficient between phenotypic values ​​and genome-estimated breeding values ​​(i.e., predicted values). The Pearson correlation coefficient provides a measure of linear consistency between predicted and actual values ​​when evaluating the model's predictive performance. Unbiasedness is the regression coefficient of phenotypic values ​​to individual genome-estimated breeding values, indicating whether the regression coefficient between predicted and actual values ​​is close to 1.0. The difference between unbiasedness and 1.0 determines whether the predicted values ​​have systematic bias. It is important to note that the simulated dataset contains both phenotypic values ​​and actual breeding values. Therefore, when making predictions on the simulated data, this embodiment uses phenotypic values ​​for prediction and uses actual breeding values ​​and genome-estimated breeding values ​​(i.e., predicted values) to calculate accuracy and unbiasedness.

[0194] 1.2 Results

[0195] 1.2.1 Prediction performance of the Ukin method on simulated datasets

[0196] Figure 2Tables 1-3 below illustrate the performance of the Ukin method (including UKin+Yang and UKin+VanRaden) and other genome relation matrix construction methods in GBLUP predictions. For each box plot, the midline represents the median, the middle "×" symbol represents the mean, and the top and bottom of each box represent the maximum and minimum values. In the simulated dataset, the prediction accuracy for all traits ranged from 0.466 to 0.629. The prediction accuracy for trait T1 ranged from 0.592 to 0.616, for trait T2 from 0.613 to 0.629, for trait T3 from 0.538 to 0.545, for trait T4 from 0.505 to 0.535, for trait T5 from 0.563 to 0.572, and for trait T6 from 0.466 to 0.472. Overall, the Yang method and the UKin+VanRaden method performed best on this dataset, with prediction accuracies of 0.616, 0.629, 0.545, 0.535, 0.572, 0.472 and 0.608, 0.627, 0.545, 0.525, 0.572, 0.471 for the six traits, respectively. The prediction accuracies of the other three methods were relatively similar.

[0197] Table 1-3 GBLUP prediction accuracy using different genome relation matrices in the simulated dataset.

[0198]

[0199]

[0200] Figure 3 Tables 1-4 show a comparison of the unbiasedness of GBLUP predictions using different genomic relation matrices. The unbiasedness of predictions for all traits ranged from 0.953 to 1.598, indicating that the overall GBLUP predictions were overestimated relative to the true values. The unbiasedness of predictions for trait T1 ranged from 1.038 to 1.107, for trait T2 from 1.012 to 1.042, for trait T3 from 0.953 to 0.955, for trait T4 from 1.501 to 1.589, for trait T5 from 1.004 to 1.020, and for trait T6 from 1.070 to 1.087. Except for trait T3, the unbiasedness of all traits was greater than 1.0. The mean unbiased prediction scores for all traits by each method were 1.11, 1.11, 1.13, 1.12, and 1.11, respectively, indicating that the unbiased prediction scores of all methods were very similar.

[0201] Table 1-4 Unbiasedness of GBLUP predictions using different genome relation matrices in the simulated dataset.

[0202]

[0203] 1.2.2 Prediction performance of the UKin method on real datasets

[0204] The real datasets include wheat datasets, Holstein cattle datasets, and Large White pig datasets. Figure 4 , Figure 5 Tables 1-5 show the comparison of prediction accuracy results in GBLUP using the UKin method (including UKin+Yang and UKin+VanRaden) and other genome relation matrix construction methods in these datasets. In the wheat dataset, the prediction accuracy for all traits ranged from 0.384 to 0.513. The prediction accuracy for trait W1 ranged from 0.490 to 0.513, for trait W2 from 0.476 to 0.487, for trait W3 from 0.384 to 0.391, and for trait W4 from 0.450 to 0.466. Overall, the Yang method and the UKin+Yang method performed relatively well, with prediction accuracies of 0.513, 0.479, 0.391, 0.450 and 0.513, 0.478, 0.391, 0.451 for the four traits, respectively. In the Holstein cattle dataset, the prediction accuracy for all traits ranges from 0.738 to 0.816. The prediction accuracy for trait MY ranges from 0.768 to 0.772, for trait MFP from 0.805 to 0.816, and for trait SCS from 0.738 to 0.740. The results for trait MF and trait SCS are very close among the five methods, while for trait MFP, the prediction accuracy of VanRaden, UKin+Yang, and Weir methods is 0.816, and the prediction accuracy of Yang and UKin+VanRaden methods is 0.805, showing a significant difference. In the Large White pig dataset, the prediction accuracy for trait AGE100 ranges from 0.742 to 0.746, and the prediction accuracy for trait BF100 ranges from 0.823 to 0.824. The accuracy results of the methods for constructing the genome relationship matrix for these two traits were almost equal. In contrast, the UKin+VanRaden method had a slightly higher average accuracy of 0.785.

[0205] Figure 6Tables 1-6 show a comparison of the unbiasedness of GBLUP predictions using the UKin method and other genomic relation matrices across all real-world datasets. For the wheat dataset, the unbiasedness of predictions for all traits ranged from 0.954 to 1.322. The prediction accuracy for trait W1 ranged from 1.018 to 1.234, for trait W2 from 1.049 to 1.322, for trait W3 from 0.954 to 1.069, and for trait W4 from 1.002 to 1.157. Overall, except for the VanRaden method, the average unbiasedness of predictions for the other four methods was approximately 1.0, indicating relatively stable prediction performance. For the Holstein cattle dataset, the unbiasedness of predictions for all traits ranged from 0.996 to 1.007, indicating that all methods performed relatively stably on this dataset. In the Large White pig dataset, the unbiased prediction accuracy of all methods ranged from 0.988 to 0.991, and all methods performed relatively stably. For these three datasets, the unbiased prediction accuracy for all traits ranged from 0.954 to 1.322. Except for the VaRaden method, which achieved an unbiased prediction accuracy of 1.322 ± 0.237 for the W2 trait in the wheat dataset, the unbiased prediction accuracy for the other traits and methods was very close to 1.0. This indicates that the UKin method provides stable unbiased prediction results in real-world datasets, comparable to traditional methods.

[0206] Table 1-5 GBLUP prediction accuracy using different genome relation matrices in real datasets

[0207]

[0208] Table 1-6 Unbiasedness of GBLUP predictions using different genome relation matrices in real datasets

[0209]

[0210] Traditional genomic relationship matrices often deviate from the true kinship of individuals, affecting subsequent genetic analyses. This embodiment proposes an unbiased estimation method, the UKin method, which corrects this bias by integrating genetic information from the entire population. The method's ability to improve heritability estimation accuracy is demonstrated. Specifically, the UKin method's predictive performance in the GBLUP model was evaluated on simulated and multiple real-world datasets. The results show varying degrees of variation in both prediction accuracy and unbiasedness. Overall, the UKin method exhibits similar predictive accuracy to traditional methods on both simulated and real-world datasets. These results demonstrate that the UKin method addresses the biased estimation problem of traditional methods while maintaining the accuracy and unbiasedness of genomic predictions, showing valuable application potential in predicting complex traits. This superior performance may stem from two reasons. First, it more accurately reflects the genetic relationships between individuals. Unbiased estimation ensures that the estimated kinship values ​​do not deviate from the true values ​​due to systematic bias, which is particularly important for predicting breeding values ​​in genomic prediction. Secondly, the UKin method effectively handles negative values ​​in the genome relation matrix, which can be a cause of bias in traditional methods. By optimizing the handling of these negative values, the UKin method can more accurately capture genetic variations, thereby ensuring the accuracy and unbiasedness of predictions.

[0211] Constructing the genome relation matrix is ​​a crucial step in genome selection. This example explores the UKin-based genome relation matrix construction method and its application and performance in genome prediction. Through studies on simulated and real datasets, the prediction performance of the UKin method and traditional methods in GBLUP is systematically compared. The results show that the UKin method not only solves the biased estimation problem of traditional methods, but also achieves similar prediction accuracy and unbiasedness to traditional methods on both simulated and three real datasets. The UKin method corrects the bias problem in traditional genome relation matrices, providing a new approach and means to improve the accuracy and unbiasedness of genome prediction, and offering strong support for precision genetic improvement and breeding decisions.

[0212] 1.2.3 Prediction performance of kernel function method on simulated datasets

[0213] In the analysis of the simulated dataset, different kernel functions were used to construct the genome relation matrix to evaluate the prediction performance of GBLUP. A total of four kernel function methods were analyzed: polynomial, cosine, Gaussian (Rbf), and hyperbolic tangent kernel.

[0214] From Table 1-7 and Figure 7From the results of the accuracy, the prediction accuracy of all traits ranges from 0.460 to 0.637. The prediction accuracy of trait T1 ranges from 0.587 to 0.618, the prediction accuracy of trait T2 ranges from 0.607 to 0.637, the prediction accuracy of trait T3 ranges from 0.528 to 0.554, the prediction accuracy of trait T4 ranges from 0.503 to 0.537, the prediction accuracy of trait T5 ranges from 0.558 to 0.584, and the prediction accuracy of trait T6 ranges from 0.460 to 0.489. Generally speaking, the accuracy of the GBLUP-Polynomial and GBLUP-Rbf methods is slightly better than other methods. Specifically, the average prediction accuracy of the GBLUP-Polynomial method is 0.570, followed by the GBLUP-Rbf method with an average prediction accuracy of 0.569. In contrast, the average prediction accuracies of the GBLUP-Cosine and GBLUP-Sigmoid methods are 0.550 and 0.541 respectively. This indicates that in the simulated dataset, the methods based on polynomial kernel and Gaussian kernel are relatively more effective in constructing genomic relationship matrices.

[0215] Table 1-7 GBLUP prediction accuracies for constructing genomic relationship matrices using different kernel functions in the simulated dataset

[0216]

[0217] The results of the prediction unbiasedness are as Figure 8 shown in Table 1-8. The prediction unbiasedness of all traits ranges from 0.945 to 1.659. The prediction unbiasedness of trait T1 ranges from 1.099 to 1.114, the prediction unbiasedness of trait T2 ranges from 1.024 to 1.038, the prediction unbiasedness of trait T3 ranges from 0.945 to 0.960, the prediction unbiasedness of trait T4 ranges from 1.641 to 1.659, the prediction unbiasedness of trait T5 ranges from 1.004 to 1.020, and the prediction unbiasedness of trait T6 ranges from 1.064 to 1.077. Sorted by the average prediction unbiasedness from low to high, it is GBLUP-Sigmoid < GBLUP-Rbf < GBLUP-Polynomial < GBLUP-Cosine. The average prediction unbiasedness of the GBLUP-Sigmoid method is 1.14, which is the lowest among the four methods. The unbiasedness of GBLUP-Cosine is 1.69, which is the highest among the four methods. This indicates that the prediction results of the Sigmoid kernel function are closer to the true values compared to other methods, while the prediction results of the cosine kernel function are relatively the farthest from the true values.

[0218] Table 1-8 GBLUP prediction unbiasedness for constructing genomic relationship matrices using different kernel functions in the simulated dataset

[0219]

[0220] 1.2.4 Prediction performance of kernel function methods on real datasets

[0221] Figure 9 , Figure 10 Tables 1-9 show the prediction accuracy of GBLUP in constructing genomic relation matrices using kernel functions on real datasets. In the wheat dataset, the prediction accuracy for all traits ranges from 0.378 to 0.548. Specifically, the prediction accuracy for trait W1 ranges from 0.474 to 0.548, for trait W2 from 0.481 to 0.491, for trait W3 from 0.378 to 0.401, and for trait W4 from 0.451 to 0.497. In this dataset, the GBLUP-Polynomial method performs best overall, followed by the GBLUP-Cosine method, while the GBLUP-Sigmoid method performs worst. This indicates that in the wheat dataset, polynomial and Gaussian kernels provide better genomic relation matrices, thus enhancing the accuracy of the GBLUP prediction model. In the Holstein cattle dataset, the prediction accuracy for all traits ranges from 0.737 to 0.816. The prediction accuracy for trait MY ranges from 0.768 to 0.775, for trait MFP from 0.802 to 0.816, and for trait SCS from 0.737 to 0.740. For the MFP trait, the GBLUP-Cosine method performs slightly better than other methods. Overall, the differences between the methods are not significant, indicating that these kernel function methods have relatively balanced prediction accuracy in the Holstein cattle dataset. In the Large White pig dataset, the prediction accuracy for trait AGE100 ranges from 0.742 to 0.782, and for trait BF100 from 0.823 to 0.843. In this dataset, the GBLUP-Cosine and GBLUP-Polynomial methods provide relatively higher prediction accuracy, especially for the BF100 trait.

[0222] Table 1-9 GBLUP prediction accuracy using different genome relation matrices in real datasets

[0223]

[0224] Figure 11Table 1-10 shows the unbiasedness of GBLUP predictions when constructing genome relation matrices using kernel functions on real datasets. In the wheat dataset, the unbiasedness of predictions for all traits ranges from 0.929 to 1.077. The unbiasedness ranges for trait W1 from 1.010 to 1.032, for trait W2 from 1.052 to 1.077, for trait W3 from 0.929 to 0.972, and for trait W4 from 1.015 to 1.024. Overall, the average unbiasedness of predictions for all methods is close to or equal to 1.0, showing good predictive performance. For the Holstein cattle dataset, the unbiasedness of predictions for all traits ranges from 0.997 to 1.156. The unbiasedness range for trait MY is 0.997–1.066, for trait MFP it is 0.998–1.156, and for trait SCS it is 0.999–1.067. The GBLUP-Cosine method shows unbiasedness very close to 1.0 for all three traits, indicating that its predictions are very close to the true values. The GBLUP-Polynomial and GBLUP-Rbf methods show slightly higher unbiasedness for some traits. This result highlights the excellent unbiasedness performance of the GBLUP-Cosine method in the Holstein cattle dataset, especially for the MFP trait. For the Large White pig dataset, the unbiasedness range for all traits is 0.989–1.027. The unbiasedness range for trait AGE100 is 0.992–1.015, and for trait BF100 it is 0.991–1.027. In this dataset, all methods performed well overall. Considering the unbiased prediction results of the wheat, Holstein cattle, and Large White pig datasets, the GBLUP-Cosine method demonstrated excellent unbiased prediction across all datasets, particularly the Holstein cattle dataset. Similar to the results of the UKin method, except for the MFP trait, almost all methods achieved an unbiasedness of around 1.0 across all traits.

[0225] Table 1-10 Unbiasedness of GBLUP predictions using different genome relation matrices in real datasets

[0226]

[0227] While traditional methods for constructing genomic relation matrices have been widely used in genomic selection, their potential for further improving selection performance has become relatively limited. This embodiment introduces the concept of kernel functions from machine learning to investigate methods for constructing genomic relation matrices between individuals. Machine learning kernel functions are used to construct genomic relation matrices between individuals, replacing the traditional relation matrices in GBLUP, and this approach is studied for the first time on different types of datasets. The information required for constructing genomic relation matrices using kernel functions is the same as traditional methods, requiring only the genotype data of the population. By applying kernel functions in GBLUP, the model can better simulate the complex genetic relationships between genes and the nonlinear combinations of gene effects, thereby improving the explanatory power of genetic variation. Furthermore, the kernel function method offers a variety of kernel function options, allowing the model to flexibly select the most suitable kernel function based on the characteristics of the dataset and the research objectives.

[0228] Kernel functions map the input space to a high-dimensional feature space through nonlinear transformations, making input data easier to separate and thus improving the accuracy and stability of regression and classification problems. In this embodiment, multinomial kernels, cosine kernels, Gaussian kernels, and hyperbolic tangent kernels are used, each with its specific application scenarios and advantages. This embodiment applies the GBLUP method using kernel functions to different datasets. The results show that multinomial and Gaussian kernels outperform cosine and hyperbolic tangent kernels on both simulated and real datasets overall, except for the Holstein cow dataset, where the Gaussian kernel shows lower accuracy. These results also highlight the applicability of different kernel function methods to different datasets and traits, and the importance of selecting the most suitable kernel function for improving the performance of the GBLUP prediction model. Comparing the results of all kernel functions with the Yang method in traditional genome relation matrix construction, we found that the prediction accuracy of the polynomial kernel improved by 0.4%, 1.2%, 2.1%, 0.0%, 2.0%, 3.7%, 8.1%, 0.5%, 4.2%, 7.7%, 0.3%, -1.7%, 0.0%, 5.2%, and 2.3% for all traits, respectively, while the prediction accuracy of the Gaussian kernel improved by 0.5%, 0.9%, 1.6%, 0.3%, 1.7%, 2.7%, 7.5%, 0.9%, 4.2%, 6.8%, 0.1%, 0.5%, 0.2%, 4.2%, and 1.8% for all traits. Among these traits, both kernel functions showed advantages in improving genome prediction accuracy, particularly in the wheat and Large White pig datasets. This means that polynomial kernels and Gaussian kernels have unique advantages in constructing genome relation matrices. These advantages can solve some scientific problems in genetic assessment and prediction in theory and practice, especially providing new perspectives and methods for analyzing and predicting the genetic mechanisms of complex traits.

[0229] II. Utilizing machine learning strategies to capture more non-additive effects in genome prediction

[0230] With the rapid development of bioinformatics and statistical genetics, machine learning has become an important tool for analyzing complex biological data. Especially in genome selection, the application of machine learning algorithms provides new avenues for improving the efficiency and accuracy of genetic breeding. Compared with traditional genome selection methods, kernel ridge regression with polynomial (K) P The KRR (Ridge Regression) machine learning strategy combines the regularization advantages of ridge regression (RR) with the ability of kernel functions to handle nonlinear data, making it a powerful tool for processing high-dimensional genetic data. By introducing kernel functions, K... P RR can map the original feature space to a higher-dimensional space, thereby capturing complex nonlinear relationships in genotype data, which is particularly important for understanding the multigene control mechanisms of hereditary traits. In this embodiment, K will be introduced. P This paper discusses the theoretical basis, implementation methods, and application of RR to genomic data from diverse genetic backgrounds, including additive, dominant, and epistatic effects. By comparing it with traditional genomic selection methods such as GBLUP and BayesB, it demonstrates the advantages of K... P The prediction performance and efficiency of RR on multiple datasets are presented. This embodiment aims to provide a more comprehensive perspective to demonstrate the application potential of machine learning methods in genomic selection and genetic breeding, with the goal of providing new tools and ideas for plant and animal breeding.

[0231] 2.1 Materials and Methods

[0232] 2.1.1 Test Methods

[0233] 2.1.1.1 Machine Learning Strategy: Multinomial Kernel Ridge Regression

[0234] Ridge regression is a variant of linear regression that addresses the problem of collinearity among independent variables and avoids overfitting by adding a regularization term to the loss function. The loss function L(w) of ridge regression is defined as:

[0235] L(w) = ||Xw - y|| 2 +λ‖w‖ 2 (2-1)

[0236] Where X is the data matrix; w is the weight vector; y is the target variable; and λ is the regularization parameter.

[0237] The multinomial kernel trick is a method that uses a kernel function to map data to a high-dimensional feature space, allowing linear algorithms to implicitly handle nonlinear relationships in the original space. Kernel ridge regression addresses nonlinear problems by combining ridge regression and the kernel trick. The goal of kernel ridge regression is to minimize the following loss function:

[0238] L(w)=‖φ(X)wy‖ 2 +λ‖w‖ 2 (2-2)

[0239] Here, φ(X) represents the result of mapping the original data X to a high-dimensional space through a kernel function. The optimal solution to this equation can be obtained by solving the following system of linear equations:

[0240] (K+λI)α=y (2-3)

[0241] Where K is the kernel matrix, and the elements K ij =K(x) i ,x j α represents the kernel function value between training samples i and j; I is the identity matrix; α is the coefficient vector. For a new input x, its predicted value can be calculated using the following formula:

[0242]

[0243] Where, α i These are the coefficients obtained from the above solution, x i These are training samples. The machine learning strategy used in this embodiment is kernel ridge regression with polynomial (K) P RR), is based on the above process and incorporates the polynomial kernel function. The results were obtained through integration. The optimal parameters of the polynomial kernel were found using a grid search method, and the Ki algorithm was implemented using the Python package sklearn. P The genomic prediction process for RR.

[0244] 2.1.1.2 Machine Learning Strategies: Multinomial Kernel Support Vector Machine Regression

[0245] Support Vector Machine (SVM) regression is an extension of SVM for regression problems. Its main idea is to minimize the error and maximize the margins, where some error is tolerable. Assume that when f(x) and y... i The difference between them is greater than ò If the loss is calculated only at a specific time, then the loss function can be expressed as:

[0246]

[0247] Based on the loss function described above, the goal of support vector machine regression is to find a function f(x) = w T φ(x)+b, such that the following objective function is minimized:

[0248]

[0249] The following conditions must be met:

[0250]

[0251] Here, φ(x) represents the function that maps the input x to a high-dimensional space, w and b are the model's weights and biases, respectively (w and b are solved iteratively), C is the regularization parameter (C is solved using a grid search method), and ξ i and It is a slack variable (ξ) i and The solution method is the difference method, used to handle unavoidable errors. Using the Lagrange multiplier method, the original optimization problem can be transformed into its dual problem, making the solution more efficient, especially when using kernel functions to process high-dimensional data. After transformation and derivation, the prediction model of support vector machine regression can be expressed as:

[0252]

[0253] Where K is the kernel function; and α i These are the positive weights of each observation, obtained during the Lagrange multiplier transformation process. Based on the above process, and by integrating the polynomial kernel function with support vector machine regression, S is obtained. P VR strategy. Similarly, this method also uses a grid search approach to find the optimal parameters, and implements S using the Python package sklearn. P The process of genome prediction for VR.

[0254] 2.1.1.3 BayesB Method

[0255] The Bayes-B method assumes that most markers have no effect, only a few markers have an effect, and that the effect follows a gamma or exponential distribution (MEUWISSEN et al., 2009). Its statistical model can be expressed as:

[0256]

[0257] Where y is the phenotypic value vector; m j It is the genotype of the j-th locus; α j It is the effect value of the j-th label, which follows the law of effect. e is a random residual that follows a set... αj The prior depends on the labeling effect method and the prior probability π, which determines the proportion of labels with effect values ​​out of the total number of labels, and is determined empirically. The genome prediction process is implemented using the R package "BGLR".

[0258] 2.1.1.4 Optimal Linear Unbiased Prediction of the Genome

[0259] Same as 1.1.1.1, no further details will be provided.

[0260] 2.1.1.5 Optimal Linear Unbiased Prediction of Genomic Dominance

[0261] The GDBLUP method treats additive and dominant genetic effects as random effects. The statistical model is as follows:

[0262] y = 1μ + Z a g a +D d g d +e (2-10)

[0263] Among them, D d It is g d The correlation matrix; Defined as dominant inheritance, y represents the variance of dominant inheritance; μ represents the marker-based dominant genome relation matrix; g represents the observed value of the target trait; μ represents the population mean; g represents the variance of dominant inheritance. a It is a sports value; Z a It is a sports species value g a The correlation matrix. The construction of the D matrix and the implementation of GDBLUP are both achieved using the R package "sommer".

[0264] The method for constructing the D matrix is ​​briefly described below:

[0265] Among them W ik =1-||x ik ||,x ik This represents the genotype of the k loci of the i-th individual.

[0266] Z a The construction method is briefly described as follows: Za represents the correlation matrix with breeding values, and its dimension is m×n, where m represents m phenotypes and n represents n individuals.

[0267] 2.1.1.6 Optimal Linear Unbiased Prediction of Genomic Epistasis

[0268] The GEBLUP method treats additive and epistatic genetic effects as random effects. The statistical model is as follows:

[0269] y = 1μ + Z a g a+D ep g ep +e (2-11)

[0270] Where D ep It is g ep The correlation matrix; Defined as epistatic effect, E represents the epistatic variance, E is the epistatic genomic relationship matrix based on SNP markers, y is the observed value of the target trait; μ is the mean; g is the epistatic variance. a It is a sports value; Z a It is a sports species value g a The correlation matrix. The construction of the E matrix and the implementation of GEBLUP are both achieved using the R package "sommer".

[0271] 2.1.2 Test Materials

[0272] 2.1.2.1 Additive Effects Dataset

[0273] The additive effects dataset uses the wheat dataset and the Holstein cattle dataset, which have been introduced in section 1.1.2 Experimental Materials and will not be repeated here.

[0274] 2.1.2.2 Dominant Effect Dataset

[0275] (1) PIC Core Pig Dataset

[0276] This dataset comes from the core pig line of the Pig Improvement Company (PIC) and includes 3,534 individuals and 5 traits. The quality control criteria for genotype data were: removal of loci without location information, loci located on sex chromosomes, and duplicate loci; Hardy-Weinberg equilibrium <10. -4 SNPs with a minimum allele frequency >0.05 and a locus detection rate >0.95 were retained; individuals without the phenotype were deleted. A total of 38,891 SNPs and 2,314 individuals were retained. Based on the heritability of these five traits, we selected the trait T2(h)... 2 =0.16), T3(h 2 =0.38) and T4(h 2 The analysis was performed with a variance of 0.58. The dominance variances for these traits were 2%, 7%, and 1%, respectively. This dataset has been used in several studies related to dominance effects. The phenotypes used in the analysis have been corrected for either fixed effects or weighted progeny averages.

[0277] (2) Scottish Blackface Sheep Dataset

[0278] This dataset consists of 752 Scottish Blackface sheep. Paternal and maternal lineages were recorded at birth for all animals, along with complete pedigrees from the establishment of the flock in 1988. The three live body weight traits considered in this example are body weight at 16 weeks (BW16), body weight at 20 weeks (BW20), and body weight at 24 weeks (BW24). Phenotypic values ​​were corrected for sex, year, age, and group. After quality control, 37,243 SNPs were retained for further analysis. The dominance variances of the three traits accounted for 38%, 6%, and 30% of the total phenotypic variance, respectively.

[0279] 2.1.2.3 Episode Effects Dataset

[0280] (1) Back-crossing simulation dataset

[0281] This dataset contains 600 individuals, 1 trait, and 121 molecular markers evenly distributed on one chromosome. Among them, 9 molecular markers overlap with the main effect QTL. Of the possible marker pairs, 13 pairs exhibited interaction effects. The fixed effects only included the population mean, with a phenotypic variance of 106, of which 90 came from all marker effects (main effects and epistatic effects), and the remainder from the covariance term.

[0282] (2) Rice dataset

[0283] This dataset includes four traits: yield, number of pages per plant, number of grains per ear, and grain weight per kilogram (KWG). After quality control, 278 hybrid combinations and 1,619 genotypes remained. Phenotypic values ​​were adjusted for year. The epistatic effect variances for the four traits were estimated to be 95%, 69%, 52%, and 23%, respectively.

[0284] Table 2-1 below lists the statistical summary of the above dataset.

[0285] Table 2-1 Statistical summary of the dominant effect dataset and the epistatic effect dataset

[0286]

[0287] 2.1.3 Evaluation Indicators

[0288] Similar to section 1.1.2, a 5-fold cross-validation procedure with 5 repetitions was used for analysis. The dataset for each analysis was randomly divided into 5 parts: 80% of the individuals were used as the training set, and 20% as the validation set. The performance of each method was evaluated using prediction accuracy and computation time. The method for calculating accuracy has been described in section 1.1.3, the evaluation metrics. All computations were performed on the same research server (Ubuntu 18.04.6, 72 CPUs @ 2.30GHz), ensuring the same number of CPU threads and memory.

[0289] 2.2 Results

[0290] 2.2.1 Prediction results in an additive genetic background

[0291] Figure 11 , Figure 12 Table 2-2 presents the prediction accuracy results for the two datasets under the additive genetic background, applying K... P RR, S P Four methods, VR, GBLUP, and BayesB, were used to predict the breeding values ​​of seven traits based on genomic estimation. The midline of the box plot represents the median, the top and bottom of each box represent the maximum and minimum values, and the "×" symbol represents the mean. In the wheat dataset, the prediction accuracy for all traits ranged from 0.354 to 0.544. The prediction accuracy for trait W1 ranged from 0.502 to 0.544, for trait W2 from 0.425 to 0.491, for trait W3 from 0.354 to 0.389, and for trait W4 from 0.440 to 0.483. The prediction accuracy for each trait varied within the ranges of 0.042, 0.066, 0.035, and 0.043, indicating relatively stable prediction results. For all traits, K... P RR showed the highest prediction accuracy. In the Holstein cattle dataset, the prediction accuracy for trait MY ranged from 0.756 to 0.791, for trait MFP from 0.759 to 0.870, and for trait SCS from 0.717 to 0.742. The mean prediction accuracies for each trait were 0.773, 0.814, and 0.736, respectively. The results from different methods varied little across the three traits, except that BayesB showed significantly higher accuracy for trait MFP compared to other methods. Overall, BayesB achieved the highest prediction accuracy for traits MY and MFP. P RR and GBLUP showed similar and stable performance for traits MY and MKG. However, for trait SCS, the prediction accuracy of all methods was almost consistent.

[0292] Table 2-2 Prediction accuracy of each trait in the additive genetic background dataset

[0293]

[0294] 2.2.2 Prediction Results under the Background of Superordinate Effects

[0295] Figure 13 Tables 2 and 3 present the prediction accuracy results for the two datasets under the epistatic effect genetic background. Ki was used. P RR, S P The VR, BayesB, GBLUP, and GEBLUP methods were compared and analyzed for all five traits. It can be seen that the prediction accuracy of the simulated traits ranged from 0.530 to 0.768, and the prediction accuracy of each method was ranked K. P RR>GEBLUP>BayesB>GBLUP>S P VR. In the rice dataset, the prediction accuracy for the trait Yield ranged from 0.372 to 0.417, for Tiller from 0.418 to 0.509, for Grain from 0.505 to 0.626, and for KGW from 0.759 to 0.830. For simulated traits, Yield, and Tiller, K... P RR performed best. For the Grain and KWG traits, BayesB performed best. Overall, S always... P VR performed the worst.

[0296] Besides S P Aside from VR and GBLUP, the prediction accuracy of other methods did not differ significantly for the same trait. Comparing the results of GBLUP and GBLUP, it was found that adding epistatic effects to the prediction model changed the prediction accuracy for different traits. The prediction accuracy for simulated traits, trait Yield, and trait Grain improved by 24.75%, 6.64%, and 2.24%, respectively, while the prediction accuracy for trait Tiller decreased by 1.83%, and the prediction accuracy for trait KWG remained unchanged. Meanwhile, K... P The prediction accuracy of RR was improved by 25.42%, 12.01%, 2.42%, 3.07%, and 0.23% compared to GBLUP, respectively. (Compared to K) P The results of RR and GEBLUP showed that K P The predictive accuracy of RR was improved by 0.54%, 5.04%, 4.33%, 0.81%, and -0.19% respectively compared to the latter. These comparisons indicate that K P RR does not require prior calculation of the epistatic effect matrix, and can directly incorporate epistatic factors into the model while achieving good predictive performance.

[0297] Table 2-3 Predictive accuracy of epistatic effect genetic background traits

[0298]

[0299] 2.2.3 Prediction results under the background of dominant effects

[0300] Figure 14 Tables 2 and 4 present the prediction accuracy results for the two datasets under the dominant effect genetic background. Ki was used. P RR, S P Five methods, VR, BayesB, GBLUP, and GDBLUP, were used to predict the genome sequence of all six traits. In the PIC core pig dataset, the prediction accuracy for trait T2 ranged from 0.465 to 0.497, and the prediction accuracy of each method was ranked K. P RR>GBLUP>GDBLUP>BayesB>S P VR. The prediction accuracy of trait T3 ranged from 0.320 to 0.323, and the prediction accuracy of each method was ranked as follows: S P VR>BayesB=GBLUP>GDBLUP>K P RR. The prediction accuracy for trait T4 ranged from 0.425 to 0.450, and the prediction accuracy of each method was ranked as follows: BayesB = GBLUP > GDBLUP > K P RR>S P VR. Combining GDBLUP with GBLUP, K... P RR and GBLUP, K P Comparing RR and GDBLUP revealed that the differences between them were all within ±0.01, indicating that the prediction accuracy of different methods on this dataset differed only slightly.

[0301] In the Scottish Blackface sheep dataset, the prediction accuracy of the BW16 trait ranged from 0.189 to 0.261, and the prediction accuracy of each method was ranked as follows: S P VR>K P RR > BayesB > GDBLUP > GBLUP. The prediction accuracy for trait BW20 ranged from 0.297 to 0.360, and the prediction accuracy of each method was ranked as follows: S P VR>K P RR > BayesB > GBLUP > GDBLUP. The prediction accuracy for trait BW24 ranged from 0.133 to 0.212, and the prediction accuracy of each method was ranked as K. P RR>BayesB>S P VR > GDBLUP > GBLUP. Overall, K... P RR and S PVR showed the highest prediction accuracy on this dataset, while GBLUP showed the worst. Comparing the results of GBLUP and GBLUP, it was found that adding the dominance effect matrix to the model improved the prediction accuracy of the three traits by 5.78%, 0.60%, and 30.99%, respectively. (Compared to K...) P The results of RR and GDBLUP showed that K was present in all three traits. P The accuracy of RR was improved by 29.10%, 16.67%, and 21.56% compared to GDBLUP, respectively.

[0302] Table 2-4 Predictive accuracy of dominant effect genetic background traits

[0303]

[0304] 2.2.4 Comparison of running times

[0305] Figure 15 Runtime for all traits and methods on the same device was recorded. The x-axis represents different traits, and the y-axis represents the time consumed by different methods, in seconds, with logarithms to base 10 calculated. Runtime for traits in the wheat dataset ranged from 1.01 to 318.37, for the Holstein cattle dataset from 154.19 to 33071.24, for simulated traits in the simulated dataset from 0.50 to 1081.00, for the rice dataset from 0.40 to 1099.30, for the PIC core pig dataset from 149.30 to 227856.00, and for the Scottish Blackface sheep dataset from 8.00 to 31985.00. It can be observed that the runtime of methods increases with the number of individuals or markers. Among all traits, BayesB took the longest time, K... P RR takes the shortest time. In some cases, such as rice traits, K... P RR was observed to be approximately 2500 times faster than BayesB. In both the simulated dataset and the rice dataset, S... P VR is faster than GBLUP, and in the pig and sheep datasets, S P VR has a slower computing speed.

[0306] 2.3 Discussion

[0307] 2.3.1K P Predictive performance of RR on additive datasets

[0308] The performance of genomic selection methods depends heavily on the genetic structure of traits and population structure. While no single method consistently outperforms others across all datasets, the results of this embodiment reveal the unique advantages of different approaches. In this embodiment, the machine learning strategy K was first tested. P Performance of RR on datasets controlled solely by additive effects, including the wheat and Holstein cattle datasets. Compare with K. P RR, S P The predicted breeding values ​​for genome estimation in VR, BayesB, and GBLUP datasets revealed that K... P The RR method performed best across all four traits in the wheat dataset, and showed similar prediction accuracy to GBLUP in the dairy cow dataset. Previous studies using GBLUP for genome prediction on the wheat dataset yielded accuracies of 0.512, 0.483, 0.401, and 0.463, while the results in this example are 0.507, 0.487, 0.385, and 0.461, showing a high degree of consistency. In previous studies using the GBLUP method to estimate the prediction accuracy of three traits in the German Holstein cattle dataset, the results were 0.774, 0.816, and 0.738, while the results in this example are 0.772, 0.816, and 0.740, also showing a high degree of consistency.

[0309] 2.3.2K P Predictive performance of RR on non-additive datasets

[0310] In the early stages of genomic selection, many methods only consider additive effects in genomic data. In reality, non-additive effects play a central role in hybrid vigor, polymorphism, and genomic selection. Some studies have shown that models integrating non-additive effects achieve better predictive performance compared to genomic selection models that only consider additive effects. For example, in existing studies, the predictive performance of two-step and non-additive effect models was evaluated in seven traits of a cassava population, and the results showed that the latter improved predictive performance by an average of 10%. Therefore, considering non-additive effects in genomic selection models is essential. Thus, in this embodiment, the epistatic effect matrix and the dominance epistatic matrix are integrated into the GBLUP model, respectively, resulting in the GEBLUP and GDBLUP models. The results show that in datasets containing non-additive effects, the predictive accuracy of both the GEBLUP and GDBLUP models is significantly improved compared to GBLUP, which also supports the above viewpoint.

[0311] K P RR, S PThe VR, BayesB, GBLUP, and GEBLUP methods were first applied to datasets containing epistatic effects. In the simulated dataset, 13 possible label pairs had interaction effects, and there was a high proportion of additive and epistatic variance. The method with the best prediction accuracy was K... P RR, followed by GEBLUP, BayesB, GBLUP, and S P VR. Nine labels in this dataset overlap with the main effect's QTL, which may explain why BayesB outperforms GBLUP and S. P The reason for VR is that the BayesB method has been observed to perform well in predicting traits controlled by large-effect genes or QTLs, which also explains the excellent performance of BayesB in the MFP of the Holstein cattle dataset. For the rice dataset, the epistatic effect proportions of its four traits were 95%, 69%, 52%, and 23%, respectively. The results of this dataset indicate that when there is more epistatic variance in a trait, the machine learning strategy K... P RR has better predictive power.

[0312] After that, K P RR, S P The VR, BayesB, GBLUP, and GDBLUP methods were applied to datasets containing dominant effects. For the PIC core pig dataset, the proportions of dominant variance for the three traits are relatively small, at 2%, 7%, and 1%, respectively. This low proportion of dominant variance may affect K... P The result of RR, i.e., K P RR does not outperform other methods on this dataset, which is why GDBLUP does not show a significant improvement over GBLUP. In the Scottish Blackface sheep dataset, the dominance variance of the three traits accounts for 38%, 6%, and 30% of the total phenotypic variance, respectively. P RR consistently outperformed other methods. Furthermore, it can be observed that as the proportion of non-additive effects increases, K... P The RR prediction performance also showed an increasing trend. From the above results and discussion, it can be seen that the machine learning strategy K... P RR can effectively capture non-additive effects in genomic data and enable them to contribute to genomic prediction models.

[0313] 2.3.3 Model computational efficiency

[0314] In this embodiment, the runtime of all methods and traits on the same device was statistically analyzed. The results show that BayesB typically requires the longest computation time because it uses a specific Markov chain Monte Carlo algorithm for solving. In contrast, machine learning strategy K... PRR has the shortest computation time across all analyzed datasets, sometimes even thousands of times faster than other methods. This significant speed advantage may be related to K... P The characteristics of the RR algorithm itself are relevant. First, kernel ridge regression can use closed-form solutions to estimate model parameters, meaning the model can directly obtain the optimal solution without iterative optimization. Second, the regularization term in kernel ridge regression makes the model more inclined to use low-dimensional feature subspaces, thereby reducing computational complexity. In the results, we also note that S... P VR runtime is always longer than K. P RR takes longer. Although both are based on kernel function training data, they solve optimization problems in different ways. P VR (Reverse Vector Machine) is based on the principle of Support Vector Machines (SVMs). It minimizes prediction error by constructing an optimal hyperplane. The process of finding this hyperplane involves solving a convex quadratic programming problem, that is, optimizing the loss function while satisfying a series of inequality constraints. Solving convex quadratic programming problems is usually quite complex, so when dealing with large-scale datasets, S... P VR runtimes can be very long.

[0315] Ignoring non-additive genetic effects can bias multiple genetic analysis processes. In this embodiment, we investigated kernel ridge regression and multinomial kernel functions based on machine learning algorithms and proposed a machine learning strategy K. P RR. Using this strategy in conjunction with traditional genomic selection methods, genomic prediction studies were conducted on datasets with different genetic backgrounds. The results showed that (1) in an additive genetic background, the machine learning strategy K... P RR has similar genomic estimation breeding value prediction accuracy to traditional genomic selection methods GBLUP and BayesB. When the trait is controlled by a large-effect gene, the BayesB method has better prediction accuracy. (2) In the context of non-additive inheritance, adding non-additive effects to the GBLUP model significantly improved prediction accuracy. Compared with traditional genomic selection methods GBLUP and BayesB, the machine learning strategy K P RR has higher or similar prediction accuracy, which illustrates the advantages of machine learning in transforming high-dimensional feature spaces and capturing non-additive effects. (3) Machine learning strategy K P The RR method has the shortest overall runtime, while the BayesB method has the longest.

[0316] It should be stated that the above-described specific embodiments are merely preferred embodiments of the present invention and the technical principles employed. Those skilled in the art should understand that various modifications, equivalent substitutions, and variations can be made to the present invention. However, such variations, as long as they do not depart from the spirit of the present invention, should be within the scope of protection of the present invention. Furthermore, some terminology used in this specification and claims is not limiting, but merely for ease of description.

Claims

1. A genomic selection method for machine learning of additive and non-additive effects, characterized in that, Including the following steps: S1, construct the genome relationship matrix using the UKin unbiased kinship estimation method; S2, adding non-additive effects to the linear model and constructing a non-additive genomic relationship matrix; S3 uses machine learning to capture the additive and non-additive effects of traits, and uses a machine model that has completed learning to predict individual traits and output the prediction results; In step S2, the statistical model that adds non-additive effects to the linear model is expressed by the following formula (2-10): ; In formula (2-10), observed value representing the target trait; represents the population average; represents a sports breed value; represents a sports breed value of the association matrix; denotes the dominance effect, is defined as the normal distribution of dominance inheritance; is the variance of dominance inheritance; is the marker-based dominance genomic relationship matrix; is the incidence matrix of the graph represents the remaining residual; In step S1, the genomic relationship matrix constructed using the UKin unbiased kinship estimation method is expressed by the following expression: (1-4) (1-5) (1-6) As can be seen from equations (1-4) - (1-6), The expected value of the score is independent of the chosen SNP, therefore, The estimate of the score can be expressed as: ;(1-7) in Individual and The true kinship coefficient between them; yes The estimated value; express The actual calculated value; and Including individuals and Other individuals besides and This refers to the number of minor alleles; It is the average minor allele frequency of this SNP. It is the variance of genotypes; Indicates the first One SNP; Indicates the expected value; This represents a group of n individuals; It is the number of SNPs; This represents the j-th SNP locus of the k-th individual; This represents the j-th SNP locus of the l-th individual; express Expected value; In step S3, genome prediction is performed using multinomial kernel ridge regression, specifically including the following steps: A1, construct a multinomial kernel ridge regression prediction model as expressed in the following formula (2-1): ; In equation (2-4), represents the predicted value of the genome of the individual input; Indicates the first The coefficients corresponding to each training sample; denotes the number of training samples; a kernel matrix representing kernel function values of the training samples; represents a kernel function value between the i-th training sample and​​ A2, the individual input to the polynomial kernel ridge regression prediction model, the model prediction output pair genomic prediction result.

2. The machine learned genomic selection method for additive and non-additive effects of claim 1, wherein, The loss function of the polynomial kernel ridge regression prediction model is expressed by the following formula (2-2): ; In Equation (2-2), denotes a loss function; representing the original data the result of mapping to a high-dimensional space by a kernel function; is the weight vector; Target variable; is a regularization parameter; The goal of kernel ridge regression is to minimize the loss function expressed by formula (2-2).

3. The machine learned genomic selection method for additive and non-additive effects of claim 2, wherein, The optimal solution of the loss function expressed by formula (2-2) is obtained by solving the linear equation expressed by formula (2-3): ; In equation (2-3), is the identity matrix; is a coefficient vector.

4. The machine learned genomic selection method for additive and non-additive effects of claim 1, wherein, Genome prediction using multinomial kernel support vector machine regression includes the following steps: B1, construct a support vector machine regression prediction model as expressed by the following formula (2-8): ;(2-8) In equation (2-8), represents the predicted value of the genome of the individual input; represents the kernel function value between the i-th training sample and the input x.​ represents the positive weight value of the individual prediction value; represents a positive weight in the training set; represents the number of training samples; Paranoia; B2, the individual input to the support vector machine regression prediction model, the model prediction output pair genomic prediction result.

5. The machine learned genomic selection method for additive and non-additive effects of claim 4, wherein, The loss function of the support vector machine regression prediction model is expressed by the following formula (2-5): ;(2-5) In equation (2-5), denotes a loss function; representative individual corresponding genomic truth; is a threshold value.

6. The machine learned genomic selection method for additive and non-additive effects of claim 5, wherein, The optimal solution of the loss function expressed by formula (2-5) is constrained by the following formulas (2-6) and (2-7): ;(2-6) The following conditions must be met: ;(2-7) In formulas (2-6), (2-7), is a slack variable; , For the first Relaxation variables for each training sample; is a regularization parameter; is a weight; To be paranoid; To train the sample mapping to a high-dimensional space; corresponding to the i-th training sample; and corresponding to the i-th training sample; and is the transpose of the weight ​