Multi-omics genome selection method and system based on combination of PCA dimension reduction and SVR

By combining PCA dimensionality reduction and dynamic multi-source SVR loss function in genome selection prediction, the problems of insufficient accuracy of genome selection prediction and poor robustness of model in the prior art are solved, and higher prediction accuracy and model stability are achieved.

CN120148641APending Publication Date: 2025-06-13FOSHAN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510210886.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing machine learning-based genome selection prediction methods are insufficiently accurate in the prediction of complex traits, and the model is poorly robust, making it difficult to adapt to diverse sample data.

Method used

Using a multiomic genome selection method based on the combination of PCA dimensionality reduction and SVR, we can adapt to the diversity of sample data and improve prediction accuracy by constructing dynamic multi-source SVR loss function and PCA dimensionality reduction processing.

Benefits of technology

Effectively reduce data dimensions, improve computational efficiency, avoid dimension explosions and gradient disappearances, and improve the accuracy of multiomics data genome selection and the biological rationality and predictive robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148641A_ABST
    Figure CN120148641A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and provides a multi-omics genome selection method and system based on PCA dimension reduction and SVR, and the method comprises the steps: constructing a dynamic multi-source SVR loss function based on multi-omics data and environmental factors associated with the multi-omics data, and constructing an SVR model; the SVR model is trained; pCA dimension reduction processing is carried out on the target object multi-omics data, and the target character phenotype is predicted according to the target object multi-omics data after PCA dimension reduction processing. According to the method, the dynamic multi-source SVR loss function is constructed, the SVR model is constructed according to the SVR loss function, multiple factors possibly influencing the accuracy of character phenotype prediction of the target object are considered, and the method adapts to the diversity of sample data, has better adaptability and has higher biological rationality and prediction robustness in complex character prediction compared with traditional SVR.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a multi-omics genome selection method and system based on the combination of PCA dimensionality reduction and SVR. Background Art

[0002] For decades, animal breeding technologies have greatly improved the production performance of animals, and genomic selection plays an increasingly important role in modern animal breeding. Existing technologies such as GBLUP and Bayesian methods mainly focus on additive effects and ignore the epistatic effects between genes and non-linear prediction models. In recent years, with the rapid development of machine learning theory and practical application technologies, machine learning has also been successfully used for the analysis and interpretation of various omics data such as genomic data.

[0003] The Chinese invention patent with the application number "CN202411179964.7" relates to the technology of realizing genomic selection prediction based on machine learning. By combining the advantages of polynomial kernel function and kernel ridge regression methods, it can greatly improve the genomic prediction effect. Another Chinese invention patent with the application number "CN202311218507.X" discloses a whole-genome prediction method, which solves the technical problems of inaccurate prediction and unstable model in the prior art during whole-genome prediction. For example, the Chinese invention patent with the application number "201910060405.7" discloses a whole-genome prediction method and device, which can infer the genotype of hybrids based on the genotypes of parents and predict their phenotypic data. Thus, for genomic prediction based on machine learning, there are still many technical problems to be solved and many technical solutions to be proposed in its practical application. Summary of the Invention

[0004] Based on this, in order to improve the accuracy of gene selection prediction for multi-omics data, the present invention provides a multi-omics genome selection method and system based on the combination of PCA dimensionality reduction and SVR. By constructing a dynamic multi-source SVR loss function based on multi-omics data and environmental factors associated with the multi-omics data, it can adapt to the diversity of sample data, has better adaptability, and has stronger biological rationality and prediction robustness in complex trait prediction compared with traditional SVR; by performing PCA dimensionality reduction processing on the multi-omics data of the target object, it can effectively reduce the data dimension, improve the calculation efficiency, avoid dimensional explosion and gradient disappearance, and at the same time improve the accuracy of multi-omics data genome selection. The specific technical solutions are as follows:

[0005] A multi-omics genome selection method based on the combination of PCA dimensionality reduction and SVR, comprising the following steps:

[0006] Construct a dynamic multi-source SVR loss function based on multi-omics data and environmental factors associated with the multi-omics data, and construct an SVR model according to the SVR loss function;

[0007] Obtain a training sample set, perform PCA dimensionality reduction processing on the training sample set, and train the SVR model according to the training sample set after PCA dimensionality reduction processing;

[0008] Obtain multi-omics data of a target object, perform PCA dimensionality reduction processing on the multi-omics data of the target object, and predict the target trait phenotype according to the multi-omics data of the target object after PCA dimensionality reduction processing and the trained SVR model;

[0009] Among them, the multi-omics data of the target object includes target genomic data and target transcriptomic data, and the training sample set includes genomic sample data, transcriptomic sample data, and trait phenotype sample data.

[0010] The multi-omics genome selection method based on the combination of PCA dimensionality reduction and SVR constructs a dynamic multi-source SVR loss function, and constructs an SVR model according to the SVR loss function, taking into account multiple factors that may affect the prediction accuracy of the target object's trait phenotype, adapting to the diversity of sample data, having better adaptability, and having stronger biological rationality and prediction robustness in complex trait prediction than traditional SVR. In addition, before predicting the target trait phenotype using the multi-omics data of the target object and the trained SVR model, performing PCA dimensionality reduction processing on the multi-omics data of the target object can effectively reduce the data dimension, improve the calculation efficiency, avoid dimensional explosion and gradient disappearance, and at the same time improve the accuracy of multi-omics data genome selection.

[0011] Preferably, before performing PCA dimensionality reduction processing on the multi-omics data of the target object, preprocess the multi-omics data of the target object.

[0012] Preferably, the SVR loss function

[0013]

[0014] Among them, L = λ g ||w g || 2 + λ t ‖‖w t || 2 + λ ef ‖|w ef || 2 represents the regularization task loss, k represents the main task loss weight, (1 - k) represents the regularization task loss weight, N represents the number of trait phenotype sample data, yi represents the true value of the phenotypic sample data of the i-th trait, y' i represents the predicted value of the phenotypic sample data of the i-th trait, ε i = ε 0 × e -γE , E represents an intermediate variable, wg, wt, and wef represent the eigenvectors of genomic data, transcriptomic data, and environmental factors respectively, λg, λt, λef, α, β, γ represent hyperparameters, |||| represents the norm, C represents the regularization parameter, ξ i respectively represent the slack variables of the i-th target genomic data, n represents the number of target genomic data, e represents the natural constant, ε 0 represents the preset standard insensitive error loss value, ε i represents the insensitive error loss value corresponding to the i-th environmental factor, E i represents the actual value of the i-th environmental factor, E i ' represents the standard value of the i-th environmental factor, m represents the number of environmental factors, || represents the absolute value function, W g 、W t 、 respectively represent the genomic data weight matrix, the transcriptomic data weight matrix, and the transpose of the genomic data weight matrix, and Q represents the preset regulatory matrix.

[0015] Preferably, the specific method for performing PCA dimensionality reduction on the preprocessed multi-omics data of the target object includes the following steps:

[0016] Calculate the covariance matrix for the preprocessed multi-omics data of the target object;

[0017] Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors;

[0018] Select the principal components and retain a preset proportion of the variance;

[0019] Perform PCA transformation on the reference population data and perform PCA transformation on the validation population data to project it into the principal component space of the reference population data.

[0020] A multi-omics genomic selection system based on the combination of PCA dimensionality reduction and SVR, used to implement the multi-omics genomic selection method based on the combination of PCA dimensionality reduction and SVR, includes:

[0021] A construction module, used to construct a dynamic multi-source SVR loss function based on multi-omics data and environmental factors associated with the multi-omics data, and construct an SVR model according to the SVR loss function;

[0022] A training module, configured to obtain a training sample set, perform PCA dimensionality reduction processing on the training sample set, and train the SVR model according to the training sample set after PCA dimensionality reduction processing to obtain a trained SVR model;

[0023] A dimensionality reduction module, configured to perform PCA dimensionality reduction processing on multi-omics data of a target object;

[0024] Wherein, the SVR model is used to predict a target trait phenotype according to the multi-omics data of the target object after PCA dimensionality reduction processing, the multi-omics data of the target object includes target genomic data and target transcriptomic data, and the training sample set includes genomic sample data, transcriptomic sample data, and trait phenotype sample data.

[0025] Preferably, the multi-omics genomic selection system based on the combination of PCA dimensionality reduction and SVR further includes:

[0026] A preprocessing module, configured to preprocess the multi-omics data of the target object before performing PCA dimensionality reduction processing on the multi-omics data of the target object.

[0027] Preferably, the SVR loss function

[0028]

[0029] Where L = λ g ||w g || 2 + λ t ||w t || 2 + λ ef ||w ef || 2 represents the regularization task loss, k represents the main task loss weight, (1 - k) represents the regularization task loss weight, N represents the number of trait phenotype sample data, y i represents the true value of the i-th trait phenotype sample data, y' i represents the predicted value of the i-th trait phenotype sample data, ε i = ε 0 × e -γE , E represents an intermediate variable, wg, wt, and wef respectively represent the feature vectors of genomic data, transcriptomic data, and environmental factors, λg, λt, λef, α, β, and γ represent hyperparameters, |||| represents the norm, C represents the regularization parameter, ξ i respectively represent the relaxation variables of the i-th target genomic data, n represents the number of target genomic data, e represents the natural constant, ε 0 represents a preset standard insensitive error loss value, εi Denote the insensitive error loss value corresponding to the \(i\)-th environmental factor as \(E\). i Denote the actual value of the \(i\)-th environmental factor as \(E\). i Denote the standard value of the \(i\)-th environmental factor as \(m\), where \(m\) represents the number of environmental factors, and \(||\) represents the absolute value function. \(W\). g and \(W\). t and respectively represent the genomic data weight matrix, the transcriptomic data weight matrix, and the transpose of the genomic data weight matrix, and \(Q\) represents a preset regulatory matrix.

[0030] Preferably, the dimensionality reduction module includes:

[0031] A calculation unit for calculating the covariance matrix of the preprocessed multi-omics data of the target object;

[0032] A decomposition unit for performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors;

[0033] A selection unit for selecting principal components and retaining a preset proportion of variance;

[0034] A conversion unit for performing PCA conversion on the reference population data and performing PCA conversion on the validation population data to project it into the principal component space of the reference population data.

[0035] A multi-omics genomic selection device based on the combination of PCA dimensionality reduction and SVR, which includes:

[0036] A controller;

[0037] A memory storing executable instructions;

[0038] Wherein, the executable instructions can run on the controller and implement the multi-omics genomic selection method based on the combination of PCA dimensionality reduction and SVR.

[0039] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the multi-omics genomic selection method based on the combination of PCA dimensionality reduction and SVR as described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The present invention can be further understood from the following description in conjunction with the drawings. The components in the drawings are not necessarily drawn to scale, but the emphasis is placed on showing the principles of the embodiments. In different views, the same reference numerals designate corresponding parts.

[0041] Figure 1 is the overall flow schematic diagram of the multi-omics genomic selection method in an embodiment of the present invention;

[0042] Figure 2 It is a schematic flowchart of the specific method for PCA dimensionality reduction processing in an embodiment of the present invention;

[0043] Figure 3 It is a schematic flowchart of the method for obtaining the main task loss weight in an embodiment of the present invention;

[0044] Figure 4 It is a schematic diagram of the overall structure of the multi-omics genomic selection system in an embodiment of the present invention. Detailed implementation manners

[0045] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with its embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0046] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are only for the purpose of illustration and do not represent the only implementation manner.

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used herein in the description of the present invention are only for the purpose of describing specific implementation manners and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0048] The "first" and "second" in the present invention do not represent specific quantities and orders, but are only used for name distinction.

[0049] Multi-omics data refers to the data integrated at different omics levels (such as genomics, transcriptomics, proteomics, etc.), which comprehensively reveals the biological characteristics of organisms and their dynamic changes under different conditions. It is an interdisciplinary research method that helps to analyze complex relationships such as genes and the environment, genes and phenotypes, metabolism and physiology, and is widely used in fields such as animal science, precision medicine, plant science, and microbiology.

[0050] Genotype and environment are two important factors that affect the morphological manifestation of an individual organism. Genotype refers to the genetic makeup carried by an individual, while the environment refers to the external environmental conditions in which the individual is located. The interaction between the two jointly determines the phenotypic characteristics of the individual organism. In biology, phenotype refers to the morphological characteristics presented by an individual organism in the external environment. Genotype, on the other hand, refers to the genetic composition carried by an individual. The genotype of an individual determines its genetic information, while the phenotype is the result of the combined action of genotype and environment. That is to say, multi-omics data such as genomics, epigenomics, transcriptomics, proteomics, metabolomics, and microbiomics jointly affect the phenotypic traits of the living system.

[0051] In existing machine learning-based genomic selection prediction methods, they often rely on strong assumptions and mainly perform linear regression analysis. This method limits the accuracy of its prediction because it fails to fully capture the deep interactions between genotype and phenotype and the impact of environmental factors on phenotypic traits. Therefore, the existing machine learning-based genomic selection prediction methods have poor robustness and may be somewhat lacking in accuracy in predicting complex traits, and there is room for improvement.

[0052] To solve the above technical problems, as Figure 1 shown, the present invention provides a multi-omics genomic selection method based on PCA dimensionality reduction and SVR combination, which includes the following steps:

[0053] S1, construct a dynamic multi-source SVR loss function based on multi-omics data and environmental factors associated with the multi-omics data, and construct an SVR model according to the SVR loss function.

[0054] Specifically, the environmental factors associated with the multi-omics data include but are not limited to environmental temperature and humidity, light, atmospheric pressure, oxygen concentration, and PM2.5 concentration during data measurement.

[0055] For the SVR model, the SVR model can be constructed using the Sklearn (V1.3.2) software package of Python. After constructing the SVR model, preferably, hyperparameter optimization is performed on the SVR model. Specifically, combined with five-fold cross-validation at one time, Grid Search is used for hyperparameter optimization to find the optimal parameter combination of the model by systematically traversing various combinations of hyperparameters.

[0056] S2, obtain a training sample set, perform PCA dimensionality reduction processing on the training sample set, and train the SVR model according to the training sample set after PCA dimensionality reduction processing.

[0057] The training sample set includes genomic sample data, transcriptomic sample data, and trait phenotype sample data. Preferably, before training the SVR model according to the training sample set, preprocessing and PCA dimensionality reduction processing are first performed. The preprocessing is mainly used to check and process missing values and outliers, and perform quality control to ensure the quality of genomic and transcriptomic data. The PCA dimensionality reduction processing includes, but is not limited to, covariance matrix calculation, eigenvalue decomposition, principal component selection, and PCA transformation of reference population data and validation population data.

[0058] After completing the PCA dimensionality reduction processing of the training sample set, the principal components of the genomic data and the transcriptomic data are concatenated together. For the training sample set, the reference population and the validation population can be divided according to a preset ratio to better complete the deep learning work of the SVR model.

[0059] S3. Obtain the multi-omics data of the target object, perform PCA dimensionality reduction processing on the multi-omics data of the target object, and predict the target trait phenotype according to the multi-omics data of the target object after PCA dimensionality reduction processing and the trained SVR model. Among them, the multi-omics data of the target object includes target genomic data and target transcriptomic data.

[0060] Specifically, PCA dimensionality reduction is implemented by importing the PCA class from the sklearn.decomposition module using Python3. Before performing PCA dimensionality reduction processing on the multi-omics data of the target object, the multi-omics data of the target object can also be preprocessed first. The preprocessing is used to check and process missing values and outliers, and perform quality control to ensure the quality of genomic and transcriptomic data.

[0061] Preferably, as Figure 2 shown, in step S3, the specific method for performing PCA dimensionality reduction processing on the preprocessed multi-omics data of the target object includes the following steps:

[0062] S31. Calculate the covariance matrix for the preprocessed multi-omics data of the target object. Specifically, the covariance matrix is automatically calculated inside the PCA class.

[0063] S32. Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors. Specifically, the PCA class automatically performs eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors.

[0064] S33. Select the principal components and retain a preset proportion of the variance. Specifically, set n_components = 0.95 to retain 95% of the variance.

[0065] S34. Perform PCA transformation on the reference population data and perform PCA transformation on the validation population data to project it into the principal component space of the reference population data. Use the fit_transform method to perform PCA transformation on the reference population data, and use pca.transform to project the validation population data into the principal component space of the reference population data.

[0066] After that, use the concat method of pandas to splice the principal components of the genomic data and the transcriptomic data in the multi-omics data of the target object together for the target genomic data and the target transcriptomic data.

[0067] Preferably, the multi-omics genomic selection method based on PCA dimensionality reduction and SVR combination further includes index evaluation. The index evaluation specifically includes: using the Pearson correlation coefficient between the estimated breeding value (GEBV) and the trait phenotypic value (P) as the evaluation index, and selecting the average value of the accuracy of 5 times of five-fold cross-validation as the final index to measure the prediction accuracy.

[0068] That is, the model accuracy cov(GEBV,P) represents the covariance between the estimated breeding value (GEBV) and the trait phenotypic value (P), and var(GEBV) and var(P) respectively represent the variances of the estimated breeding value (GEBV) and the trait phenotypic value (P).

[0069] In summary, the multi-omics genomic selection method based on PCA dimensionality reduction and SVR combination constructs a dynamic multi-source SVR loss function, and constructs an SVR model according to the SVR loss function, taking into account multiple factors that may affect the prediction accuracy of the trait phenotype of the target object, adapting to the diversity of sample data, having better adaptability, and having stronger biological rationality and prediction robustness in predicting complex traits than traditional SVR.

[0070] In the prior art, both the GBLUP (Genomic best linear unbiased prediction) and Bayes (Bayes A, Bayes B, Bayes C, Bayes Cπ, Bayesian methods based on prior information) methods simply assume that only additive effects can be stably inherited while ignoring the epistatic effects between genes. Moreover, both the GBLUP and Bayes methods estimate the GEBV (Genomic estimated breeding value) by constructing a linear model while ignoring the non-linear prediction model. Machine learning improves the prediction performance of the model by learning "experience" from data, without the need to establish assumptions and can construct non-linear prediction models. However, due to the progress of high-throughput genotyping technology and the continuous reduction of the detection cost of single nucleotide polymorphisms (SNPs), a large amount of genomic data has been generated. These data are usually high-dimensional, and the application of multi-omics further reduces the computational efficiency of current genomic selection methods and overloads the available computational resources, resulting in overfitting of model parameters in machine learning models, leading to dimensional explosion and gradient disappearance. At the same time, the information of hundreds of thousands of potentially irrelevant markers in the multi-omics genomic prediction model may have a negative impact on the prediction accuracy of target traits.

[0071] Before predicting the target trait phenotype using the multi-omics genomic selection method based on PCA dimensionality reduction and SVR combination of the present invention with the multi-omics data of the target object and the trained SVR model, the multi-omics data of the target object is first subjected to PCA dimensionality reduction processing, which can effectively reduce the data dimension, improve the computational efficiency, avoid dimensional explosion and gradient disappearance, and at the same time improve the accuracy of genomic selection of multi-omics data.

[0072] As a preferred technical solution, the SVR loss function

[0073]

[0074] where L = λ g ||w g || 2 + λ t ||w t || 2 + λ ef ||w ef || 2 represents the regularization task loss, k represents the main task loss weight, (1 - k) represents the regularization task loss weight, N represents the number of trait phenotype sample data, y iRepresents the true value of the phenotypic sample data of the i-th trait, y' i Represents the predicted value of the phenotypic sample data of the i-th trait, ε i = ε 0 × e -γE , E represents an intermediate variable, wg, wt, and wef respectively represent the feature vectors of genomic data, transcriptomic data, and environmental factors, λg, λt, λef, α, β, γ represent hyperparameters, |||| represents the norm, C represents the regularization parameter, which can be understood as a penalty factor, ξ i respectively represent the slack variables of the i-th target genomic data, n represents the number of target genomic data, e represents the natural constant, ε 0 Represents the preset standard insensitive error loss value, ε i Represents the insensitive error loss value corresponding to the i-th environmental factor, E i Represents the actual value of the i-th environmental factor, E i ' represents the standard value of the i-th environmental factor, m represents the number of environmental factors, || represents the absolute value function, W g 、W t 、 respectively represent the genomic data weight matrix, the transcriptomic data weight matrix, and the transpose of the genomic data weight matrix, and Q represents the preset regulatory matrix.

[0075] Specifically, Q represents the preset eQTL (expression quantitative trait locus) regulatory matrix. The main task loss weight, the weights of genomic data, transcriptomic data, and environmental factors, the regularization parameter, the slack variable of the i-th target genomic data, the preset standard insensitive error loss value, and the standard value of the i-th environmental factor can all be set by technicians according to experience.

[0076] L = λ g ‖‖ w g ‖‖ 2 + λ t ‖‖ w t ‖‖ 2 + λ ef ‖| w ef || 2 Represents the regularization task loss, which is a multi-source regularization term that combines the weights of genomic data, transcriptomic data, and environmental factors, applies different regularizations to the weight vectors of different sample data, and thus controls the complexity of the SVR model to prevent overfitting. Through the formula Integrating the actual values of multiple different environmental factor values with the corresponding standard values, taking into account the influence of environmental factors on the prediction accuracy of trait phenotypes, making the SVR model more adaptable, thus allowing a larger error, measuring environmental factors such as temperature and humidity during multi-omics data determination, expanding the tolerance of the insensitive error loss ε in a harsh environment. At the same time, through the formula (ε i -ε 0 ) 2 it can also control its adjustment amplitude to avoid over-relaxation.

[0077] Through the formula cross-omics regulation constraints can be achieved, forcing the genomic data weight matrix and the transcriptomic data weight matrix to conform to the known regulatory relationships. The role of setting the γ hyperparameter is to characterize the adjustment efficiency of the dynamic changes of environmental factors on the model's fault tolerance, while β is used to quantify the correction intensity of the known regulatory network on the SVR model.

[0078] Generally speaking, the SVR loss function organically integrates the adjustment of the dynamic insensitive error loss ε, multi-source data regularization, cross-omics constraints, and the main task loss, realizing adaptive modeling of multi-dimensional biological sample data, and having stronger biological rationality and prediction robustness in complex trait prediction than traditional SVR.

[0079] Specifically, the value range of the main task loss weight is preferably 0.6 - 0.9. As a preferred technical solution, as Figure 3 shown, the method for obtaining the main task loss weight includes the following steps:

[0080] The first step is to obtain the complexity of the genomic sample data the complexity of the transcriptomic sample data and the complexity of the trait phenotype sample data

[0081] The second step is to calculate the main task loss weight k = cpx 1 ×k 1 +cpx 2 ×k 2 +cpx 3 ×k 3 .

[0082] Among them, n1 represents the number of genomic sample data, x 1i represents the value of the i-th genomic sample data, u 1 represents the average value of the genomic sample data, n2 represents the number of transcriptomic sample data, x 2i represents the value of the i-th transcriptomic sample data, u2 represents the average value of transcriptome sample data, n3 represents the number of trait phenotype sample data, and x 3i represents the value of the i-th trait phenotype sample data, and u 3 represents the average value of the i-th trait phenotype sample data, and k 1 、k 2 、k 3 respectively represent the weight coefficients of genomic sample data, transcriptome sample data, and trait phenotype sample data.

[0083] Through the formula k = cpx 1 ×k 1 +cpx 2 ×k 2 +cpx 3 ×k 3 , it can dynamically obtain the main task loss weight according to the complexity of genomic sample data, transcriptome sample data, and trait phenotype sample data, so as to adapt to different changes in individual identities, improving flexibility and the scope of application.

[0084] The present invention also provides a multi-omics genomic selection system based on the combination of PCA dimensionality reduction and SVR, which is used to implement the multi-omics genomic selection method based on the combination of PCA dimensionality reduction and SVR, as Figure 4 shown, and it includes a construction module, a training module, and a dimensionality reduction module.

[0085] The construction module is used to construct a dynamic multi-source SVR loss function based on multi-omics data and environmental factors associated with the multi-omics data, and construct an SVR model according to the SVR loss function; the training module is used to obtain a training sample set, perform PCA dimensionality reduction processing on the training sample set, and train the SVR model according to the training sample set after PCA dimensionality reduction processing to obtain a trained SVR model; the dimensionality reduction module is used to perform PCA dimensionality reduction processing on the target object multi-omics data.

[0086] The multi-omics genomic selection system further includes a data acquisition module for obtaining the target object multi-omics data of the training sample set.

[0087] Among them, the SVR model is used to predict the target trait phenotype according to the target object multi-omics data after PCA dimensionality reduction processing. The target object multi-omics data includes target genomic data and target transcriptome data, and the training sample set includes genomic sample data, transcriptome sample data, and trait phenotype sample data.

[0088] Specifically, the dimensionality reduction module includes a calculation unit, a decomposition unit, a selection unit, and a conversion unit.

[0089] The calculation unit is used to calculate the covariance matrix of the multi-omics data of the target object after preprocessing; the decomposition unit is used to perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors; the selection unit is used to select the principal components and retain a preset proportion of the variance; the conversion unit is used to perform PCA conversion on the reference population data and perform PCA conversion on the validation population data to project it into the principal component space of the reference population data.

[0090] The multi-omics genome selection system based on PCA dimensionality reduction and SVR combination further includes a preprocessing module for preprocessing the multi-omics data of the target object before performing PCA dimensionality reduction processing on the multi-omics data of the target object.

[0091] Preprocessing is mainly used to check and process missing values and outliers, and perform quality control to ensure the quality of genome and transcriptome data. PCA dimensionality reduction processing includes, but is not limited to, covariance matrix calculation, eigenvalue decomposition, principal component selection, and PCA conversion of reference population data and validation population data.

[0092] After completing the PCA dimensionality reduction processing of the training sample set, the principal components of the genome data and the transcriptome data are spliced together. For the training sample set, the reference population and the validation population can be divided according to a preset proportion to better complete the deep learning work of the SVR model.

[0093] The multi-omics genome selection system based on PCA dimensionality reduction and SVR combination constructs a dynamic multi-source SVR loss function, and constructs an SVR model according to the SVR loss function, taking into account multiple factors that may affect the prediction accuracy of the target object's trait phenotype, adapting to the diversity of sample data, having better adaptability, and having stronger biological rationality and prediction robustness in complex trait prediction compared to traditional SVR.

[0094] In addition, before predicting the target trait phenotype using the multi-omics data of the target object and the trained SVR model, the multi-omics data of the target object is first subjected to PCA dimensionality reduction processing, which can effectively reduce the data dimension, improve the calculation efficiency, avoid dimensional explosion and gradient disappearance, and at the same time improve the accuracy of multi-omics data genome selection.

[0095] As a preferred technical solution, the SVR loss function

[0096]

[0097] where L = λ g ||w g || 2 + λ t ||w t || 2 + λef |‖w ef ‖‖ 2 represents the regularization task loss, k represents the main task loss weight, (1 - k) represents the regularization task loss weight, N represents the number of trait phenotype sample data, and y i represents the true value of the i-th trait phenotype sample data, and y' i represents the predicted value of the i-th trait phenotype sample data, ε i = ε 0 ×e -γE , E represents an intermediate variable, wg, wt, and wef respectively represent the feature vectors of genomic data, transcriptomic data, and environmental factors, λg, λt, λef, α, β, and γ represent hyperparameters, ‖‖‖‖ represents the norm, C represents the regularization parameter, and ξ i respectively represent the slack variables of the i-th target genomic data, n represents the number of target genomic data, e represents the natural constant, and ε 0 represents the preset standard insensitive error loss value, and ε i represents the insensitive error loss value corresponding to the i-th environmental factor, and E i represents the actual value of the i-th environmental factor, and E i ' represents the standard value of the i-th environmental factor, m represents the number of environmental factors, || represents the absolute value function, and W g 、W t 、 respectively represent the genomic data weight matrix, the transcriptomic data weight matrix, and the transpose of the genomic data weight matrix, and Q represents the preset regulation matrix.

[0098] In summary, the SVR loss function organically integrates the dynamic insensitive error loss ε adjustment, multi-source data regularization, cross-omics constraints, and the main task loss, realizes the adaptive modeling of multi-dimensional biological sample data, and has stronger biological rationality and prediction robustness in complex trait prediction than traditional SVR.

[0099] The present invention also provides a multi-omics genomic selection device based on PCA dimensionality reduction and SVR combination, which includes: a controller; a memory storing executable instructions; wherein, the executable instructions can run on the controller and implement the multi-omics genomic selection method based on PCA dimensionality reduction and SVR combination described above.

[0100] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the multi-omics genomic selection method based on PCA dimensionality reduction and SVR combination as described above.

[0101] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0102] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A multi-omics genome selection method based on PCA dimensionality reduction combined with SVR, characterized in that: The multi-omics genome selection method based on PCA dimensionality reduction combined with SVR comprises the following steps: constructing a dynamic multi-source SVR loss function based on multi-omics data and environmental factors associated with the multi-omics data, and constructing an SVR model according to the SVR loss function; Acquire a training sample set, perform PCA dimensionality reduction processing on the training sample set, and train the SVR model according to the training sample set after the PCA dimensionality reduction processing; Acquire multi-omics data of the target object, perform PCA dimensionality reduction processing on the multi-omics data of the target object, and predict the target trait phenotype according to the multi-omics data of the target object after the PCA dimensionality reduction processing and the trained SVR model; The target object multi-omics data includes target genome data and target transcriptome data, and the training sample set includes genome sample data, transcriptome sample data and trait phenotype sample data.

2. A multi-omics genome selection method based on PCA dimensionality reduction combined with SVR as claimed in claim 1, characterized in that: Before performing PCA dimensionality reduction processing on the target object multi-omics data, the target object multi-omics data is first preprocessed.

3. A multi-omics genome selection method based on PCA dimensionality reduction combined with SVR as claimed in claim 2, characterized in that: The SVR loss function Where L = λ g ||w g || 2 +λ t ||w t || 2 +λ ef ||w ef || 2 represents the regularized task loss, k represents the main task loss weight, (1-k) represents the regularized task loss weight, N represents the number of trait phenotype sample data, y i Represents the true value of the phenotypic sample data of the i-th trait, y' i represents the predicted value of the phenotype sample data of the i-th trait, ε i =ε0×e -γE , E represents the intermediate variable, w g 、w t 、w ef Represent the feature vectors of genomic data, transcriptomic data and environmental factors, respectively, g , t , ef , α, β, γ represent hyperparameters, |||| represents the norm, C represents the regularization parameter, ξ i They represent the slack variables of the i-th target genome data, n represents the number of target genome data, e represents a natural constant, ε0 represents the preset standard insensitive error loss value, and ε i represents the insensitive error loss value corresponding to the i-th environmental factor, E i represents the actual value of the i-th environmental factor, E i ' represents the standard value of the i-th environmental factor, m represents the number of environmental factors, || represents the absolute value function, W g , W t , They represent the genome data weight matrix, the transcriptome data weight matrix and the transpose of the genome data weight matrix respectively, and Q represents the preset regulatory matrix.

4. A multi-omics genome selection method based on PCA dimensionality reduction combined with SVR as claimed in claim 3, characterized in that: The specific method of performing PCA dimensionality reduction processing on the pre-processed multi-omics data of the target object comprises the following steps: Calculating the covariance matrix of the preprocessed multi-omics data of the target object; Performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues ​​and corresponding eigenvectors; Select principal components, retaining a preset proportion of variance; The reference group data were PCA transformed, and the validation group data were PCA transformed to project into the principal component space of the reference group data.

5. A multi-omics genome selection system based on PCA dimensionality reduction combined with SVR, used to implement the multi-omics genome selection method based on PCA dimensionality reduction combined with SVR as described in any one of claims 1 to 4, characterized in that: The multi-omics genome selection system based on PCA dimensionality reduction combined with SVR includes: A construction module, for constructing a dynamic multi-source SVR loss function based on multi-omics data and environmental factors associated with the multi-omics data, and constructing an SVR model according to the SVR loss function; A training module, used for acquiring a training sample set, performing PCA dimensionality reduction processing on the training sample set, and training the SVR model according to the training sample set after the PCA dimensionality reduction processing to obtain a trained SVR model; Dimensionality reduction module, used to perform PCA dimensionality reduction on multi-omics data of target objects; Among them, the SVR model is used to predict the target trait phenotype based on the target object multi-omics data after PCA dimensionality reduction processing, the target object multi-omics data includes target genome data and target transcriptome data, and the training sample set includes genome sample data, transcriptome sample data and trait phenotype sample data.

6. A multi-omics genome selection system based on PCA dimensionality reduction combined with SVR as claimed in claim 5, characterized in that: The multi-omics genome selection system based on PCA dimensionality reduction combined with SVR also includes: The preprocessing module is used to preprocess the target object multi-omics data before performing PCA dimensionality reduction processing on the target object multi-omics data.

7. A multi-omics genome selection system based on PCA dimensionality reduction combined with SVR as claimed in claim 6, characterized in that: The SVR loss function Where L = λ g ||w g || 2 +λ t ||w t || 2 +λ ef ||w ef || 2 represents the regularized task loss, k represents the main task loss weight, (1-k) represents the regularized task loss weight, N represents the number of trait phenotype sample data, y i Represents the true value of the phenotypic sample data of the i-th trait, y' i represents the predicted value of the phenotype sample data of the i-th trait, ε i =ε0×e -γE , E represents the intermediate variable, w g 、w t 、w ef Represent the feature vectors of genomic data, transcriptomic data and environmental factors, respectively, g , t , ef , α, β, γ represent hyperparameters, |||| represents the norm, C represents the regularization parameter, ξ i They represent the slack variables of the i-th target genome data, n represents the number of target genome data, e represents a natural constant, ε0 represents the preset standard insensitive error loss value, and ε i represents the insensitive error loss value corresponding to the i-th environmental factor, E i represents the actual value of the i-th environmental factor, E i ' represents the standard value of the i-th environmental factor, m represents the number of environmental factors, || represents the absolute value function, W g , W t , They represent the genome data weight matrix, the transcriptome data weight matrix and the transpose of the genome data weight matrix respectively, and Q represents the preset regulatory matrix.

8. A multi-omics genome selection system based on PCA dimensionality reduction combined with SVR as claimed in claim 7, characterized in that: The dimension reduction module comprises: A calculation unit, used for calculating the covariance matrix of the preprocessed multi-omics data of the target object; A decomposition unit, used for performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues ​​and corresponding eigenvectors; Selection unit, used to select principal components, retaining a preset proportion of variance; The transformation unit is used to perform PCA transformation on the reference group data and perform PCA transformation on the validation group data to project them into the principal component space of the reference group data.

9. A multi-omics genome selection device based on PCA dimensionality reduction combined with SVR, characterized in that: The multi-omics genome selection device based on PCA dimensionality reduction combined with SVR includes: Controller; A memory storing executable instructions; The executable instructions can be run on the controller and implement the multi-omics genome selection method based on the combination of PCA dimensionality reduction and SVR as described in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multi-omics genome selection method based on the combination of PCA dimension reduction and SVR as described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Whole genome predicting method based on random forest model and device

    CN109727642A

  • Whole genome prediction method based on deep learning

    CN116959585A

  • Genome selection method based on kernel function and machine learning

    CN119108013A