A method and apparatus for predicting genome-wide breeding values based on a virtual expression matrix driven by genetic variation.
By constructing a virtual expression matrix driven by genetic variation and utilizing an elastic network regression model and kernel matrix, the problem of insufficient virtual expression feature extraction in existing breeding prediction models is solved, achieving efficient and low-cost prediction of complex traits.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-17
AI Technical Summary
Existing breeding prediction models lack technical solutions for accurately extracting and effectively integrating virtual expression features based on DNA sequence information, resulting in low efficiency in predicting complex traits and high costs for acquiring multi-omics data, which limits the application of whole-genome selection.
By acquiring transcriptome expression data from a small number of training samples, a virtual expression matrix driven by genetic variation is constructed. An elastic network regression model is used to calculate the predictive weight of genetic variation sites on gene expression levels. By combining the genetic relationship kernel matrix and the virtual expression level relationship kernel matrix, an optimal linear unbiased prediction model is constructed, which reduces detection costs and improves prediction accuracy.
While maintaining model simplicity, it significantly improves the genetic capture ability of complex traits, reduces breeding testing costs, and achieves efficient prediction of complex traits.
Smart Images

Figure CN122417142A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics, and in particular to a method and apparatus for predicting whole-genome breeding values based on a virtual expression matrix constructed by genetic variation. Background Technology
[0002] As global climate change and population growth increasingly pressure food security, traditional crop breeding models are facing severe efficiency bottlenecks. Throughout the history of crop improvement, from early phenotypic selection to molecular marker-assisted selection (MAS), breeders have been seeking more efficient breeding methods. However, while MAS has achieved significant success in regulating relatively well-defined qualitative traits and simple traits controlled by a few major genes, most important agronomic traits are complex quantitative traits driven by a few major genes and a large number of minor genes. Because current technologies struggle to fully explore and utilize these minor, diffusely influential loci, MAS faces significant challenges in improving the selection efficiency of such complex traits. Against this backdrop, Genomic Selection (GS), a technique that can use whole-genome marker information to predict the early genomic estimated breeding value (GEBV) of candidate individuals, is considered a major revolution in breeding technology.
[0003] While genetic engineering (GS) technology has shown great potential in plant and animal breeding, its core assumption—the existence of linkage disequilibrium between markers and trait-controlling genes—is often affected by population structure, marker density, and environmental factors in practical applications. More importantly, current GS models mainly rely on single nucleotide polymorphism (SNP) markers, often treating the complex mapping process from genotype to phenotype as a "black box." This limits the model's ability to capture non-additive effects and gene-environment interactions (G×E), leading to the phenomenon of "loss of heritability."
[0004] The transcriptome, as a bridge connecting the genome and phenotype, reflects gene expression levels at specific times and spaces and is considered the most direct "intermediate phenotype," making it an effective way to recover lost heritability. However, existing multi-omics integration methods still have shortcomings in data fusion strategies. Traditional "early fusion" strategies are prone to overfitting due to the extremely high dimensionality of omics data and the presence of significant environmental noise. Furthermore, from an economic perspective in breeding practice, although multi-omics prediction has significant theoretical advantages, the cost of acquiring dynamic omics data such as transcriptomics and metabolomics is far higher than that of static DNA sequence data. For breeding companies and research institutions, conducting large-scale omics data analysis on the entire population sample at the model application stage is extremely uneconomical. The high detection costs and cumbersome experimental procedures greatly limit the large-scale implementation of multi-omics models in frontline breeding.
[0005] Studies have shown that expression changes involve various genetic variations, including structural variations (SVs). In complex genomes such as cotton, where the scope and intensity of genomic variations vary, expression variations driven by genetic variations are key regulatory signals affecting yield and quality. However, existing breeding prediction models lack a technical solution that can accurately extract and effectively integrate virtual expression features (GReX) based solely on DNA sequence information. Furthermore, how to maximize the capture of complex traits while maintaining the simplicity of prediction models to reduce breeding testing costs remains a pressing technical challenge in the field of genome-wide selection. Summary of the Invention
[0006] This application provides a method and apparatus for predicting the whole-genome breeding value by constructing a virtual expression matrix driven by genetic variation. This method only measures the transcriptome expression data of a small number of training samples and calculates the importance of different genetic variation sites to gene expression levels to estimate the breeding value of the target sample.
[0007] In a first aspect, embodiments of this application provide a method for predicting genome-wide breeding values by constructing a virtual expression matrix driven by genetic variation, the method comprising:
[0008] A population to be predicted, comprising multiple samples, is obtained, along with whole-genome genetic variation data and gene annotation data of the population to be predicted. The samples include training samples and target samples. Transcriptome expression data for each training sample is obtained, including the gene expression level of each gene in the training sample. A weight prediction model is trained based on whole-genome genetic variation data, gene annotation data, and transcriptome expression data of each training sample. The predicted weight of each genetic variation site on the corresponding gene expression level is obtained based on the trained weight prediction model, resulting in a variation site-gene weight matrix. A genetic relationship kernel matrix is constructed based on the genotype coding value of each sample in the whole genome genetic variation data. A virtual expression level relationship kernel matrix is calculated based on the whole genome genetic variation data and the variation site-gene weight matrix. The genetic relationship kernel matrix represents the genetic similarity between any two samples in the same gene. The virtual expression level relationship kernel matrix represents the virtual expression level similarity between any two samples in the same gene. The virtual expression level is the component of the gene expression level of the sample driven by genetic variation. For each trait, a corresponding optimal linear unbiased prediction model is constructed. The optimal linear unbiased prediction model uses the genetic variation relation kernel matrix and the virtual expression relation kernel matrix as the covariance structure of random effects, and predicts the genetic variation effect vector and virtual expression effect vector of each target sample in the population to be predicted based on the phenotypic data of the corresponding trait of each training sample. The genetic variation effect vector and the virtual expression effect vector of each target sample under the same trait in the population to be predicted are summed to obtain the breeding value of each target sample under the corresponding trait.
[0009] Secondly, embodiments of this application provide a genome-wide breeding value prediction device for constructing a virtual expression matrix driven by genetic variation, comprising: The acquisition module is used to acquire a population to be predicted that includes multiple samples, and to acquire whole-genome genetic variation data and gene annotation data of the population to be predicted, wherein the samples include training samples and target samples. The weight prediction module is used to obtain the transcriptome expression data of each training sample. The transcriptome expression data includes the gene expression level of each gene in the training sample. The weight prediction model is trained based on whole genome genetic variation data, gene annotation data and transcriptome expression data of each training sample. Based on the trained weight prediction model, the predicted weight of each genetic variation site on the corresponding gene expression level is obtained, and the variation site-gene weight matrix is obtained. The virtual expression calculation module constructs a genetic relationship kernel matrix based on the genotype coding value of each sample in the whole genome genetic variation data, and calculates a virtual expression level relationship kernel matrix based on the whole genome genetic variation data and the variation site-gene weight matrix. The genetic relationship kernel matrix represents the genetic similarity between any two samples in the same gene, and the virtual expression level relationship kernel matrix represents the virtual expression level similarity between any two samples in the same gene. The virtual expression level is the component of the sample gene expression level driven by genetic variation. The prediction module constructs a corresponding optimal linear unbiased prediction model for each trait. The optimal linear unbiased prediction model uses the genetic variation relation kernel matrix and the virtual expression relation kernel matrix as the covariance structure of random effects, and predicts the genetic variation effect vector and virtual expression effect vector of each target sample in the population to be predicted based on the phenotypic data of the corresponding trait of each training sample. The breeding value calculation module sums the genetic variation effect vector and the virtual expression effect vector for the same trait of each target sample in the population to be predicted, and obtains the breeding value of each target sample under the corresponding trait.
[0010] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform a whole-genome breeding value prediction method for constructing a virtual expression matrix driven by genetic variation.
[0011] The main contributions and innovations of this invention are as follows: This application embodiment measures transcriptome expression data only on a small number of training samples, uses elastic network regression to train a weight prediction model and constructs a variant site-gene weight matrix, eliminating the need for transcriptome testing on the entire population, significantly reducing breeding testing costs and experimental complexity, and overcoming the economic bottleneck of multi-omics breeding implementation. This application embodiment calculates a virtual expression level relationship matrix driven by genetic variation based on whole-genome genetic variation data and the variant site-gene weight matrix, transforming discrete genetic markers into biologically meaningful continuous intermediate phenotypic signals, eliminating environmental and non-genetic noise, and retaining only heritable regulatory components. This application embodiment uses the genetic variation relationship matrix and the virtual expression level relationship matrix as the covariance structure of the optimal linear unbiased prediction model with two random effects, simultaneously capturing DNA... Sequence-level genetic effects and transcriptional regulation-level effects effectively recover "missing heritability" and enhance the genetic capture ability of complex quantitative traits. The elastic network regression model in this application combines L1 and L2 regularized training weights to achieve sparse screening of key genetic variation sites, handle strong correlations between variation sites, ensure the robustness of the model in complex genetic backgrounds, and avoid overfitting. This application's embodiment sums the genetic variation effect vector and the virtual expression effect vector to calculate the breeding value, combining two layers of core genetic information to comprehensively analyze the trait regulation mechanism, significantly improving the prediction accuracy of key agronomic traits such as yield and fiber quality.
[0012] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description
[0013] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of a genome-wide breeding value prediction method based on a genetic variation-driven virtual expression matrix according to an embodiment of this application; Figure 2 This is a comparison chart of the performance of the single genetic variation method (G) and the integrated prediction model (G+E) of this scheme on multiple traits according to the embodiments of this application; Figure 3 This is a structural block diagram of a genome-wide breeding value prediction device constructed from a genetic variation-driven virtual expression matrix according to an embodiment of this application; Figure 4 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0014] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0015] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; and multiple steps described in this specification may be combined into a single step in other embodiments.
[0016] Example 1 This application provides a genome-wide breeding value prediction method based on a virtual expression matrix constructed using genetic variation. This method estimates the breeding value of the target sample by measuring transcriptome expression data from a small number of training samples and calculating the importance of different genetic variation sites for gene expression levels. Specifically, refer to... Figure 1 The method includes: A population to be predicted, comprising multiple samples, is obtained, along with whole-genome genetic variation data and gene annotation data of the population to be predicted. The samples include training samples and target samples. Transcriptome expression data for each training sample is obtained, including the gene expression level of each gene in the training sample. A weight prediction model is trained based on whole-genome genetic variation data, gene annotation data, and transcriptome expression data of each training sample. The predicted weight of each genetic variation site on the corresponding gene expression level is obtained based on the trained weight prediction model, resulting in a variation site-gene weight matrix. A genetic relationship kernel matrix is constructed based on the genotype coding value of each sample in the whole genome genetic variation data. A virtual expression level relationship kernel matrix is calculated based on the whole genome genetic variation data and the variation site-gene weight matrix. The genetic relationship kernel matrix represents the genetic similarity between any two samples in the same gene. The virtual expression level relationship kernel matrix represents the virtual expression level similarity between any two samples in the same gene. The virtual expression level is the component of the gene expression level of the sample driven by genetic variation. For each trait, a corresponding optimal linear unbiased prediction model is constructed. The optimal linear unbiased prediction model uses the genetic variation relation kernel matrix and the virtual expression relation kernel matrix as the covariance structure of random effects, and predicts the genetic variation effect vector and virtual expression effect vector of each target sample in the population to be predicted based on the phenotypic data of the corresponding trait of each training sample. The genetic variation effect vector and the virtual expression effect vector of each target sample under the same trait in the population to be predicted are summed to obtain the breeding value of each target sample under the corresponding trait.
[0017] In the current embodiment, the whole-genome genetic variation data is a set of DNA sequence differences between samples in the population to be predicted. The whole-genome genetic variation data is the fundamental source of genetic variation that determines traits. The gene annotation data marks the location information of each gene in the species to which the population to be predicted belongs.
[0018] In other words, gene annotation data marks the location of each gene, while genome-wide genetic variation data marks each genetic variation site on each gene.
[0019] In the current embodiment, the whole-genome genetic variation data is one or more of single nucleotide polymorphisms, insertion / deletion variations, and structural variations.
[0020] Furthermore, the acquired whole-genome genetic variation data were preprocessed. The preprocessing steps included using bioinformatics software to filter out data with a minimum allele frequency of less than 0.05 and a deletion rate greater than 0.1.
[0021] In the current embodiment, only a small portion of the training samples in the population to be predicted are obtained to predict the breeding value of the target sample with only whole-genome genetic variation data. Therefore, the number of training samples in the population to be predicted is less than that of the target sample.
[0022] Specifically, since the cost of obtaining transcriptome expression data is high, if the transcriptome expression data of each sample in the population to be predicted is obtained to calculate the breeding value, the time and economic costs are very high. Therefore, this scheme adopts a combination of weighted prediction model and optimal linear unbiased prediction model, so that only a small number of training samples in the population to be predicted need to be obtained to accurately predict the breeding value of other target samples.
[0023] In the current embodiment, the transcriptome expression data of each training sample is standardized to eliminate the systematic bias caused by batch effect during the detection of expression levels of different samples, and to ensure that the expression levels of the same gene are comparable among different samples.
[0024] In the step of training a weighted prediction model based on whole-genome genetic variation data, gene annotation data, and transcriptome expression data of each training sample, the gene location of each gene in each training sample is defined according to the gene annotation data. The upstream and downstream preset regions centered on each gene location are obtained as cis-regulatory windows. The gene expression level of the corresponding cis-regulatory window is obtained based on the corresponding transcriptome expression data. The genetic variation sites in each cis-regulatory window are obtained based on the whole-genome genetic variation data. An elastic network regression model is trained based on the genetic variation sites in each cis-regulatory window and the corresponding gene expression levels. The trained elastic network regression model is used as the weighted prediction model.
[0025] In this scheme, a 100Kb region upstream and downstream of the center of each gene is defined as a cis-regulatory window using gene annotation data. The cis-regulatory window contains all gene fragments that affect the expression level of the corresponding gene, thus ensuring coverage of all genetic variation sites that affect the expression level of the corresponding gene, while avoiding the inclusion of too many irrelevant variations that would increase the computational burden of the model, and ensuring the efficiency and accuracy of weight prediction.
[0026] In other words, this approach first defines a cis-regulatory window for each gene in the training samples, then obtains the genetic variation sites within each cis-regulatory window based on whole-genome genetic variation data, thereby determining which gene segments affect the corresponding gene expression level. Finally, by training an elastic network regression model, it predicts the influence weights of different genetic variation sites on the gene expression level of a particular gene, thus obtaining a variation site-gene weight matrix applicable to the entire population.
[0027] In this scheme, the elastic net regression model used is the Elastic Net Regression model. The elastic net regression model is used to train the prediction weights of different genetic variation sites on the corresponding gene expression levels, that is, to train the degree of influence of different genetic variation sites on gene expression levels.
[0028] Furthermore, when training the resilient network regression model, the sum of the mean squared error term and the regularization penalty term is used as the total loss function. The formula for the mean squared error term is as follows:
[0029] in, Let be the mean squared error term for the j-th gene, N be the total number of training samples, and i be the index of the training sample. Let be the expression level of the j-th gene in the i-th training sample. This represents the total number of genetic variation sites within the cis-regulatory window of the target gene, where k is the index of the genetic variation site. Let be the genotype encoding value of the i-th training sample at the k-th genetic variation site. This represents the predicted weight of the k-th genetic variation site on the expression level of the j-th gene in the variation site-gene weight matrix. This is the intercept term of the elastic network regression model.
[0030] The formula for the regularization penalty term is expressed as follows:
[0031] in, For regularization penalty terms, For regularization parameters, This is the mixing ratio parameter.
[0032] Specifically, in the regularization penalty term, the L1 part The L2 part is used to make the prediction weights of insignificant genetic variation sites tend to 0, thereby achieving sparse screening of features. This is used to handle the strong correlations that may exist between genetic variation sites, thereby ensuring the robustness of the model in complex genetic contexts.
[0033] Specifically, regularization parameters The overall strength of the penalty is used to control its intensity. A higher intensity indicates more severe compression of the weights by the model, thus eliminating more irrelevant genetic variation sites. (Mixing ratio parameter) This scheme is used to balance variable selection and model stability. It is 0.5.
[0034] In the current embodiment, the variant site-gene weight matrix records the degree of influence of each genetic variant site on the expression level of different genes.
[0035] In other words, this approach only requires actual measurement of transcriptome expression data for each training sample. Then, a weighted prediction model is used to predict the influence of different genetic variation sites on the expression levels of different genes. The weighted prediction model can then obtain the influence of the target sample's genetic variation sites on the expression levels of different genes based on the target sample's genetic variation sites recorded in the whole genome genetic variation data, without having to spend a lot of time measuring transcriptome expression data for each sample.
[0036] In the current embodiment, the whole genome genetic variation data is numerically encoded to obtain a genetic variation matrix, which contains the genetic variation information of each sample on each gene. A genetic relationship kernel matrix is constructed based on the centered genetic variation matrix using a linear kernel function, and the genetic variation relationship kernel matrix is represented by X.
[0037] Specifically, based on whole-genome genetic variation data, each genetic variation site of each sample is numerically encoded as 0, 1, or 2. This invention employs an additive model for numerical encoding of genetic variation sites: 0 represents that the sample does not contain the mutant allele (homozygous wild type), 1 represents that it contains one mutant allele (heterozygous type), and 2 represents that it contains two mutant alleles (homozygous mutant type). Through this encoding, the genotyping data is transformed into a computable linear numerical matrix.
[0038] In the current embodiment, a virtual expression matrix is obtained by calculating the virtual expression level of each gene in each sample based on the genetic variation matrix. A virtual expression level relation kernel matrix is then constructed based on the centered virtual expression matrix using a linear kernel function. The formula for calculating the virtual expression level is as follows:
[0039] in, For the i-th sample Virtual expression levels of individual genes This represents the total number of genetic variation sites in the sample. An index of genetic variation sites. For the i-th sample in the genetic variation matrix Numerical coding results of individual genetic variation sites This represents the predicted weight of the k-th genetic variation site on the expression level of the j-th gene in the variation site-gene weight matrix. This is the intercept term.
[0040] Specifically, the virtual expression levels of all genes are integrated column-wise to ultimately construct a virtual expression level relationship matrix for the entire population.
[0041] Specifically, the virtual expression level relational matrix represents the portion of expression variation directly driven by genetic variation (DNA level). In this way, the present invention successfully transforms discrete genetic marker information into continuous, biologically regulatory "intermediate phenotype" signals; reduces environmental noise: because the matrix is generated based on genetic weight prediction, it eliminates noise in the transcriptome measurement data caused by experimental operations, environmental fluctuations, or non-genetic factors, thus allowing the model to focus more on heritable regulatory components; achieves full population coverage: using this method, for a large number of samples in the breeding population that have not undergone transcriptome sequencing, their virtual expression levels can be inferred solely from their existing whole-genome genetic variation data. This effectively completes high-dimensional omics information without requiring large-scale experimental measurement of transcriptome data from all samples, greatly reducing the detection costs and technical implementation barriers in actual breeding.
[0042] Specifically, the genetic variation matrix and the virtual expression matrix are centered to eliminate the influence of dimensions and ensure that the mean of the features is 0. The formula for the centered processing is expressed as follows:
[0043] in, These are the eigenvalues after centering. These are the original matrix values. Let be the average value of the k-th genetic variation site in the entire population. This is used when centering the genetic variation matrix. This serves as an index for genetic variation sites; during the centering of the virtual expression matrix, For gene indexing.
[0044] The formula for a linear kernel function is as follows:
[0045] in, From a data perspective, when The time indicates that the genetic variation matrix is being calculated, when Time indicates that the virtual representation matrix is being calculated. This represents the centered genetic variation relationship matrix / virtual expression level relationship matrix. This is a genetic relationship kernel matrix / virtual expression kernel matrix. The total number of features in the genetic variation relationship matrix / virtual expression level relationship matrix. As a denominator, it serves a standardization function, ensuring that the numerical scale of the kernel matrix is comparable.
[0046] Specifically, the genetic relationship matrix represents the standard genomic kinship between samples at the DNA sequence level, reflecting the most basic genetic similarity between samples.
[0047] Specifically, the virtual expression kernel matrix represents the similarity between samples at the transcriptional regulation level. That is, the closer the gene expression patterns of two samples driven by genetic variations are, the higher their similarity value in the virtual expression kernel matrix.
[0048] In the current embodiment, using the phenotypic data of the corresponding traits of the training samples as the dependent variable, an optimal linear unbiased prediction model is constructed, including a genetic variation effect vector and a virtual expression effect vector. The genetic relation kernel matrix is the covariance structure of the genetic variation effect, and the virtual expression kernel matrix is the covariance structure of the virtual expression effect vector. The optimal linear unbiased prediction model is trained using the restricted maximum likelihood method to obtain the genetic variation variance components and the virtual expression variance components. Based on the genetic variation variance components and the genetic relation kernel matrix, the genetic variation effect vector of each target sample on the corresponding trait is calculated, and based on the virtual expression variance components and the virtual expression kernel matrix, the virtual expression effect vector of each target sample on the corresponding trait is calculated.
[0049] Specifically, using the phenotypic data of the corresponding traits in the training samples as the dependent variable, an optimal linear unbiased prediction model is constructed. The expression for the optimal linear unbiased prediction model is:
[0050] in, Phenotypic data corresponding to the traits of the training samples, This represents the average level of the phenotypic data corresponding to the training samples. This is the genetic variation effect vector of the training samples. This is the virtual expression effect vector of the training samples. It is a random residual vector.
[0051] Specifically, assuming that the genetic variation effect vector and the genetic relationship kernel matrix satisfy a normal distribution, and that the genetic relationship kernel matrix is the covariance structure of the genetic variation effect, then we can obtain:
[0052] in, This represents the variance component of genetic variation.
[0053] Assuming that the virtual expression effect vector and the virtual expression kernel matrix follow a normal distribution, and that the virtual expression kernel matrix is the covariance structure of the virtual expression effect vector, then we can obtain:
[0054] in, This represents the virtual expression of variance components.
[0055] The residual random residual vector also follows a normal distribution, i.e. ,in It is the identity matrix. This represents the residual variance components.
[0056] Specifically, the restricted maximum likelihood method is used to train the optimal linear unbiased prediction model to estimate different variance components during training. The specific calculation process utilizes the `mmer()` function from the `sommer` package in R for fitting. Through iterative optimization, the optimal linear unbiased prediction model decomposes the total variance of the corresponding phenotypic data into components of genetic variation variance, dummy expression variance, and residual variance.
[0057] Specifically, the formula for summing the genetic variation effect vector and the virtual expression effect vector for the same trait in each target sample within the population to be predicted is expressed as follows:
[0058] in, These are the breeding values for the corresponding traits. This represents the genetic variation effect vector of the target sample. This is the virtual expression effect vector of the target sample.
[0059] Specifically, in order to assess the degree of influence of different genetic levels on the target trait, this embodiment calculates the proportion of the variance component corresponding to each nucleus to the total variance of the phenotypic data. Quantitative analysis was performed on the genetic contribution rate at the level of genetic variation and virtual expression. The calculation formula is as follows:
[0060] in, As a virtual expression of variance components, The variance component of genetic variation The residual variance components, This represents the virtual expression variance component / genetic variation variance component.
[0061] Specifically, quantitative analysis of the genetic contribution rate at the level of genetic variation and virtual expression helps to elucidate the regulatory mechanisms of traits and verify the technical contribution of introducing a virtual expression matrix to recover "missing heritability".
[0062] In some embodiments, the performance of the integrated prediction model combining the single genetic variation method (G) and the proposed method (G+E) on multiple traits is compared as follows: Figure 2 As shown, by Figure 2 It can be seen that the proposed method achieves a significant improvement in prediction performance for the following core traits: Yield characteristics: single boll weight, seed cotton weight, lint percentage, lint weight, etc.; Quality characteristics: micronaire value, fiber strength, fiber uniformity, fiber length, etc.
[0063] Cross-validation results show that, compared with a single genetic variation prediction model, the integrated prediction method proposed in this invention significantly improves the ability to capture traits, and its PCC gain demonstrates the technical superiority of constructing a virtual matrix by driving intermediate regulatory signals through genetic variation.
[0064] Example 2 Based on the same concept, referencing Figure 3 This application also proposes a genome-wide breeding value prediction device for constructing a virtual expression matrix driven by genetic variation, comprising: The acquisition module is used to acquire a population to be predicted that includes multiple samples, and to acquire whole-genome genetic variation data and gene annotation data of the population to be predicted, wherein the samples include training samples and target samples. The weight prediction module is used to obtain the transcriptome expression data of each training sample. The transcriptome expression data includes the gene expression level of each gene in the training sample. The weight prediction model is trained based on whole genome genetic variation data, gene annotation data and transcriptome expression data of each training sample. Based on the trained weight prediction model, the predicted weight of each genetic variation site on the corresponding gene expression level is obtained, and the variation site-gene weight matrix is obtained. The virtual expression calculation module constructs a genetic relationship kernel matrix based on the genotype coding value of each sample in the whole genome genetic variation data, and calculates a virtual expression level relationship kernel matrix based on the whole genome genetic variation data and the variation site-gene weight matrix. The genetic relationship kernel matrix represents the genetic similarity between any two samples in the same gene, and the virtual expression level relationship kernel matrix represents the virtual expression level similarity between any two samples in the same gene. The virtual expression level is the component of the sample gene expression level driven by genetic variation. The prediction module constructs a corresponding optimal linear unbiased prediction model for each trait. The optimal linear unbiased prediction model uses the genetic variation relation kernel matrix and the virtual expression relation kernel matrix as the covariance structure of random effects, and predicts the genetic variation effect vector and virtual expression effect vector of each target sample in the population to be predicted based on the phenotypic data of the corresponding trait of each training sample. The breeding value calculation module sums the genetic variation effect vector and the virtual expression effect vector for the same trait of each target sample in the population to be predicted, and obtains the breeding value of each target sample under the corresponding trait.
[0065] Example 3 This embodiment also provides an electronic device, see reference. Figure 4 It includes a memory 404 and a processor 402, wherein the memory 404 stores a computer program and the processor 402 is configured to run the computer program to perform the steps in any of the above method embodiments.
[0066] Specifically, the processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0067] Memory 404 may include a mass storage device for data or instructions. For example, and not limitingly, memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 404 may include removable or non-removable (or fixed) media. Where appropriate, memory 404 may be internal or external to a data processing device. In a particular embodiment, memory 404 is non-volatile memory. In a particular embodiment, memory 404 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable read-only memory (PROM), an erasable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), an electrically alterable read-only memory (EAROM), or flash memory, or a combination of two or more of these. Where appropriate, the RAM can be Static Random-Access Memory (SRAM) or Dynamic Random-Access Memory (DRAM). DRAM can be Fast Page Mode Dynamic Random-Access Memory (FPMDRAM), Extended Data Out Dynamic Random-Access Memory (EDODRAM), Synchronous Dynamic Random-Access Memory (SDRAM), etc.
[0068] The memory 404 can be used to store or cache various data files that need to be processed and / or communicated, as well as possible computer program instructions executed by the processor 402.
[0069] The processor 402 reads and executes computer program instructions stored in the memory 404 to implement any of the whole-genome breeding value prediction methods for constructing a virtual expression matrix driven by genetic variation in the above embodiments.
[0070] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408, wherein the transmission device 406 is connected to the processor 402, and the input / output device 408 is connected to the processor 402.
[0071] Transmission device 406 can be used to receive or send data via a network. Specific examples of the network described above may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, transmission device 406 may be a radio frequency (RF) module used for wireless communication with the Internet.
[0072] Input / output device 408 is used to input or output information. In this embodiment, the input information may be whole-genome genetic variation data, gene annotation data, etc., and the output information may be the breeding value of each target sample under the corresponding trait.
[0073] Optionally, in this embodiment, the processor 402 can be configured to perform the following steps via a computer program: A population to be predicted, comprising multiple samples, is obtained, along with whole-genome genetic variation data and gene annotation data of the population to be predicted. The samples include training samples and target samples. Transcriptome expression data for each training sample is obtained, including the gene expression level of each gene in the training sample. A weight prediction model is trained based on whole-genome genetic variation data, gene annotation data, and transcriptome expression data of each training sample. The predicted weight of each genetic variation site on the corresponding gene expression level is obtained based on the trained weight prediction model, resulting in a variation site-gene weight matrix. A genetic relationship kernel matrix is constructed based on the genotype coding value of each sample in the whole genome genetic variation data. A virtual expression level relationship kernel matrix is calculated based on the whole genome genetic variation data and the variation site-gene weight matrix. The genetic relationship kernel matrix represents the genetic similarity between any two samples in the same gene. The virtual expression level relationship kernel matrix represents the virtual expression level similarity between any two samples in the same gene. The virtual expression level is the component of the gene expression level of the sample driven by genetic variation. For each trait, a corresponding optimal linear unbiased prediction model is constructed. The optimal linear unbiased prediction model uses the genetic variation relation kernel matrix and the virtual expression relation kernel matrix as the covariance structure of random effects, and predicts the genetic variation effect vector and virtual expression effect vector of each target sample in the population to be predicted based on the phenotypic data of the corresponding trait of each training sample. The genetic variation effect vector and the virtual expression effect vector of each target sample under the same trait in the population to be predicted are summed to obtain the breeding value of each target sample under the corresponding trait.
[0074] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.
[0075] Generally, various embodiments can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention can be implemented in hardware, while others can be implemented by firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, by way of non-limiting example, these blocks, apparatuses, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0076] Embodiments of the present invention can be implemented by computer software, which may be executable by a data processor of a mobile device, such as a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets, and / or macros can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product may include one or more computer-executable components configured to perform the embodiments when the program is run. The one or more computer-executable components may be at least one piece of software code or a portion thereof. Additionally, it should be noted in this respect that, as Figure 4Any box in the logical flow can represent a program step, or interconnected logic circuits, boxes and functions, or a combination of program steps and logic circuits, boxes and functions. Software can be stored on physical media such as memory chips or blocks of storage implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs, etc. The physical medium is a non-transient medium.
[0077] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0078] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for predicting genome-wide breeding values by constructing a virtual expression matrix driven by genetic variation, characterized in that, Includes the following steps: A population to be predicted, comprising multiple samples, is obtained, along with whole-genome genetic variation data and gene annotation data of the population to be predicted. The samples include training samples and target samples. Transcriptome expression data for each training sample is obtained, including the gene expression level of each gene in the training sample. A weight prediction model is trained based on whole-genome genetic variation data, gene annotation data, and transcriptome expression data of each training sample. The predicted weight of each genetic variation site on the corresponding gene expression level is obtained based on the trained weight prediction model, resulting in a variation site-gene weight matrix. A genetic relationship kernel matrix is constructed based on the genotype coding value of each sample in the whole genome genetic variation data. A virtual expression level relationship kernel matrix is calculated based on the whole genome genetic variation data and the variation site-gene weight matrix. The genetic relationship kernel matrix represents the genetic similarity between any two samples in the same gene. The virtual expression level relationship kernel matrix represents the virtual expression level similarity between any two samples in the same gene. The virtual expression level is the component of the gene expression level of the sample driven by genetic variation. For each trait, a corresponding optimal linear unbiased prediction model is constructed. The optimal linear unbiased prediction model uses the genetic variation relation kernel matrix and the virtual expression relation kernel matrix as the covariance structure of random effects, and predicts the genetic variation effect vector and virtual expression effect vector of each target sample in the population to be predicted based on the phenotypic data of the corresponding trait of each training sample. The genetic variation effect vector and the virtual expression effect vector of each target sample under the same trait in the population to be predicted are summed to obtain the breeding value of each target sample under the corresponding trait.
2. The method for predicting whole-genome breeding values by constructing a virtual expression matrix driven by genetic variation according to claim 1, characterized in that, The whole-genome genetic variation data includes one or more of single nucleotide polymorphisms, insertion / deletion variations, and structural variations.
3. The method for predicting whole-genome breeding values by constructing a virtual expression matrix driven by genetic variation according to claim 1, characterized in that, In the step of training the weighted prediction model based on whole-genome genetic variation data, gene annotation data, and transcriptome expression data of each training sample, the gene position of each gene in each training sample is defined according to the gene annotation data, the upstream and downstream preset regions are obtained as cis-regulatory windows with each gene position as the center, the gene expression level of the corresponding cis-regulatory window is obtained based on the corresponding transcriptome expression data, and the genetic variation sites in each cis-regulatory window are obtained based on the whole-genome genetic variation data. Based on the genetic variation sites in each cis-regulatory window and the corresponding gene expression levels, an elastic network regression model is trained, and the trained elastic network regression model is used as a weight prediction model.
4. The method for predicting whole-genome breeding values by constructing a virtual expression matrix driven by genetic variation according to claim 1, characterized in that, When training a resilient network regression model, the sum of the mean squared error term and the regularization penalty term is used as the total loss function. The formula for the mean squared error term is as follows: The formula for the regularization penalty term is expressed as follows: in, For regularization penalty terms, For regularization parameters, For mixing ratio parameters, Let be the mean squared error term for the j-th gene, N be the total number of training samples, and i be the index of the training sample. Let be the expression level of the j-th gene in the i-th training sample. This represents the total number of genetic variation sites within the cis-regulatory window of the target gene, where k is the index of the genetic variation site. Let be the genotype encoding value of the i-th training sample at the k-th genetic variation site. This represents the predicted weight of the k-th genetic variation site on the expression level of the j-th gene in the variation site-gene weight matrix. This is the intercept term of the elastic network regression model.
5. The method for predicting whole-genome breeding values by constructing a virtual expression matrix driven by genetic variation according to claim 1, characterized in that, The genetic variation matrix is obtained by numerically encoding the whole genome genetic variation data. The genetic variation matrix contains the genetic variation information of each sample on each gene. A genetic relationship kernel matrix is constructed based on the centered genetic variation matrix using a linear kernel function.
6. The method for predicting whole-genome breeding values by constructing a virtual expression matrix driven by genetic variation according to claim 5, characterized in that, The virtual expression matrix is obtained by calculating the virtual expression level of each gene in each sample based on the genetic variation matrix. A virtual expression level relation kernel matrix is then constructed based on the centered virtual expression matrix using a linear kernel function. The formula for calculating the virtual expression level is as follows: in, For the i-th sample Virtual expression levels of individual genes This represents the total number of genetic variation sites in the sample. An index of genetic variation sites. For the i-th sample in the genetic variation matrix Numerical coding results of individual genetic variation sites This represents the predicted weight of the k-th genetic variation site on the expression level of the j-th gene in the variation site-gene weight matrix. This is the intercept term.
7. The method for predicting whole-genome breeding values by constructing a virtual expression matrix driven by genetic variation according to claim 1, characterized in that, Using the phenotypic data of the corresponding traits in the training samples as the dependent variable, an optimal linear unbiased prediction model is constructed, including a genetic variation effect vector and a virtual expression effect vector. The genetic relation kernel matrix is the covariance structure of the genetic variation effect, and the virtual expression kernel matrix is the covariance structure of the virtual expression effect vector. The optimal linear unbiased prediction model is trained using the restricted maximum likelihood method to obtain the genetic variation variance components and the virtual expression variance components. Based on the genetic variation variance components and the genetic relation kernel matrix, the genetic variation effect vector of each target sample on the corresponding trait is calculated, and based on the virtual expression variance components and the virtual expression kernel matrix, the virtual expression effect vector of each target sample on the corresponding trait is calculated.
8. A genome-wide breeding value prediction device based on the construction of a virtual expression matrix driven by genetic variation, characterized in that, include: The acquisition module is used to acquire a population to be predicted that includes multiple samples, and to acquire whole-genome genetic variation data and gene annotation data of the population to be predicted, wherein the samples include training samples and target samples. The weight prediction module is used to obtain the transcriptome expression data of each training sample. The transcriptome expression data includes the gene expression level of each gene in the training sample. The weight prediction model is trained based on whole genome genetic variation data, gene annotation data and transcriptome expression data of each training sample. Based on the trained weight prediction model, the predicted weight of each genetic variation site on the corresponding gene expression level is obtained, and the variation site-gene weight matrix is obtained. The virtual expression calculation module calculates a genetic variation relationship matrix based on whole-genome genetic variation data; and calculates a virtual expression level relationship matrix based on whole-genome genetic variation data and a variation site-gene weight matrix. The genetic variation relationship matrix includes genetic variation information for each genetic variation site of each sample in the population to be predicted, and the virtual expression level relationship matrix includes genetic regulation prediction values for each genetic variation site of each sample in the population to be predicted, wherein the genetic regulation prediction values are the heritability of the corresponding genetic variation site. The prediction module constructs a corresponding optimal linear unbiased prediction model for each trait. The optimal linear unbiased prediction model uses the genetic variation relationship matrix and the virtual expression level relationship matrix as the covariance structure of random effects, and predicts the genetic variation effect vector and virtual expression effect vector of each target sample in the population to be predicted based on the phenotypic data of the corresponding trait of each training sample. The breeding value calculation module sums the genetic variation effect vector and the virtual expression effect vector for the same trait of each target sample in the population to be predicted, and obtains the breeding value of each target sample under the corresponding trait.
9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform a whole-genome breeding value prediction method for constructing a virtual expression matrix driven by genetic variation, as described in any one of claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements a whole-genome breeding value prediction method for constructing a virtual expression matrix driven by genetic variation as described in any one of claims 1-7.