GAN model construction method, and data fitting, phenotype prediction, sample augmentation and breeding method based on GAN model
By constructing a GAN model, using the adversarial training of generators and discriminators, fit multi-omic data characteristics and perform phenotypic predictions, the problems of insufficient data volume and low genome selection accuracy in the existing technology are solved, and higher accuracy of breeding models are achieved.
Patent Information
- Application Number
- PCT/CN2023/136547
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-05
AI Technical Summary
The amount of data available for breeding is insufficient and the accuracy of genomic selection is not high enough, especially when dealing with complex nonlinear associations and multiomic data.
Using the construction method of the GAN model, the multiomic data characteristics are fitted and phenotypic prediction is performed through adversarial training of generators and discriminators. The method involves obtaining real multiomics data and fitting genotype data, and training through the cross loss entropy loss function of the generator and discriminator until the discriminator cannot correctly distinguish between the real and fitted data.
In the case of insufficient data volume, we can give full play to the advantages of deep learning algorithms, improve the accuracy of genome selection, effectively capture complex nonlinear associations, and generate data that meets the characteristics of multi-omic data.
Smart Images

Figure CN2023136547_05062025_PF_FP_ABST
Abstract
Description
GAN model construction method and data fitting, phenotype prediction, sample expansion and breeding methods based on GAN model Priority information
[0001] This application claims priority to Chinese invention patent application CN2023116017459, filed on November 27, 2023, entitled "Method for constructing a GAN model and methods for data fitting, phenotype prediction, sample expansion, and breeding based on a GAN model," the contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of biotechnology, and more particularly to a method for constructing a GAN model and methods for data fitting, phenotype prediction, sample expansion, and breeding based on the GAN model. A GAN (Generative Adversarial Network) is a deep generative model based on adversarial learning. Background Art
[0003] Genomic selection implements a data-driven scientific breeding program. Based on all available genetic information of individual genomes and the relationship between genotype and phenotype of the reference population, a breeding model is constructed to estimate the effect value of single nucleotide polymorphisms, and then estimate the breeding value of the candidate population. By screening individuals with higher breeding value, actual breeding is carried out to achieve the goal of rapidly improving the breeding population. Currently, the widely used breeding models can be divided into two categories according to the use of data: models based on genomic data and models based on multi-omics data; according to the use of algorithmic models, they can also be divided into two categories: models based on statistical methods and models based on machine learning methods. These four different types of breeding models are of great significance to the development of genomic selection, especially the development of intelligent breeding systems, and have effectively promoted the progress of contemporary biological breeding. However, as research continues to deepen, many problems still exist, mainly in the following aspects:
[0004] First, in the early days of genomic selection research, statistical methods were used to model genomic data, achieving excellent results in practical applications. This was particularly true for predicting phenotypes with high heritability and significant main effects. However, many important economic traits are quantitative traits. These traits are regulated not only by major loci but also by a large number of minor loci, and they interact closely with the environment. Furthermore, the intrinsic interaction patterns of individual phenotypes, in addition to additive effects, also include dominance effects and epistatic effects. This means that the association between genotype and phenotype is not a simple linear relationship but rather involves complex nonlinear relationships. These complex nonlinear relationships cannot be well captured by statistical models.
[0005] Secondly, while addressing the aforementioned shortcomings, the application of multi-omics data has gradually expanded: not only encompassing as many potential causal molecular markers as possible at the genomic level, but also providing more functional information about SNPs through the use of multi-omics data. Deep learning algorithms are also being applied to autonomously learn associations between key breeding data and phenotypes, encompassing not only linear but also complex nonlinear relationships. These two applications, to some extent, address the shortcomings of genomic data and statistical algorithms and have been proven to effectively improve the accuracy and efficiency of genomic selection, meeting the needs of smart breeding systems in the big data era. However, acquiring comprehensive multi-omics data is costly, and in practice, measuring multi-omics data for every candidate individual is challenging. Therefore, systematic and in-depth research is needed to indirectly obtain multi-omics data for candidate individuals through multi-omics data from training populations. Furthermore, deep learning models require extensive training data to achieve optimal performance. However, current genomic selection faces the challenge of estimating the effect sizes of tens of thousands or even tens of millions of SNPs across hundreds or thousands of individuals, which prevents the full potential of deep learning models. The specific manifestations are that although many works have shown that the introduction of deep learning algorithms, especially the integration of multi-omics data, can improve the accuracy of phenotypic prediction, the improvement is not as high as expected.
[0006] Therefore, we need a more complete intelligent breeding system built with deep learning algorithms to make up for the lack of data while giving full play to the advantages of deep learning algorithms and comprehensively improving the accuracy of genome selection. Technical issues
[0007] One purpose of this application is to provide a method for constructing a GAN model and methods for data fitting, phenotype prediction, sample expansion, and breeding based on the GAN model, at least to solve the problems of insufficient data volume and insufficient accuracy of genomic selection for breeding. Technical Solutions
[0008] To achieve the above objectives, some embodiments of the present application provide the following aspects:
[0009] In a first aspect, some embodiments of the present application provide a method for constructing a GAN model: the GAN model includes a first generator G1, a second generator G2, a first discriminator D1, and a second discriminator D2;
[0010] The construction method comprises: obtaining real multi-omics data and obtaining fitted genotype data, wherein the real multi-omics data comprises genotype data and selectively comprises at least one of epigenomic data, transcriptomic data, proteomic data, metabolomic data, and functional group data of the target species, as well as real phenotypic data y realThe fitted genotype data is a set of random values with the same dimension as the genotype data; the label value of the real multi-omics data is True, and the label value of the fitted genotype data is False;
[0011] The real multi-omics data is input into the first generator G1, and the data features aggregated by the first generator G1 are input into the first discriminator D1 to obtain the predicted phenotypic data y pre , based on the predicted phenotypic data y pre and the real phenotypic data y real The difference between the two constructs the loss function training to obtain the trained first generator G1 and the first discriminator D1;
[0012] The network structure and basic parameters of the second generator G2 are initially constructed to be consistent with those of the trained first generator G1; the real multi-omics data are input into the second discriminator D2 through the aggregation features of the first generator G1 to obtain a first judgment result, the fitted genotype data are input into the second discriminator D2 through the aggregation features of the second generator G2 to obtain a second judgment result, the parameters of the second discriminator D2 are updated based on the loss function value when the first judgment result is True and the second judgment result is False, and the parameters of the second generator G2 are updated based on the loss function value when the second judgment result is True, and adversarial training is performed until the second discriminator D2 cannot correctly distinguish whether the input is true or false, thereby obtaining the trained second generator G2 and second discriminator D2.
[0013] In a preferred embodiment, the loss function training is constructed based on the difference between the predicted phenotypic data and the true phenotypic data, and the loss function is defined by the mean absolute error, loss function L(x) = |D1(x|G1) - y real |.
[0014] In a preferred embodiment, the loss function used in the adversarial training is defined by cross-loss entropy, which is defined as: ;
[0015] When receiving the input of the first generator G1 and judging it as true, the loss function ;
[0016] Receive the input of the second generator G2 and judge it to be False, the loss function ;
[0017] Receive the input of the second generator G2 and judge it to be True, the loss function ;
[0018] The loss function of the discriminator D2 is the minimax valuation function V(G2, D2):
[0019] The loss function of the generator G2 is defined as minimizing the valuation function V(G2):
[0020] in, The output of G1 processing the real multi-omics data, The output of the fitted genotype data is processed for the G2.
[0021] In a preferred embodiment, the adversarial training steps are:
[0022] Step A: Randomly select a set of random variables with the same dimensions as the true genotype data from the specified data distribution;
[0023] Step B: Use G2 to receive the random variable generated in step A, fit the data features, and label it as False;
[0024] Step C: Select a certain number of samples from the real data, use G1 to obtain the real data features, and label them as True;
[0025] Step D: using B and C to train D2 according to the cross-loss entropy loss function V(G2, D2);
[0026] Step E: regenerate a set of random variables according to A, define the label as True, and train G2 according to the cross-loss entropy loss function V(G2);
[0027] Step F: Repeat the above AE steps according to the specified number of steps until the set conditions are met and the training stops.
[0028] In a second aspect, some embodiments of the present application also provide a method for data fitting, which uses a GAN model constructed by the construction method described above to fit multi-omics data, including the following steps: inputting the genotype data of the candidate population into the G2, obtaining estimated values of the parameters of each layer in the G2, extracting the estimated values of the parameters of each layer in the G2, corresponding to the position where the multi-omics data is input in the G1 model, to achieve fitting of the multi-omics data.
[0029] In a third aspect, some embodiments of the present application further provide a method for phenotype prediction, wherein the GAN model constructed using the construction method described above is used for phenotype prediction, comprising the following steps:
[0030] Inputting the candidate population multi-omics data into the G1, and inputting the data features aggregated by the G1 into the D1 to obtain the predicted phenotype of the candidate population;
[0031] Alternatively, the candidate population genotype data is input into the G2 to obtain the estimated values of the parameters of each layer in the G2, the estimated values of the parameters of each layer in the G2 are extracted, corresponding to the position of the multi-omics data input in the G1 model, input into the G1 to obtain the aggregated features, and then the predicted phenotype of the candidate population is obtained through the D1;
[0032] Alternatively, the candidate population genotype data is input into the G2 to obtain aggregated features, and then input into the D1 to obtain the predicted phenotype of the candidate population.
[0033] In a fourth aspect, some embodiments of the present application further provide a method for sample expansion, wherein the GAN model constructed using the above method is used to perform sample expansion, comprising the following steps:
[0034] Obtain fitted genotype data, input the fitted genotype data into G2, obtain estimated values of parameters at each layer in G2, extract estimated values of parameters at each layer in G2, corresponding to the position where multi-omics data is input in the G1 model, input into G1 to obtain aggregated features, and then obtain the predicted phenotype corresponding to the fitted genotype data through D1;
[0035] Alternatively, the obtained fitted genotype data is input into the G2 to obtain aggregated features, and then into the D1 to obtain the predicted phenotype corresponding to the fitted genotype data.
[0036] In a fifth aspect, some embodiments of the present application further provide a breeding method, which uses the GAN model constructed by the method described above to perform phenotype prediction, and uses the obtained predicted phenotype for breeding; the predicted phenotype is obtained by the method described above.
[0037] In a sixth aspect, some embodiments of the present application further provide a computer device comprising: one or more processors; and a memory storing computer program instructions, wherein the computer program instructions, when executed, cause the processor to execute the method described above.
[0038] In a seventh aspect, some embodiments of the present application further provide a computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method as described above. Beneficial effects
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] First, GAN focuses on the genome selection model, where the generator continuously fits the characteristics of real data, and the discriminator continuously distinguishes true from false data. Through the continuous adversarial learning between the generator and the discriminator, the generator achieves the goal of generating data that meets the expected data characteristics.
[0041] Second, to address the drawbacks of high cost of measuring multi-omics data and difficulty in obtaining them in actual breeding experiments, the generative adversarial genomic selection system of the present invention can utilize a training population with measured multi-omics data, including genome, epigenome, transcriptome, proteome, metabolome, phenotype, etc., to construct a multi-omics data generator, autonomously learn and fit multi-omics data features, and generate multi-omics data features for candidate populations.
[0042] 3. To address the defect that the number of training populations is far lower than the number of molecular markers, the generative adversarial genomic selection system of the present invention can use the multi-omics data of the training population to construct a population data generator, autonomously learn and fit the association between multi-omics data features and phenotypes, and expand the candidate population. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] FIG1 is a schematic diagram of the architecture of a GAN model provided in an embodiment of the present application;
[0044] FIG2 is the experimental data of phenotype prediction provided in the examples of this application. Modes for Carrying Out the Invention
[0045] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0046] The following terms are used in this article: GAN (Generative Adversarial Network), a generative adversarial network, is a deep generative model based on adversarial learning; LASSO (Least Absolute Shrinkage and Selection Operator); SVR (Support Vector Regression); SVC (Support Vector Classification); RFC (Random Forest Classification); PCC (Pearson correlation coefficient); SNP (Single Nucleotide Polymorphism); AA: represents the AA allele; AT: represents the AT allele; TT: represents the TT allele.
[0047] Example 1
[0048] A method for constructing a GAN model, the GAN model includes a first generator G1, a second generator G2, a first discriminator D1, and a second discriminator D2. As shown in Figure 1, it is a schematic diagram of the architecture of the GAN model of the present invention. The GAN model consists of four parts: a first generator G1, a second generator G2, a first discriminator D1, and a second discriminator D2. The first generator G1 is responsible for multi-omics data feature extraction, mainly including receiving real multi-omics data as input, and outputting the extracted aggregated features through processing of a multi-layer neural network. The second generator G2 is responsible for multi-omics data feature fitting, mainly including receiving fitted genotype data as input, and outputting the fitted aggregated features through processing of a multi-layer neural network. Among them, the fitted genotype data is a random value with the same dimension as the real genotype data, and the value space of the random value is the value space of the real genotype. In addition, the structure and basic parameters of the multi-layer neural network of G2 are consistent with those of G1. The first discriminator D1 is responsible for predicting the phenotype using the extracted information features. This mainly includes receiving the aggregated features extracted from the first generator G1 or the aggregated features fitted from the second generator G2 as input, processing them through a multi-layer neural network, and outputting the predicted phenotypic data. The second discriminator D2 is responsible for determining whether the data comes from the first generator G1 (real data) or the second generator G2 (fitted data). This mainly includes receiving the aggregated features extracted from the first generator G1 or the aggregated features fitted from the second generator G2 as input, processing them through a multi-layer neural network, and outputting the discriminant information.
[0049] The method for constructing the GAN model of this embodiment includes the following steps:
[0050] Step S101, obtaining real multi-omics data, said real multi-omics data including genotype data, optionally including multi-omics data and real phenotypic data real The multi-omics data includes at least one of the epigenomic data, transcriptomic data, proteomic data, metabolomic data, and functional group data of the target species.
[0051] Obtain fitted genotype data, where the fitted genotype data is a set of random values with the same dimension as the genotype data; the label value of the real multi-omics data is True, and the label value of the fitted genotype data is False;
[0052] Step S102: input the real multi-omics data into the first generator G1, and input the data features aggregated by the first generator G1 into the first discriminator D1 to obtain the predicted phenotype data y pre , based on the predicted phenotypic data y pre and the real phenotypic data y real The difference between the two constructs the loss function training to obtain the trained first generator G1 and the first discriminator D1;
[0053] Step S103, initially constructing the network structure and basic parameters of the second generator G2 to be consistent with the trained first generator G1; inputting the real multi-omics data into the second discriminator D2 through the aggregation features of the first generator G1 to obtain a first judgment result, inputting the fitted genotype data into the second discriminator D2 through the aggregation features of the second generator G2 to obtain a second judgment result, updating the second discriminator D2 parameters based on the loss function value when the first judgment result is True and the second judgment result is False, and updating the second generator G2 parameters based on the loss function value when the second judgment result is True, and performing adversarial training until the second discriminator D2 cannot correctly distinguish whether the input is true or false, thereby obtaining the trained second generator G2 and second discriminator D2.
[0054] Specifically, in step S101, the currently commonly used and widely used multi-omics databases include cbioPortal, ISwine, FAANG, GWAS Atlas, IAnimal, iSMOD, CottonMD, Teabase, SoyOmics and other databases.
[0055] In some embodiments of the present application,
[0056] Real multi-omics data generator G1:
[0057] Based on the structure of multi-omics data, a deep neural network is constructed, starting from the bottom-level genotype data, extracting data features layer by layer until the output of the top-level aggregated features.
[0058] Generally speaking, there should be at least one input layer and one output layer to meet the input requirements of the minimum genotype data.
[0059] The real multi-omics data generator, that is, the first generator G1, must input genotype data, defined as real_g1; followed by multi-omics data that can be optionally input, such as epigenomic data real_e1, transcriptomic data real_t1, proteomic data real_p1, metabolomic data real_m1, and functional group data real_f1.
[0060] The output function for receiving input data (genotype data is required, other data can be input selectively or all data can be input) can be: F(x) = G1(x| real_g1),
[0061] Or F(x) = G1(x| real_g1, real_e1), F(x) = G1(x| real_g1, real_e1, real_t1, real_p1, real_m1, real_f1), etc.
[0062] The main purpose of F(x) in this step is to cooperate with the D1 generator below to confirm the multi-omics features that need to be fitted, which can also be understood as quantitative multi-omics data features.
[0063] Fitting the multi-omics data generator G2:
[0064] Fitting the multi-omics data generator, that is, the input data of the second generator G2 is a set of random values fake_g1 of the same dimension as real_g1, and its random space is the value space of real_g1;
[0065] There are several ways to encode the real genotype data, such as [-1, 0, 1] or [0, 1, 2], etc.
[0066] If for a sample, real_g1 has 10 SNPs, for the first encoding case [-1, 0, 1], the 10 values of real_g1 are these three values, and therefore fake_g1 is also randomly selected from these three values.
[0067] The network structure and basic parameters of G2 are consistent with those of the real multi-omics data generator G1. This is primarily to ensure that the generated multi-omics data conform to the same data distribution as the real multi-omics data. This design ensures that, after the entire process is complete, the data generated by G2 will conform to the characteristics of the G1 data. Since each layer of G1 is designed for each layer of multi-omics data, each layer of G2 also corresponds to a layer of fitted multi-omics data.
[0068] The network structure and basic parameters of G2 are consistent with those of the real multi-omics data generator G1. This is primarily to ensure that the generated multi-omics data conform to the same data distribution as the real multi-omics data. This design ensures that, after the entire process is complete, the data generated by G2 will conform to the characteristics of the G1 data. Since each layer of G1 is designed for each layer of multi-omics data, each layer of G2 also corresponds to a layer of fitted multi-omics data.
[0069] Once the model training is completed, the model can be backtracked to extract the value of each layer to correspond one-to-one to the multi-omics data of each layer.
[0070] Because the random genotype data input by G2 is actually a hypothetical sample, multi-omics data can be generated based on this sample.
[0071] Because the outputs of G1 and G2 are consistent, and the model frameworks of G1 and G2 are consistent, for G1, the received multi-omics data are real multi-omics data, so the data of each layer obtained by G2 can correspond to the multi-omics data of G2 samples.
[0072] The main purpose of F(x) in this step is to cooperate with the previous first generator G1 and the following second discriminator D2 to generate multi-omics data that conforms to the real feature distribution of real data.
[0073] The output function for receiving input data (only random noise, which can be regarded as the genotype data of a certain sample) is: F(x) = G2(x| fake_g1)
[0074] The F(x) in this step is to cooperate with the D2 discriminator below to achieve the purpose of generating the same distribution as the real data.
[0075] True phenotype discriminator D1
[0076] Based on the aggregated features extracted by the real multi-omics data generator, a deep neural network is constructed to summarize the features layer by layer until the top-level phenotypic values are output.
[0077] Generally speaking, it should include at least 1 input layer and 1 output layer.
[0078] The true phenotype discriminator, that is, the input data of the first generator D1 is the aggregated data features generated by G1, and the output is defined as: F(x) = D1(x| G1)
[0079] The main purpose of F(x) in this step is to fix the real multi-omics data characteristics.
[0080] The specific process is called phenotype prediction, which means that for the multi-omics data input into G1, the sample phenotype (y real ) output. Therefore, the goal is to get as close to the real phenotype as possible.
[0081] The loss function is defined by the mean absolute error: L(x) = |D1(x|G1) - y real |
[0082] Fitting the multi-omics data discriminator D2
[0083] Fitting multi-omics data discrimination, that is, the second discriminator D2 constructs a deep neural network based on the aggregated features extracted by the fitting multi-omics data generator, and summarizes the features layer by layer until the top-level discriminant information is output.
[0084] Generally speaking, it should include at least 1 input layer and 1 output layer.
[0085] The discriminator D2 in this step is the discriminator of the classic GAN. Its main purpose is to ensure that the output of G2 is consistent with the output of G1. That is, the discriminator cannot determine whether the data comes from real data or fitted data, thus achieving a fake-real effect, that is, generating data that conforms to the characteristics of real data.
[0086] The received data comes from the input of G1 and the input of G2, and the output data is the probability of judging whether it is true or not:
[0087] F(x) = D2(x|G1), the data probability range is [0,1]
[0088] F(x) = D2(x|G2), the data probability range is [0,1]
[0089] The main process is as follows:
[0090] Receive the input of the discriminator G1 (the data distribution is replaced by P) and judge it to be true. The loss function is defined by the cross loss entropy:
[0091] The cross loss entropy is defined as:
[0092] The loss function is defined as:
[0093] Receive the input of the discriminator G2 (the data distribution is replaced by P) and judge it to be false. The loss function is defined by the cross loss entropy:
[0094] The loss function is defined as:
[0095] Receive the input of the discriminator G2 (the data distribution is replaced by P) and judge it to be true. The loss function is defined by the cross loss entropy:
[0096] The loss function is defined as:
[0097] The loss function of the entire discriminator D2 can be defined as the minimization and maximization function V(G2, D2):
[0098] The loss function of the entire generator G2 can be defined as the minimization valuation function V(G2):
[0099] in, The output of G1 processing the real multi-omics data, The output of the fitted genotype data is processed for the G2.
[0100] The training process is as follows:
[0101] Step A: Randomly select a set of random variables with the same dimensions as the true genotype data from the specified data distribution;
[0102] Step B: Use G2 to receive the random variables generated by A, fit the data features, and label them as False;
[0103] Step C: Select a certain number of samples from the real data, use G1 to obtain the real data features, and label them as True;
[0104] Step D: Use steps B and C to train D2 according to the loss function defined by V(G2, D2);
[0105] Step E: Generate a set of random variables again according to step A, define the label as True, and train G2 according to the loss function defined by V(G2);
[0106] Step F: Repeat the above process according to the specified number of steps.
[0107] Generally speaking, when G2 is updated, D2 is fixed, and when D2 is updated, G2 is fixed.
[0108] Example 2
[0109] The GAN model constructed using the above construction method is used to fit multi-omics data, including the following steps: inputting the genotype data of the candidate population into the G2, obtaining the estimated values of the parameters of each layer in the G2, extracting the estimated values of the parameters of each layer in the G2, corresponding to the position where the multi-omics data is input in the G1 model, and realizing the fitting of the multi-omics data.
[0110] Example 3
[0111] The GAN model constructed using the above construction method is used for phenotype prediction, including the following methods: inputting the multi-omics data of the candidate population into the G1, and inputting the data features aggregated by the G1 into the D1 to obtain the predicted phenotype of the candidate population; or, inputting the genotype data of the candidate population into the G2, obtaining the estimated values of the parameters of each layer in the G2, extracting the estimated values of the parameters of each layer in the G2, corresponding to the position where the multi-omics data is input in the G1 model, inputting the aggregated features into the G1, and then obtaining the predicted phenotype of the candidate population through the D1; or, inputting the genotype data of the candidate population into the G2 to obtain the aggregated features and entering them into the D1 to obtain the predicted phenotype of the candidate population.
[0112] Figure 2 shows the experimental data for phenotype prediction provided in the examples of this application. The left figure is a bar chart showing the relative improvement (%) of the GAN model compared to four other models, and the right figure is a scatter plot of the GAN model's predicted and fitted phenotypes. The left figure shows that for the same sample batch, the inventors evaluated the performance by calculating the Pearson correlation coefficient (PCC) between the predicted phenotypes obtained by the model and the true phenotypes. Among all correlation coefficient calculation methods, the most common in the biological field is the Pearson correlation coefficient, also known as the Pearson product-moment correlation coefficient. This is a linear correlation coefficient that measures the degree of linear correlation between two variables. The performance evaluation results show that compared to the least absolute shrinkage and selection operator (LASSO), support vector regression (SVR), support vector classification (SVC), and random forest classification (RFC), the GAN model improves on the PCC metric by 28%, 10%, 54%, and 20%, respectively, leveraging the advantages of deep learning algorithms to improve the accuracy of genomic selection. The right figure shows that for the same set of genotype data, the correlation coefficient (PCC) between the predicted phenotype obtained by the first discriminator D1 receiving the aggregated features extracted from the first generator G1 and the predicted phenotype obtained by the first discriminator D1 receiving the aggregated features fitted from the second generator G2 can reach 0.5, which can make up for the lack of data volume.
[0113] Example 4
[0114] The GAN model constructed by the above method is used for sample expansion: fitting genotype data is obtained, the fitting genotype data is input into the G2, the estimated values of the parameters of each layer in the G2 are obtained, the estimated values of the parameters of each layer in the G2 are extracted, the positions corresponding to the multi-omics data input in the G1 model are input into the G1 to obtain the aggregated features, and then the predicted phenotype corresponding to the fitting genotype data is obtained through the D1; alternatively, the fitting genotype data is obtained and input into the G2 to obtain the aggregated features and enter into the D1 to obtain the predicted phenotype corresponding to the fitting genotype data.
[0115] Example 5
[0116] The predicted phenotype obtained by the above method is used for breeding: the data based on the predicted phenotype can be used for breeding and provide a reference for breeding.
[0117] The predicted phenotype is obtained by a phenotype prediction method, including but not limited to the following methods: inputting the multi-omics data of the candidate population into the G1, inputting the data features aggregated by the G1 into the D1, and obtaining the predicted phenotype of the candidate population; or, inputting the genotype data of the candidate population into the G2, obtaining the estimated values of the parameters of each layer in G2, extracting the estimated values of the parameters of each layer in the G2, corresponding to the position where the multi-omics data is input in the G1 model, inputting the aggregated features into the G1, and then obtaining the predicted phenotype of the candidate population through the D1; or, inputting the genotype data of the candidate population into the G2 to obtain the aggregated features and entering them into the D1 to obtain the predicted phenotype of the candidate population.
[0118] In summary, this application builds on the aforementioned GAN model. Based on the input of real multi-omics data, it enters G1, extracts features, and then enters D1 to train G1 and D1. Next, a random noise stream is generated and input to G2, along with real data input to G1. The outputs of both streams enter D2, training G2 and D2. This approach can at least compensate for insufficient data while fully leveraging the advantages of deep learning algorithms to comprehensively improve the accuracy of genomic selection.
[0119] Specifically, the following technical effects can be achieved through the embodiments of the present invention:
[0120] 1. Fitting of Real Multi-omics Data
[0121] Through the real multi-omics data generator G1 and the real phenotype discriminator D1, the multi-omics data of the training population and the phenotype association are learned. Subsequently, the real multi-omics data is fitted by fitting the multi-omics data generator G2 and fitting the multi-omics data discriminator D2.
[0122] Input a set of data with known genotypes but unknown multi-omics data, obtain the estimation of the parameters of each layer of G2 in the model through G2, extract the parameters of each layer in G2, correspond to the position of multi-omics data input in the G1 model, and realize the fitting of multi-omics data.
[0123] 2. Prediction of Candidate Groups
[0124] On the basis of implementation one, by inputting the genotype data of the candidate population, the multi-omics data is fitted by fitting the multi-omics data generator, and the phenotype of the candidate population is predicted by the true phenotype discriminator.
[0125] This implementation scheme can measure across the multi-omics data of the candidate population, and through the characteristics of the multi-omics data of the candidate population and according to the genotype of the candidate population, fit and generate a set of outputs that conform to the aggregated characteristics of the multi-omics data.
[0126] If the multi-omics data of this sample are known: the input multi-omics data passes through G1 and enters D1 to realize phenotype prediction.
[0127] If the multi-omics data of this sample is unknown, but the genotype is definitely known: the input genotype data passes through G2 and enters D1 to realize the phenotypic prediction.
[0128] 3. Expansion of the training group
[0129] Based on the first implementation, we input random noise and fit the multi-omics data generator to fit the multi-omics data. Then, we use the true phenotype discriminator to fill in the expanded population phenotype. By extracting the specific parameters of each layer of the multi-omics data generator network, we can extract the multi-omics data.
[0130] A random noise set is randomly generated (see the description in Section 2. Fitting the Multi-omics Data Generator G2) and fed into G2. G2 then estimates the parameters of each layer of the model. The parameters of each layer in G2 are extracted, corresponding to the position of the multi-omics data input in the G1 model. These parameters are then fed into G1 to obtain aggregated features, which are then fed into D1 to obtain predicted phenotypic values. Alternatively, a random noise set is randomly generated and fed into G2 to obtain aggregated features, which are then fed into D1 to obtain predicted phenotypic values. This provides both the sample genotypes and phenotypic values, thus expanding the training population.
[0131] 4. The above-mentioned predicted phenotypic data can be used for breeding.
[0132] It is obvious to those skilled in the art that the present application is not limited to the details of the above-mentioned exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the present application is defined by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and scope of the equivalent elements of the claims are included in the present application. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim can also be implemented by one unit or device through software or hardware. Words such as first and second are used to indicate names and do not indicate any particular order.
Claims
1. A method for constructing a GAN model, characterized in that, the GAN model includes a first generator G1, a second generator G2, a first discriminator D1, and a second discriminator D2; The construction method includes: obtaining real multi-omics data and obtaining fitted genotype data, where the real multi-omics data includes genotype data and optionally includes at least one of epigenomic data, transcriptomic data, proteomic data, metabolomic data, and functional group data of the target species, as well as real phenotype data y real ; the fitted genotype data is a set of random values with the same dimension as the genotype data; the label value of the real multi-omics data is True, and the label value of the fitted genotype data is False; Input the real multi-omics data into the first generator G1, and input the data features aggregated by the first generator G1 into the first discriminator D1 to obtain the predicted phenotypic data y pre , based on the predicted phenotypic data y pre and the real phenotypic data y real construct a loss function for training to obtain the trained first generator G1 and first discriminator D1; The network structure and basic parameters of the initially constructed second generator G2 are kept consistent with those of the trained first generator G1; the real multi-omics data is input into the second discriminator D2 through the aggregated features of the first generator G1 to obtain a first determination result, and the fitted genotype data is input into the second discriminator D2 through the aggregated features of the second generator G2 to obtain a second determination result. The parameters of the second discriminator D2 are updated based on the loss function values where the first determination result is True and the second determination result is False. The parameters of the second generator G2 are updated based on the loss function value where the second determination result is True. After adversarial training until the second discriminator D2 cannot correctly distinguish true or false of the input, the trained second generator G2 and second discriminator D2 are obtained.
2. The method according to claim 1, characterized in that, Training is performed by constructing a loss function based on the difference between the predicted phenotypic data and the true phenotypic data. The loss function is defined by the mean absolute error, and the loss function L(x) = |D1(x|G1) - y real |.
3. The method according to claim 1, characterized in that, The loss function used in the adversarial training is defined by the cross-entropy loss, and the cross-entropy loss is defined as: ; When receiving the input of the first generator G1 and determining it to be true, the loss function ; Receives the input of the second generator G2 and determines it to be False, and the loss function ; Receives the input of the second generator G2 and determines it to be True, the loss function ; The loss function of the discriminator D2 is to minimize the minimax evaluation function V(G2, D2): ; The loss function of the generator G2 is defined as minimizing the valuation function V(G2): ; Among them, is the output of processing the real multi-omics data for the G1, is the output of G2 processing the fitted genotype data.
4. The method according to claim 3, characterized in that, the adversarial training steps are as follows: Step A: Randomly select a set of random variables with the same dimension as the real genotype data from a specified data distribution; Step B: Use G2 to receive the random variables generated in Step A, fit the data features, and the label is False; Step C: Select a certain number of samples from the real data, and use G1 to obtain the real data features, and the label is True; Step D: Use B and C to train D2 according to the cross-entropy loss function V(G2, D2); Step E: Regenerate a set of random variables according to A, and define the label as True, and train G2 according to the cross-entropy loss function V(G2); Step F: Repeat the above steps A - E according to the specified number of steps until the set condition is met and stop training.
5. A method for data fitting, characterized in that, using the GAN model constructed by the construction method described in any one of claims 1 - 4 for multi-omics data fitting, including the following steps: The candidate population genotype data is input into G2 to obtain the estimated values of the parameters of each layer in G2, and the estimated values of the parameters of each layer in G2 are extracted, corresponding to the positions where the multi-omics data is input in the G1 model, to achieve the fitting of the multi-omics data.
6. A method for phenotype prediction, characterized in that, using the GAN model constructed by the construction method described in any one of claims 1 - 4 for phenotype prediction, including the following steps: Input the candidate population multi-omics data into G1, and input the data features aggregated by G1 into D1 to obtain the predicted phenotype of the candidate population; Alternatively, input the candidate population genotype data into the G2 to obtain the estimated values of the parameters of each layer in the G2. Extract the estimated values of the parameters of each layer in the G2, corresponding to the positions where the multi-omics data is input in the G1 model, input them into the G1 to obtain aggregated features, and then obtain the predicted phenotypes of the candidate population through the D1. Alternatively, input the candidate population genotype data into the G2 to obtain aggregated features, and then enter the D1 to obtain the predicted phenotypes of the candidate population.
7. A method for sample augmentation Characterized in that Use the GAN model constructed by the method according to any one of claims 1-4 for sample augmentation, including the following steps: Obtain fitted genotype data, input the fitted genotype data into the G2 to obtain the estimated values of the parameters of each layer in the G2. Extract the estimated values of the parameters of each layer in the G2, corresponding to the positions where the multi-omics data is input in the G1 model, input them into the G1 to obtain aggregated features, and then obtain the predicted phenotypes corresponding to the fitted genotype data through the D1. Alternatively, input the obtained fitted genotype data into the G2 to obtain aggregated features, and enter the D1 to obtain the predicted phenotypes corresponding to the fitted genotype data.
8. A breeding method Characterized in that Use the GAN model constructed by the method according to any one of claims 1-4 for phenotype prediction, and use the obtained predicted phenotypes for breeding; the predicted phenotypes are obtained by the method according to claim 6 or 7.
9. A computer device Characterized in that The device includes: One or more processors; and A memory storing computer program instructions, and when the computer program instructions are executed, the processors execute the method according to any one of claims 1-7.
10. A computer-readable medium having computer program instructions stored thereon, and the computer program instructions can be executed by a processor to implement the method according to any one of claims 1-7.
Citation Information
Patent Citations
Method for judging planting value of seed group
CN111627495A
Method for estimating breeding value by fitting genome with non-additive effect
CN111883206A
Breeding method, device and equipment based on multi-omics data and deep learning
CN114743601A
Soft measurement virtual modeling data generation method based on double adversarial learning
CN115935785A
Cited By
Joint angle prediction method, system and equipment for infant crawling and medium
CN120600226A