Plant breeding method and device based on two-channel convolutional neural network
The gene prediction model constructed using a dual-channel convolutional neural network solves the efficiency and accuracy problems of random regression models in processing high-dimensional genotype data, enabling early and accurate prediction of the content of specified compounds and improving the efficiency and accuracy of plant breeding.
Patent Information
- Application Number
- CN202511108857.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-12-19
AI Technical Summary
In existing technologies, the optimal linear unbiased prediction model of random regression is inefficient in processing high-dimensional genotype data and has low prediction accuracy for complex traits. It is difficult to improve the prediction efficiency and accuracy of the content of a specified compound, which affects the feasibility of whole-genome selection breeding in plants.
By employing a dual-channel convolutional neural network approach, a general feature covering the entire genome and prior features corresponding to different specified compounds are constructed. Feature extraction and fusion are then performed using a pre-built gene prediction model to capture nonlinear genetic effects and improve the prediction accuracy of the content of specified compounds.
It achieves accurate prediction in the early stages of target plants, reduces experimental cycle and resource consumption, improves the prediction efficiency and accuracy of specified compound content, and provides a feasible strategy for whole-genome selection breeding of plants.
Smart Images

Figure CN121171332A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of biological breeding, and in particular to a plant breeding method and device based on a double-channel convolutional neural network. BACKGROUND
[0002] Genomic Selection (GS) refers to constructing a genetic prediction model by using whole genome markers to predict the genetic potential of a target trait. Moreover, since the core hypothesis of genomic selection is that there is linkage disequilibrium (LD) between single nucleotide polymorphisms and quantitative trait nucleotide (QTL), it is suitable for predicting complex quantitative traits. Further, since the traditional breeding cycle is long and inefficient, genomic selection can significantly accelerate genetic gain through early prediction in the seedling stage, so this method is widely used in plant and animal breeding.
[0003] To improve the prediction performance of the genetic prediction model, various modeling methods have been proposed. Among them, random regression best linear unbiased prediction model (rrBLUP) is the most typical traditional statistical model in genetic prediction model. Although the best linear unbiased prediction model has many advantages such as high efficiency and stability, the random regression best linear unbiased prediction model also has some disadvantages. For example, the random regression best linear unbiased prediction model is not suitable for processing high-dimensional genotype data, and the prediction accuracy for complex traits is low.
[0004] In summary, how to improve the prediction efficiency and accuracy of the content of the specified compound, and then provide a feasible strategy for plant genomic selection breeding related problems need to be solved.
[0005] In view of the technical problems of the prior art that how to improve the prediction efficiency and accuracy of the content of the specified compound, and then provide a feasible strategy for plant genomic selection breeding related problems need to be solved, no effective solution has been proposed so far. SUMMARY
[0006] Embodiments of the present disclosure provide a plant breeding method and device based on a double-channel convolutional neural network to at least solve the technical problems of the prior art that how to improve the prediction efficiency and accuracy of the content of the specified compound, and then provide a feasible strategy for plant genomic selection breeding related problems need to be solved.
[0007] According to an aspect of the embodiments of the present disclosure, there is provided a plant breeding method based on a double-channel convolutional neural network, comprising: determining a target plant, a plurality of specified compounds in the target plant, and a plurality of SNP sites in the specified compounds; from the plurality of SNP sites, screening a plurality of first tag SNP sites for covering a whole genome, and determining each first tag SNP site as a general feature for inputting into a general information channel of each gene prediction model, the model architecture of the general information channel being a convolutional neural network; from the plurality of SNP sites, screening first candidate SNP sites corresponding to different specified compounds, and determining the first candidate SNP sites corresponding to different specified compounds as prior features for inputting into a prior information channel of each gene prediction model, wherein each gene prediction model corresponds to a different compound, and the model architecture of the prior information channel is a convolutional neural network; and based on the general features, each prior feature, and each gene prediction model, predicting the content corresponding to the specified compounds, and performing breeding feasibility analysis on the target plant according to the predicted content.
[0008] According to another aspect of the embodiments of the present disclosure, there is also provided a storage medium comprising a stored program, wherein the program is executed by a processor when the program is run.
[0009] According to another aspect of the embodiments of the present disclosure, there is also provided a plant breeding device based on a double-channel convolutional neural network, comprising: a target plant determination module for determining a target plant, a plurality of specified compounds in the target plant, and a plurality of SNP sites in the specified compounds; a general feature determination module for screening a plurality of first tag SNP sites for covering a whole genome from the plurality of SNP sites, and determining each first tag SNP site as a general feature for inputting into a general information channel of each gene prediction model, the model architecture of the general information channel being a convolutional neural network; a prior feature determination module for screening first candidate SNP sites corresponding to different specified compounds from the plurality of SNP sites, and determining the first candidate SNP sites corresponding to different specified compounds as prior features for inputting into a prior information channel of each gene prediction model, wherein each gene prediction model corresponds to a different compound, and the model architecture of the prior information channel is a convolutional neural network; and a content prediction module for predicting the content corresponding to the specified compounds based on the general features, each prior feature, and each gene prediction model, and performing breeding feasibility analysis on the target plant according to the predicted content.
[0010] According to another aspect of the embodiments of the present disclosure, a plant breeding device based on a double-channel convolutional neural network is also provided, comprising: a processor; and a memory connected with the processor, configured to provide the processor with instructions to process the following processing steps: determining a target plant, a plurality of specified compounds in the target plant, and a plurality of SNP sites in the specified compounds; screening a plurality of first tag SNP sites from the plurality of SNP sites for covering a whole genome, and determining each first tag SNP site as a general feature for inputting into a general information channel of each gene prediction model, the model architecture of the general information channel being a convolutional neural network; screening first candidate SNP sites corresponding to different specified compounds from the plurality of SNP sites, and determining the first candidate SNP sites corresponding to different specified compounds as prior features for inputting into a prior information channel of each gene prediction model, wherein each gene prediction model corresponds to a different compound, and the model architecture of the prior information channel is a convolutional neural network; and predicting the content corresponding to the specified compounds based on the general features, each prior feature, and each gene prediction model, and performing breeding feasibility analysis on the target plant according to the predicted content.
[0011] The present application provides a plant breeding method based on a double-channel convolutional neural network. Unlike the existing genome selection method based on traditional statistical models, in the present application, the general features covering the whole genome and the prior features corresponding to different specified compounds are extracted and fused by using the pre-constructed gene prediction model, so as to effectively model the nonlinear genetic effect, capture genetic signals at different levels, and thus improve the accuracy of predicting the content of the specified compounds in the target plant.
[0012] Therefore, since the present application can accurately predict the content of the specified compounds in the target plant by using the gene prediction model, even in the early stage when the target plant has not been fully phenotyped, the target plant individuals meeting the requirements can be selected according to the content of the specified compounds. The selected target plant individuals are used to breed offspring or form an excellent core population, so as to realize phenotype selection in advance and reduce the test cycle and resource consumption.
[0013] Therefore, the present application can improve the prediction efficiency and accuracy of the content of the specified compounds. Thus, the technical problem of how to improve the prediction efficiency and accuracy of the content of the specified compounds in the prior art is solved, and a feasible strategy for plant whole genome selection breeding is provided. BRIEF DESCRIPTION OF DRAWINGS
[0014] The accompanying drawings, which are included to provide a further understanding of the present disclosure and constitute a part of this application, illustrate certain illustrative embodiments of the present disclosure and together with the description, serve to explain the present disclosure. In the drawings:
[0015] Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to Embodiment 1 of the present disclosure;
[0016] Figure 2 is a schematic diagram of a system for plant breeding based on a double-channel convolutional neural network according to Embodiment 1 of the present disclosure;
[0017] Figure 3 is a modular schematic diagram of a gene prediction model according to Embodiment 1 of the present disclosure;
[0018] Figure 4 is a flowchart of a method for plant breeding based on a double-channel convolutional neural network according to Embodiment 1 of the present disclosure;
[0019] Figure 5 is a visual diagram of 12 chromosomes of Litsea cubeba according to Embodiment 1 of the present disclosure;
[0020] Figure 6 is a specific architecture diagram of a gene prediction model according to Embodiment 1 of the present disclosure;
[0021] Figure 7 is a schematic diagram of the distribution of the contents of 8 terpene compounds according to Embodiment 1 of the present disclosure;
[0022] Figure 8 is a principal component analysis diagram of a plurality of first plant samples and a plurality of second plant samples according to Embodiment 1 of the present disclosure;
[0023] Figure 9 is a distribution density diagram of second tag SNP site samples under different marker densities according to Embodiment 1 of the present disclosure;
[0024] Figure 10 is a schematic diagram of the prediction ability of the rrBLUP model corresponding to different orders of magnitude of second tag SNP site samples according to Embodiment 1 of the present disclosure;
[0025] Figure 11 is a comparison diagram of the prediction accuracy between the gene prediction model and the rrBLUP model on various specified compounds according to Embodiment 1 of the present disclosure;
[0026] Figure 12 is a schematic diagram of a plant breeding device based on a double-channel convolutional neural network according to Embodiment 2 of the present disclosure; and
[0027] Figure 13 is a schematic diagram of a plant breeding device based on a dual-channel convolutional neural network according to Embodiment 3 of the present application. DETAILED DESCRIPTION
[0028] In order to enable persons skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, not all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by persons skilled in the art without creative labor should fall within the scope of protection of the present disclosure.
[0029] It should be noted that the terms "first", "second", and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0030] Embodiment 1
[0031] According to the present embodiment, a method embodiment of plant breeding based on a dual-channel convolutional neural network is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that herein.
[0032] The method embodiment provided in the present embodiment can be executed in a mobile terminal, a computer terminal, a server, or a similar computing device. Figure 1 A hardware structure block diagram of a computing device for implementing a method of plant breeding based on a dual-channel convolutional neural network is shown. As Figure 1As shown, the computing device can include one or more processors (which can include, but are not limited to, processing devices such as microprocessors, MCUs, or programmable logic devices, FPGAs, etc.), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, the transmission device, and the input / output interface are connected to the processor through a bus. In addition, it can also include a display, a keyboard, and a cursor control device connected to the input / output interface. Those skilled in the art can understand that Figure 1 The structure shown is only a schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computing device can include more or fewer components than those shown in Figure 1 or have a different configuration than that shown in Figure 1 .
[0033] It should be noted that the one or more processors and / or other data processing circuits described above can be referred to herein generally as "data processing circuits". The data processing circuits can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuits can be a single independent processing module, or all or part of any one of the other elements incorporated into the computing device. As referred to in the embodiments of the present disclosure, the data processing circuits serve as a processor to control, for example, the selection of the variable resistance terminal path connected to the interface.
[0034] The memory can be used to store software programs and modules of application software, such as program instructions / data storage means corresponding to the plant breeding method based on a double-channel convolutional neural network in the embodiments of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, i.e., implements the plant breeding method based on a double-channel convolutional neural network of the application program described above. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely disposed relative to the processor, which can be connected to the computing device through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0035] The transmission device is configured to receive or send data via a network. The network can include, for example, a wireless network provided by a communication provider of the computing device. In one example, the transmission device includes a network interface controller (NIC) that can connect to other network devices through a base station to communicate with the Internet. In one example, the transmission device can be a radio frequency (RF) module that is configured to communicate with the Internet via wireless communication.
[0036] The display can be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with the user interface of the computing device.
[0037] It should be noted that, in some optional embodiments, the above-mentioned Figure 1 The computing device can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or a combination of both hardware and software elements. It should be noted that Figure 1 is merely one example of a particular implementation and is intended to illustrate the types of components that can be present in the computing device.
[0038] Figure 2 is a schematic diagram of a system for plant breeding based on a dual-channel convolutional neural network according to embodiments of the present application. Referring to Figure 2 As shown in the figure, the system includes a terminal device 100 and a processor 200. The terminal device 100 is in communication connection with the processor 200, so that a breeder can send an instruction for analyzing the breeding feasibility of a target plant to the processor 200 through the terminal device 100. The processor 200 can respond to the instruction by predicting the content corresponding to various specified compounds using various gene prediction models, and analyzing the breeding feasibility of the target plant and providing a feasibility analysis strategy according to the predicted content corresponding to various specified compounds.
[0039] Figure 3 is a modular schematic diagram of a gene prediction model according to embodiments of the present application. Referring to Figure 3 As shown in the figure, the gene prediction model includes a feature extraction module, a feature fusion module, a feature integration module, and an output module.
[0040] The feature extraction module is in communication connection with the feature fusion module, and is configured to extract features of general features in the general information channel by using a first convolutional neural network, and is further configured to extract features of prior features in the prior information channel by using a second convolutional neural network. The feature fusion module is in communication connection with the feature integration module, and is configured to fill and splice the first feature map corresponding to the general information channel and the second feature map corresponding to the prior information channel output by the feature extraction module, and generate a third feature map. The feature integration module is in communication connection with the output module, and is configured to compress the number of channels and integrate the general information and the prior information in the third feature map, and generate a feature vector. The output module is configured to perform layer-by-layer nonlinear mapping on the feature vector by using a multi-layer fully connected network, and output a predicted content corresponding to the specified compound.
[0041] It should be noted that the terminal device 100 and the processor 200 in the system can be applicable to the hardware structure described above.
[0042] Under the above operating environment, according to a first aspect of the embodiment, a plant breeding method based on a double-channel convolutional neural network is provided, which is realized by the processor 200 shown in Figure 2 . Figure 4 The flowchart of the method is shown, and the method includes: Figure 4
[0043] S402: determining a target plant, a plurality of specified compounds in the target plant, and a plurality of SNP sites in each specified compound;
[0044] S404: selecting a plurality of first tag SNP sites for covering the whole genome from the plurality of SNP sites, and determining each first tag SNP site as a general feature for inputting into a general information channel of each gene prediction model, and the model architecture of the general information channel is a convolutional neural network;
[0045] S406: selecting a first candidate SNP site corresponding to different specified compounds from the plurality of SNP sites, and determining the first candidate SNP site corresponding to different specified compounds as a prior feature for inputting into a prior information channel of each gene prediction model, wherein each gene prediction model corresponds to a different compound, and the model architecture of the prior information channel is a convolutional neural network; and
[0046] S408: predicting the content corresponding to each specified compound based on the general features, each prior feature, and each gene prediction model, and performing breeding feasibility analysis on the target plant according to the predicted content.
[0047] Specifically, first, the breeder determines a target plant, a plurality of specified compounds in the target plant, and a plurality of first SNP sites of the various specified compounds through the terminal device 100 (S402). Among them, since the fruits of Litsea cubeba contain rich terpenes, especially monoterpenes, which account for more than 90% of the volatile components in fresh fruits, the lemon-like aroma rich in myrcia and neral in monoterpenes makes Litsea cubeba have important application value in essential oils, cosmetics, and food flavorings. However, Litsea cubeba also contains a small amount of caryophyllene with a pungent odor, which is not conducive to the quality of essential oils, so in this embodiment, Litsea cubeba germplasm resources are taken as the target plant, and the breeding purpose includes reducing the content of caryophyllene and increasing the content of citral. Those skilled in the art should note that the above is only an example of Litsea cubeba to illustrate the technical solutions disclosed in the present application, and the actual application situation is not limited thereto.
[0048] Thus, in the case where the breeder determines to take Litsea cubeba germplasm resources as the target plant, the breeder determines 8 main terpenes in Litsea cubeba germplasm resources that affect the content of essential oils through the terminal device 100. Among them, the 8 main terpenes (i.e., a plurality of specified compounds) include citral, myrcia, neral, limonene, a-pinene, β- phellandrene, β-caryophyllene, and β-pinene. And among them, the 8 main terpenes in the above Litsea cubeba germplasm resources that affect the content of essential oils can be pre-stored to the terminal device 100, thereby facilitating calling.
[0049] Figure 5 is a visual map of 12 chromosomes of Litsea cubeba according to the embodiments of the present application. As shown in Figure 5 The breeder extracts the genome of the Litsea cubeba germplasm resources using the improved cetyltrimethylammonium bromide method (CTAB method). Then, the breeder performs sequencing through the terminal device 100 and using the DNBSEQ-T7 platform, and the average sequencing depth is set to 10x. The raw sequencing data is subjected to quality control using Trimmomatic, and after alignment to the Litsea cubeba reference genome by BWA software, single nucleotide polymorphism detection (SNP) and InDel detection are performed using GATK software, and filtering parameters (such as QD<2.0, MQ<40, etc.) are set to eliminate low-quality sites. Thus, the breeder can determine a plurality of SNP sites in the 8 main terpenes (i.e., specified compounds) through the terminal device 100. Among them, Figure 5 From the outside in, they are SNP site density, InDel density, and gene density.
[0050] Then, the breeder divides the LD blocks by performing LD analysis on the whole genome through the terminal device 100 and using the PLINK v1.9 software. Then, the second tag SNP sites that can represent the haplotype structure of the region are selected in each LD block. Since the number of the first tag SNP sites has been determined when the gene prediction model is trained, the breeder can select the optimal number of the first tag SNP sites (for example, 2000 tag SNP sites) from the plurality of SNP sites. The optimal number of the first tag SNP sites contains key genetic information and can meet the accuracy of prediction.
[0051] In the case where the terminal device 100 determines the plurality of first tag SNP sites, the plurality of first tag SNP sites are sent to the processor 200. The processor 200 determines the plurality of first tag SNP sites as general features of a general information channel for inputting into the plurality of gene prediction models (S404). The model architecture of the general information channel is a convolutional neural network.
[0052] Meanwhile, the breeder selects the first candidate SNP sites corresponding to different specified compounds from the plurality of SNP sites through the terminal device 100. Specifically, first, the breeder performs whole genome association analysis on 8 major terpene compounds in Litsea cubeba germplasm resources (i.e., target plants) through the terminal device 100 and using the GEMMA software, and determines the association analysis result. The association analysis result is used to indicate a plurality of second candidate SNP sites (for example, 886 significant sites) that reach a pre-set significance threshold in the plurality of SNP sites. Then, in order to improve the quality of prior information and reduce the false positive rate of the gene prediction model, the breeder reselects the plurality of second candidate SNP sites through the terminal device 100 to determine the first candidate SNP sites (for example, 49 high-quality candidate functional SNPs) corresponding to different specified compounds. The above will be described in detail later, and therefore will not be described here.
[0053] In the case where the terminal device 100 determines the plurality of first candidate SNP sites, the plurality of first candidate SNP sites are sent to the processor 200. The processor 200 determines the plurality of first candidate SNP sites as prior features of a prior information channel for inputting into the plurality of gene prediction models (S406). The model architecture of the prior information channel is a convolutional neural network.
[0054] Finally, the processor 200 predicts the content corresponding to various specified compounds based on the general features, the prior features corresponding to the specified compounds, and the respective gene prediction models. Specifically, Figure 6 is a specific architecture diagram of the gene prediction model according to the embodiments of the present application. Referring to Figure 6As shown, in order to improve the whole genome selection prediction accuracy of the content of terpene compounds in Litsea cubeba, the application provides a double-channel convolutional neural network prediction framework (i.e., a gene prediction model). The gene prediction model is realized based on a PyTorch platform and integrates two different types of information channels: one is a general information channel for processing general features covering the whole genome, and the other is a prior information channel for processing important candidate functional sites.
[0055] In terms of input data, the input of the general information channel is a one-dimensional vector with a length of 2000, representing the genotype information (i.e., general features) of 2000 first label SNP sites. The input of the prior information channel is the candidate functional markers corresponding to the specified compound (i.e., prior features). When there is no significant SNP for the specified compound (such as β-pinene and α-terpineol), the gene prediction model will automatically close the prior information channel and only enable the general information channel to enhance the flexibility and generalization ability of the prediction model architecture.
[0056] And the processor 200 uses the first convolutional neural network corresponding to the communication information channel to extract features from the general features and outputs a first feature map; uses the second convolutional neural network corresponding to the prior information channel to extract features from the prior information and outputs a second feature map. Then the processor 200 uses the feature fusion module to fill, splice and convolve the first feature map and the second feature map, and generates a third feature map. Finally, the processor 200 inputs the feature vector corresponding to the third feature map into the fully connected network, thereby outputting the predicted content corresponding to the specified compound. The above content will be described in detail later, so it will not be described here.
[0057] In addition, it is worth noting that due to the fundamental conflict between the genetic mechanisms and content levels of various specified compounds, it is impossible to predict the eight terpene compounds by only one gene prediction model. Each specified compound has exclusive neural network parameters (weight matrix), and eight independent gene prediction models are packaged as files ending with.pth and do not interfere with each other when running. And each gene prediction model only loads the prior features of the specified compound when predicting, without loading the prior features of other specified compounds in the same target plant (such as the gene prediction model corresponding to β-pinene being unaware of the citral value).
[0058] And the skilled person in the art needs to note that the general features of the general information channel in the above gene prediction model are not limited to 2000 first label SNP sites. The application only determines 2000 first label SNP sites as the best choice for general features when taking Litsea cubeba germplasm resources as an example, and the actual situation is not limited to this. That is, for different plants, the order of magnitude of the first label SNP sites of the general features is also different.
[0059] Thus, in a case where the processor 200 determines to predict the content corresponding to each of the specified compounds by using the plurality of genetic prediction models, the breeding feasibility of the target plant can be analyzed according to the predicted content, and a breeding feasibility analysis report can be generated (S408). That is, a breeder can check the content corresponding to each of the specified compounds and the feasibility analysis report corresponding to the target plant through the terminal device 100.
[0060] As described in the background, in order to improve the prediction performance of the genetic prediction model, a plurality of modeling methods have been proposed in succession. Among them, the random regression best linear unbiased prediction model (rrBLUP) is the most typical traditional statistical model in the genetic prediction model. However, although the best linear unbiased prediction model has many advantages such as high efficiency and stability, the random regression best linear unbiased prediction model also has some disadvantages. For example, the random regression best linear unbiased prediction model is not suitable for processing high-dimensional genotype data, and the prediction accuracy for complex traits is low.
[0061] In summary, how to improve the prediction efficiency and accuracy of the content of the specified compound, and then provide a feasible strategy for plant whole genome selection breeding is a problem to be solved.
[0062] Therefore, the present application provides a plant breeding method based on a double-channel convolutional neural network. And unlike the existing genome selection method based on the traditional statistical model, in the present application, the pre-constructed genetic prediction model is used to extract and fuse the general features covering the whole genome and the prior features corresponding to different specified compounds, so as to effectively model the nonlinear genetic effect and capture genetic signals at different levels, thereby improving the accuracy of predicting the content of the specified compound in the target plant.
[0063] Therefore, since the present application can accurately predict the content of the specified compound in the target plant by using the genetic prediction model, even in the early stage when the target plant has not been fully phenotyped, the target plant individuals that meet the requirements can be selected according to the content of the specified compound. And the selected target plant individuals are used to breed offspring or form an excellent core population, so as to realize phenotype selection in advance and reduce the test cycle and resource consumption.
[0064] Thus, the present application can improve the prediction efficiency and accuracy of the content of the specified compound. And it solves the technical problem of how to improve the prediction efficiency and accuracy of the content of the specified compound, and then provides a feasible strategy for plant whole genome selection breeding.
[0065] Optionally, from the plurality of SNP sites, screening the first candidate SNP site corresponding to the different specified compound comprises: performing whole genome association analysis on various specified compounds in the target plant, and determining the association analysis result, wherein the association analysis result is used to indicate a plurality of second candidate SNP sites reaching a pre-set significance threshold in the plurality of SNP sites; and screening the first candidate SNP site corresponding to the different specified compound from the plurality of second candidate SNP sites.
[0066] Specifically, in the case that the breeder determines the plurality of SNP sites through the terminal device 100, the GEMMA software is further used to perform whole genome association analysis on the eight terpenes of the Litsea cubeba germplasm resources by using the Mixed Linear Model (MLM), and the association analysis result is determined. The association analysis result is used to indicate a plurality of second candidate SNP sites (for example, 886 candidate SNP sites reaching the significance threshold) reaching a pre-set significance threshold in the plurality of SNP sites. And the significance threshold is determined to be P<8.37×10 -7 .
[0067] In addition, when performing the whole genome association analysis on the eight terpenes, the kinship matrix is used as a random effect to correct the influence of the population structure. And the first three principal components (PCs) are used as fixed effect covariates to further control the systematic bias.
[0068] Further, in order to improve the quality of prior information and reduce the false positive rate of the gene prediction model, the plurality of second candidate SNP sites (that is, 886 candidate SNP sites reaching the significance threshold) obtained after the whole genome association analysis are further strictly filtered. Therefore, first, the processor 200 retains the gene sites related to the key metabolic pathways (MVA and MEP) of terpene synthesis, and further requires that the associated genes must have effective transcription expression (FPKM>5) in the Litsea cubeba germplasm resources in combination with the transcriptome data.
[0069] Then, in order to determine the reliability of the markers, the plurality of second candidate SNP sites screened by the above preliminary screening are re-verified in an independent test population, and the statistically insignificant markers (P>1×10 -5 ) are removed.
[0070] Finally, redundancy was removed from the numerous significant second candidate SNP sites using LD analysis, retaining only representative and independent second candidate SNP sites. After the three screening processes described above, 49 first candidate SNP sites (i.e., 49 candidate functional SNP sites) were obtained, corresponding to different specified compounds. These included 11 citral sites, 8 geraniol sites, 3 neraldehyde sites, 8 β-phellandrene sites, 4 β-caryophyllene sites, and 15 limonene sites. Furthermore, no significantly associated sites meeting all criteria were detected for the two traits, β-pinene and α-terpineol.
[0071] The processor 200 then uses the aforementioned 49 high-quality candidate functional SNP sites as prior features for input to the prior information channel.
[0072] Therefore, by using important SNP sites identified in genome-wide association analysis as prior information to input into a gene prediction model based on a dual-channel convolutional neural network, prediction efficiency and accuracy can be effectively improved.
[0073] Optionally, the gene prediction model includes a first convolutional neural network, a second convolutional neural network, a feature fusion module, and a fully connected network. Based on general features, various prior features, and various gene prediction models, the operation of predicting the content corresponding to various specified compounds includes: extracting features from the general features using the first convolutional neural network and outputting a first feature map, wherein the size of each convolutional kernel in the first convolutional neural network is 11; extracting features from the prior features using the second convolutional neural network and outputting a second feature map, wherein the size of each convolutional kernel in the second convolutional neural network is 7; performing filling, concatenation, and convolution operations on the first and second feature maps using the feature fusion module to generate a third feature map; and determining the feature vector corresponding to the third feature map and inputting the feature vector into the fully connected network to output the predicted content corresponding to the specified compound.
[0074] Specifically, refer to Figure 3 and Figure 6 As shown, the feature extraction module includes a first convolutional neural network corresponding to the general information channel and a second convolutional neural network corresponding to the prior information channel. The first convolutional neural network corresponding to the general information channel receives a one-dimensional vector of length 2000, representing the genotypic information of 200 tag SNPs (i.e., multiple first tag SNP sites). The second convolutional neural network corresponding to the prior information channel receives several candidate functional markers (i.e., first candidate SNP sites) corresponding to a specified compound.
[0075] In addition, the use structure of the gene prediction model has flexibility. When there is no significant candidate SNP for the specified compound (i.e., there is no first candidate SNP site), the prior information channel will be automatically closed, and the gene prediction model only enables the general information path to make a prediction.
[0076] And the first convolutional neural network corresponding to the general information channel and the second convolutional neural network corresponding to the prior information channel are both composed of three layers of one-dimensional convolutional networks in parallel. Among them, the size of each layer of convolutional kernel in the first convolutional neural network corresponding to the general information channel is 11, so as to effectively extract the global and long-distance genetic structure features. The second convolutional neural network corresponding to the prior information channel adaptively adjusts the number of convolutional layers, kernel size, and pooling strategy according to the number of first candidate SNP sites. For example, when the number of first candidate SNP sites is small (less than 10), the gene prediction model only uses one or two layers of small convolutional kernels (1 or 3) to capture local features. When the number of first candidate SNP sites is large, up to three layers of convolution are used, and the pooling operation is increased for dimension compression. In addition, the output channel number of the prior information channel is automatically aligned to be consistent with the output channel number of the general information path, so as to facilitate subsequent fusion.
[0077] Further, after each convolutional layer of the first convolutional neural network and each convolutional layer of the second convolutional neural network, Batch Normalization (Batch Normalization) is followed to accelerate model convergence, and a ReLU activation function is used to introduce nonlinearity, and finally the feature map is reduced in dimension through a max-pooling operation (pooling kernel size is 2, step is 2). Thus, the first convolutional neural network of the final feature extraction module outputs a first feature map, and the second convolutional neural network outputs a second feature map.
[0078] Then the processor 200 fills the first feature map and the second feature map using a feature fusion module to ensure that the length dimensions of the first feature map and the second feature map are consistent. Further, the feature fusion module concatenates the first feature map and the second feature map in the channel dimension and forms a fused feature map.
[0079] Further, the feature integration module performs a 1x1 convolution operation on the fused feature map, which can effectively compress the channel number and perform deep feature integration of general and prior information, generate a third feature map, and flatten the integrated third feature map into a one-dimensional feature vector.
[0080] Finally, the feature integration module transmits the one-dimensional feature vector to an output module composed of three layers of fully connected networks. Among them, the number of neurons of the three layers of fully connected networks is 128, 64 and 1 in turn, and through layer-by-layer nonlinear mapping, the final output is the accurate prediction value of the specified compound.
[0081] Therefore, the gene prediction model based on the double-channel convolutional neural network can not only effectively utilize a small number of high-impact SNP sites (i.e., the first candidate SNP site), but also fully extract potential signals in the genome range, balance specificity and comprehensiveness, and be suitable for plant population prediction tasks with complex genetic regulation and limited sample size. Furthermore, the content of the specified compound can be efficiently and accurately predicted.
[0082] In addition, the gene prediction model constructed in the present application adopts modular design, has good scalability and portability, is suitable for breeding research of other plants or complex number shapes, and can provide efficient and reliable technical support for molecular breeding and early selection of plants such as Litsea cubeba. Therefore, the breeding cycle is shortened, the selection efficiency is improved, and it has broad application prospect and popularization value.
[0083] Optionally, the device further comprises: constructing each gene prediction model in advance and training each gene prediction model, wherein the specific operation comprises: obtaining a plurality of first plant samples, and determining a plurality of specified compounds in the first plant samples and the content of each specified compound in the first plant samples; based on the genetic distance optimization principle, screening the plurality of first plant samples, and determining a plurality of second plant samples in the plurality of first plant samples; determining a general feature sample corresponding to each second plant sample and a prior feature sample corresponding to a different specified compound in each second plant sample; based on the plurality of general feature samples, the plurality of prior feature samples corresponding to different specified compounds, and the content of each compound in the second plant sample, constructing a plurality of training sample pairs corresponding to different specified compounds; and training the gene prediction model corresponding to different specified compounds using the plurality of training sample pairs corresponding to different specified compounds.
[0084] Specifically, before predicting the content of the corresponding specified compound using each gene prediction model, the gene prediction model corresponding to each specified compound needs to be constructed in advance, and the plurality of gene prediction models are trained respectively. Hereinafter, taking Litsea cubeba germplasm resources as an example, the gene prediction model corresponding to 8 terpene compounds is trained.
[0085] The specific training method is as follows: first, the breeding personnel collect 945 samples of 8-year-old Litsea cubeba germplasm resources (i.e., a plurality of first plant samples). Then, 310 female individuals with normal fruit development are selected from the 945 samples of Litsea cubeba germplasm resources, and 100 g of mature fruits are collected from each sample of Litsea cubeba germplasm resources. Then, the essential oil is extracted by water vapor distillation, and the relative content of 8 main terpene compounds in the essential oil is analyzed by gas chromatography and mass spectrometry (GC-MS). The 8 main terpene compounds include citral, geranial, neral, limonene, alpha-terpineol, beta-phellandrene, beta-caryophyllene, and beta-pinene.
[0086] Figure 7 is a schematic diagram of the distribution of the content of the 8 terpene compounds according to the embodiments of the present application. As shown in Figure 7 , the black line represents the fitted distribution curve, so that the content of each specified compound in the population is approximately normally distributed. In addition, the phenotypic statistics of the 8 terpene compounds are shown in Table 1:
[0087] Table 1
[0088]
[0089] After that, in order to determine the generalization ability and prediction accuracy of the gene prediction model, the breeding personnel construct a core sample set with maximum genetic diversity through the terminal device 100 and using CoreHunter3 software. Specifically, the genetic distance between the 310 samples of Litsea cubeba germplasm resources (i.e., each first plant sample) is calculated multiple times, and each sample of Litsea cubeba germplasm resources is sorted according to the number of times it is selected into the core set in the calculation. The average genetic distance between any sample of Litsea cubeba germplasm resources in the core set and the adjacent sample is maximized.
[0090] After that, according to the pre-set sample selection ratio, and according to the sorting result, a plurality of second plant samples are selected from a plurality of first plant samples. The sample selection ratio may, for example, be pre-set by the breeding personnel according to experience, or be set by the terminal device 100 based on historical data. The above content will be described in detail later, and therefore will not be described here.
[0091] Therefore, based on the above method, the breeding personnel may, for example, select the top 70% of the 310 samples of Litsea cubeba germplasm resources (i.e., 217 samples of Litsea cubeba germplasm resources) in the number of hits as a plurality of second plant samples for training the gene prediction model through the terminal device 100.
[0092] Further, in the case that the breeder determines a plurality of second plant samples for training the gene prediction model through the terminal device 100, the general feature sample corresponding to each second plant sample is further determined, and the prior feature sample corresponding to different specified compounds in each second plant sample is determined.
[0093] The specific method for determining the general feature sample corresponding to each second plant sample is as follows: first, the breeder determines a plurality of SNP site samples of various specified compounds in the plurality of second plant samples through the terminal device 100. Then, the breeder performs LD analysis on the whole genome of each second plant sample through the terminal device 100, and determines the first tag SNP site sample within each LD block in the whole genome. The first tag SNP site sample can represent the haplotype structure of the corresponding LD block. Further, the breeder selects second tag SNP site samples of different orders of magnitude through the terminal device 100 based on the location of the first tag SNP site sample on the whole genome. Then, the processor 200 constructs an rrBLUP model, and compares the prediction ability of the rrBLUP model for the specified compound based on second tag SNP site samples of different orders of magnitude. Finally, the processor 200 determines the third tag SNP site sample corresponding to the optimal order based on the prediction ability, and takes the third tag SNP site sample as the general feature sample. The above content will be described in detail later, and therefore will not be described here.
[0094] And the specific method for determining the prior feature sample corresponding to different specified compounds in each second plant sample is as follows: first, the breeder performs whole genome association analysis on various specified compounds in the plurality of second plant samples through the terminal device 100, and determines the association analysis result. The association analysis result is used to indicate a plurality of first candidate SNP site samples that reach a pre-set significance threshold in the plurality of SNP site samples. Then, the breeder filters out the second candidate SNP site sample corresponding to different specified compounds from the first candidate SNP site sample through the terminal device 100, and takes the second candidate SNP site sample as the prior feature sample. The above content will be described in detail later, and therefore will not be described here.
[0095] Further, the processor 200 constructs a plurality of training sample pairs corresponding to different specified compounds based on the plurality of general feature samples, the plurality of prior feature samples corresponding to different specified compounds, and the content of various specified compounds in the second plant sample.
[0096] Finally, the processor 200 trains the gene prediction model corresponding to various specified compounds respectively by using the plurality of training sample pairs corresponding to different specified compounds.
[0097] Optionally, based on the genetic distance optimization principle, the operation of screening the plurality of first plant samples and determining the plurality of second plant samples in the plurality of first plant samples comprises: calculating the genetic distance between each first plant sample multiple times, and sorting each first plant sample according to the number of hits of each first plant sample in the calculation of being selected into the core set, wherein the average genetic distance between any first plant sample in the core set and the adjacent first plant sample is maximized; and according to the pre-set sample selection ratio, and according to the sorting result, the plurality of second plant samples are selected from the plurality of first plant samples.
[0098] Specifically, Figure 8 is a principal component analysis diagram of the plurality of first plant samples and the plurality of second plant samples according to the embodiments of the present application. Referring to Figure 8 It is shown that, in order to determine the generalization ability and prediction accuracy of the gene prediction model, the CoreHunter3 software is used to construct a core sample set with maximum genetic diversity. This process is based on the genetic distance optimization principle. First, the breeder inputs the genotype data filtered by LD through the terminal device 100. Then, the "entry-to-nearest-entry" algorithm is used to calculate the genetic distance between each first plant sample.
[0099] Further, the above optimization process is repeated 100 times, and each first plant sample is sorted according to the "number of hits" of each first plant sample in the calculation of being selected into the core set. Finally, the first plant samples with the top 70% of the number of hits (i.e., the sample selection ratio) are selected from the 310 samples, and the above first plant samples are used as the plurality of second plant samples for training each gene prediction model.
[0100] In addition, the processor 200 also uses the remaining 30% of the samples as a test set for testing the performance of each gene prediction model, for calculating the Pearson correlation coefficient between the predicted value and the true observed value.
[0101] Thus, by screening the plurality of first plant samples based on the genetic distance optimization principle and determining the plurality of second plant samples, it can be ensured that the plurality of second plant samples selected have high diversity, and in the case of training each gene prediction model using the plurality of second plant samples, the stability and generalization ability of the gene prediction model can be enhanced.
[0102] Optionally, the operation of determining the common feature sample corresponding to each second plant sample comprises: determining a plurality of SNP site samples of each designated compound in the plurality of second plant samples; performing LD analysis on the whole genome of each second plant sample, and determining a first tag SNP site sample in each LD block in the whole genome, wherein the first tag SNP site sample can represent the haplotype structure of the corresponding LD block; selecting second tag SNP site samples of different orders based on the positions of the first tag SNP site samples on the whole genome; constructing an rrBLUP model, and comparing the prediction ability of the rrBLUP model for the designated compound based on the second tag SNP site samples of different orders; and determining a third tag SNP site sample corresponding to an optimal order based on the prediction ability, and taking the third tag SNP site sample as the common feature sample.
[0103] Specifically, Figure 9 is a distribution density diagram of the second tag SNP site sample under different marker densities according to the embodiments of the present application. As shown in Figure 9 In the training of each gene prediction model, too many markers will introduce background noise and increase the computational burden, and too few markers may lose key genetic information. Therefore, it is crucial to screen common feature samples of appropriate quantity and covering the whole genome.
[0104] Therefore, first, the breeder extracts the genome of the Litsea cubeba germplasm resource sample (i.e., a plurality of second plant samples) using the improved cetyltrimethylammonium bromide method (CTAB method). Then, the breeder performs sequencing through the terminal device 100 and by using the DNBSEQ-T7 platform, and the average sequencing depth is set to 10x. The raw sequencing data is subjected to quality control using Trimmomatic, and after alignment to the Litsea cubeba reference genome by BWA software, single nucleotide polymorphism detection (SNP) and InDel detection are performed using GATK software, and filtering parameters (such as QD<2.0, MQ<40, etc.) are set to eliminate low-quality sites. Therefore, the breeder can determine 7790030 SNP sites (i.e., a plurality of SNP site samples) of 8 major terpene compounds (i.e., designated compounds) through the terminal device 100.
[0105] Then, the breeder performs LD analysis on the whole genome by using the PLINK v1.9 software through the terminal device 100, and divides the LD blocks. Then, a first tag SNP site sample capable of representing the haplotype structure of the region is selected in each LD block.
[0106] And in order to determine the optimal number of markers, a Python script is used to uniformly select different orders of magnitude of second tag SNP site samples according to the physical positions of the first tag SNP site samples on the whole genome. For example, the orders of magnitude can be 100, 300, 500, 1,000, 2,000, 4,000, 6,000, 8,000, and 10,000.
[0107] Further, based on these different density marker sets, the mixed.solve function in the rrBLUP package of R language is used to construct a traditional rrBLUP model, and different orders of magnitude of second tag SNP site samples are input into the rrBLUP model to compare the prediction ability of the rrBLUP model for the eight terpenes.
[0108] Figure 10 is a schematic diagram of the prediction ability of the rrBLUP model corresponding to different orders of magnitude of second tag SNP site samples according to the embodiments of the present application. As shown in Figure 10 , with the increasing orders of magnitude, the prediction ability of the rrBLUP model generally shows an upward trend, and tends to be stable when the order of magnitude is 2000 (i.e., the optimal order of magnitude), so that 2000 third tag SNP site samples are determined as the universal feature samples.
[0109] Alternatively, the operation of determining the prior feature sample corresponding to the different specified compound in each second plant sample includes: performing whole genome association analysis on various specified compounds in the plurality of second plant samples, and determining the association analysis result, wherein the association analysis result is used to indicate that a plurality of first candidate SNP site samples reaching a pre-set significance threshold are in the plurality of SNP site samples; from the plurality of first candidate SNP site samples, the second candidate SNP site sample corresponding to the different specified compound is screened out, and the second candidate SNP site sample is taken as the prior feature sample.
[0110] Specifically, first, the breeder determines the plurality of SNP site samples through the terminal device 100, and further uses the mixed linear model (MLM) of the GEMMA software to perform whole genome association analysis on the eight terpenes of the Litsea cubeba germplasm resources, and determines the association analysis result. Among them, the association analysis result is used to indicate that a plurality of first candidate SNP site samples (for example, 886 candidate SNP sites reaching the significance threshold) reaching a pre-set significance threshold are in the plurality of SNP site samples. And wherein the significance threshold is determined to be P<8.37x10 -7 .
[0111] In addition, in the whole genome association analysis of the 8 terpene compounds, the kinship matrix was taken as a random effect to correct the influence of population structure. And the first three principal components (PCs) were taken as fixed effect covariates to further control systematic bias.
[0112] Finally, in order to improve the quality of prior information and reduce the false positive rate of gene prediction model, the present application further strictly filters the plurality of first candidate SNP site samples (i.e., 886 candidate SNP sites reaching the significance threshold) obtained after the whole genome association analysis. Therefore, first, the processor 200 retains the gene sites related to the key metabolic pathways (MVA and MEP) of terpene synthesis, and further requires that the associated genes have effective transcription expression (FPKM>5) in the Litsea cubeba germplasm resources in combination with the transcriptome data.
[0113] Then, in order to determine the reliability of the markers, the plurality of first candidate SNP site samples screened by the above preliminary screening are re-verified in an independent test population, and statistically insignificant markers (P>1×10 -5 ) are removed.
[0114] Finally, the significant plurality of first candidate SNP site samples are de-redundant by LD analysis, and only representative and independent first candidate SNP site samples are retained, so that after the above three screenings, the second candidate SNP site samples corresponding to different specified compounds (i.e., 49 candidate functional SNP sites) are obtained. Among them, 11 for citral, 8 for geranial, 3 for neral, 8 for β-phellandrene, 4 for β-caryophyllene, and 15 for limonene. In addition, for the two traits of β-pinene and α-terpineol, no significant associated sites meeting all the conditions were detected.
[0115] In addition, in order to evaluate the effectiveness of each trained gene prediction model, the traditional rrBLUP method is used as a baseline model. When constructing the rrBLUP model, the input features include the same 2000 third tag SNP site samples as the gene prediction model. The rrBLUP model uses the rrBLUP package in R language, and introduces a fixed effect term to include the second candidate SNP site samples corresponding to different specified compounds. Specifically, for each specified compound, the second candidate SNP site samples screened out by the whole genome association analysis are set as fixed effects, and the remaining 2000 third tag SNP site samples are taken as random effects, and the "mixed.solve" function is used for fitting.
[0116] Further, the Pearson correlation coefficient (the correlation between the predicted value and the observed value of the phenotype) is taken as an index to analyze the prediction performance corresponding to the gene prediction model of each specified compound. Figure 11is a comparison chart of prediction accuracy between the genetic prediction model and the rrBLUP model on various specified compounds according to the embodiments of the present application. Referring to Figure 11 As shown in the figure, in the prediction of the eight major terpenes, the prediction performance of the genetic prediction model on most of the specified compounds is better than that of the rrBLUP model. Among them, limonene has the highest prediction ability, which can reach 0.78. The prediction abilities of citral and geranial are 0.72 and 0.71 respectively, both of which have high prediction accuracy.
[0117] Compared with the rrBLUP model (prediction ability is 0.61), the prediction ability of the genetic prediction model on geranial (0.71) is the most significant. For traits without significant prior markers, such as β-pinene, the PKDP can reach a prediction ability of 0.65 even only relying on the general path, which is higher than the rrBLUP model.
[0118] Therefore, the genetic prediction model in the present application can not only effectively use a small number of high-impact SNP sites, but also fully extract potential signals in the genome, balance specificity and comprehensiveness, and is suitable for plant population prediction tasks with complex genetic regulation and limited sample size.
[0119] Therefore, according to the first aspect of the present embodiment, the technical effect of improving the prediction efficiency and accuracy of the content of the specified compound can be achieved.
[0120] In addition, referring to Figure 1 As shown in the figure, according to the second aspect of the present embodiment, a storage medium is provided. The storage medium comprises a stored program, wherein when the program is executed by a processor, the method described in any one of the above embodiments is executed.
[0121] Therefore, according to the present embodiment, the technical effect of improving the prediction efficiency and accuracy of the content of the specified compound can be achieved.
[0122] It should be noted that for the above-mentioned method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the order of the described actions, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.
[0123] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0124] Example 2
[0125] Figure 12 A plant breeding device 1200 based on a dual-channel convolutional neural network according to this embodiment is shown, which corresponds to the method described in Embodiment 1. Reference Figure 12 As shown, the device 1200 includes: a target plant determination module 1210, used to determine the target plant, multiple specified compounds in the target plant, and multiple SNP sites in the various specified compounds; a general feature determination module 1220, used to screen multiple first tag SNP sites from the multiple SNP sites to cover the entire genome, and to determine each first tag SNP site as a general feature for input into a general information channel in each gene prediction model, wherein the model architecture of the general information channel is a convolutional neural network; a prior feature determination module 1230, used to screen first candidate SNP sites corresponding to different specified compounds from the multiple SNP sites, and to determine the first candidate SNP sites corresponding to different specified compounds as prior features for input into a prior information channel in each gene prediction model, wherein each gene prediction model corresponds to different compounds, and the model architecture of the prior information channel is a convolutional neural network; and a content prediction module 1240, used to predict the content corresponding to various specified compounds based on the general features, each prior feature, and each gene prediction model, and to perform a breeding feasibility analysis on the target plant based on the predicted content.
[0126] Optionally, the prior feature determination module 1230 includes: a first association analysis module, used to perform genome-wide association analysis on various specified compounds in the target plant and determine the association analysis results, wherein the association analysis results are used to indicate multiple second candidate SNP sites that reach a pre-set significance threshold among multiple SNP sites; and a first screening module, used to screen out first candidate SNP sites corresponding to different specified compounds from multiple second candidate SNP sites.
[0127] Optionally, the gene prediction model comprises a first convolutional neural network, a second convolutional neural network, a feature fusion module, and a fully connected network, and the content prediction module 1240 comprises: a first feature extraction module configured to perform feature extraction on the general features by using the first convolutional neural network and output a first feature map; a second feature extraction module configured to perform feature extraction on the prior features by using the second convolutional neural network and output a second feature map; a feature map generation module configured to perform padding, splicing, and convolution operations on the first feature map and the second feature map by using the feature fusion module, and generate a third feature map; and a predicted content output module configured to determine a feature vector corresponding to the third feature map, and input the feature vector into the fully connected network, so as to output the predicted content corresponding to the specified compound.
[0128] Optionally, the device 1200 further comprises a model training module configured to pre-construct each gene prediction model and train each gene prediction model respectively, wherein the model training module comprises: a plant sample acquisition module configured to acquire a plurality of first plant samples, and determine a plurality of specified compounds in the first plant samples and contents of various specified compounds in the first plant samples; a plant sample screening module configured to screen the plurality of first plant samples based on a genetic distance optimization principle, and determine a plurality of second plant samples in the plurality of first plant samples; a feature sample determination module configured to determine general feature samples corresponding to each second plant sample and prior feature samples corresponding to different specified compounds in each second plant sample; a training sample pair construction module configured to construct a plurality of training sample pairs corresponding to different specified compounds based on a plurality of general feature samples, a plurality of prior feature samples corresponding to different specified compounds, and contents of various specified compounds in the second plant samples; and a model training submodule configured to train corresponding gene prediction models by using the plurality of training sample pairs corresponding to different specified compounds respectively.
[0129] Optionally, the plant sample screening module comprises: a second screening module configured to calculate genetic distances between each first plant sample multiple times, and sort each first plant sample according to a hit number of each first plant sample being selected into a core set in the calculation, wherein an average genetic distance between any first plant sample in the core set and an adjacent first plant sample is maximized; and a third screening module configured to select a plurality of second plant samples from the plurality of first plant samples according to a pre-set sample selection ratio and according to a sorting result.
[0130] Optionally, the characteristic sample determining module comprises: a SNP site sample determining module, configured to determine a plurality of SNP site samples of each specified compound in the plurality of second plant samples; an LD analysis module, configured to perform LD analysis on a whole genome of each second plant sample, and determine a first tag SNP site sample in each LD block in the whole genome, wherein the first tag SNP site sample is capable of representing a haplotype structure of the corresponding LD block; a SNP site sample selecting module, configured to select second tag SNP site samples of different orders of magnitude based on positions of the first tag SNP site samples on the whole genome; a prediction ability comparing module, configured to construct an rrBLUP model, and compare prediction abilities of the rrBLUP model for the specified compound based on the second tag SNP site samples of different orders of magnitude; and a universal characteristic sample determining module, configured to determine a third tag SNP site sample corresponding to an optimal order based on the prediction abilities, and take the third tag SNP site sample as a universal characteristic sample.
[0131] Optionally, the characteristic sample determining module comprises: a fourth screening module, configured to perform whole genome association analysis on each specified compound in the plurality of second plant samples, and determine an association analysis result, wherein the association analysis result is used to indicate a plurality of first candidate SNP site samples reaching a pre-set significance threshold in the plurality of SNP site samples; and a fifth screening module, configured to screen, from the first candidate SNP site samples, second candidate SNP site samples corresponding to different specified compounds, and take the second candidate SNP site samples as prior characteristic samples.
[0132] According to the embodiment, the technical effect of improving the prediction efficiency and accuracy for the content of the specified compound can be achieved.
[0133] Embodiment 3
[0134] Figure 13 A plant breeding device 1300 based on a double-channel convolutional neural network according to the embodiment is shown, which corresponds to the method according to Embodiment 1. Referring to Figure 3As shown, the apparatus 1300 includes a processor 1310 and a memory 1320 connected with the processor 1310, for providing the processor 1310 with instructions to process the following processing steps: determining target plants, a plurality of specified compounds in the target plants, and a plurality of SNP sites in the various specified compounds; from the plurality of SNP sites, screening a plurality of first tag SNP sites for covering a whole genome, and determining each first tag SNP site as a general feature for inputting into a general information channel of each gene prediction model, a model architecture of the general information channel being a convolutional neural network; from the plurality of SNP sites, screening first candidate SNP sites corresponding to different specified compounds, and determining the first candidate SNP sites corresponding to the different specified compounds as prior features for inputting into a prior information channel of each gene prediction model, wherein each gene prediction model corresponds to a different compound, and a model architecture of the prior information channel being a convolutional neural network; and based on the general features, each prior feature, and each gene prediction model, predicting contents corresponding to the various specified compounds, and performing breeding feasibility analysis on the target plants according to the predicted contents.
[0135] Optionally, the operation of screening, from the plurality of SNP sites, the first candidate SNP sites corresponding to the different specified compounds includes: performing whole genome association analysis on the various specified compounds in the target plants, and determining an association analysis result, wherein the association analysis result is used to indicate a plurality of second candidate SNP sites reaching a pre-set significance threshold in the plurality of SNP sites; and screening, from the plurality of second candidate SNP sites, the first candidate SNP sites corresponding to the different specified compounds.
[0136] Optionally, the gene prediction model includes a first convolutional neural network, a second convolutional neural network, a feature fusion module, and a fully connected network, and the operation of predicting, based on the general features, each prior feature, and each gene prediction model, the contents corresponding to the various specified compounds includes: performing feature extraction on the general features by using the first convolutional neural network, and outputting a first feature map; performing feature extraction on the prior features by using the second convolutional neural network, and outputting a second feature map; performing padding, splicing, and convolution operations on the first feature map and the second feature map by using the feature fusion module, and generating a third feature map; and determining a feature vector corresponding to the third feature map, and inputting the feature vector into the fully connected network, so as to output the predicted contents corresponding to the specified compounds.
[0137] Optionally, the apparatus 1300 further includes: constructing each gene prediction model in advance and training each gene prediction model, wherein the specific operations include: obtaining a plurality of first plant samples, and determining a plurality of specified compounds in the first plant samples and the content of each specified compound in the first plant samples; filtering the plurality of first plant samples based on a genetic distance optimization principle, and determining a plurality of second plant samples in the plurality of first plant samples; determining a general feature sample corresponding to each second plant sample and a prior feature sample corresponding to a different specified compound in each second plant sample; constructing a plurality of training sample pairs corresponding to different specified compounds based on the plurality of general feature samples, the plurality of prior feature samples corresponding to different specified compounds, and the content of each compound in the second plant sample; and training the corresponding gene prediction model using the plurality of training sample pairs corresponding to different specified compounds.
[0138] Optionally, the operation of filtering the plurality of first plant samples based on a genetic distance optimization principle, and determining a plurality of second plant samples in the plurality of first plant samples includes: calculating the genetic distance between each first plant sample multiple times, and sorting each first plant sample according to the number of times it is selected into a core set in the calculation, wherein the average genetic distance between any first plant sample in the core set and the adjacent first plant sample is maximized; and selecting the plurality of second plant samples from the plurality of first plant samples according to a pre-set sample selection ratio and according to the sorting result.
[0139] Optionally, the operation of determining a general feature sample corresponding to each second plant sample includes: determining a plurality of SNP site samples of each specified compound in the plurality of second plant samples; performing LD analysis on the whole genome of each second plant sample, and determining a first tag SNP site sample in each LD block in the whole genome, wherein the first tag SNP site sample can represent the haplotype structure of the corresponding LD block; selecting second tag SNP site samples of different orders of magnitude based on the position of the first tag SNP site sample on the whole genome; constructing an rrBLUP model, and comparing the prediction ability of the rrBLUP model for the specified compound based on the second tag SNP site samples of different orders of magnitude; and determining a third tag SNP site sample corresponding to the optimal order based on the prediction ability, and taking the third tag SNP site sample as the general feature sample.
[0140] Optionally, the operation of determining the prior feature sample corresponding to the different specified compound in each second plant sample comprises: performing a genome-wide association analysis on each specified compound in the plurality of second plant samples, and determining an association analysis result, wherein the association analysis result is used to indicate that a plurality of first candidate SNP site samples reach a preset significance threshold in the plurality of SNP site samples; and screening a second candidate SNP site sample corresponding to the different specified compound from the plurality of first candidate SNP site samples, and taking the second candidate SNP site sample as the prior feature sample.
[0141] Therefore, according to the embodiment, the technical effect of improving the prediction efficiency and accuracy of the content of the specified compound can be achieved.
[0142] The above-mentioned serial numbers of the embodiments of the application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0143] In the above-mentioned embodiments of the application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0144] In several embodiments provided by the present application, it should be understood that the disclosed technology can be implemented in other ways. Of course, the unit embodiment described above is only illustrative, and for example, the division of units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, unit or module, and can be electrical or other forms.
[0145] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment.
[0146] In addition, each functional unit in each embodiment of the application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0147] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0148] The above is only the preferred embodiment of the present application, it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, can make a number of improvements and refinements, these improvements and refinements should also be considered as the protection scope of the present application.
Claims
1. A method for plant breeding based on a dual-channel convolutional neural network, characterized in that, The method comprises the following steps: determining a target plant, a plurality of specified compounds in the target plant, and a plurality of SNP sites in the specified compounds; selecting a plurality of first tag SNP sites from the plurality of SNP sites for covering the whole genome, and determining each first tag SNP site as a general feature for inputting into a general information channel of each gene prediction model, the model architecture of the general information channel being a convolutional neural network; selecting first candidate SNP sites corresponding to different specified compounds from the plurality of SNP sites, and determining the first candidate SNP sites corresponding to the different specified compounds as prior features for inputting into a prior information channel of the each gene prediction model, wherein the each gene prediction model corresponds to a different compound, and the model architecture of the prior information channel is a convolutional neural network; and predicting the content corresponding to the specified compounds based on the general features, the each prior feature, and the each gene prediction model, and performing breeding feasibility analysis on the target plant according to the predicted content.
2. The method of claim 1, wherein, The operation of selecting first candidate SNP sites corresponding to different specified compounds from the plurality of SNP sites comprises: performing whole genome association analysis on the specified compounds in the target plant, and determining an association analysis result, wherein the association analysis result is used to indicate a plurality of second candidate SNP sites reaching a pre-set significance threshold in the plurality of SNP sites; and selecting the first candidate SNP sites corresponding to the different specified compounds from the plurality of second candidate SNP sites.
3. The method of claim 1, wherein, The gene prediction model comprises a first convolutional neural network, a second convolutional neural network, a feature fusion module, and a fully connected network, and the operation of predicting the content corresponding to the specified compounds based on the general features, the each prior feature, and the each gene prediction model comprises: performing feature extraction on the general features by using the first convolutional neural network, and outputting a first feature map; performing feature extraction on the prior features by using the second convolutional neural network, and outputting a second feature map; performing padding, splicing, and convolution operations on the first feature map and the second feature map by using the feature fusion module, and generating a third feature map; and determining a feature vector corresponding to the third feature map, and inputting the feature vector into the fully connected network, so as to output the predicted content corresponding to the specified compounds.
4. The method of claim 1, wherein, The method further comprises the following steps: pre-constructing the each gene prediction model, and training the each gene prediction model respectively, wherein the specific operation comprises: obtaining a plurality of first plant samples, and determining a plurality of specified compounds in the first plant samples and the content of the specified compounds in the first plant samples; selecting a plurality of second plant samples from the plurality of first plant samples based on a genetic distance optimization principle; and determining a plurality of SNP sites in the second plant samples, and determining a plurality of first tag SNP sites from the plurality of SNP sites for covering the whole genome, and determining each first tag SNP site as a general feature for inputting into a general information channel of each gene prediction model, the model architecture of the general information channel being a convolutional neural network. determining a general feature sample corresponding to each of the second plant samples and a prior feature sample corresponding to a different specified compound in the second plant sample; constructing a plurality of training sample pairs respectively corresponding to the different specified compounds based on the plurality of general feature samples, the plurality of prior feature samples corresponding to the different specified compounds and contents of the different specified compounds in the second plant samples; and training the gene prediction model corresponding to the different specified compound respectively by using the plurality of training sample pairs corresponding to the different specified compound.
5. The method of claim 4, wherein, The operation of screening the plurality of first plant samples and determining a plurality of second plant samples in the plurality of first plant samples based on a genetic distance optimization principle, comprising: calculating the genetic distance between the plurality of first plant samples multiple times, and sorting each first plant sample according to the number of times it is selected into a core set in the calculation, wherein the average genetic distance between any first plant sample in the core set and the adjacent first plant sample is maximized; and selecting the plurality of second plant samples from the plurality of first plant samples according to the preset sample selection ratio and the sorting result.
6. The method of claim 4, wherein, The operation of determining a general feature sample corresponding to each of the second plant samples, comprising: determining a plurality of SNP site samples of different specified compounds in the plurality of second plant samples; performing LD analysis on the whole genome of the second plant sample, and determining a first tag SNP site sample in each LD block in the whole genome, wherein the first tag SNP site sample can represent the haplotype structure of the corresponding LD block; selecting second tag SNP site samples of different orders of magnitude based on the location of the first tag SNP site sample on the whole genome; constructing an rrBLUP model and comparing the prediction ability of the rrBLUP model for the specified compound based on second tag SNP site samples of different orders of magnitude; and determining a third tag SNP site sample corresponding to an optimal order based on the prediction ability, and taking the third tag SNP site sample as the general feature sample.
7. The method of claim 6, wherein, The operation of determining a prior feature sample corresponding to a different specified compound in the second plant sample, comprising: performing whole genome association analysis on different specified compounds in the plurality of second plant samples, and determining an association analysis result, wherein the association analysis result is used to indicate a plurality of first candidate SNP site samples reaching a preset significance threshold in the plurality of SNP site samples; and screening second candidate SNP site samples corresponding to different specified compounds from the first candidate SNP site samples, and taking the second candidate SNP site samples as the prior feature sample.
8. A storage medium, characterized by The storage medium comprises a stored program, wherein the program is executed by a processor when the program is run to perform the method of any one of claims 1 to 7. 9.A plant breeding device based on a dual-channel convolutional neural network, characterized by, comprising: a target plant determining module configured to determine a target plant, a plurality of specified compounds in the target plant, and a plurality of SNP sites in each of the specified compounds; a general feature determining module configured to filter a plurality of first tag SNP sites from the plurality of SNP sites, the first tag SNP sites being used to cover a whole genome, and determine each of the first tag SNP sites as a general feature used to input into a general information channel of each of the gene prediction models, the general information channel having a model architecture of a convolutional neural network; a prior feature determining module configured to filter a plurality of first candidate SNP sites corresponding to different specified compounds from the plurality of SNP sites, and determine the first candidate SNP sites corresponding to the different specified compounds as prior features used to input into a prior information channel of each of the gene prediction models, wherein each of the gene prediction models corresponds to a different compound, and the prior information channel has a model architecture of a convolutional neural network; and a content predicting module configured to predict contents corresponding to the different specified compounds based on the general features, the prior features, and the gene prediction models, and perform a breeding feasibility analysis on the target plant based on the predicted contents. 10.A plant breeding device based on a dual-channel convolutional neural network, characterized by, comprise: a processor; and a memory connected to the processor, configured to provide the processor with instructions to process the following processing steps: determine a target plant, a plurality of specified compounds in the target plant, and a plurality of SNP sites in each of the specified compounds; filter a plurality of first tag SNP sites from the plurality of SNP sites, the first tag SNP sites being used to cover a whole genome, and determine each of the first tag SNP sites as a general feature used to input into a general information channel of each of the gene prediction models, the general information channel having a model architecture of a convolutional neural network; filter a plurality of first candidate SNP sites corresponding to different specified compounds from the plurality of SNP sites, and determine the first candidate SNP sites corresponding to the different specified compounds as prior features used to input into a prior information channel of each of the gene prediction models, wherein each of the gene prediction models corresponds to a different compound, and the prior information channel has a model architecture of a convolutional neural network; and predict contents corresponding to the different specified compounds based on the general features, the prior features, and the gene prediction models, and perform a breeding feasibility analysis on the target plant based on the predicted contents.