Corn whole genome association analysis method, device and equipment and readable storage medium

By combining convolutional neural networks and member contribution models, maize genotype and phenotypic data are integrated and processed to construct a genome-wide association analysis model. This solves the problems of high computational cost and low efficiency in existing models, and achieves more efficient and accurate assessment of maize breeding traits.

CN121237214APending Publication Date: 2025-12-30BEIJING RES CENT FOR INFORMATION TECH & AGRI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511201491.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing genome-wide association analysis models are computationally intensive and inefficient when quantitatively assessing the impact of each SNP locus on breeding traits, and they are insufficient in expressing higher-order complex effects.

Method used

A method combining convolutional neural networks and member contribution models was adopted to integrate genotypic and phenotypic data of maize samples. The data matrix was processed through a multi-level quality control strategy, and a convolutional neural network was trained to construct a genome-wide association analysis model to quantitatively evaluate the impact of each SNP locus on breeding traits.

Benefits of technology

This improves the accuracy and efficiency of whole-genome analysis of maize, enabling more accurate prediction of maize phenotypes and quantification of the impact of each SNP locus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237214A_ABST
    Figure CN121237214A_ABST
Patent Text Reader

Abstract

The invention relates to the field of gene data analysis, and provides a corn whole genome association analysis method, device and equipment and a readable storage medium, the method comprises the following steps: integrating a genotype data matrix and a phenotype data matrix of a corn sample to obtain a whole genome association analysis data matrix; training a convolutional neural network based on the whole genome association analysis data matrix; combining the trained convolutional neural network with the member contribution model to construct a whole genome association analysis model; and based on the whole genome association analysis model, obtaining an influence result of each gene locus of the to-be-analyzed corn on the corn phenotype. According to the corn whole genome correlation analysis method, the whole genome correlation analysis model is constructed through the combination of the convolutional neural network and the member contribution model, correlation analysis is performed on the corn whole genome, and the precision and efficiency of corn whole genome analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of genetic data analysis, and in particular to a corn whole genome association analysis method, device, equipment and readable storage medium. BACKGROUND

[0002] Mining single nucleotide polymorphism (SNP) sites that have a significant impact on corn traits is of great importance to corn breeding, and its significance lies in quantitatively evaluating the influence of each SNP site on breeding traits (yield and resistance, etc.) to provide a basis for breeding corn varieties with excellent traits. At present, the whole genome association analysis model for quantitatively evaluating the influence of each SNP site on breeding traits mainly includes two kinds of general linear model and linear mixed model. However, the existing whole genome association analysis model has great limitations, mainly including large amount of calculation and low efficiency; the existing model focuses more on additive effect and is insufficient in expressing high-order complex effect (such as epistasis). SUMMARY

[0003] The present application provides a corn whole genome association analysis method, device, equipment and readable storage medium to solve the limitation problem of the existing whole genome association analysis model for quantitatively evaluating the influence of each SNP site on breeding traits.

[0004] The present application provides a corn whole genome association analysis method, comprising the following steps: Integrating the genotype data matrix and the phenotype data matrix of the corn sample to obtain a whole genome association analysis data matrix; Training a convolutional neural network based on the whole genome association analysis data matrix; Combining the trained convolutional neural network and the member contribution model to construct a whole genome association analysis model; Based on the whole genome association analysis model, obtaining the influence result of each genetic site of the corn to be analyzed on the corn phenotype.

[0005] According to the corn whole genome association analysis method provided by the present application, the genotype data matrix and the phenotype data matrix of the corn sample are integrated to obtain a whole genome association analysis data matrix, which includes the following steps: Collecting corn genotype data and corn phenotype data of the corn sample; Arranging the corn phenotype data by a best linear unbiased prediction (BLUP) model to obtain a phenotype data matrix; Standardizing the corn genotype data by a multi-level quality control strategy to obtain a genotype data matrix.

[0006] According to the corn whole genome correlation analysis method provided by the application, the multi-level quality control strategy comprises individual quality control strategy, gene quality control strategy and population quality control strategy; the genotype data of the corn sample is filtered by the individual missing rate threshold in the individual quality control strategy; The genotype data of the corn sample is filtered by the individual missing rate threshold in the individual quality control strategy; The genotype data of the corn sample is filtered by the gene missing rate threshold, the minor allele frequency threshold and the linkage disequilibrium degree in the gene quality control strategy; The genotype data of the corn sample is filtered by the gene missing rate threshold, the minor allele frequency threshold and the linkage disequilibrium degree in the gene quality control strategy; The genotype data of the corn sample is filtered by the gene missing rate threshold, the minor allele frequency threshold and the linkage disequilibrium degree in the gene quality control strategy.

[0007] According to the corn whole genome correlation analysis method provided by the application, the genotype data of the corn sample is filtered by the individual missing rate threshold in the individual quality control strategy; The genotype data of the corn sample is filtered by the individual missing rate threshold in the individual quality control strategy;

[0008] According to the corn whole genome correlation analysis method provided by the application, the genotype data of the corn sample is filtered by the individual missing rate threshold in the individual quality control strategy; The genotype data of the corn sample is filtered by the individual missing rate threshold in the individual quality control strategy; The genotype data of the corn sample is filtered by the individual missing rate threshold in the individual quality control strategy.

[0009] According to the corn whole genome correlation analysis method provided by the application, the genotype data of the corn sample is filtered by the individual missing rate threshold in the individual quality control strategy; The genotype data of the corn sample is filtered by the individual missing rate threshold in the individual quality control strategy; The genotype data of the corn sample is filtered by the individual missing rate threshold in the individual quality control strategy.

[0010] The application also provides a corn whole genome correlation analysis device, comprising the following modules: The whole genome correlation analysis data matrix determination module is used for integrating the genotype data matrix and the phenotype data matrix of the corn sample to obtain the whole genome correlation analysis data matrix; a training module configured to train a convolutional neural network based on the whole genome association analysis data matrix; a whole genome association analysis model construction module configured to combine the trained convolutional neural network and the member contribution model to construct a whole genome association analysis model; a whole genome association analysis module configured to obtain an influence result of each genetic locus of a corn to be analyzed on a corn phenotype based on the whole genome association analysis model.

[0011] The application also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the corn whole genome association analysis method according to any one of the above when executing the computer program.

[0012] The application also provides a non-transitory computer readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the corn whole genome association analysis method according to any one of the above.

[0013] The application also provides a computer program product comprising a computer program, and the computer program is executed by a processor to implement the corn whole genome association analysis method according to any one of the above.

[0014] The corn whole genome association analysis method, device, equipment and readable storage medium provided by the application integrate the genotype data matrix and the phenotype data matrix of the corn sample to obtain a whole genome association analysis data matrix for training a convolutional neural network and a member contribution model; after the convolutional neural network for predicting the corn phenotype and the member contribution model for quantitatively evaluating the influence of each SNP locus on breeding traits are trained, the two are combined to construct a whole genome association analysis model, which is used for whole genome association analysis of the corn to be analyzed to obtain an influence result of each genetic locus of the corn to be analyzed on the corn phenotype. The whole genome association analysis model constructed by combining the convolutional neural network and the member contribution model is used for whole genome association analysis of the corn, and the precision and efficiency of the corn whole genome analysis are improved. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0016] Figure 1 is one of the flowcharts of the corn whole genome association analysis method provided by the application.

[0017] Figure 2 is a network framework schematic diagram of the whole genome association analysis model provided by the application.

[0018] Figure 3 is a second flowchart of the whole genome association analysis method for corn provided by the application.

[0019] Figure 4 is a structural schematic diagram of the whole genome association analysis device for corn provided by the application.

[0020] Figure 5 is a structural schematic diagram of the electronic device provided by the application. DETAILED DESCRIPTION

[0021] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described below in connection with the drawings in the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0022] The corn whole genome association analysis method, device, equipment and readable storage medium of the present application will be described below in connection with Figures 1-5

[0023] is a first flowchart of the whole genome association analysis method for corn provided by the application, as shown in the figure, the method comprises the following steps: Figure 1 Figure 1 Step 100, integrating the genotype data matrix and the phenotype data matrix of the corn sample to obtain a whole genome association analysis data matrix. Specifically, the specific implementation steps of the corn whole genome association analysis method provided by the present application are as follows. Step S1, collecting the genotype data of the corn sample, constructing a three-level quality control strategy, filtering and standardizing the genotype data of the corn sample from three levels of individual level, SNP level and population structure level, and arranging the genotype data of the corn sample into an m*n genotype data matrix of m varieties and n SNP sites.

[0024] Step S2, collecting the phenotype data of the corn sample, eliminating the influence of environmental factors on the phenotype data through the Best Linear Unbiased Prediction (BLUP) model, and arranging the phenotype data of the corn sample into a phenotype data matrix of m varieties and y trait types.

[0025]

[0026] ​​Step S3: Based on the maize variety, merge the phenotypic data and genotypic data into a genome-wide association analysis (GWAA) data matrix. The GWAA data matrix is ​​arranged with maize samples as rows and genotypic and phenotypic data as columns. The matrix dimension of the GWAA data matrix is ​​m×(n+y).

[0027] Step 200: Train a convolutional neural network based on the genome-wide association analysis data matrix; Specifically, in step S4, based on the number of trait types y, construct CNN models respectively (e.g., Figure 2 As shown in the figure, a convolutional neural network (CNN) is constructed.

[0028] Step S5: Convolutional Neural Network Training. The genome-wide association analysis data matrix is ​​used as input data to train the convolutional neural network until the phenotypic prediction accuracy of the convolutional neural network reaches the expected value.

[0029] Step 300: Combine the trained convolutional neural network and member contribution model to construct a genome-wide association analysis model; based on the genome-wide association analysis model, obtain the influence results of each gene locus of maize on the maize phenotype.

[0030] Specifically, by combining convolutional neural networks and member contribution models, the CNN-SHAP model in this embodiment is constructed. This model is used for genome-wide association analysis (GWAA) of maize phenotypes, and the input data for the GWAA model is the standardized GWAA data matrix mentioned above. SHAP (Shapley additive explanations) is a model interpretation method based on Shapley values ​​in game theory, used to quantify the contribution of features to the predictions of machine learning models.

[0031] Step S6: Calculate the SHAP value of each SNP locus (i.e., gene locus) in the genotype data using the member contribution model, which is used as the significance of the influence of each SNP locus on the maize phenotype, i.e., the result of the influence of each gene locus on the maize phenotype.

[0032] This embodiment integrates the genotype and phenotypic data matrices of maize samples to obtain a genome-wide association analysis (GWAS) data matrix for training convolutional neural networks and member contribution models. After the convolutional neural network for predicting maize phenotypes and the member contribution model for quantifying the impact of each SNP locus on breeding traits are trained, they are combined to construct a GWAS model for association analysis of the entire maize genome, yielding the impact of each gene locus on the maize phenotype. This invention, through the GWAS model constructed by combining convolutional neural networks and member contribution models, improves the accuracy and efficiency of maize genome-wide analysis.

[0033] Figure 3 This is the second schematic diagram of the genome-wide association analysis method for maize provided by this invention, as shown below. Figure 3 As shown, the method may further include: Step 10: Collect maize genotype and phenotypic data from maize samples; Step 20: Organize the maize phenotypic data using the optimal linear unbiased prediction BLUP model to obtain the phenotypic data matrix; Step 30: Standardize the maize genotype data using a multi-level quality control strategy to obtain a genotype data matrix.

[0034] Specifically, taking a maize variety count of m=70 as an example, the quality-controlled genotype data of m×n obtained are shown in Table 1. A, T, C, and G are the four bases of the DNA molecule, which are the basic coding units of genetic information. By encoding the values ​​of SNP sites numerically, a standardized genotype data matrix is ​​obtained, which is Table 2 obtained after numericalizing Table 1.

[0035] Table 1

[0036] Table 2

[0037] Table 3

[0038] Examples of maize phenotypes based on traits such as yield, plant height, and resistance are shown in Table 3. Yield (per mu) can be expressed in kilograms; plant height in centimeters; and resistance (gray spot disease resistance) in grades. For example, maize sample 1 has a yield of 681.3 kg / mu, a plant height of 236 cm, and a gray spot disease resistance grade of 1.

[0039] In this embodiment, the collected maize genotype data and maize phenotypic data are organized to obtain the phenotypic data matrix and genotype data matrix of maize samples, respectively.

[0040] In one embodiment, the maize genome-wide association analysis method provided in this application may further include: Step 21: Process the corn phenotypic data using the BLUP model to obtain a standardized phenotypic data matrix.

[0041] Specifically, the maize phenotypic data were processed using the BLUP model to eliminate the influence of environmental factors on the maize phenotypic data. The maize phenotypic data in Table 3 were then organized into standardized phenotypic data for 70 varieties and 3 traits, as shown in Table 4.

[0042] Table 4

[0043] By using the variety name, the genotype data matrix and phenotype data matrix obtained above were merged into a standardized genome-wide association analysis data matrix, as shown in Table 5.

[0044] Table 5

[0045] This embodiment uses the BLUP model to eliminate the influence of environmental factors on maize phenotypic data.

[0046] In one embodiment, the maize genome-wide association analysis method provided in this application may further include: Step 31: Filter the maize genotype data of the maize samples using the individual missing rate threshold in the individual quality control strategy; Step 32: Filter the maize genotype data of the maize sample using the gene deletion rate threshold, minor allelic frequency threshold, and linkage disequilibrium degree in the gene quality control strategy. Step 33: Filter the maize genotype data of the maize samples using the kinship matrix in the population quality control strategy; Step 34: Standardize the filtered maize genotype data to obtain a genotype data matrix.

[0047] Specifically, multi-level quality control strategies (taking a three-level quality control strategy as an example) include individual quality control strategies, gene quality control strategies, and population quality control strategies.

[0048] Individual quality control strategy: At the individual level, the individual missing rate threshold is set to ≤10%, and corn samples with an individual missing rate higher than this threshold are removed, as are corn samples with DNA degradation or hybridization failure.

[0049] Gene quality control strategy: At the SNP level, this includes a triple filtering strategy: gene deletion rate filtering, minor allele frequency (MAF) screening, and linkage disequilibrium (LD) pruning. The SNP deletion rate threshold is set to ≤10%; the minor allele frequency threshold is set to ≥5%; and the LD severity is set to r² < 0.05, where r² is a measure of LD, a statistic that measures the degree of non-random association between alleles at two genetic loci.

[0050] Population quality control strategy: At the population structure level, population quality control is achieved by constructing a kinship matrix (K matrix, a symmetric matrix used in genetic analysis to quantify the degree of kinship between individuals). After filtering, the values ​​of SNP loci are represented numerically by encoding, resulting in a standardized genotype data matrix of 70 varieties at multiple SNP loci, as shown in Table 3.

[0051] This embodiment employs a three-level quality control strategy to filter and standardize the genotype data of maize samples at the individual, SNP, and population structure levels.

[0052] In one embodiment, the maize genome-wide association analysis method provided in this application may further include: Step 210: Use the genome-wide association analysis data matrix as input data to train the convolutional neural network; Step 220: If the phenotypic prediction accuracy of the convolutional neural network reaches the expected value, determine that the training of the convolutional neural network is complete.

[0053] Specifically, taking maize with trait number y=3, and yield, plant height, and resistance as phenotypes, three CNN-SHAP models were constructed for genome-wide association analysis of the three maize phenotypes. For example... Figure 2 As shown, the convolutional neural network of each genome-wide association analysis model includes an input layer, three convolutional layers, two pooling layers, one fully connected layer, and an output layer. The activation function of the convolutional neural network can be the Rectified Linear Unit (ReLU) activation function; the batch size can be set to 15; the learning rate can be set to 0.001; the epochs can be set to 200; the kernel size can be set to 8×8; the pooling method can be max pooling; the stride can be set to 1; and the optimizer can be the Adaptive Moment Estimation (Adam) optimizer.

[0054] (1) The contribution of each feature value to the prediction result is calculated using the member contribution model, as shown in Formula 1, where, It is a feature The contribution value; For the entire feature set; For feature subset Weights for the unordered nature of the feature arrangement; The total number of permutations of all features; The weights for the arrangement of the remaining features that were not included in the subset; It is except features A subset of features other than those in the above categories; It is a feature subset Add features Subsequent model predictions; It is a feature subset The model's predicted value.

[0055] (2) The loss function during the training of a convolutional neural network can be the mean squared error, as shown in Equation 2, where, It is the first The loss function of a convolutional neural network; This represents the total number of corn samples. It is a corn sample. No. The actual phenotypic data corresponding to individual traits; It is a convolutional neural network through corn samples The genotype data predicted the first Phenotypic data of individual traits.

[0056] In this embodiment, the genome-wide association analysis data matrix of maize samples is used as the training data for the convolutional neural network, thereby training the phenotypic prediction accuracy of the convolutional neural network to reach the expected value.

[0057] In one embodiment, the maize genome-wide association analysis method provided in this application may further include: Step 310: Input the genotype data of the maize to be analyzed into the trained convolutional neural network to obtain the predicted phenotype of the maize to be analyzed; Step 320: Based on the trained member contribution model, determine the contribution value of the gene locus to the predicted phenotype; the gene locus is determined based on the genotype data of the maize to be analyzed.

[0058] Specifically, in step S6 above, the phenotypic prediction accuracy is set as the mean absolute percentage error (MASE). As shown in Formula 3, the expected value can be set to 20%.

[0059] (3) After the phenotypic prediction accuracy of the convolutional neural network reaches the expected value, the SHAP value of each SNP locus to the prediction result (i.e., the contribution value in this embodiment) is calculated using the member contribution model and used as the output of the genome-wide association analysis model. The larger the SHAP value of an SNP locus, the greater its contribution to the prediction result, that is, the more significant its influence on the corresponding trait, as shown in Table 6.

[0060] Table 6

[0061] This embodiment achieves high accuracy in predicting maize phenotypes through training a convolutional neural network. Then, the SHAP algorithm is used to calculate the SHAP value of each SNP to determine the significance of the influence of each SNP site on the maize phenotype.

[0062] The following describes the maize genome-wide association analysis device provided by the present invention. The maize genome-wide association analysis device described below and the maize genome-wide association analysis method described above can be referred to in correspondence.

[0063] Please refer to Figure 4 The present invention also provides a maize genome-wide association study device, comprising: The genome-wide association analysis data matrix determination module 401 is used to integrate the genotype data matrix and phenotypic data matrix of maize samples to obtain the genome-wide association analysis data matrix; Training module 402 is used to train a convolutional neural network based on the genome-wide association analysis data matrix; The genome-wide association analysis model construction module 403 is used to combine the trained convolutional neural network and the member contribution model to construct a genome-wide association analysis model; The genome-wide association analysis module 404 is used to obtain the effect of each gene locus of maize on the maize phenotype based on the genome-wide association analysis model.

[0064] Optionally, the maize genome-wide association study device further includes: The data acquisition module is used to collect maize genotype data and maize phenotypic data from maize samples; The corn phenotypic data processing module is used to process the corn phenotypic data using the best linear unbiased prediction BLUP model to obtain a phenotypic data matrix. The maize genotype data standardization processing module is used to standardize the maize genotype data through a multi-level quality control strategy to obtain a genotype data matrix.

[0065] Optionally, the multi-level quality control strategy includes an individual quality control strategy, a gene quality control strategy, and a population quality control strategy; the maize genotype data standardization processing module includes: An individual quality control strategy execution unit is used to filter the maize genotype data of the maize sample by using the individual missing rate threshold in the individual quality control strategy. The gene quality control strategy execution unit is used to filter the maize genotype data of the maize sample using the gene deletion rate threshold, minor allelic frequency threshold and linkage disequilibrium degree in the gene quality control strategy. The population quality control strategy execution unit is used to filter the maize genotype data of the maize sample through the kinship matrix in the population quality control strategy; The genotype data matrix determination unit is used to standardize the filtered maize genotype data to obtain the genotype data matrix.

[0066] Optionally, the corn phenotypic data processing module includes: The phenotypic data matrix determination unit is used to process the corn phenotypic data using the BLUP model to obtain a standardized phenotypic data matrix.

[0067] Optionally, the training module includes: A convolutional neural network training unit is used to train the convolutional neural network by taking the genome-wide association analysis data matrix as input data. The phenotypic prediction accuracy determination unit is used to determine that the training of the convolutional neural network is complete when the phenotypic prediction accuracy of the convolutional neural network reaches the expected value.

[0068] Optionally, the genome-wide association analysis module includes: The predictive phenotype determination unit is used to input the genotype data of the maize to be analyzed into the trained convolutional neural network to obtain the predicted phenotype of the maize to be analyzed. The contribution value determination unit is used to determine the contribution value of gene loci to the predicted phenotype based on the trained member contribution model; the gene loci are determined based on the genotype data of the maize to be analyzed.

[0069] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a maize genome-wide association analysis (GWIA) method, which includes: integrating genotype data matrices and phenotypic data matrices of maize samples to obtain a GWIA data matrix; training a convolutional neural network based on the GWIA data matrix; combining the trained convolutional neural network with a member contribution model to construct a GWIA model; and obtaining the influence of each gene locus on the maize phenotype based on the GWIA model.

[0070] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0071] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the maize genome-wide association analysis method provided by the above methods. This method includes: integrating the genotype data matrix and phenotype data matrix of maize samples to obtain a genome-wide association analysis data matrix; training a convolutional neural network based on the genome-wide association analysis data matrix; combining the trained convolutional neural network and a member contribution model to construct a genome-wide association analysis model; and obtaining the influence results of each gene locus of the maize to be analyzed on the maize phenotype based on the genome-wide association analysis model.

[0072] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the maize genome-wide association analysis method provided by the above methods. The method includes: integrating genotype data matrices and phenotypic data matrices of maize samples to obtain a genome-wide association analysis data matrix; training a convolutional neural network based on the genome-wide association analysis data matrix; combining the trained convolutional neural network with a member contribution model to construct a genome-wide association analysis model; and obtaining the influence results of each gene locus of the maize to be analyzed on the maize phenotype based on the genome-wide association analysis model.

[0073] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0074] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for maize genome-wide association analysis, characterized in that, The method comprises the following steps: integrating a genotype data matrix and a phenotype data matrix of corn samples to obtain a whole genome association analysis data matrix; training a convolutional neural network based on the whole genome association analysis data matrix; combining the trained convolutional neural network and a member contribution model to construct a whole genome association analysis model; obtaining the influence of each genetic locus of a corn to be analyzed on the corn phenotype based on the whole genome association analysis model.

2. The maize genome-wide association analysis method of claim 1, wherein, The method for integrating the genotype data matrix and the phenotype data matrix of the corn samples to obtain the whole genome association analysis data matrix comprises the following steps: collecting corn genotype data and corn phenotype data of corn samples; arranging the corn phenotype data by a best linear unbiased prediction (BLUP) model to obtain a phenotype data matrix; standardizing the corn genotype data by a multi-level quality control strategy to obtain a genotype data matrix.

3. The maize genome-wide association analysis method of claim 2, wherein, The multi-level quality control strategy comprises an individual quality control strategy, a gene quality control strategy and a population quality control strategy. The method for standardizing the corn genotype data by the multi-level quality control strategy to obtain the genotype data matrix comprises the following steps: filtering the corn genotype data of the corn samples by an individual missing rate threshold in the individual quality control strategy; filtering the corn genotype data of the corn samples by a gene missing rate threshold, a minor allele frequency threshold and a linkage disequilibrium degree in the gene quality control strategy; filtering the corn genotype data of the corn samples by a kinship matrix in the population quality control strategy; standardizing the filtered corn genotype data to obtain the genotype data matrix.

4. The maize genome-wide association analysis method of claim 2, wherein, The method for arranging the corn phenotype data by the best linear unbiased prediction (BLUP) model to obtain the phenotype data matrix comprises the following steps: processing the corn phenotype data by the BLUP model to obtain a standardized phenotype data matrix.

5. The maize genome-wide association analysis method of claim 1, wherein, The method for training the convolutional neural network based on the whole genome association analysis data matrix comprises the following steps: training the convolutional neural network by taking the whole genome association analysis data matrix as input data; determining that the training of the convolutional neural network is completed when the phenotype prediction accuracy of the convolutional neural network reaches an expected value.

6. The maize genome-wide association analysis method of claim 1, wherein, The method for obtaining the influence of each genetic locus of a corn to be analyzed on the corn phenotype based on the whole genome association analysis model comprises the following steps: inputting the genotype data of the corn to be analyzed into the trained convolutional neural network to obtain a predicted phenotype of the corn to be analyzed; determining a contribution value of a genetic locus to the predicted phenotype based on the trained member contribution model; the genetic locus is determined based on the genotype data of the corn to be analyzed.

7. A maize whole genome association analysis apparatus, comprising: The method comprises the following steps: a whole genome association analysis data matrix determination module is configured to integrate a genotype data matrix and a phenotype data matrix of corn samples to obtain a whole genome association analysis data matrix; a training module is configured to train a convolutional neural network based on the whole genome association analysis data matrix; A whole genome association analysis model construction module, configured to combine the trained convolutional neural network and the member contribution model to construct a whole genome association analysis model; A whole genome association analysis module, configured to obtain an influence result of each genetic locus of the corn to be analyzed on a corn phenotype based on the whole genome association analysis model.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The computer program is executed by the processor to implement the whole genome association analysis method of corn according to any one of claims 1 to 6. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the whole genome association analysis method of corn according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the whole genome association analysis method of corn according to any one of claims 1 to 6.