A method for mining excellent genes of sorghum
By using genetic algorithms in artificial intelligence to process sorghum gene expression data, the problem of unstable gene mining in traditional breeding methods has been solved, enabling efficient mining of superior sorghum genes, providing a candidate gene set for sorghum breeding, and supporting high-yield and high-quality breeding.
Patent Information
- Application Number
- CN202211397185.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-09
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2042-11-09
AI Technical Summary
Existing methods for discovering superior genes in sorghum mainly rely on traditional breeding techniques, which cannot reliably obtain and transfer superior genes, and AI breeding has not been widely applied in the field of sorghum.
Using genetic algorithms from artificial intelligence, we can mine superior genes in sorghum through gene expression data preprocessing and feature selection. This includes processing the gene expression matrix of quality traits, initializing the genetic algorithm, and evaluating it with a random forest algorithm to determine the set of superior genes.
It enables efficient and accurate discovery of superior sorghum genes, providing a candidate gene set for sorghum breeding and supporting the breeding of high-yield, high-quality, and high-efficiency new varieties.
Smart Images

Figure CN116130006B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of biological information, and particularly relates to a method for mining excellent genes of sorghum. BACKGROUND
[0002] The process of cultivating excellent crops is called crop breeding or variety improvement. High yield, stable yield, high quality and high efficiency are the goals of breeding. Excellent crop agronomic traits are mainly controlled by excellent key functional genes, and identifying and applying these genes to crop genetic improvement is one of the central tasks of germplasm research.
[0003] The existing method for obtaining / mining excellent genes of sorghum is the traditional breeding technology method, which usually discovers and obtains excellent genes through natural variation selection breeding method and hybridization breeding method. However, these methods are subject to the occurrence of excellent variation in nature or the method itself, and there is no way to obtain specific gene or genes that will also appear trait separation and cannot be stably inherited to the next generation. With the development of modern biological technology, molecular assisted breeding has emerged, which has the advantages of being fast, accurate and not affected by environmental conditions. Thanks to the vigorous development of biological omics and artificial intelligence, and the core strategic demand of the state for "species source", AI breeding has become a hot word in recent years. AI breeding is to use artificial intelligence technology to help breeders accelerate the process of screening breeding materials, which includes both genotype big data analysis and prediction, and phenotype big data analysis and prediction. In essence, it is hoped that various algorithms of artificial intelligence can accelerate the process of "picking one out of millions" and "finding a needle in a haystack". However, at present, AI breeding technology is mainly used for examination, and there is no report on using AI combined with biological omics technology to mine excellent genes of sorghum. SUMMARY
[0004] The purpose of the present application is to provide a method for mining excellent genes of sorghum, which uses genetic algorithms in artificial intelligence algorithms to mine excellent genes of sorghum and provides a set of candidate excellent genes for sorghum breeding.
[0005] In order to achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0006] One purpose of the present application is to provide a method for mining excellent genes of sorghum, which comprises the following steps:
[0007] 1) Determine the quality traits required for breeding sorghum varieties according to literature or experience, and select sorghum varieties with and without the required quality traits;
[0008] 2) Obtain gene expression of the transcriptome of the sorghum variety by sequencing;
[0009] 3) Pretreatment of gene expression data of the transcriptome: including: deleting duplicate data; using normal distribution to determine that gene expression values greater than 3δ are abnormal values, replacing the abnormal values with the maximum value other than the abnormal values; taking logarithmic processing on the data; quantile standardization is performed on the gene expression matrix to make the samples comparable;
[0010] 4) Feature selection on the pretreated data for each quality trait using a genetic algorithm: including: the original feature set is the gene set obtained by sequencing; N genes are randomly selected as the initial feature gene set after genetic algorithm initialization; the feature gene set is evaluated using a random forest algorithm, the feature genes are sorted according to the importance given by the random forest, the importance is used as the fitness value, and the final feature gene set is the excellent gene set of a certain trait of sorghum;
[0011] 5) The excellent gene set of all traits selected in 4) is used as the excellent gene set of sorghum;
[0012] The quality traits include one or more of the following: plump grains, high yield, low tannin content, high protein content, high starch content, and high trace element content.
[0013] Select sorghum varieties with desired quality traits as positive samples for feature gene selection, and select sorghum varieties without desired quality traits as negative samples for feature gene selection, and perform transcriptome sequencing on the positive and negative samples.
[0014] As an implementable way, in step 1), search the database with sorghum and desired quality traits as a combination field to determine sorghum varieties with and without desired quality traits.
[0015] Preferably, the sorghum varieties are as consistent as possible in other traits in addition to having and not having desired quality traits.
[0016] As an implementable way, the coding method of the genetic algorithm is 0 / 1 coding, that is, if gene i is selected, it is 1, and if the gene is not selected, it is 0, and thus the chromosome coding is a 0 / 1 sequence, and the population is multiple 0 / 1 sequences; the selection method uses fitness proportionate method to calculate the selection probability; the crossover probability is set to 0.5, single-point crossover is used, and the mutation rate is 0.0002.
[0017] Further, the fitness proportionate method to calculate the selection probability is: Where i is a certain gene, f i is the fitness, p si is the probability of being selected.
[0018] As an implementable way, the random forest algorithm adopts 10-fold cross-validation to evaluate the feature set with average classification accuracy.
[0019] As an implementable way, the genes of each selected excellent quality trait are taken and set as a final set of excellent quality genes of sorghum.
[0020] Compared with the prior art, the present application has the following advantages and effects:
[0021] The method for mining excellent genes of sorghum provided by the present application can efficiently and accurately mine excellent genes of sorghum through the genetic algorithm in the artificial intelligence algorithm, and provides a candidate excellent gene set for sorghum breeding. After important agronomic traits are obtained through the method of the present application, relevant scientists can perform function analysis, which not only meets the needs of theoretical research and international gene resource competition, but also provides a gene candidate source basis for obtaining excellent new sorghum varieties with high yield, high quality and good stress resistance through molecular breeding. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 A specific flowchart of the method for mining excellent genes of sorghum is provided.
[0023] Figure 2 A specific flowchart of the genetic algorithm for modeling and mining excellent genes is provided. DETAILED DESCRIPTION
[0024] The present application will be further described below in combination with specific embodiments, but the present application is not limited to the following embodiments. Unless otherwise specified, the experimental methods used in the embodiments are conventional methods, and the materials, reagents and the like used are commercially available.
[0025] Embodiment 1
[0026] A method for mining excellent genes of sorghum comprises the following steps in sequence:
[0027] 1) Determine the quality traits to be selected and bred through literature and the experience of breeding experts, such as: grain plumpness, high yield, low tannin content, high protein content, high starch content and high trace element content. Select sorghum varieties with and without the required quality traits. The selected varieties are as consistent as possible in other traits except for the excellent traits. For example, sorghum A has the characteristic of grain plumpness, and sorghum B is basically the same as sorghum A in other traits except for grain plumpness.
[0028] Specifically as follows:
[0029] Through the pubmed database, search the database by combining the English representation of sorghum (Broomcorn, Red Sorghum, Red Shum) with the English word group of quality traits, read the literature, and determine the selection of sorghum varieties with and without related quality traits. Here, through the retrieval of relevant data, the variety A with full grains is selected as the positive sample and the variety B (with basically the same other traits as sorghum A) without full grains is selected as the negative sample.
[0030] 2) Select sorghum with and without the quality traits described in 1) to obtain gene expression of the transcriptome through biological experiments and high-throughput transcriptome sequencing technology (RNA-Seq).
[0031] Specifically as follows:
[0032] Select the seeds of sorghum variety A with the required quality traits as the positive sample for feature gene selection, and select the seeds of sorghum variety B without the required quality traits as the negative sample for feature gene selection, and try to select an equal number of positive and negative samples for transcriptome sequencing. The transcriptome sequencing method is referred to in Hufford MB, Seetharam AS, Woodhouse MR, et al. De novo assembly, annotation, and comparative analysis of 26 diverse maize genomes. Science. 2021; 373(6555): 655-662. doi: 10.1126 / science.abg5289.
[0033] 3) Data preprocessing is performed on the transcriptome expression data obtained in step 2);
[0034] Specifically as follows:
[0035] The transcriptome sequencing data obtained in step 2) is the gene expression matrix G ij (G ij , which is the expression value of gene i in sample j), redundant data is deleted, that is, data with repeated gene names is removed; the variance δ of the gene expression matrix is calculated, and the gene expression value greater than 3δ is determined as an abnormal value, and the abnormal value is replaced by the maximum value other than the abnormal value; The data is logarithmically processed G' ij = log2(1+G ij ); The gene expression matrix is quantile standardized to make the samples comparable; thus, the gene expression matrix after preprocessing is obtained.
[0036] 4) For each quality trait, use a genetic algorithm to select features from the preprocessed data;
[0037] The original feature set is the gene set obtained by all sequencing; 30 genes are randomly selected as the initial feature gene set; the feature gene set is evaluated using the random forest algorithm.
[0038] Specifically: A) Select n samples from the sample set as a training set using the bootstrap method with sampling with replacement; generate a decision tree using the sampled sample set. B) At each node generated: randomly select d features (genes) without repetition, and divide the sample set using the d features respectively to find the best division feature (using the Gini coefficient); C) Repeat steps A) to B) k times, where k is the number of decision trees in the random forest; D) use the trained random forest to predict the test samples, and sort the feature genes according to the importance given by the random forest. The importance is used as the fitness value; the set number of iterations is used as the termination condition of the algorithm; the final feature gene set is the excellent gene set S of the certain quality trait of sorghum k , where k is a certain quality trait;
[0039] The coding method of the genetic algorithm is binary coding, that is, if the gene i is selected, it is 1, and if the gene is not selected, it is 0, so the chromosome code is a 0 / 1 sequence, and the population is multiple 0 / 1 sequences. The selection method uses the fitness proportion method Calculate the selection probability. Set the crossover probability P c = 0.5, use single-point crossover, and the mutation rate P m = 0.002.
[0040] Random forest, 10-fold cross-validation, average classification accuracy to evaluate the feature set.
[0041] The union of the gene sets selected for each excellent quality trait is S = S1∪S2∪…∪S n , as the final excellent quality gene set of sorghum.
[0042] Example 2
[0043] 1) Determine the quality traits needed for selection and breeding: plump grains; through literature review and query of variety identification information.
[0044] 2) Select sorghum with the required quality traits as positive samples for feature gene selection, and select sorghum without the required quality traits as negative samples for feature gene selection, and perform transcriptome sequencing on the positive and negative samples together.
[0045] The transcriptome sequencing adopts the method reported in the literature “Stark R, Grzelak M, Hadfield J. RNA sequencing: the teenage years. Nat Rev Genet. 2019; 20(11): 631-656. doi: 10.1038 / s41576-019-0150-2”, and the sequencing results are one fastq file for each sample. The data analysis and interpretation of all fastq files adopt the method described in the literature Sarah Djebali, Valentin Wucher, Sylvain Foissac, Christophe Hitte, Erwan Corre, Thomas Derrien. Bioinformatics Pipeline for Transcriptome Sequencing Analysis. Methods Mol Biol. 2017; 1468: 201-19. doi: 10.1007 / 978-1-4939-4035-6_14. Through the analysis and interpretation of this method, a gene expression matrix is obtained, in which the rows are gene names, the columns are samples, and the values are gene expression values.
[0046] 3) Data preprocessing of transcriptome expression data
[0047] The specific processing method for the transcriptome sequencing data obtained in step 2) is as follows: check the gene expression matrix obtained in the previous step for duplicate data, and delete duplicate data if any; all expression data uses log2 logarithm as the new gene expression value, and when processing, in order to avoid logarithmic errors caused by missing values and 0, add 1 to each expression value. Detect abnormal gene expression values in the matrix, calculate the standard deviation of all gene expression values in all samples, and use normal distribution to determine gene expression values greater than 3 standard deviations as abnormal values. Replace the abnormal values with the maximum value other than the abnormal values.
[0048] Therefore, the “data after preprocessing” obtained in this step 3) is specifically: the gene expression data matrix after preprocessing, which has no change in row name and column name compared to the gene expression matrix of step 2, and the expression values in it are processed according to step 3).
[0049] 4) Feature selection for each quality trait using genetic algorithm on the data after preprocessing
[0050] The genetic algorithm is used for feature selection for the pretreated data obtained in step 3), and specifically, since the original feature set is a gene set obtained by all sequencing, a binary coded gene list is used, that is, the coding mode is 0 / 1 coding, the selected gene is coded as 1, and the unselected gene is coded as 0, so that the chromosome coding is a 0 / 1 sequence, and the population is a plurality of 0 / 1 sequences. The genetic algorithm is initialized, 30 genes are randomly selected, that is, the 30 genes are coded as 1, and the rest are 0, as an initial feature gene set, the feature gene set is evaluated using the random forest algorithm, the feature genes are sorted according to the importance given by the random forest, and the importance is used as the fitness value. The selection mode adopts the fitness ratio mode The selection probability is calculated, wherein f i is the fitness value of gene i. The fixed crossover probability is set as P c = 0.5, single-point crossover is adopted, and the fixed mutation rate is P m = 0.002. The set iteration number 1000 is used as the algorithm termination condition. The finally determined feature gene set is a high-quality gene set S k of a certain trait of sorghum, wherein k is the selected grain fullness quality trait.
[0051] The result obtained in step 4) is a gene set with excellent grain fullness.
[0052] 5), all the genes selected in 4) are collected as a high-quality gene set of sorghum.
[0053] The genes obtained in step 4) are specifically a collection of all genes coded as 1.
[0054] The credibility of the result set is verified by using the gene function annotation (http: / / www.geneontology.org) of bioinformatics, verifying that the functions of the selected genes are all related to grain fullness, and verifying from the metabolic pathway (https: / / www.genome.jp / kegg / pathway.html) that the above-mentioned genes participate in the process of grain fullness, so that the selected gene set is a high-quality gene set.
[0055] The above embodiment is the best implementation manner of the present application, but the implementation manner of the present application is not limited to the above embodiment, and any change, modification, substitution, combination, simplification made without departing from the spirit and principle of the present application should be an equivalent replacement manner, and all are included in the protection scope of the present application.
Claims
1. A method for mining elite genes of Sorghum, characterized by, The method comprises the following steps: 1) determining the quality traits required for breeding a sorghum variety according to literature or experience, and selecting sorghum varieties with and without the required quality traits; 2) obtaining gene expression of the transcriptome of the sorghum varieties by sequencing; 3) preprocessing the gene expression data of the transcriptome, and performing quantile normalization on the gene expression matrix to make the samples comparable: including deleting duplicate data; using normal distribution to determine that gene expression values greater than 3δ are abnormal values, and replacing the abnormal values with the maximum value other than the abnormal values; and taking logarithmic processing on the data; 4) performing feature selection on the preprocessed data using a genetic algorithm for each quality trait: including that the original feature set is the entire set of genes obtained by sequencing; N genes are randomly selected as the initial feature gene set after genetic algorithm initialization; the feature gene set is sorted according to the importance given by the random forest algorithm using the random forest algorithm to evaluate the feature gene set, the importance is used as the fitness value, and the final feature gene set is the excellent gene set of the sorghum trait; 5) taking the excellent gene set of all traits selected in 4) as the excellent gene set of sorghum; The quality traits include one or more of the following: plump grains, high yield, low tannin content, high protein content, high starch content, and high trace element content; Sorghum varieties with the required quality traits are selected as positive samples for feature gene selection, and sorghum varieties without the required quality traits are selected as negative samples for feature gene selection, and the positive and negative samples are subjected to transcriptome sequencing.
2. The method of claim 1, wherein the excellent genes are selected from the group consisting of SEQ ID NOs: 1 to 22. In step 1), search the database with sorghum and the required quality traits as the combination field to determine sorghum varieties with and without the required quality traits.
3. The method of claim 2, wherein the excellent genes are selected from the group consisting of SEQ ID NOs: 1 to 6. The sorghum varieties are as consistent as possible in other traits except for having and not having the required quality traits.
4. The method of claim 1, wherein the excellent genes are selected from the group consisting of SEQ ID NOs: 1 to 22. The coding method of the genetic algorithm is 0 / 1 coding, that is, selecting gene i is 1, and not selecting the gene is 0, and thus the chromosome coding is a 0 / 1 sequence, and the population is a plurality of 0 / 1 sequences; the selection method adopts a fitness proportion method to calculate the selection probability; the crossover probability is set to 0.5, single-point crossover is adopted, and the mutation rate is 0.0002.
5. The method of claim 4, wherein the excellent genes are selected from the group consisting of SEQ ID NOs: 1 to 20. The fitness-proportional method calculates the selection probability as: where i is a certain gene, f i is the fitness, p si is the probability of being selected.
6. The method of claim 1, wherein the excellent genes are selected from the group consisting of SEQ ID NOs: 1 to 6. The random forest algorithm adopts 10-fold cross-validation to evaluate the feature set by the average classification accuracy.
7. The method of claim 1, wherein the excellent genes are selected from the group consisting of SEQ ID NOs: 1 to 6. The selected gene sets of each excellent quality trait are taken as the final excellent quality gene set of sorghum.
Citation Information
Patent Citations
Method and application of pan-tumor targeted drug sensitivity state evaluation model constructed based on high-throughput sequencing data and clinical phenotypes
CN111640508A
Mutant gene classification method based on NLP
CN115186769A