Method for quickly exploring quantitative trait genes of rice

By scoring the matching degree between QTL mapping results of multiple subpopulations and candidate gene variant sites in rice NAM populations, and combining gene function and expression information, rapid and accurate mapping of quantitative trait genes in rice was achieved, solving the problems of long cycle and low efficiency in existing technologies and improving gene discovery efficiency.

CN116052771BActive Publication Date: 2026-01-13SHANGHAI ZKW BREEDING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310058245.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-13
Publication Date
2026-01-13
Estimated Expiration
2043-01-13

AI Technical Summary

Technical Problem

Existing technologies are time-consuming and inefficient in the process of locating quantitative trait genes in rice, making it difficult to quickly discover important genes.

Method used

By cross-validating QTL mapping results from multiple subpopulations in the rice NAM population and combining information such as gene expression levels, histone modifications, and -log10P values ​​of associated QTL sites for weighted scoring, valuable candidate genes were screened out, and a rapid mapping method was established.

Benefits of technology

It significantly improved the efficiency of quantitative trait gene discovery in rice, shortened the mapping time from more than 2 years to less than 1 year, and reduced the need to construct secondary mapping populations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052771B_ABST
    Figure CN116052771B_ABST
Patent Text Reader

Abstract

The application discloses a method for quickly exploring quantitative trait genes of rice, and the method comprises the following steps: performing whole genome association analysis on a rice nested linkage mapping population and linkage analysis on each subpopulation, and exploring quantitative trait loci (QTL) which can be identified by both methods; comprehensively analyzing the function, variation and expression of candidate genes in the QTL; performing weighted scoring and sorting on the number of keyword matches in the annotation information of the candidate genes, the influence of the variation sites of the candidate genes on the gene function, the expression amount of the candidate genes in specific tissues and the like, and selecting three genes with the highest scores as the candidate genes of the target quantitative trait.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biology, in particular to the field of plant genetics and genetic engineering. BACKGROUND

[0002] Rice is one of the most important food crops in China, and rice is the staple food of more than half of the population. Important agronomic traits of rice, such as yield, quality, and resistance, are quantitative traits, which are generally controlled by multiple quantitative trait genes. Cloning of rice quantitative trait genes and analysis of the genetic mechanism of rice quantitative traits are important foundations for rice molecular breeding.

[0003] Some cloned rice quantitative trait genes have been successfully applied to rice molecular breeding, such as the rice amylose content control gene Waxy and the rice blast resistance gene Pigm. Studies have found that the Waxy gene encodes a granular starch synthase enzyme. In japonica rice, a base mutation occurs in the intron splicing site of the Waxy gene, causing the intron to fail to splice normally, weakening the gene function, reducing the amylose content of rice, and making the rice taste soft. The Pigm gene is a NBS-LRR gene cloned from the local variety "Gu Mei 4", which can form a homodimer and has strong broad-spectrum resistance to rice blast. Waxy and Pigm, and other rice quality and resistance-related genes, have been applied to rice molecular breeding, significantly improving the quality and disease resistance of target varieties.

[0004] Currently, rice quantitative trait genes are mainly located by linkage analysis and association analysis, and further fine mapping is performed using map-based cloning methods. For example, the rice heat tolerance gene TT1 was fine mapped using map-based cloning. The F2 population was constructed by selfing the cross between heat-tolerant and heat-intolerant parents, and TT1 was initially located on chromosome 3 by linkage analysis. Subsequently, a BC4F2 population containing 6721 single plants was constructed by multiple backcrosses, and TT1 was located within a 12.69 kb range containing 2 candidate genes using the secondary population. The map-based cloning method can accurately locate the target gene within a few tens of kb, laying a foundation for further cloning and verification of candidate genes. However, this method requires a long time to construct a secondary mapping population containing a large number of lines, resulting in a long research cycle and limiting the rapid discovery of important rice quantitative trait genes. Therefore, there is an urgent need to develop a rapid rice quantitative trait gene mapping method.

[0005] The present application innovates the evaluation method of rice gene function and variation, cross- validates multiple population QTL mapping results, and combines gene expression, histone modification, and the -log 10The P value and the like are weighted and scored to establish a rapid quantitative trait gene positioning method of the rice NAM population, rapid and accurate positioning of the quantitative trait gene of rice is realized, and the efficiency of the important quantitative trait gene exploration of rice is improved. SUMMARY

[0006] The purpose of the present application is to provide a method capable of rapidly exploring quantitative trait genes of rice. Through a large amount of research, the present application proposes a method of scoring the matching degree of QTL positioning results of multiple subpopulations in a rice NAM population and candidate gene variation sites, and developing a variation scoring standard according to the position, size and effect of the candidate gene variation site. In addition, the function of the candidate gene is scored by searching the gene annotation information according to the target trait keywords, and the candidate gene expression amount, histone modification information and QTL site-log 10 P value and the like are weighted and scored to rapidly explore quantitative trait genes of rice, and the efficiency of the quantitative trait gene exploration of rice is significantly improved.

[0007] The present inventors have found, after a large amount of research, that for a candidate gene, the valuable candidate gene can be screened by scoring and sorting the information such as variation position, size, genetic effect, the number of times of matching the keywords and gene annotation information of the homologous gene related literature, the ratio of the expression amount of the candidate gene in the target tissue to that in other tissues, the matching of the parent variation and the QTL site, and whether there is histone modification.

[0008] Specifically, the present application provides a method for rapidly exploring quantitative trait genes of rice, which comprises the following steps:

[0009] Step 1: Based on the genotypes and phenotypes of the rice NAM population, perform whole genome association analysis to explore quantitative trait loci significantly related to the quantitative trait phenotype, denoted as association QTL loci; and perform linkage analysis on each subpopulation of the mapping population to explore quantitative trait loci linked to the quantitative trait phenotype in each subpopulation, denoted as linkage QTL loci; based on the proximity of the positions of the association QTL loci and the linkage QTL loci, select candidate QTL loci, and all genes within all candidate QTL loci are used as candidate genes;

[0010] Step 2: Whole genome sequencing, assembly and alignment of the population parents of the rice NAM population, mining of variation sites different from the reference genome Nipponbare, and determination of the location and variation type of the variation sites on the genes; collection of published gene information and related research paper keywords in rice and Arabidopsis thaliana, and annotation of the gene functions in the entire genome according to the homologous genes; transcriptome sequencing of the population parents and collection of rice gene expression data in public databases to obtain gene expression data; collection of eChIP-Seq data of rice tissues to analyze the H3K27me3 histone modification information of the candidate genes of rice;

[0011] Step 3: Exclusion of candidate genes without base difference between the parents of the population in which the linked QTL is located; weighted scoring according to the genetic effect size of the variation of the candidate genes, the number of matching times of the target phenotype-specific keywords in the annotation information of the candidate genes, the expression amount of the candidate genes in the specific tissues of the target traits, whether the variation of the candidate genes in different parents completely matches the QTL, the significant association level of the QTL gene in which the candidate gene is located, and the histone modification information of the candidate gene, and selection of a number of genes with the highest scores as the candidate genes to be verified of the target quantitative trait.

[0012] In a preferred implementation manner,

[0013] The step 1 comprises:

[0014] Step 1.1: selecting a plurality of rice materials with obvious differences in the target quantitative trait as parents, wherein 1 rice variety is used as a common parent to cross with other materials and then self-crossed for several generations to construct a rice nested linkage mapping NAM population comprising a plurality of sets of recombinant inbred line populations, and performing whole genome association analysis on the entire NAM population to mine QTL sites significantly associated with the target trait;

[0015] Step 1.2: performing linkage analysis on each recombinant inbred line population in the NAM population to mine QTL sites significantly linked to the target trait;

[0016] Step 1.3: selecting QTL sites with the same or similar positions in the two types of sites as candidate QTL sites according to the positions of the significantly associated and linked QTL sites, and selecting all genes in the QTL sites as candidate genes.

[0017] In another preferred implementation manner,

[0018] The step 2 comprises:

[0019] Step 2.1: Whole genome sequencing and RNA sequencing of the population parents using second or third generation sequencing technology, comparing the genomes of all parents with the reference genome to determine the difference between the two variation sites, according to the location of the variation site, screening the variation sites in the coding region and non-coding region of the candidate gene; compare the variation sites of the common parent with the variation sites of other parents, and find the variation sites different from the common parent from the non-common parent;

[0020] Step 2.2: Collect all gene information and related research papers published in rice and Arabidopsis, according to the annotation of homologous genes and related paper information, the function of each candidate gene is annotated, and the functional annotation text of the candidate gene is formed;

[0021] Step 2.3: Collecting transcriptome data of different tissues of rice from public database to obtain rice gene expression information;

[0022] Step 2.4: Collecting eChIP-Seq data of rice tissues to analyze H3K27me3 histone modification information of rice candidate genes.

[0023] In another preferred implementation, the reference genome is the rice variety "Nipponbare" genome (MSU7.0), and the variation sites include single nucleotide variation sites (SNP), insertion and deletion (InDel) and structural variation (SV).

[0024] In another preferred implementation, the step 3 includes:

[0025] Step 3.1: Scoring the effect of the variation on the function of the candidate gene according to the information of the variation, the size of the variation fragment, the position of the variation, etc.

[0026] Step 3.2: Collecting published papers related to target traits in rice, Arabidopsis, corn and other species, and summarizing the keywords of a certain trait-related gene; using these keywords to search the annotation text of the candidate gene, and scoring the gene function of the candidate gene according to the number of keyword matches;

[0027] Step 3.3: Comparing the expression amount of the candidate gene in the target tissue with the average expression amount in other tissues, and scoring the gene expression amount of the candidate gene according to the ratio;

[0028] Step 3.4: Scoring the genetic effect and reliability of the candidate gene according to the -log 10 P value of the significantly associated QTL site;

[0029] Step 3.5. Scoring the matching degree between the candidate gene variation and the population QTL according to whether the variation of the candidate gene in different parents matches the QTL of the sub-population;

[0030] Step 3.6. Scoring whether the candidate gene is involved in the regulation of agronomic traits based on whether the candidate gene is a transposon gene;

[0031] Step 3.7. Adding or weighted adding the scores obtained in steps 3.1-3.6, and ranking the candidate genes according to the total scores of the candidate genes, and selecting a plurality of candidate genes based on the high-low scores.

[0032] In another preferred implementation, the scoring rules include comprehensive scoring of the candidate gene variation according to the size, position and effect of the variation site of the candidate gene, scoring of the function of the candidate gene according to the matching number of different keywords summarized in a large number of literatures for different traits and the gene annotation information, scoring of the expression level of the candidate gene by selecting key tissues for different traits and comparing the expression levels of the key tissues and other tissues, cross-validation scoring of multiple populations by using whether the QTL sites from different populations and the variation sites of the candidate gene in different parents are the same, scoring of the importance of the gene by using histone modification information, and scoring of the reliability of the candidate gene according to the-log 10 P value of the associated QTL site.

[0033] The application further provides an application of the method, and the application is to quickly mine quantitative trait genes in the rice NAM population.

[0034] The application further provides a gene mined by using the method, and the gene sequence is shown in the sequence table SEQ ID No. 1.

[0035] Technical effects

[0036] Before the method of the application is proposed, the rice quantitative trait gene mining mainly depends on the method of map-based cloning, and the cycle of mining the quantitative trait gene is long and the efficiency is low, and after the method of the application is used, the time for fine mapping of the QTL can be reduced from more than 2 years to less than 1 year.

[0037] The method of the application is suitable for the discovery of quantitative trait genes of rice or other plants. By matching the QTL positioning results of multiple subpopulations with candidate gene variation sites and scoring the functions of the variation sites according to the positions, sizes and effects of the candidate gene variation sites, and by searching the gene annotation information according to the target trait keywords to score the functions of the candidate genes, the target genes can be scored more accurately, so that the genes most closely related to quantitative traits in plants, especially rice, can be found, the number of genes to be verified can be reduced, and the secondary crop population does not need to be constructed, thereby improving the efficiency of the discovery of quantitative trait genes of rice.

[0038] The method of the application has the advantage of shortening the period of discovering candidate genes. Compared with the published cloning of quantitative trait genes of rice, the method of the application does not need to construct a secondary mapping population for fine mapping, thereby saving a large amount of time.

[0039] Preferably, the method of the application can be applied to various crops with reference genome sequences and capable of constructing a primary artificial population. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 Development process of the method for rapid discovery of quantitative trait genes of rice;

[0041] Figures 2-4 Functional verification process of the OsFTL1 gene of rice, wherein, Figure 2 Whole-genome association analysis result of the heading date; Figure 3 Whole-genome linkage analysis result of the heading date; Figure 4 Result of the functional verification of the genes, the left graph is the result of knocking out the OsFTL1 gene in Huanghuazhan, and the right graph is the result of complementing the OsFTL1 gene of Huanghuazhan in Kasalath. The results show that knocking out the OsFTL1 gene in Huanghuazhan makes Huanghuazhan heading 3-4 days earlier, and complementing the OsFTL1 gene of Huanghuazhan in Kasalath makes Kasalath delay heading 2-3 days. DETAILED DESCRIPTION

[0042] The application will be described in further detail below with reference to the embodiments and the accompanying drawings, but the embodiments of the application are not limited thereto.

[0043] The application will be described below with the rapid discovery of quantitative trait genes of the NAM population of rice as an example.

[0044] Example 1: Development of the method for rapid discovery of quantitative trait genes of the NAM population of rice

[0045] In this embodiment, the development process of the rice quantitative trait gene rapid exploration technology will be described. In general, first, linkage and association analysis of quantitative traits are performed on the rice NAM population respectively, and linkage QTL sites and association QTL sites are obtained respectively. The association QTL sites with the same position as the linkage QTL sites are taken as candidate QTL sites. The genes contained in the QTL sites are scored according to the degree of the influence of the variation on the function of the candidate gene, the number of times of matching the target trait keywords with the candidate gene annotation, and the expression amount of the candidate gene in the target tissue, and so on. The candidate genes are sorted according to the total score, and the top three genes with the highest scores are taken as the candidate genes. Preferably, when scoring, the score is comprehensively scored based on the weight of different subjects and the influence of the gene on the subject.

[0046] Specifically, the rice quantitative trait gene rapid exploration technology is realized through the following steps:

[0047] Step 1.1: 16 rice materials with obvious differences in target quantitative traits are selected as parents, and one rice variety “Huanghuazhan” is used as a common parent to cross and self-cross for more than 6 generations to construct a rice NAM population containing 15 sets of recombinant inbred lines. Three generations of Nanopore sequencing are performed on the 16 parents, and after assembly, high-precision genome sequences at the chromosome level are obtained. Using Nipponbare as the reference genome, Mummer software is used for pairwise alignment of Nipponbare genome and parent genome to obtain high-accuracy variation information (including SNP, Inde, SV); 0.05-fold low-coverage sequencing is performed on the 15 sets of recombinant inbred lines. Continue to use Nipponbare genome as the reference genome, and use Segmap process to obtain the genotype information of each set of population. Then integrate the genotypes of the 15 sets of recombinant inbred lines, unify the position information, and obtain the genotype matrix of the whole offspring population. Finally, based on this, the identification of the target quantitative trait is realized (in this embodiment, the heading date is taken as an example).

[0048] Step 1.2: Perform whole genome association analysis of the target trait on the whole NAM population using whole genome association analysis method (e.g. FarmCPU model of rMVP package), obtain SNP sites significantly associated with the target trait of heading date of the whole NAM population, remove nearby highly linked SNP signal sites. After deduplication, obtain significant SNP sites strongly associated with the target trait of heading date. Cycle all significant SNP sites, take the position of the SNP site as the interval of 200 kb before and after the associated QTL site, at the same time, combine whole genome linkage analysis (e.g. composite interval mapping method of WinQTLcart) to perform linkage analysis on 15 subpopulations in the whole population respectively, and explore the linkage QTL sites related to heading date. Screen from the associated QTL sites, select the associated QTL sites meeting the following conditions as candidate QTL sites: there is a linkage QTL site within 500 kb before and after the associated QTL site. All genes within the candidate QTL sites are selected as candidate genes.

[0049] Step 1.3: Obtain high-quality variation sites different between the other parent and the common parent by step 1.1. Judge each candidate gene, remove the candidate gene without variation sites within the coding region and the promoter within 2 kb. Compare the positions of QTL sites in different subpopulations, and explore the QTL sites with the same position. Compare the parents with the same QTL sites, whether there is at least one parent with variation in the candidate gene, if there is no variation, remove the candidate gene. If the same candidate interval is located in multiple linkage populations, it is proved that the interval has high credibility, and the existence of correct candidate gene also has high possibility, so the gene with coding region difference between parents in the interval also has high credibility.

[0050] Step 1.4: For each candidate gene, search the same gene and homologous gene information from public databases such as RAPDB, MSU, TAIR, etc., annotate the candidate gene based on the functional description contained in the same gene and homologous gene information to form an annotation text. Then, according to the research target (target trait), obtain the papers related to the research, and extract the keywords therein. Use the extracted keywords to search the annotation of the candidate gene and the related research on the candidate gene, and score the gene function according to the number of matches between the keywords and the annotation of the candidate gene and the related research content of the candidate gene. Taking the heading date research as an example, collect and sort the published QTL genes related to the heading date of rice or the flowering period of plants in rice, Arabidopsis and other species, and extract the keywords therein, including flowering, photoperiod, light receptor, circadian rhythms, diurnal rhythm, biological clock, heading date, florigen, CCT, CONSTANS, FT, MADS, flowering plant, etc. Use the above keywords to search the annotation of the candidate gene and the title and abstract of the related papers of the candidate gene, and score the gene function according to the number of matches between the keywords and the two contents. The more the number of matches, the higher the score. The scoring in this step and the scoring in each of the following steps can construct a corresponding mapping table or a relationship function. For example, a mapping table between the number of matches and the score is constructed (Table 1), and the scoring based on the matching degree is performed based on the mapping table.

[0051] Table 1 Scoring table for matching target trait keywords with candidate gene annotation

[0052]

[0053]

[0054] Step 1.5: Collect 17 types of rice tissues in NCBI, such as leaves, branch buds, flowers, roots, etc., take Nipponbare as the reference genome, align and normalize the transcriptome data, and calculate the expression amount (FPKM) of the candidate gene in different tissues. Compare the ratio of the expression amount of the candidate gene in the key tissue related to the target trait to the average expression amount in other tissues, and score the gene expression amount according to the ratio. The higher the ratio, the higher the score. When the expression amount of the candidate gene in the target tissue is less than 0.05, the expression score is 0; when the expression amount of the candidate gene in the target tissue is greater than 0.05, if its expression amount is more than 10 times, 5 times, 2 times or 1 time of the average expression amount in non-target tissues, the score is 22, 20, 18 or 15, respectively; and if the expression amount is less than or equal to 1 time of the average expression amount in non-target tissues, the score is 10.

[0055] Step 1.6: Variations in coding and non-coding regions are annotated using methods such as SNPeffect, SIFT, NCBI-CDD, etc. to determine the impact of variations in candidate genes on gene function. In the coding region, scoring is performed according to whether it causes loss of gene function, the size of the SIFT value of missense mutation, whether it is located in an important module, the length of structural variation, etc. In the non-coding region, scoring is performed according to the length of the structural variation, whether it is in the 5'UTR and 3'UTR region, whether it is an intron splicing site, whether it is in the chromatin open region, whether it is in the promoter regulatory element, etc. Each condition corresponds to a different score, which is pre-set through a mapping table. The greater the impact of the variation on the gene function, the higher the score (Table 2). The score with the higher score in the two groups of scoring results according to the coding region and the non-coding region is taken as the effect score of the candidate gene variation site.

[0056] Table 2: Candidate gene variation scoring table

[0057]

[0058]

[0059]

[0060] Step 1.7: According to the significant association QTL site given by the FarmCPU model, the -log 10 P value of the candidate gene is scored, and the score is scored according to the number of sites determined, -log 10 P value. The higher the -log 10 P value, the higher the score. If there are more than five signal peaks, the score is divided into five levels, each level is 2 points, if there are less than five, the total score is 10 points, and the score is divided according to the number of peaks. The score of each peak is determined according to which gradient it is located in, and the corresponding score is added. If it is close to the significant association site given by the MLM model, additional points are added, -log 10 P greater than 20, 10 points; 10 to 20, 5 points; less than 10, 0 points. The score of the genetic effect or reliability of the candidate gene is determined.

[0061] Step 1.8: Research has shown that genes with inhibitory histone modifications usually have a greater impact on traits, and are judged to be important genes. eChIP-Seq data of four types of tissues in rice are collected to analyze the H3K27me3 histone modification information of the candidate genes in rice, which is used to determine whether the candidate gene is an important gene. If the candidate gene is modified by H3K27me3 histone, it is considered that the expression of the gene is inhibited, and it may be an important gene, and a higher score is given, otherwise a lower score is given.

[0062] Step 1.9: For the candidate genes with different variation sites in different parents, multiple alleles are combined. Each candidate gene is scored at the gene level, and whether the variation of the multiple alleles matches the QTL site obtained from the subpopulation linkage analysis is scored. If it is completely matched, an additional 5 points are added, and if it is not matched, no points are added.

[0063] Step 1.10: According to the functional annotation of the rice genome on the Rice Genome Annotation Project, it is judged whether the candidate gene is a transposon gene. If it is a transposon gene, all scores are cleared, otherwise, the scores obtained in other steps are retained.

[0064] Step 1.11: The scores of steps 1.4 to 1.10 are combined, the total score of all candidate genes in the QTL interval is added or weighted, and the results are sorted, and the top 3 genes are selected as the candidate genes of the rice target quantitative trait Figure 1 ).

[0065] Example 2 Rapid mining and verification of rice heading date genes

[0066] In this embodiment, the rice quantitative trait gene rapid mining technology is used to mine the rice heading date genes. In general, the rice quantitative trait gene rapid mining method developed in Example 1 is used to mine the heading date related genes in the rice NAM population, and it is analyzed whether the cloned major heading date genes can be located, whether the located heading date genes belong to the top 3 candidate genes in the corresponding QTL site, and one of the QTL sites is cloned and verified.

[0067] Specifically, the rice quantitative trait gene rapid mining technology is realized by the following steps:

[0068] Step 2.1: The method in Example 1 is used to rapidly mine the heading date genes of the rice NAM population, and 26 heading date related QTL sites are obtained, including 78 candidate genes.

[0069] Step 2.2: The variation of the cloned rice heading date QTL genes in the rice parents is analyzed, and it is found that Ghd7, Ghd7.1, Ghd8, OsSOC1, Ehd1, Hd1, RFT1, Hd3a, Hd16, OsMADS51, OsMADS56, and Ef-cd genes have variations in the NAM population parents. The analysis of QTL sites and found that all these genes are located in the mined QTL sites. The analysis of the candidate genes found that 91.7% of them belong to the top 3 candidate genes, indicating that the rice quantitative trait gene rapid mining method can be well applied to the mining of rice heading date QTL genes.

[0070] Step 2.3: A main-effect QTL site for heading date was found near 6.5 Mb of rice chromosome 1. According to the scoring, LOC_Os01g11940 (the sequence of which is shown in the sequence listing), LOC_Os01g12240, and LOC_Os01g11960 were the top three candidate genes, with total scores of 67, 55, and 52, respectively. Among them, LOC_Os01g11940 is the homologous gene OsFTL1 of the rice florigen gene Hd3a. Analysis found that OsFTL1 in the common parent "Huanghuazhan" of the NAM population was wild type, while there was a frameshift mutation in the coding region of OsFTL1 gene in the parents Kasalath and Basmati.

[0071] Step 2.4: Two CRISPR / Cas9 target sites were designed on the first exon of OsFTL1 (OsU3T1: 5'-aCAGCTGCTGTACCCTCGCCG-3'; OsU6aT2: 5'-gTAGGGCGCATGAGCGATGAG-3'), and an OsFTL1 gene editing binary vector was constructed using the pYLCRISPR / Cas9Pubi-H vector as the backbone and transformed into "Huanghuazhan" through Agrobacterium tumefaciens to knock out the OsFTL1 gene. An 8Kb OsFTL1 gene in the "Huanghuazhan" background was amplified by PCR (FTL1 prog F: 5'-GATATCCAGATCCAGTGGGAAAGTAGCAGCCACCAACTACCG-3'; FTL1 prog R: 5'-AGCGGCCGCACTAGTAAGCATTCTTCTTCCTCCGGTTCCAG-3') and integrated into the pRHEcMyc vector by homologous recombination to construct an OsFTL1 complementary binary vector (OsFTL1 pro HHZ :OsFTL1 g HHZ -cMyc), and the "Huanghuazhan" OsFTL1 gene was introduced into "Kasalath" through Agrobacterium tumefaciens. Homozygous gene knockout lines (HHZ-osftl1) and positive complementary lines (Kas-OsFTL1 HHZ ) were identified from the knockout and complementary lines, respectively, and the heading dates of T2 generation knockout and complementary plants were counted. It was found that the gene knockout material was 3.8 days earlier than "Huanghuazhan" in heading date, and the complementary line was 2.6 days later than "Kasalath" in heading date Figure 4 ).

[0072] In summary, OsFTL1 was found to be a candidate gene of heading date QTL by QTL mapping, and the gene function was verified for the first time by gene editing and functional complementation. The results show that OsFTL1 gene is a QTL gene for controlling the heading date of rice. The above results show that the method of rapid exploration of quantitative trait genes of rice can quickly explore the heading date QTL gene of rice.

[0073] Although the principles of the present application have been described in a detailed manner above in connection with preferred embodiments thereof, it is to be understood that many modifications, equivalents, and alternatives to the above-described embodiments will be apparent to those of ordinary skill in the art, and the above-described embodiments should not be construed to limit the scope of the present application. The details in the embodiments do not constitute a limitation on the scope of the present application, and any obvious changes, simple replacements, etc. based on the technical solutions of the present application, without departing from the spirit and scope of the present application, fall within the protection scope of the present application.

Claims

1. A method for rapid discovery of quantitative trait genes in rice, characterized in that, The method includes the following steps: Step 1: Based on the genotype and phenotype of the rice NAM population, perform genome-wide association analysis to identify quantitative trait loci that are significantly associated with the quantitative trait phenotype, and denote them as associated QTL loci; and perform linkage analysis on each subpopulation of the mapping population to identify quantitative trait loci linked to the quantitative trait phenotype in each subpopulation, and denote them as linked QTL loci. Based on the similarity of sites among linked QTL sites, candidate QTL sites are selected, and all genes within all candidate QTL sites are selected as candidate genes. Step 2: Perform whole-genome sequencing, assembly, and alignment on the parents of the rice NAM population to identify variant sites that differ from the reference genome of Nipponbare rice, and determine the location and type of these variant sites on the genes; collect published gene information and keywords from relevant research papers in rice and Arabidopsis thaliana, and annotate gene functions throughout the genome based on homologous genes; perform transcriptome sequencing on the parents and collect rice gene expression data from public databases to obtain gene expression levels; collect eChIP-Seq data from rice tissues and analyze H3K27me3 histone modification information of rice candidate genes; Step 3: Exclude candidate genes that show no base differences between parents in the linked QTL population; weighted scores are applied based on the magnitude of the genetic effect of candidate gene variations, the number of matches between target phenotype-specific keywords in the candidate gene annotation information, the expression level of the candidate gene in the target trait-specific tissue, whether the variation of the candidate gene between different parents completely matches the QTL, the significant association level of the candidate gene with the QTL gene, and the histone modification information of the candidate gene. The genes with the highest scores are selected as candidate genes to be validated for the target quantitative trait. Step 3 includes: Step 3.1 Score the variation effect of candidate genes based on the impact of the variation on gene function, the size of the variation fragment, and the location of the variation. Step 3.2 Collect published papers related to the target trait from species such as rice, Arabidopsis thaliana, and maize, and summarize the keywords of genes related to a certain trait; use these keywords to search the annotation texts of candidate genes, and score the gene function of candidate genes based on the number of keyword matches; Step 3.3 Compare the expression levels of candidate genes in the target tissue with the average expression levels in other tissues, and score the gene expression levels of candidate genes based on the ratios; Step 3.4 Based on the -log of significantly associated QTL sites 10 The P-value scores the genetic effect and reliability of candidate genes; Step 3.5 Score the matching degree between candidate gene variants and population QTLs based on whether the variants of candidate genes match the QTLs of the subpopulation in different parents; Step 3.6 Based on whether the candidate gene is a transposon gene, score the candidate gene to determine whether the candidate gene is involved in the regulation of agronomic traits; Step 3.7 Add up or weighted sum the scores obtained in steps 3.1-3.6, sort the candidate genes according to their total scores, and select a number of candidate genes based on their scores.

2. The method for rapid discovery of quantitative trait genes in rice according to claim 1, characterized in that, Step 1 includes: Step 1.1: Select multiple rice materials with significant differences in the target quantitative trait as parents. One rice variety is used as a common parent and hybridized with other materials. After several generations of continuous self-pollination, a rice nested composite mapping (NAM) population containing multiple sets of recombinant inbred lines is constructed. Genome-wide association analysis is performed on the entire NAM population to identify QTL loci that are significantly associated with the target trait. Step 1.2: Perform linkage analysis on each recombinant inbred line population in the NAM population to identify QTL sites that are significantly linked to the target trait; Step 1.3: Based on the location of significantly associated and linked QTL sites, select QTL sites that are in the same or similar locations in both types of sites as candidate QTL sites, and select all genes within the QTL sites as candidate genes.

3. The method for rapid discovery of quantitative trait genes in rice according to claim 1, characterized in that, Step 2 includes: Step 2.1: Use second- or third-generation sequencing technology to perform whole-genome sequencing and RNA sequencing on the parent population. Compare the genomes of all parents with the reference genome to identify the variant sites that differ between the two. Based on the location of the variant sites, screen for variant sites in the coding and non-coding regions of candidate genes. Compare the variant sites of the common parents with the variant sites of other parents to discover variant sites that differ from the common parents from the non-common parents. Step 2.2: Collect all published gene information and related research papers in rice and Arabidopsis thaliana. Based on the annotation of homologous genes and related paper information, annotate the function of each candidate gene to form a functional annotation text of the candidate genes; Step 2.3: Collect transcriptome data from different rice tissues from public databases to obtain rice gene expression information; Step 2.4: Collect eChIP-Seq data from rice tissues and analyze the H3K27me3 histone modification information of rice candidate genes.

4. The method for rapid discovery of quantitative trait genes in rice according to claim 1, characterized in that, The reference genome is the genome of the rice variety "Nipponbare" (MSU7.0), and the variant sites include single nucleotide variants (SNPs), insertions / deletions (InDel), and structural variations (SVs).

5. The application of the method according to any one of claims 1-4, wherein the application is to rapidly discover quantitative trait genes in rice NAM populations.